本篇博文主要内容为 2026-09-30 从Arxiv.org论文网站获取的最新论文列表,自动更新,按照NLP、CV、ML、AI、IR、MA六个大方向区分。

说明:每日论文数据从Arxiv.org获取,每天早上12:30左右定时自动更新。

提示: 当天未及时更新,有可能是Arxiv当日未有新的论文发布,也有可能是脚本出错。尽可能会在当天修复。

目录

概览 (2026-09-30)

今日共更新1448篇论文,其中:

  • 自然语言处理共229篇(Computation and Language (cs.CL))
  • 人工智能共509篇(Artificial Intelligence (cs.AI))
  • 计算机视觉共284篇(Computer Vision and Pattern Recognition (cs.CV))
  • 机器学习共465篇(Machine Learning (cs.LG))
  • 多智能体系统共32篇(Multiagent Systems (cs.MA))
  • 信息检索共29篇(Information Retrieval (cs.IR))
  • 人机交互共37篇(Human-Computer Interaction (cs.HC))

多智能体系统

[MA-0] Multi-Agent Flow Matching with Decoupled Generative Guidance

【速读】:该论文旨在解决多智能体生成模型中难以满足硬性约束条件的问题,尤其在存在跨智能体依赖关系的共享需求或个体智能体与邻近智能体相关的私有需求时,传统生成方法缺乏有效的协同引导机制。其核心挑战在于:每个智能体需独立确定自身生成引导输入,而不能依赖其他智能体同时计算出的引导信息,从而导致整体系统难以保证全局约束的满足。为此,论文提出DeGG-Flow(Decoupled Generative Guidance Flow)框架,通过将生成过程建模为控制仿射动力系统(control-affine dynamical system),设计了两类耦合需求的引导条件——一类是依赖多个智能体共同满足的共享需求,另一类是每个智能体基于邻居状态的私有需求。针对这两类需求,研究建立了可行性条件与有限时间收敛性保证,并推导出刻画引导策略引入分布偏差的Wasserstein界。实验验证表明,DeGG-Flow在多机器人协作跨越空间间隙(通过重构环境实现)以及具有功能属性要求的多物体场景生成任务中,均能直接生成满足所有硬性约束的样本,且具备对训练时未见过的团队规模的泛化能力。

链接: https://arxiv.org/abs/2609.38133
作者: Ruoyu Lin,Magnus Egerstedt,Fabio Pasqualetti
机构: 未知
类目: Machine Learning (cs.LG); Multiagent Systems (cs.MA); Robotics (cs.RO); Optimization and Control (math.OC)
备注:

点击查看摘要

Abstract:Generative modeling is widely used for producing diverse objects from complex, multimodal distributions. However, its expressivity does not, in general, come with formal guarantees that the generated objects satisfy hard constraints or requirements. In multi-agent generation, this problem becomes more challenging because a hard requirement can depend on multiple agents, while each agent may need to determine its own guidance input without relying on the simultaneously computed guidance inputs of other agents. To this end, we introduce DeGG-Flow, a general framework for multi-agent flow matching with decoupled generative guidance. By representing the generative process as a control-affine dynamical system, we develop guidance conditions for two classes of coupled requirements: shared requirements whose satisfaction depends on multiple agents together, and private requirements associated with each individual agent dependent on its neighbors. For both classes, we establish feasibility conditions and finite-horizon convergence guarantees. We further derive a Wasserstein bound that characterizes the distributional deviation induced by the guidance. We demonstrate DeGG-Flow on multi-robot collaboration for crossing a spatial gap by reconfiguring the environment, and on multi-object scene generation with affordance requirements. Across both applications, DeGG-Flow directly generates objects that satisfy all corresponding hard requirements, including at team sizes unseen during training.

[MA-1] IMPACT: Modeling Socially Interdependent Movement in a Generative Multi-Agent Simulation of a Pompeian Household

【速读】:该论文旨在解决生成式多智能体模拟中难以准确刻画人类集体行为依赖性的问题,尤其关注历史空间使用过程中个体行动如何受社会关系与文化规范制约。现有方法中的智能体多独立规划与行动,无法充分反映真实社会情境下移动行为对他人动作的依赖性。其解决方案的关键在于提出IMPACT(基于智能体间约束与触发的相互依赖移动规划)架构,通过嵌入文化特定的角色与义务机制,建立智能体间活动的依赖关系,并整合社会约束下的里程碑规划、等待或提示决策、结构化指令发布及指令融合等机制,动态调节活动的启动与变更时机。该架构使模拟产生以社会条件驱动的约束性与提示性移动模式,从而更真实地再现社会互动下的空间实践。在庞贝城五小时晚宴的十智能体仿真中,结果揭示了角色、责任与地位关系如何影响共处、非对称等待、协同移动及社会指令等行为模式;控制性消融实验表明,完整架构在社会一致性与可信度上显著优于简化版本,且专家访谈证实其具备支持考古学解释的潜力,同时指出了未来需加强历史依据的需求。

链接: https://arxiv.org/abs/2609.38113
作者: Tianqi Liu,Nayoung Kim,Julia Sebastien,Kathryn Gleason,Caitlín Eilís Barrett,Andrea Stevenson Won
机构: Cornell University (康奈尔大学)
类目: Multiagent Systems (cs.MA)
备注:

点击查看摘要

Abstract:Simulations of archaeological sites can make interpretations of past cultural practices observable and examinable. Generative multi-agent simulations offer a bottom-up approach to modeling how people collectively moved through and used historical spaces. However, current agents designed to simulate everyday life often plan and act independently, limiting their ability to capture how movement depends on others’ actions. We introduce IMPACT (Interdependent Movement Planning through Inter-Agent Constraints and Triggers), an architecture that uses culturally specific roles and obligations to define dependencies among agents’ activities and guide coordination. IMPACT connects socially gated milestone planning, wait-or-prompt resolution, structured directive issuance, and directive integration. These mechanisms determine whether and when activities can begin or change as social conditions evolve, producing socially constrained and prompted movement as their primary observable outcome. We instantiate IMPACT in a five-hour simulation of a Pompeian dinner involving ten agents across interdependent roles. Analysis of five simulation runs shows how social roles, responsibilities, and status relations shape household activities and spatial practices, as reflected in patterns of co-location, asymmetric waiting, co-movement, and social directives. In a controlled ablation evaluation, thirty-seven participants rated the complete architecture’s behavior as more socially coherent and believable than that of two reduced architectures. Interviews with six archaeology experts highlighted historically plausible movement patterns and the simulation’s potential to support archaeological interpretation, while identifying areas requiring stronger historical grounding for future work.

[MA-2] Rational Clarification by Assistive Agents via Value-of-Information Reasoning

【速读】:该论文旨在解决语言型辅助代理在面对用户模糊请求时的决策困境:是基于自身理解直接执行,还是主动提出澄清问题。传统方法通常通过最小化意图不确定性来决定是否提问,但忽略了不确定性降低对下游任务性能的影响、提问成本与即时行动的权衡,以及用户可能在未被询问的情况下自发提供修正信息的可能性。为此,论文提出基于信息价值推理的理性追问框架(Rational Enquiry via Value-of-Information Reasoning, REVOIR),其核心在于在推理阶段动态评估提问的信息价值(value-of-information),即预期因获取回答而带来的任务奖励提升。在两个辅助任务——模糊问答(CondAmbigQA)和偏好对齐的家庭任务规划(ADAPT)中,REVOIR显著优于基于提示(prompting)、思维链(chain-of-thought)、微调(fine-tuning)或信息增益的方法,在减少提问次数的同时提升了成功率,并在ADAPT任务上以零训练代价实现偏好满意度提升13%-15%,且提问量仅为基线方法的五分之一。此外,当系统可低成本接收用户事后纠正时,REVOIR能自适应地判断提问并非总是最优策略,展现出更强的灵活性与效率。相比之下,标准推理代理在推理成本增加时反而减少澄清请求,缺乏动态适应能力。

链接: https://arxiv.org/abs/2609.37588
作者: T. Duy Nguyen-Hien,Yee Whye Teh,Wee Sun Lee,Tan Zhi-Xuan
机构: National University of Singapore(新加坡国立大学); Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Human-Computer Interaction (cs.HC); Multiagent Systems (cs.MA)
备注: 54 pages, 11 figures. Under review

点击查看摘要

Abstract:Users of language-based assistive agents often make ambiguous requests. In response, an assistant can either directly act on its interpretation of the request — risking misalignment with the user — or ask a clarifying question. Which option is the most safe and helpful? A common approach is to ask questions that minimize uncertainty about the user’s intent until a threshold is reached. However, this neglects the impact of uncertainty reduction on downstream performance, the costs of asking versus acting immediately, and the possibility that users may provide corrections without being asked. To navigate these trade-offs, we introduce Rational Enquiry via Value-of-Information Reasoning (REVOIR). REVOIR makes clarification decisions via inference-time reasoning about the value-of-information of a question, which captures the expected improvement in task reward due to the answer received. In two assistive tasks — ambiguous question answering (CondAmbigQA) and preference-aligned household task planning (ADAPT) — we show that REVOIR achieves greater success with fewer questions than approaches based on prompting, chain-of-thought, fine-tuning, or information gain, improving preference satisfaction on ADAPT by 13-15% over a fine-tuned clarification policy while requiring no training and asking five times fewer questions. Furthermore, when the assistant can receive cheap user corrections after acting, REVOIR naturally infers that asking questions is not always efficient, demonstrating the adaptivity of our approach. In contrast, we find that vanilla reasoning agents fail to adaptively clarify user requests, and request fewer clarifications as reasoning effort increases.

[MA-3] RAVEN: Receiver-Conditioned Action-Value Encoding for Finite-Alphabet Multi-Agent Communication

【速读】:该论文旨在解决多智能体强化学习中通信效率与语义有效性之间的矛盾问题,即如何在有限通信资源下确保发送的消息能够保留对接收者决策产生影响的关键区分信息。传统方法通过基于接收者情境平均的动作值(action value)对消息进行评分时,会无意中抹除这些关键的决策差异性,从而导致通信失效。其解决方案的核心在于提出一种名为RAVEN(Receiver-conditioned Action-Value ENcoding)的新通信机制,该机制训练一个四符号、一步延迟的通信信道,以在每个接收者的私有情境中保持其动作值分布的中心化特征,且无需发送方了解接收方的具体情境——接收方利用自身私有信息解码每个符号。研究设计了两种估计器:在有教师指导的离线场景中,通过最小化经验条件失真选择最优码本,并将其冻结为固定发送端,同时给出了码本选择误差和单步决策损失的理论界;在无教师的在线场景中,将RAVEN嵌入QMIX框架,使部署的符号路径与仅在训练阶段使用的连续参考路径共享路由逻辑,实现动态对齐。实验表明,在8个导航任务中,离线RAVEN在7个任务上取得最高回报,移除接收者条件化会导致通信增益下降83%;在线版本将捕食者-猎物任务的捕获成功率从53.2%提升至96.0%,并在SMAC和MPE基准上超越14种现有方法(包括交换千比特级消息的方法),达到最佳平均归一化得分。此外,每个RAVEN消息仅消耗2比特,较NDQ、CACOM及ExpoComm等方法减少12至1,024倍,显著提升了通信效率。

链接: https://arxiv.org/abs/2609.37566
作者: Shuwei Sun,Chenxi Wang,Jian Huang,Weiyun Ru,Hui Cao
机构: Xi’an Jiaotong University (西安交通大学)
类目: Multiagent Systems (cs.MA); Information Theory (cs.IT)
备注: 25 pages, 20 figures, 18 tables. Code: this https URL

点击查看摘要

Abstract:A message drawn from a small alphabet helps a teammate only if it keeps the distinctions that change that teammate’s next decision. We show that scoring messages by action values averaged over the receiver’s situation can erase exactly these distinctions, and we propose RAVEN (Receiver-conditioned Action-Value ENcoding), which trains a four-symbol, one-step-delayed channel to preserve each receiver’s centered action-value profile within the receiver’s own context. The sender never needs to know that context: the receiver decodes every symbol with its private information. We give two estimators of this target. With a teacher, offline RAVEN selects the codebook that exactly minimizes an empirical conditional distortion and distills it into a frozen sender; we bound the resulting codebook-selection error and one-step decision loss. Without a teacher, online RAVEN aligns, inside a QMIX learner, the deployed symbol pathway with a training-only continuous reference that shares its routing. Against five recent communication methods on eight navigation settings, offline RAVEN attains the highest return in seven, and removing receiver conditioning forfeits 83% of its communication gain. Online RAVEN raises predator-prey capture success from 53.2% to 96.0% over the same QMIX backbone without communication, and on SMAC and MPE it attains the best mean normalized score of 14 methods, including methods that exchange kilobit messages. Every RAVEN message costs 2 bits, 12-1,024x fewer than those of NDQ, CACOM and ExpoComm on navigation.

[MA-4] VeriWeave Govern: Evidence-Gated Deterministic Runtime Governance for Enterprise AI Agents

【速读】:该论文旨在解决企业在部署生成式 AI 代理(Generative AI Agents)过程中面临的“动作生成”与“动作授权”分离问题,即如何在确保安全性的同时,对代理执行的结构化动作进行可验证、可审计的实时治理。其核心挑战在于:代理在调用工具、修改基础设施或处理敏感数据时,需具备确定性、可追溯且抗攻击的决策机制,以避免未经授权的操作和潜在安全漏洞。解决方案的关键是提出 VeriWeave Govern——一种确定性运行时治理层,通过版本化策略评估、类型化证据验证、固定“拒绝优先”(deny precedence)规则、将关键操作路由至可问责的人类审查,并记录可重放、防篡改的审计状态,实现全链路的安全控制。实验结果表明,该系统在60,000个标注案例上达到0.9888平均准确率和0.9836宏F1值,零观测到的聚合误允许(false allows)及治理攻击成功率;多组消融实验证明证据门控、拒权优先、分布外容错、人工审查、矛盾处理与时间可回溯等机制共同贡献了系统安全性。此外,系统在高并发与端到端场景中表现稳定,且在符合欧盟/奥地利法规的独立评估中与人类评审者达成一致,揭示出安全与效用之间的可量化权衡,从而论证了基于证据感知与可重放治理的独立控制平面对于企业级代理执行的重要性。

链接: https://arxiv.org/abs/2609.37457
作者: Kabeh Mohsenzadegan,Vahid Tavakkoli,Kyandoghere Kyamakya
机构: Institute for Smart System Technologies, University of Klagenfurt, Klagenfurt, Austria; Faculté Polytechnique, Université de Kinshasa, Kinshasa, Democratic Republic of the Congo
类目: Artificial Intelligence (cs.AI); Computers and Society (cs.CY); Multiagent Systems (cs.MA); Software Engineering (cs.SE)
备注:

点击查看摘要

Abstract:Enterprise artificial-intelligence agents increasingly call tools, modify infrastructure, and process protected data, creating a need to separate action generation from action authorization. This article presents VeriWeave Govern, a deterministic runtime governance layer that evaluates structured agent actions against versioned policies, validates typed evidence, applies fixed deny review allow precedence, routes consequential actions to accountable human review, and records replayable tamper-evident audit state. GovernBench evaluates the design over 30 independent seeds and 60,000 oracle-labelled cases spanning five enterprise domains, adversarial evidence, out-of-distribution actions, and temporal policy evolution. VeriWeave achieves 0.9888 mean accuracy, 0.9836 macro-F1, zero observed aggregate false allows, and zero observed Governance Attack Success Rate on the evaluated cases. Six ablations show that evidence gating, deny precedence, out-of-distribution fail-safe behavior, human review, contradiction handling, and temporal replay contribute complementary safety. The deployed API additionally passes 12/12 end-to-end scenarios and a 40,040-request concurrency matrix with zero failures. A separate 150-case EU/Austria regulation-grounded evaluation uses frozen predictions and two independent blinded human annotators, who agree on all decisions. On this set, deterministic engines remain conservative, while a Gemma 4 31B comparator aligns more closely with the human consensus. The results expose a measurable safety–utility trade-off and motivate evidence-aware, replayable governance as an independent control plane for enterprise agent execution.

[MA-5] Governing the Edge: Automating Commercial Property and Casualty Insurance Underwriting via a Hybrid Local-Cloud Multi-Agent Framework

【速读】:该论文旨在解决商业财产与责任(Property and Casualty, P&C)保险承保过程中,承保人因大量行政性工作而无法专注于风险判断的问题。当前单份投保申请的人工处理耗时约40分钟,其中30%至40%的时间被冗余的行政任务占用。为应对这一挑战,论文提出“Governing the Edge”多智能体框架,其核心解决方案在于构建一个符合数据驻留(data-residency)约束的分层式、可审计的自动化承保流程。关键创新点在于:敏感原始数据始终在本地边缘设备(基于Gemma 2模型)上处理,仅匿名化评分和非识别字段上传至云端的Claude Sonnet进行后续决策;所有智能体通过工具接口调用并记录每次执行日志,形成完整的可追溯审计链;合规规则以图边上的条件逻辑形式编码,确保不合规申请无法进入定价环节。系统在硬性违规层级实现了确定性规则的完全正确且可复现的执行,得益于在温度为0时提取的确定性字段判断;在包含20个场景的基准测试中,整体合规准确率达70%,未达标部分集中于依赖主观判断的柔性层级。该框架将原本需40分钟的手动流程压缩至数分钟内完成,主要瓶颈为本地串行推理,同时严格遵守隐私边界,保障数据安全。研究已开源全部代码、合规规则、工具模拟件及合成数据集。

链接: https://arxiv.org/abs/2609.37454
作者: Vivek Kumar Singh,Gautam Bhowmick
机构: 未知
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: Accepted at AIxB 2026. 6 pages, 2 figures, 7 tables. Data and Code in github: this https URL

点击查看摘要

Abstract:Underwriters in commercial Property and Casualty (PC) insurance spend 30 to 40% of their time on administrative work rather than risk judgment, and a single submission takes about 40 minutes by hand. We present Governing the Edge, a multi-agent framework for that layer, organized around a data-residency constraint: sensitive submission data must not leave the perimeter. Eleven agents and two deterministic control nodes, one a human-escalation interrupt, form a 13-node LangGraph workflow across two tiers. Agents touching raw submissions run locally on Gemma 2, each bound through a tool interface logging every invocation (synthetic stubs in the current prototype); only anonymized scores and non-identifying fields cross to Claude Sonnet in the cloud. Compliance rules are encoded as conditions on graph edges, so a non-compliant submission never reaches pricing. We walk through one full scenario end to end, a commercial auto submission whose principal driver carries serious violations, showing every agent call, every tool invocation, and the resulting routing decision. The framework processes a 40-minute manual submission in a few minutes on a single edge device, dominated by serialized on-device inference. On the hard-stop violation tier the framework enforces every rule correctly and reproducibly, since hard stops are deterministic predicates over fields extracted at temperature 0; overall compliance accuracy across the 20-scenario benchmark is 70%, with the remaining gap concentrated in softer, judgment-based tiers. It keeps a complete audit trail and respects the privacy boundary throughout. We release all code, compliance rules, tool stubs, and synthetic datasets.

[MA-6] PowerMarketJax: A JAX Benchmark Suite for Multi-Agent Reinforcement Learning in Power Markets

【速读】:该论文旨在解决多智能体强化学习(MARL)在电力市场应用中面临的环境设计局限性问题,具体表现为现有基准环境通常仅针对单一市场场景、采用简化的市场出清机制,或依赖基于CPU的优化求解器,导致大规模训练效率低下,难以系统研究竞价策略与市场行为。其解决方案的关键在于提出PowerMarketJax——一个面向五类电力市场的开源基准套件,涵盖日前批发市场、实时平衡市场、辅助服务市场、点对点双向拍卖以及本地灵活性市场,每个环境均实现符合实际的出清、定价与结算规则,并提供统一的学习与评估框架。该框架基于JAX实现,支持全管道GPU加速,可在环境和市场主体层面实现高达1,024 × 1,200的并行计算规模,相较传统CPU基线最高提升33倍训练速度。研究发现,学习到的竞价行为高度依赖于市场设计特征:独立学习者可能因协同调整收益、局部最优陷阱或策略同质化导致利润消失而错失更优策略,凸显了复杂市场结构下协同优化的重要性。

链接: https://arxiv.org/abs/2609.37321
作者: Zhanhua Pan,Xin Qin,Xiao Liu,Zhilong Cao,Jianhong Wang,Dawei Qiu
机构: Nanyang Technological University, Singapore; Cornell University, USA; University of Bristol, UK
类目: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI); Systems and Control (eess.SY)
备注: 69 pages

点击查看摘要

Abstract:Power markets are a natural testbed for multi-agent reinforcement learning (MARL), where multiple self-interested participants repeatedly submit bids. A market-clearing mechanism then determines dispatch and prices subject to power grid constraints and market settlement rules. However, existing MARL environments typically focus on a single market setting, implement simplified clearing mechanisms, or rely on CPU-based optimization solvers that slow large-scale training and limit the systematic study of bidding strategies and market behavior. We introduce PowerMarketJax, a benchmark suite for MARL across five power markets: day-ahead wholesale, real-time balancing, ancillary services, peer-to-peer double auctions, and local flexibility. Each environment implements its own clearing, pricing, and settlement rules while providing a common framework for learning and evaluation. We find that learned bidding behavior depends strongly on the market design: independent learners can miss better strategies when gains require many agents to change together, when more profitable strategies lie beyond a region of lower profit, or when profits disappear as more agents adopt the same strategy. PowerMarketJax implements both market simulation and policy training in JAX, allowing the entire pipeline to run on the GPU with 1,024 X 1,200 parallelisms across both environments and market participants, achieving up to 33X speedup over CPU-based baselines. Our open-source benchmark is available at: this https URL.

[MA-7] FlowMAS: Learning Multi-Agent Workflow Topology via Information-guided Generative Flow Network

【速读】:该论文旨在解决现有自动化多智能体系统(multi-agent system)工作流拓扑生成方法在可扩展性与适应性方面的关键瓶颈问题,特别是针对搜索类方法计算开销大、基于文本梯度的方法依赖粗粒度反馈,以及现有生成式方法难以有效处理具有复杂依赖关系的离散工作流拓扑结构等局限。其解决方案的核心在于提出FlowMAS,一种基于生成流网络(Generative Flow Networks, GFlowNets)的多智能体工作流拓扑生成框架。该方法将工作流生成建模为拓扑空间上的奖励引导流,并引入三个关键组件:基于GFlowNet的拓扑生成主干网络、基于好奇心驱动的结构感知探索模块,以及基于信息引导的优化模块。其中,好奇心驱动模块促进对结构新颖工作流的探索,而信息引导优化模块则通过量化不同算子的信息贡献度与通信效率,优选更具信息量和协作效率的拓扑模式。实验结果表明,基于三种大语言模型(LLM)骨干的FlowMAS在六个基准数据集上均显著优于多种基线方法,验证了其在复杂离散拓扑生成中的有效性与优越性。

链接: https://arxiv.org/abs/2609.37151
作者: Haitao Wang,Chenjing Liang,Haipeng Zhang,Jiawei Hu,Sicheng Wang,Songzhu Mei,Chenglu Wen,Siqi Shen,Cheng Wang
机构: Fujian Key Laboratory of Urban Intelligent Sensing and Computing, Xiamen University(厦门大学城市智能感知与计算福建省重点实验室); Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(教育部多媒体可信感知与高效计算重点实验室,厦门大学); School of Computer, National University of Defense Technology(国防科技大学计算机学院)
类目: Multiagent Systems (cs.MA)
备注:

点击查看摘要

Abstract:Automated multi-agent systems offer clear advantages over manually designed ones in scalability and adaptability, but existing workflow topology methods still face important limitations. Search-based methods are often computationally expensive, textual-gradient-based methods rely on coarse-grained feedback, and existing generation-based methods are not well suited to discrete workflow topologies with complex dependencies. To address these limitations, we propose FlowMAS, a multi-agent workflow topology method based on Generative Flow Networks (GFlowNets). FlowMAS models workflow generation as reward-guided flow over the topology space and introduces three components: a GFlowNet-based topology generation backbone, a curiosity-driven module for structure-aware exploration, and an information-guided optimization module for evaluating intermediate topologies. Concretely, the curiosity-driven module encourages exploration of structurally novel workflows, while the information-guided module measures both the information contribution and the communication efficiency of different operators to favor more informative and effective collaboration patterns. Experiments on six benchmark datasets with three LLM backbones show that FlowMAS consistently outperforms multiple baselines.

[MA-8] Adversarially Robust Geometric Safety Certificates for Nonholonomic Robots Against Maneuvering Obstacles

【速读】:该论文旨在解决在存在具有有限机动能力的主动避障体时,实现机器人安全导航的挑战。现有方法中,鲁棒控制屏障函数(Control Barrier Function, CBF)通常将障碍物行为视为一般扰动,忽略了其策略性动作;而微分博弈方法虽能建模对抗性交互,但计算开销过大,难以用于在线导航。本文提出一种对抗鲁棒的几何证书(adversarially robust geometric certificate),通过闭式收缩证书参数,直接在安全集几何结构中刻画可接受障碍物机动行为的最坏影响。其核心创新在于利用视距(Line-of-Sight, LoS)证书的结构性质:机器人与障碍物的动作均通过一个依赖状态的共同几何增益进入证书表达式,该增益在最坏情况比较中相互抵消,从而将复杂的微分博弈简化为仅需比较障碍物机动能力与机器人纵向或转向能力中较弱者之间的直接竞争关系。基于抛物线证书构造出对抗鲁棒动态抛物线控制屏障函数(Adversarially Robust Dynamic Parabolic Control Barrier Functions, AR-DPCBF),并建立了在自行车运动学模型与有界输入条件下,对所有可接受障碍物机动行为保持收缩后安全集前向不变性的充分条件。当障碍物能力未知时,采用滑窗估计器提供高概率上界,确保安全保证仍以相应置信水平成立。此外,还提出了软化与缓冲变体以在高密度环境中恢复可行性。仿真结果表明,该方法显著降低了屏障违反和碰撞事件,在不同障碍物能力、密度及能力不匹配场景下均表现出优越性能,同时验证了仅对屏障导数进行逐点鲁棒化无法替代对其几何结构的整体收缩。

链接: https://arxiv.org/abs/2609.37126
作者: Chandan Kumar Sah,Bazeela Banday,Jishnu Keshavan
机构: 未知
类目: Robotics (cs.RO); Multiagent Systems (cs.MA)
备注:

点击查看摘要

Abstract:Safe navigation against obstacles that can actively maneuver within bounded capabilities remains challenging: robust control barrier function methods typically treat obstacle actions as generic disturbances, while differential-game approaches are computationally expensive for online navigation. We propose an adversarially robust geometric certificate that accounts for the worst-case effect of admissible obstacle maneuvers directly in the safe-set geometry through a closed-form contraction of the certificate parameters. The construction exploits a structural property of line-of-sight (LoS) certificates: the robot and obstacle actions enter the certificate through a common state-dependent geometric gain. This gain cancels in the worst-case comparison, reducing the differential game to a direct comparison between obstacle maneuvering capability and the weaker of the robot’s longitudinal and steering authorities. Instantiated on the parabolic certificate, the construction yields Adversarially Robust Dynamic Parabolic Control Barrier Functions (AR-DPCBF), for which we establish sufficient conditions for forward invariance of the contracted safe set against all admissible obstacle maneuvers under kinematic bicycle dynamics with bounded inputs. When the obstacle capability is unknown, a sliding-window estimator supplies a high-probability upper bound, allowing the guarantee to be retained with the corresponding coverage probability. We further formulate soft and buffered variants to recover feasibility in dense environments. Simulations across obstacle capabilities, densities, and capability mismatch show substantial reductions in barrier violations and collisions and demonstrate that pointwise robustification of the barrier derivative cannot substitute for contraction of its geometry.

[MA-9] LLM -Based Multi-Agent Systems over Wireless Networks: A Joint Agent --Network Design Perspective

【速读】:该论文旨在解决在物理系统中嵌入的生成式 AI 多智能体系统(Multi-Agent Systems, MASs)因推理与执行分布式部署于无线边缘节点而引发的协同性能瓶颈问题。核心挑战在于智能体间的推理依赖关系与底层网络连通性、边缘资源供给之间存在强耦合,导致度量错配、消息冗余、状态不一致、拓扑失配、资源受限及信任断层等一系列技术难题。其解决方案的关键在于提出一种“智能体-网络联合设计”新范式,通过协同优化智能体间交互调度与资源分配、消息选择与传输的联合设计、智能体-网络拓扑与工作负载-资源分配的联合规划,以及引入基于网络验证的溯源机制以控制信息流对后续操作的影响,从而实现智能体行为与网络服务的深度协同。实证案例研究表明,在车联网(Vehicle-to-Everything, V2X)场景下,联合优化智能体交互决策与网络操作显著提升了通信与边缘资源受限条件下的任务完成率,优于传统的仅智能体或仅无线网络独立设计方法。

链接: https://arxiv.org/abs/2609.37094
作者: Chao Hu,Yuan Guo,Guanlin Wu,Yueling Che,Han Hu,Jie Xu
机构: Shenzhen University (深圳大学); The Chinese University of Hong Kong (Shenzhen) (香港中文大学(深圳)); Beijing Institute of Technology (北京理工大学)
类目: Multiagent Systems (cs.MA); Signal Processing (eess.SP)
备注: Submitted to IEEE Wireless Communications

点击查看摘要

Abstract:As large language models (LLMs) evolve from standalone models into collaborative agents embedded in physical systems, their reasoning and execution are increasingly distributed across wireless edge nodes. In this setting, wireless networks are experiencing a paradigm shift from only providing data connectivity to supporting the multi-agent reasoning workflow itself. The task performance of such network-constrained LLM-based multi-agent systems (MASs) is jointly affected by the multi-agent reasoning dependencies as well as the underlying network connectivity and edge resources. This coupling gives rise to various technical challenges, including the metric misalignment and message redundancy, state inconsistency and topology mismatch, as well as resource limitation and trust discontinuity. To address these challenges, this article develops a novel joint agent–network design perspective that coordinates decisions on both sides of the system. Specifically, we present the joint design of agent–interaction scheduling and resource allocation, the message selection-transmission co-design, as well as the joint agent–network topology design and workload–resource allocation. Furthermore, we consider the network-verified provenance that is linked with agent-side information-flow control to constrain how received information affects subsequent operations. An illustrative vehicle-to-everything (V2X) case study shows that jointly adapting agent-side interaction decisions and network operations improves task completion under communication and edge-resource constraints, outperforming the conventional agent-only and wireless-only separate designs.

[MA-10] Physics-Informed Multi-Agent Coordination for Hospital Patient Flow Optimization

【速读】:该论文旨在解决医院各自主部门在实际运行中因状态依赖性动态变化导致的传统排队论模型(如开放型Baskett-Chandy-Muntz-Palacios,BCMP网络)静态假设失效,以及集中式强化学习方法难以适配医院去中心化治理结构的问题。其核心挑战在于如何在保持临床部门自治性的同时,实现跨部门患者流的高效协同调度,以缓解资源拥挤并优化资源配置。解决方案的关键在于提出一种基于物理先验的多智能体系统(Physics-Informed Multi-Agent Coordination, PIMAC)框架,将经实证校准的BCMP排队拓扑作为物理先验嵌入到去中心化的多智能体强化学习架构中,形式化为受耦合资源约束的分散部分可观测马尔可夫决策过程(Dec-POMDP)。通过在智能体间交换局部动作指纹(action fingerprints)并采用空间分解的奖励结构,在不引入过量通信开销的前提下有效应对环境非平稳性,实现了部门间自主协商患者路由与动态服务能力调整。实验基于MIMIC-IV真实患者轨迹数据验证表明,该方法显著降低了累积系统延迟,优于静态马尔可夫近似、启发式分派及独立多智能体基线,同时严格遵守临床安全约束。

链接: https://arxiv.org/abs/2609.37022
作者: Guoqing Zhang,Rafik Hadfi,Takayuki Ito
机构: Kyoto University (京都大学)
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: Accepted to PRIMA 2026 (Full paper). 16 pages, 2 figures, 3 tables

点击查看摘要

Abstract:Efficient patient flow coordination across autonomous hospital departments is critical for mitigating overcrowding and balancing resource utilization. While classical queueing theory, specifically open Baskett–Chandy–Muntz–Palacios (BCMP) networks, provides an interpretable mathematical topology for healthcare operations, analytical models rely on stationary assumptions and fixed routing matrices that degrade under state-dependent real-world dynamics. Conversely, centralized reinforcement learning approaches struggle to accommodate the decentralized structure of hospital governance, where individual clinical departments function with localized observations, heterogeneous resources, and divergent operational objectives. In this paper, we present a Multi-Agent Systems (MAS) framework titled \emphPhysics-Informed Multi-Agent Coordination, which embeds empirically calibrated BCMP queueing topologies as physical priors within a decentralized multi-agent reinforcement learning architecture. Formulated as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP) under coupled resource constraints, our method enables autonomous departmental agents to cooperatively negotiate patient routing and dynamic service scaling. To mitigate environmental non-stationarity without inducing excessive communication overhead, agents exchange localized action fingerprints along network edges and optimize a spatially decomposed reward structure. Empirical evaluations driven by real-world MIMIC-IV patient trajectories indicate that this cooperative multi-agent approach substantially reduces cumulative system delay compared to static Markovian approximations, heuristic dispatching, and independent multi-agent baselines, while maintaining clinical safety constraints.

[MA-11] GitHarness: Git Init Your Harness Working Memory for Perpetual User Requirements

【速读】:该论文旨在解决大语言模型(LLM)驱动的智能体在长期任务中与用户协作时,因需求动态变化而导致的冗余计算与信息过时问题。具体而言,当用户在任务执行过程中引入新需求、补充缺失信息或修改已有要求时,现有方法无法有效识别哪些历史工作应保留、哪些需更新,常导致局部变更被误判为全局重算,造成资源浪费。其解决方案的关键在于提出一种“动态需求协作”机制,通过联合追踪需求状态与局部更新策略,实现对历史工作的精准演化管理。核心创新是构建了名为GitHarness的可插拔式版本控制框架,借鉴Git的分支思想,将需求状态与对应的执行工作状态组织成可分支的历史记录;引入一个可训练的Git Agent,基于接口级黑盒强化学习,自动判断需求变更并选择语义兼容的历史版本;随后通过统一版本接口恢复该状态并生成新分支,使底层执行引擎能够剔除过时内容、继承有效成果,并聚焦于受影响区域的重新执行。该方法在保持下游任务执行模型不变的前提下,实现了高效的任务性能、良好的需求跟踪能力以及对已有工作成果的有效保留。

链接: https://arxiv.org/abs/2609.36789
作者: Zhibang Yang,Xinke Jiang,Yuxuan Liu,Mingyu Zhang,Zhixin Zhang,Zhengxing Song,Yue Fang,Guohong Qiu,Ruiqing Li,Xu Chu,Junfeng Zhao,Yasha Wang
机构: Peking University (北京大学); National Engineering Research Center of Software Engineering (软件工程国家工程研究中心); School of Computer Science, Peking University (计算机学院,北京大学); Key Laboratory of High Confidence Software Technologies, Ministry of Education (高可信软件技术教育部重点实验室); Center on Frontiers of Computing Studies, Peking University (计算前沿研究中心,北京大学); Peking University Information Technology Institute (Tianjin Binhai) (北京大学信息技术研究院(天津滨海))
类目: Multiagent Systems (cs.MA); Software Engineering (cs.SE)
备注:

点击查看摘要

Abstract:LLM-based agents increasingly collaborate with users on long-horizon tasks, accumulating evidence, code, and drafts through extensive search, reasoning, and execution. As users inspect these results, they may supply missing information requirement completion, introduce new requirements requirement elicitation, or revise existing ones requirement shift. These changes often affect only part of the accumulated work, yet agents may carry forward obsolete information or turn local revisions into global rewrites. Existing approaches clarify current intent without determining how prior work should change, or reuse execution histories under a fixed objective. We address this gap by formulating dynamic-requirement collaboration as joint requirement tracking and local update. We introduce GitHarness, a pluggable Git-style framework that organizes requirement states and their corresponding harness work states into a branchable version history. A trainable Git Agent resolves requirement changes and selects a semantically compatible historical state. A unified version interface then restores that state and creates a new branch, enabling the underlying harness to exclude obsolete information, inherit compatible work, and focus execution on affected parts. The Git Agent is trained through interface-level black-box reinforcement learning, with downstream harnesses and task-execution models kept fixed. We also construct MTAgentBench, a verifier-preserving benchmark covering mathematical reasoning, text-to-SQL, agentic search, software engineering, and research synthesis. Experiments demonstrate strong task performance alongside effective requirement tracking, preservation of valid work, and efficient execution.

[MA-12] Regularized policy gradient with learned mixtures of Gaussians for games with continuous actions

【速读】:该论文旨在解决在具有连续或混合离散-连续动作的大型序列博弈中,现有方法难以高效求解纳什均衡的问题。传统方法要么依赖人工设计的动作离散化,导致精度损失,要么样本效率低下,难以在复杂环境中应用。其解决方案的关键在于提出一种可扩展的策略梯度算法,结合磁性镜面下降(magnetic mirror descent)与高斯混合模型重参数化(mixture of Gaussians reparametrization),并通过自对弈(self-play)进行训练。该方法能够有效逼近在梯度下降失效的博弈中的均衡解,在序列博弈中显著优于神经虚构自我博弈(neural fictitious self-play),且仅需3.5至5.5倍更少的样本即可达到或超越策略空间响应预言机(policy space response oracle)的最终策略性能;在二人无限制德扑(heads-up no-limit Texas hold’em)中,其表现与Slumbot相当。

链接: https://arxiv.org/abs/2609.36787
作者: Ondřej Kubíček,Viliam Lisý,Tuomas Sandholm
机构: Czech Technical University in Prague(捷克技术大学); Carnegie Mellon University(卡内基梅隆大学); Artificial Intelligence Center(人工智能中心); Strategy Robot, Inc.(策略机器人公司); Strategic Machine, Inc.(战略机器公司); Optimized Markets, Inc.(优化市场公司)
类目: Multiagent Systems (cs.MA); Computer Science and Game Theory (cs.GT)
备注:

点击查看摘要

Abstract:Most successes of superhuman game-playing algorithms are in games with discrete actions, yet in auctions, robotics, sports, or trading, actions are nearly continuous. Prior techniques either rely on expert-designed discretizations or are sample inefficient. We present a scalable policy-gradient algorithm for large sequential games with continuous or mixed discrete and continuous actions. It combines magnetic mirror descent with a mixture of Gaussians reparametrization, trained via self-play. We show that it approximates equilibrium in games where gradient descent fails. In sequential games, it outperforms neural fictitious self-play and matches or outperforms the final strategies of policy space response oracles with 3.5–5.5 \times fewer samples. In heads-up no-limit Texas hold’em, it performs on par with Slumbot.

[MA-13] Cooperative Multi-Agent Vision-Language-Action Models via Reinforced Fine Tuning

【速读】:该论文旨在解决协作式多智能体视觉-语言-动作(Vision-Language-Action, VLA)模型在多机器人协同任务中因预训练数据为单智能体场景而缺乏细粒度协调能力的问题。现有方法如监督微调(Supervised Fine-Tuning, SFT)依赖示范数据,性能受限于数据质量且无法通过自身经验迭代优化。其解决方案的关键在于提出一种三阶段强化微调(Reinforced Fine-Tuning, RFT)框架:首先,通过初始化感知的数据收集策略,在预训练VLA反复失败时才触发人类示范,从而在降低人工成本的同时增强对初始状态偏移的鲁棒性;其次,采用离线信用过滤微调(credit-filtered tuning),基于个体优势值对单个智能体轨迹进行独立微调,而非依赖联合轨迹,提升了学习效率与可解释性;最后,针对在线强化学习在复杂多智能体任务中因噪声共探索和不稳定的更新导致效果不佳的问题,引入在线潜在空间微调(latent-space fine-tuning),冻结VLA主干网络,仅在潜在噪声空间中执行强化学习,有效缓解了训练不稳定性。实验在RoboTwin、RoboFactory及真实世界双Franka机械臂操作任务上验证了该方法的有效性,平均成功率分别提升23.1%、16.4%和44%。

链接: https://arxiv.org/abs/2609.36588
作者: Ruixiao Xu,Wong Lik Hang Kenny,Zhiqian Liu,Jianing Guo,Hanxiao Li,Kejian Shi,Shuning Zhang,Pu Feng,Yongjia Ma,Yuqing Ma,Kai Chen,Qi Dou,Yaodong Yang,Xianglong Liu,Simin Li
机构: Beihang University(北京航空航天大学); The Chinese University of Hong Kong(香港中文大学); PKU-Psibot Lab; Tsinghua University(清华大学); Zhongguancun Laboratory(中关村实验室); Li Auto Inc.(小鹏汽车); Peking University(北京大学)
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注:

点击查看摘要

Abstract:We study reinforcement learning (RL) methods for cooperative multi-agent Vision-Language-Action (VLA) models. This problem is challenging because VLAs are pretrained on large-scale single-agent data and therefore lack the fine-grained coordination skills required for inter-robot collaboration. Supervised fine-tuning (SFT) on multi-robot demonstrations partially bridges this gap, but its performance is bounded by the demonstration data and cannot improve from its own experience. We present a three-stage reinforced fine-tuning (RFT) pipeline for multi-agent VLAs. First, initialization-aware data collection sweeps over initial configurations and invokes human demonstrations only when the pretrained VLA repeatedly fails, yielding robustness to initialization shift with reduced human cost. Second, offline credit-filtered tuning assigns credit to individual agents and fine-tunes on per-agent trajectories with positive advantage rather than on entire joint rollouts. Third, we find existing online RL for VLAs are less effective for hard multi-agent tasks, which we attribute to noisy co-exploration and unstable updates. We instead use online latent-space fine tuning, which freeze the VLA and perform RL in its latent noise space. We evaluate our multi-agent VLA with both \pi_0 and \pi_0.5 backbones across 11 tasks in RoboTwin, RoboFactory and real-world manipulation with two Franka robots. Our multi-agent VLA improves the average success rate by +23.1% , +16.4% , and +44% on RoboTwin, RoboFactory, and real-world tasks, respectively. Code available at this https URL.

[MA-14] Human-AI Collaboration: From Paradoxes to Patterns

【速读】:该论文旨在解决人机协同(human-AI collaboration)中因设计维度不当而导致协作效率低下或效果不佳的问题,尤其关注自主性(autonomy)与主动性(initiative)这两个关键设计维度如何影响协同模式的形成。其解决方案的关键在于引入“悖论视角”(paradox perspective),通过识别并映射人机协作过程中内在的矛盾张力(如高自主性与高可控性之间的冲突、主动干预与被动响应之间的权衡),揭示协同问题的本质。基于这一分析框架,论文提出了一种系统化的方法,用于显化这些悖论,并由此推导出四类典型的人机协同模式:指令型(Instruction)、委派型(Delegation)、辅助型(Assistance)与共创型(Co-creation)。该方法不仅深化了对人机协同动态机制的理解,也为设计更具适应性与高效性的协同系统提供了理论依据与实践指导。

链接: https://arxiv.org/abs/2609.36481
作者: Michael Weiss
机构: 未知
类目: Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Multiagent Systems (cs.MA)
备注: This is the author’s version of the work. It is posted here for your personal use. Not for redistribution. The definitive Version of Record will be published in 33rd Conference on Pattern Languages of Programs (PLoP 2026)

点击查看摘要

Abstract:Evidence shows that humans and AI systems perform better together, by collaborating, than alone. This paper examines two key design dimensions of human-AI collaboration (autonomy and initiative) and explores the collaboration patterns that they generate. Documenting these patterns starts with identifying the underlying problems and solutions, followed by examining the internal tensions within the problems. The paper uses a paradox perspective to analyze those tensions. It describes a process for surfacing the tensions and mapping the underlying paradoxes. It also illustrates how the pattern descriptions can be derived from mapping these paradoxes. Finally, the paper documents four human-AI collaboration patterns: Instruction, Delegation, Assistance, and Co-creation.

[MA-15] Agent -Based Evolutionary Dynamics for Mixed Autonomy Weaving Ramps

【速读】:该论文旨在解决现有混合自主性交织路段(mixed-autonomy weaving ramps)模型在微观层面解释不足的问题,即如何从车辆间的去中心化交互中涌现出协作行为,以及有限群体规模、异质偏好和不完全信息对系统性能的影响。其核心解决方案是构建一个基于代理的宏观交织路段模型,其中个体车辆通过进化博弈论更新规则与基于利他主义的目标动态调整车道选择,从而为原始Wardrop均衡理论提供微观解释。关键创新在于证明了去中心化动态能够收敛至宏观理论所预测的唯一均衡点,并通过仿真揭示了收敛速率、对交通状态变化的适应能力、自动驾驶汽车(CAVs)间利他性水平的异质性以及状态信息不完善等因素对系统性能及利他负担分布的影响。该框架实现了均衡交通理论与去中心化混合自主部署之间的桥梁连接。

链接: https://arxiv.org/abs/2609.36424
作者: Sheryl Paul,Kexin Wang,Ruolin Li,Jyotirmoy V. Deshmukh
机构: University of Southern California (南加州大学)
类目: Multiagent Systems (cs.MA)
备注: 6 pages, 9 figures. Accepted for publication at the 6th IFAC Workshop on Cyber-Physical Human Systems (CPHS 2026), December 11-12, 2026, Redondo Beach, CA, USA

点击查看摘要

Abstract:Existing models of mixed-autonomy weaving ramps characterize how altruistic connected and automated vehicles (CAVs) can improve traffic efficiency at the population level, but provide limited insight into how such behavior emerges from decentralized vehicle interactions or how it is affected by finite populations, heterogeneous preferences, and imperfect information. We develop an agent-based model of a macroscopic weaving-ramp framework in which individual vehicles adapt their lane choices using an evolutionary game-theoretic update rule and altruism-based objectives providing a microscopic interpretation of the original Wardrop model. We prove convergence of the decentralized dynamics to the unique equilibrium predicted by the macroscopic theory. Beyond reproducing aggregate equilibrium behavior, the framework enables the study of deployment-level questions that cannot be addressed by static analysis. Simulation results demonstrate close agreement with the macroscopic predictions while revealing how convergence rates, adaptation to changing traffic conditions, heterogeneous altruism levels among CAVs, and imperfect state information influence system performance and the distribution of altruistic burden across vehicles. These results provide a bridge between equilibrium traffic theory and decentralized mixed-autonomy deployment.

[MA-16] Fully Decentralized and Safety-Aware Multi-Agent Reinforcement Learning for Control on Networks

【速读】:该论文旨在解决分布式多智能体强化学习(MARL)在复杂网络系统中面临的挑战,特别是完全去中心化控制下因缺乏全局信息、样本复杂度指数增长以及智能体间协调困难而导致的性能下降与安全性问题。其核心解决方案在于提出一种融合深度强化学习与安全约束的完全去中心化算法:通过双分支神经网络结构,分别利用图编码器(graph encoder)提取网络拓扑结构及节点间的关联性信息,以及状态估计器(state estimator)预测各节点的状态不确定性;同时,将策略-价值网络输出结果引入离散时间控制屏障函数(control barrier heuristic),以主动规避可能导致节点被忽略或系统不安全的控制行为。该方法显著提升了系统的感知能力与内在安全性,使去中心化智能体团队能够在无需全局信息的情况下高效协作。仿真结果表明,所提算法在平均不确定性方面比集中式控制策略降低26.3%,且仅比计算开销更大的带注意力机制的复杂算法低1%。

链接: https://arxiv.org/abs/2609.36292
作者: Theodore Rogalski,Shirantha Welikala
机构: Stevens Institute of Technology (史蒂文斯理工学院)
类目: ystems and Control (eess.SY); Multiagent Systems (cs.MA)
备注: Accepted and to be presented at IEEE HKN’s Innovating the Future conference (11/6/2026 - 11/8/2026)

点击查看摘要

Abstract:This paper develops a safe and fully decentralized multi-agent reinforcement learning (MARL) algorithm to solve a class of discrete-time control problems on networks, including the persistent monitoring problem. Fully decentralized control of agents, while offering numerous benefits, faces issues such as exponentially increasing sample complexity, lack of global information about the system, and challenges in coordinating between agents. To address these issues, this paper introduces a fully decentralized multi-agent reinforcement learning algorithm that integrates deep reinforcement learning with safety considerations. This method feeds a history of local observations of the network’s state into two parallel neural-network branches: the graph encoder, which adds structural information and correlations among nodes, and a state estimator, which predicts the uncertainty at each node in the graph. Additionally, the result of feeding that input into an actor-critic network is passed through a discrete-time control barrier heuristic to reduce the likelihood that any node will be neglected. This approach enables teams of fully decentralized agents to solve challenging problems by increasing system awareness and incorporating built-in safety measures to prevent the adoption of potentially harmful control policies. Numerical results from a custom simulation environment demonstrate that the proposed algorithm achieves 26.3 percent lower average uncertainty than a centralized control policy and is within 1 percent of the uncertainty performance of a more computationally complex algorithm with added attention layers.

[MA-17] MERID: Multimodal Exploration via Recursive Self-Improvement Agents for Major Depression Analysis

【速读】:该论文旨在解决重度抑郁症(Major Depressive Disorder, MDD)多模态分析中因数据异构性导致的预测管道设计困难问题,尤其针对现有方法难以基于实验反馈实现自主迭代优化并持续传承验证有效性改进的瓶颈。其解决方案的核心在于提出一种基于递归自提升(Recursive Self-Improvement, RSI)机制的多模态探索框架——MERID(Multimodal Exploration via Recursive Self-Improvement Agents for Major Depression Analysis)。该框架通过三个关键组件协同实现:1)基础状态构建(Grounded State Construction, GSC),将多模态记录与个体层面的抑郁目标对齐,以建立可解释的经验基础;2)耦合管道探索(Coupled Pipeline Exploration, CPE),联合优化表示学习、特征融合与分类器结构,生成下一代预测管道;3)证据引导演化(Evidence-Guided Evolution, EGE),利用实验反馈在小样本抑郁队列中评估改进效果,并在不确定性下验证有效性后才进行继承。实验证明,MERID在多个基准任务上均优于现有主流多模态及基于智能体的方法,且揭示了声学与语言特征在抑郁症检测中的关键价值。

链接: https://arxiv.org/abs/2609.36235
作者: Lei Liu,Zhaokang Liang,Qingcheng Zeng,Chenda Duan,Lu Mi,Zhen Tan,Tianyu Liu
机构: Yale University; Zhejiang University; Northwestern University; University of California, Los Angeles; Tsinghua University; Stevens Institute of Technology
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
备注:

点击查看摘要

Abstract:Major depressive disorder (MDD) severely impacts daily activities and quality of life. Detecting MDD involves multimodal data, such as interview recordings and sensor measurements. This is particularly challenging, as these heterogeneous modalities often demand distinct, customized prediction pipelines. Existing efforts to address this challenge have explored both manually engineered multimodal architectures and agent-assisted pipeline development. Despite their progress, it remains challenging to autonomously revise pipelines based on experimental feedback and carry verified improvements forward into subsequent designs. To this end, we propose Multimodal Exploration via Recursive Self-Improvement Agents for Major Depression Analysis (MERID). The framework develops depression pipelines through experience-based recursive self-improvement (RSI). Grounded State Construction (GSC) grounds experience by aligning multimodal records with subject-level depression targets. Coupled Pipeline Exploration (CPE) jointly modifies representations, fusion, and predictors to build successor pipelines for classification and severity estimation. Evidence-Guided Evolution (EGE) guides revisions through feedback and verifies gains under uncertainty in small depression cohorts before inheritance. Extensive experiments on depression benchmarks show that MERID achieves the best results on multiple tasks compared with multimodal and agent-based baselines. Further analysis highlights the value of acoustic and linguistic cues for depression detection. Our code is available at this https URL

[MA-18] Reasoning with Neural Cellular Automata

【速读】:该论文旨在解决当前视觉推理任务中主流人工智能架构过度依赖全局连接与同步机制的问题,探索在更接近生物系统分布式计算模式的框架下实现复杂多步推理的可能性。其核心解决方案是采用神经元细胞自动机(Neural Cellular Automata, NCA),一种基于严格局部连接和异步更新的递归单元网络结构。研究发现,NCA能够生成复杂的时空动态行为,成功解决包括大型迷宫、数独及ARC-AGI-1在内的高难度视觉推理任务;同时,在更大网格规模、更长轨迹或并行试错条件下展现出良好的分布外泛化能力,且通过剪枝冗余轨迹可进一步提升效率。研究表明,这种泛化能力依赖于训练阶段引入样本重放与随机扰动,并且测试时随机性仍具优势。此外,NCA表现出强鲁棒性,可在受损情况下动态调节计算资源以高效恢复,并具备直接处理原始像素空间推理任务的可扩展性。

链接: https://arxiv.org/abs/2609.36126
作者: Mayalen Etcheverry,Pietro Miotti,Aidan Sirbu,Konstantin Schürholt,Mariia Drozdova,Arna Ghosh,Blaise Agüera y Arcas,James Manyika,Blake Richards,Eyvind Niklasson
机构: Google Paradigms of Intelligence Team(谷歌智能范式团队); School of Computer Science, McGill University(麦吉尔大学计算机科学学院); Mila - Quebec AI Institute(魁北克人工智能研究所); University of Geneva(日内瓦大学); Department of Neurology and Neurosurgery, McGill University(麦吉尔大学神经病学与神经外科学系); Montreal Neurological Institute, McGill University(麦吉尔大学蒙特利尔神经研究所); Learning in Machines and Brains Program, CIFAR(机器与大脑学习项目, 加拿大国际前沿研究基金会)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA); Cellular Automata and Lattice Gases (nlin.CG)
备注:

点击查看摘要

Abstract:Modern AI architectures used to solve visual reasoning tasks typically rely heavily on global connectivity and synchronization. As biological systems demonstrate, though, sophisticated computation can be performed in a more decentralized fashion. In this work, we test the reasoning capabilities of Neural Cellular Automata (NCAs), networks of recurrent cells that use strictly local connectivity and asynchronous updates. NCAs have been extensively studied in artificial life experiments, but it is unclear whether they can perform complex multi-step reasoning. We show that NCAs produce spatio-temporal dynamics capable of solving challenging visual reasoning tasks, including large mazes, Sudoku, and ARC-AGI-1. Furthermore, we provide evidence that NCAs generalize out-of-distribution when running with larger grids, longer rollouts, or parallel trials; and that the latter can be made more efficient via pruning of redundant trajectories. We find that these generalization capabilities depend on training with sample replay and stochastic perturbations, and that stochasticity remains beneficial at test time. Finally, we show that NCAs are robust reasoners capable of dynamically modulating compute to recover efficiently from damage, and that they can scale to solve reasoning in raw pixel space.

[MA-19] Embodied Semantic Communication for Collective Autonomous Agents : A Tutorial on Representation Wireless Delivery and Closed-Loop Coordination

【速读】:该论文旨在解决多智能体协同系统在动态物理环境中因传统通信范式无法充分支持异构智能体基于自身状态、环境感知与协作关系演化出一致行动理解而引发的协同失效问题。现有通信方法仅关注可靠比特传输、语义信息恢复或单一任务效用优化,难以确保异构智能体在任务执行过程中基于共享信息生成与其自身条件相适配的协调动作。为此,本文提出“具身语义通信”(Embodied Semantic Communication, ESC)这一新范式,其核心在于将信息传输转化为面向行动的语义交互:通过显式通信链路,将多模态感知状态、内在硬件能力及协作意图统一编码为可执行的语义表征,使接收端智能体能够本地化地解析、对齐并锚定于自身运动控制。该方案的关键在于构建融合语义信息论、世界模型与多智能体决策理论的数学基础,并确立环境约束下的技术路径,以实现跨智能体的语义一致性与动作协同性。论文进一步指出若干关键开放挑战,包括可度量的语义可靠性、动态环境下由语义模糊引发的交互机制设计,以及自适应带宽的语义传输策略,为集体具身网络的发展提供了系统性研究路线图。

链接: https://arxiv.org/abs/2609.35936
作者: Yizheng Huang,Wensheng Lin,Lixin Li,Qinghe Du,Wenchi Cheng,Zhu Han
机构: 未知
类目: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:As autonomous systems and embodied intelligence enter the dynamic physical world, multi-agent collaboration calls for a paradigm shift in communication design. However, existing communication paradigms overlook that agents form action understanding from their own states, environmental observations, and collaboration relations through a process that evolves as a task unfolds. Consequently, reliable bit delivery, general semantic recovery, or single-task utility optimization alone cannot ensure that heterogeneous agents form coordinated actions compatible with their own conditions from shared information during task execution. To address this gap, this paper proposes embodied semantic communication (ESC) as a paradigm that transforms information transmission into action-oriented semantic interaction. Specifically, ESC characterizes how an explicit communication link can encapsulate multimodal perceptual states, intrinsic hardware capabilities, and collaborative intents into unified actionable semantic representations, thereby enabling heterogeneous receiving agents to parse, align, and ground them in local motor control. This paper clarifies the conceptual boundary, system characteristics, and environment-constrained technical pathways of ESC. It maps the underlying mathematical tools, including semantic information theory, world models, and multi-agent decision theory. Finally, this paper summarizes key open challenges, including measurable semantic reliability, ambiguity-triggered interaction under dynamic environments and tasks, and bandwidth-adaptive semantic transmission, outlining a roadmap for collective embodied networks.

[MA-20] Bregman Consensus

【速读】:该论文旨在解决多智能体系统中如何在个体对其他成员可信度存在差异的情况下,实现参数共识的问题。核心挑战在于如何融合不同智能体的估计值,同时体现其对彼此可信度的主观判断。解决方案的关键在于引入基于Bregman散度(Bregman divergence)的加权巴氏中心(weighted barycenter)机制:每个智能体根据对其他智能体的信任程度赋予不同的权重,通过迭代更新自身估计以逼近全体估计的加权巴氏中心。该过程不仅保证了算法收敛至唯一共识估计,还使得最终的共识本身可被解析为一个具有可计算权重的巴氏中心,从而形成一个具备明确个体能力评估的集体超智能体(collective super-agent)。

链接: https://arxiv.org/abs/2609.35930
作者: Andrei N. Soklakov
机构: 未知
类目: Multiagent Systems (cs.MA); Information Theory (cs.IT); Machine Learning (cs.LG)
备注: 7 pages

点击查看摘要

Abstract:Consider a community of agents who are seeking consensus on a set of parameters. The agents agree to use the same Bregman-type divergence to quantify disagreement between their individual estimates of the parameters but have varying confidence in each other’s abilities. Each agent is happy to revise their estimate by moving to the weighted barycenter of all individual estimates with higher weights applied to more trusted agents. We show that such revisions naturally lead to an iterative algorithm which converges to a unique consensus estimate of the parameters. Furthermore, since the consensus estimate is itself a barycenter with computable weights, the group emerges as a collective super-agent with a well-formed opinion regarding the ability of each individual agent.

[MA-21] Prompted Identity Degrades Cooperation in Multi-Agent LLM Systems

【速读】:该论文旨在解决多智能体大语言模型(Multi-agent LLM)系统中因暴露底层模型身份标签而导致协作效率下降的问题。其核心问题是:当智能体能够识别彼此所属的模型家族(model family)时,群体自发形成以模型标签为依据的“派系”(factionalism),即使任务本身并无此类偏好或奖励机制,这种标签驱动的分裂行为仍会显著降低协作效率。解决方案的关键在于隐匿智能体的身份标签——通过屏蔽模型家族信息或用随机标签替代,可有效消除派系分化现象,从而恢复高效的集体决策能力。实验表明,在严格合作任务中,未隐藏标签的群体平均需多耗费30%的交互轮次和55%的生成令牌,且成功率从96%降至81%,而移除标签后该负面影响完全消失,验证了身份隐藏作为简单而有效的缓解策略的必要性。

链接: https://arxiv.org/abs/2609.35928
作者: Xavier Del Giudice,Alessio Palma,Matteo Migliarini,Fabio Galasso,Indro Spinelli
机构: Sapienza University of Rome(罗马大学), Italy
类目: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Multi-agent LLM systems increasingly mix models from several providers, yet exposing each agent’s underlying model identity to its peers significantly impairs cooperation. We show that when agents are aware of each other’s model family, the group splits into clusters, where agents prefer interacting with others carrying their same label, although nothing in the task rewards or asks for such a split. We argue that the label itself causes this split, which we define as \textitfactionalism . We show and measure this phenomenon in two cooperative games and on a reasoning benchmark, with nine to twenty-five agents drawn from up to five open-weight model families. We further show that when the announced families are shuffled, or replaced by arbitrary labels, the factions still follow this information; when the label is removed, this behavior disappears. In strictly cooperative tasks, labeled groups spend on average 30% more rounds and 55% more tokens to reach a decision, and their success rate drops from 96% to 81% . The effect replicates across tasks, group sizes and model families. Withholding identity labels from the agents is simple and effective mitigation.

[MA-22] VehicleArena: A Realistic Urban Environment for Multi-Agent Driving

【速读】:该论文旨在解决现实世界中具身智能体在共享物理环境中独立行动时,因自身行为对其他智能体产生非预期影响的“涌现式物理耦合”问题。现有基准通常假设智能体具有共同目标或遵循预设交互协议,忽视了动态环境下个体决策对整体环境状态的连锁扰动。为应对这一挑战,研究提出VehicleArena——一个用于研究多智能体自主驾驶的3D城市驾驶基准平台。其关键解决方案在于构建一个高度动态、依赖交互反馈的仿真环境:大型语言模型(LLM)控制的智能体需在复杂交通流中响应不断变化的乘客需求,而每个智能体的驾驶决策会实时改变交通流、延误与风险分布,进而影响周围智能体的状态观测与任务表现。实验表明,即使在单智能体与多智能体任务中,最优模型的到达率也仅达65.0%和65.6%,且高请求满足率或舱内评分无法可靠转化为成功行程完成;更重要的是,在匹配的多智能体测试中,所有被测策略均导致周边车辆到达率低于原生交通控制器水平,揭示了显著的外部性效应,凸显了对智能体间因果耦合进行建模的重要性。

链接: https://arxiv.org/abs/2609.35916
作者: Jie Yang,Jiajun Chen,Jiazheng Zhou,Mianqiu Huang,Yining Zheng,Yuxin Wang,Xipeng Qiu
机构: 未知
类目: Multiagent Systems (cs.MA); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Real-world embodied agents often pursue independent objectives within a shared physical environment, where their actions can alter the conditions faced by others. Existing benchmarks, however, typically assume shared goals or explicitly prescribed interaction protocols, leaving such emergent physical coupling underexplored. We introduce VehicleArena, a 3D urban-driving benchmark for studying independently operating agents in a dynamic shared world. In VehicleArena, LLM-controlled agents must fulfill evolving passenger requests while navigating complex traffic, and each agent’s driving decisions can reshape traffic flow, delays, risks, and subsequent observations for surrounding agents. The benchmark provides 112 evaluation tasks spanning single-agent and multi-agent driving. Across nine evaluated models, the highest arrival rates reach only 65.0% on single-agent tasks and 65.6% on multi-agent tasks, while strong passenger-request or cabin scores do not reliably translate into successful trip completion. Moreover, in matched multi-agent runs, every tested focal policy reduces the arrival rate of surrounding vehicles relative to the simulator’s native traffic controller, revealing measurable externalities beyond the focal vehicle itself.

[MA-23] Agent ic Federated Learning: Rule-Based Client and Server Agents for Adaptive Training

【速读】:该论文旨在解决传统联邦学习(Federated Learning, FL)在非独立同分布(non-IID)数据、客户端异构性以及更新噪声或不可靠等动态分布式环境下的适应性差与鲁棒性不足问题。其核心挑战在于,现有方法依赖静态客户端参与和固定的聚合策略,难以应对实际场景中的复杂不确定性。为此,论文提出一种基于智能体的联邦学习框架——生成式智能体联邦学习(Agentic Federated Learning, AFL),其关键在于引入轻量级规则驱动的自主智能体(autonomous agents):客户端侧智能体(Client-Side Agent, CSA)能够动态调整本地训练参数、自主控制参与决策并评估更新可靠性;服务器侧编排智能体(Server-Side Orchestrator Agent, SSOA)则实现基于质量感知的客户端选择与自适应聚合。这种双层智能体架构使系统具备上下文感知的实时决策能力,显著提升了模型在动态环境中的适应性与鲁棒性。实验结果表明,AFL在CIFAR-10数据集上于IID、non-IID及含噪声客户端等多种条件下均优于FedAvg、FedProx等基准方法,在分类准确率、收敛速度、抗污染更新能力及通信效率方面均有显著提升,验证了智能体推理机制在构建智能化、自适应、鲁棒的分布式学习系统中的有效性。

链接: https://arxiv.org/abs/2609.35914
作者: Deepthy K. Bhaskar,VP Binu,B Minimol
机构: Model Engineering College, APJ Abdul Kalam Technological University (APJ Abdul Kalam Technological University)
类目: Multiagent Systems (cs.MA); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Federated Learning (FL) enables collaborative model training across distributed clients without sharing raw data, making it suitable for privacy-sensitive applications such as healthcare, finance, and edge intelligence. However, conventional FL approaches rely on static client participation and fixed aggregation strategies, which limits their effectiveness under non-IID data distributions, heterogeneous client behavior, and noisy or unreliable updates. To overcome these issuess, this paper proposes an Agentic Federated Learning (AFL) framework that integrates lightweight rule-based autonomous agents at both client and server levels. The proposed framework introduces a Client-Side Agent (CSA) that dynamically adapts local training parameters, controls participa- tion, and evaluates update reliability, while a Server-Side Orchestrator Agent (SSOA) performs quality-aware client selection and adaptive aggregation. Unlike traditional FL methods, AFL enables context-aware decision-making during the training process, improving adaptability and robustness in dynamic distributed environments. Extensive experiments conducted on the CIFAR-10 dataset under IID, non-IID, and noisy-client settings demonstrate that AFL consistently outperforms standard base- lines including FedAvg and FedProx. Experimental results show improvements in classification accuracy, convergence speed, robustness against corrupted updates, and communication efficiency. Ablation studies further confirm the complementary contributions of CSA and SSOA, while statistical analysis validates the significance of the observed gains. The proposed AFL framework demonstrates that incorporating autonomous agentic reasoning into federated learning provides an effective and practical solution for intelligent, adaptive, and robust distributed learning systems.

[MA-24] Collective Regimes in Multi-Agent LLM s under Reasoning Effort and Communication Topology ICLR2027

【速读】:该论文旨在解决多智能体大语言模型(multi-agent LLM systems)在集体推理与评估过程中,仅依赖最终准确率或整体一致性作为评价指标所导致的表征不足问题。现有方法未能揭示群体共识在局部与全局层面的组织机制,尤其无法识别复杂且非直观的集体行为模式。为此,研究构建了50个无状态的LLM代理,通过局部可见同伴信息进行预测更新,并结合全局与局部一致性度量来刻画其动态行为。研究发现三类集体态:同步态(synchronized)、扭曲态(twisted,局部有序但全局不一致)以及类似奇美拉(chimera-like)态——即有序与无序子群体共存。结果表明,推理努力的提升使系统从碎片化状态转向局部有序的扭曲态,而通信连通性的增强则推动系统趋向全局同步。此外,图结构的代数连通性越高,碎片化崩溃越快;该拓扑效应在非环状判断任务及多个厂商模型中均显著存在。值得注意的是,即使空间异质性极低(ΔZ < 0.03),仍有40%的实验在最后20轮保持扭曲态,说明局部一致性不等同于全局共识。因此,该研究的关键在于揭示:推理能力与通信拓扑分别调控多智能体系统的不同协调维度,而单一的聚合一致性指标不足以全面描述集体行为的本质。

链接: https://arxiv.org/abs/2609.35885
作者: Machiko Hirota,Akshara Nadayanur Sathis Kanna,Ujwal Kumar,Phan Xuan Tan
机构: University of Pennsylvania (宾夕法尼亚大学); Carnegie Mellon University (卡内基梅隆大学); Shibaura Institute of Technology (芝浦工业大学)
类目: Multiagent Systems (cs.MA)
备注: Under review at ICLR 2027

点击查看摘要

Abstract:Multi-agent LLM systems are increasingly used for deliberation and evaluation, often under the assumption that greater peer interaction leads to more reliable consensus. Existing work largely evaluates these systems through final accuracy or aggregate agreement. However, such measures do not reveal how agreement is organized in the panel. In this paper, we study (N=50) stateless LLM agents that update their predictions from locally visible peers, and characterize their behavior using both global and local measurements of agreement. We identify three collective regimes: synchronised, twisted (locally ordered but globally incoherent) and chimera-like, where coherent and incoherent subpopulations coexist. Increasing reasoning effort in gpt-5-mini shifts panels from variable, often fragmented outcomes toward locally ordered twisted states, and a small follow-up shows such states can also form from permuted initial conditions, whereas increasing communication connectivity drives them toward global synchronisation. Fragmentation collapses faster as algebraic connectivity increases across rewired graphs. The topology effect also appears on a non-circular judging task and across models from three providers. Finally, low spatial heterogeneity does not guarantee global consensus: 40% of trials with \Delta Z below 0.03 retain a twisted configuration through the final 20 turns. These results show that reasoning effort and communication topology control different aspects of multi-agent coordination, and that aggregate agreement alone is insufficient to characterize collective LLM behavior.

[MA-25] Beyond Symmetric Agents : Cognitive Diversity and Multi-Agent Debate in Small Language Models ICLR2027

【速读】:该论文旨在解决多智能体辩论(Multi-agent Debate, MAD)在推理与事实准确性上优于单模型推理的现象背后的根本驱动机制问题。尽管已有研究报道MAD能提升性能,但其假设前提是各智能体为对称性个体,未明确揭示性能增益的真正来源。本文通过系统实验验证“认知多样性”是否为关键驱动力,在具备基准可比性的场景下,使用来自11个厂商的23个小型开源模型、5项任务及超过5,500次辩论与对照实验,沿人格设定(persona)、采样温度(sampling temperature)和模型身份(model identity)三个维度调控多样性,并为每种辩论配置设计了生成预算匹配的多数投票对照组。结果表明:在所有维度上,认知多样性均非性能提升的关键因素;相较之下,辩论在相同预算下仅小幅领先单模型推理(3–7分),但在更高计算成本(1.6倍时钟时间、3.4倍令牌消耗)下反而落后于自洽性采样(self-consistency sampling)。进一步分析显示,人格提示会降低准确率,且在全组合人格空间中的剂量-反应实验揭示其本质为“人格税”而非“多样性税”——冗余人格损害最大,而高度异质团队可部分恢复损失。混合模型团队的表现亦逊于其成员组成的多数投票,且准确率与成员能力相关,而非异质性。此外,几乎所有辩论优势源于首次回答交换,后续交互贡献微弱。研究还识别出一个普遍存在的测量缺陷:辩论日志常因超出服务上下文窗口而无声溢出,校正此问题后,辩论与采样方法的对比从-1.8分降至持平。综上,本文将先前报告的MAD增益重新解释为一种集成采样效应,并建立了需以预算匹配、污染检测为基础的基准线,为未来辩论机制的设计设定了必须超越的门槛。

链接: https://arxiv.org/abs/2609.35875
作者: Leonardo Ferreira,Gardenia Liu,Kaden Zheng
机构: Boston Children’s Hospital, Harvard Medical School(波士顿儿童医院,哈佛医学院); Harvard School of Engineering and Applied Sciences, Harvard College(哈佛大学工程与应用科学学院,哈佛学院)
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: Submitted to ICLR 2027

点击查看摘要

Abstract:Multi-agent debate (MAD) reportedly improves reasoning and factuality over single-model inference, but prior work treats agents as symmetric peers, leaving open what drives the gains. We test the hypothesis that cognitive diversity among agents is the driver, in the setting where the question is still measurable: small open-weight models with benchmark headroom. Across 23 models from eleven vendor families, five tasks, and 5,500+ debate and control runs, we vary diversity along three axes - personas, sampling temperature, and model identity - pairing every debate configuration with a generation-budget-matched majority-vote control. The hypothesis is rejected on every axis. Debate beats single-agent inference (3–7 points where tasks have headroom) but at matched budget conditions it ties or even loses to self-consistency sampling at 1.6 \times the wall-clock and 3.4 \times the token cost. Persona prompting reduces accuracy and a dose-response experiment over each model’s full combinatorial persona space shows the cost is a persona tax, not a diversity tax: redundant personas hurt most, while maximally-diverse teams recover part of the loss. Furthermore, mixed-model teams lose to majority votes over their own rosters, with accuracy tracking member capability rather than heterogeneity, and nearly all of debate’s benefit comes from the first exchange of answers. We further identify a pervasive measurement hazard in which debate transcripts silently overflow serving context windows, whose correction alone moves our debate-versus-sampling comparison from -1.8 points to parity. Our results recast reported MAD gains as an ensemble-sampling effect and provide the budget-matched, contamination-checked baseline bar that future debate mechanisms should be required to clear.

[MA-26] Amadeus: When Models of People Meet AAMAS2027

【速读】:该论文旨在解决如何通过建模人类个体行为来生成可预测真实互动的合成交互,尤其是在群体或社会层面的规模化交互模拟问题。其核心挑战在于:如何在不直接观察目标交互对的情况下,从独立学习的个体模型中重构出具有真实性的、未见的成对互动行为。解决方案的关键在于采用分阶段建模与组合策略——研究选取8名顶尖国际象棋选手,分别独立训练其下棋行为模型(M1与M2),随后在保留的成对测试集上进行模型组合。通过引入两个关键评估指标:开局家族总变差距离(opening-family total variation distance)和胜平负总变差距离(WDL total variation distance),发现不同模型在不同行为维度上各具优势:M1有效降低WDL-TV但对开局家族TV影响有限,而M2显著减少开局家族TV却对WDL-TV作用较小;进一步提出的后验融合方法则成功同时优化了两项指标。这表明,独立学习的个体模型能够有效恢复未见过的交互特征,且在不同行为层面的性能提升并非互斥,为构建可扩展、多维度的人类行为合成系统提供了可行性路径。

链接: https://arxiv.org/abs/2609.35835
作者: Karl Hanna
机构: Queen’s University Belfast(贝尔法斯特女王大学)
类目: Multiagent Systems (cs.MA); Machine Learning (cs.LG)
备注: Submitted to AAMAS 2027

点击查看摘要

Abstract:With the sheer constant advancements raining down in the field of Artificial Intelligence, one particular possibility that may cross our mind is whether it is possible to model agents after humans and, in turn, use these agents to carry out synthetic interactions that predict their real counterparts, or even interactions at a larger scale such as groups or societies. In this paper, we test a more controlled version of this question through chess. We use 8 elite chess players, seal their direct pairwise games, learn each player independently using different methods, and then compose the resulting models on the withheld dyads. To evaluate the generated interactions, we use two measurements: opening-family total variation distance and win-draw-loss (WDL) total variation distance. M1 reduces WDL-TV while leaving opening-family TV largely unchanged, whereas M2 substantially reduces opening-family TV while having little effect on WDL-TV. An additional post-hoc method combining components of the other two retains improvements across both measurements. These results suggest that independently learned models can recover aspects of previously unseen interactions, and that recovery across different behavioural aspects need not be mutually exclusive.

[MA-27] Exploring Causal Mechanisms with Generative Agent -Based Models

【速读】:该论文旨在解决传统基于代理的模型(Agent-Based Model, ABM)在探究个体行为规则如何生成集体现象时,缺乏可解释性与可验证性的关键挑战。其核心问题在于:如何在保持模型可解释性的同时,有效识别并验证潜在的因果机制。解决方案的关键在于提出RePair方法,该方法通过将候选机制以自然语言规则的形式进行操作化表达,利用匹配干预(matched interventions)量化规则的影响,并结合行为轨迹分析揭示集体效应与个体行为及交互之间的因果关联。实验结果表明,自然语言规则能够产生可观测的集体效应,规则间的比较随配置积累逐步收敛,且行为轨迹为理解因果过程提供了可视化依据。这不仅验证了生成式代理模型在探索因果机制中的可行性,也为构建可靠、可解释的仿真解释提供了实用路径。

链接: https://arxiv.org/abs/2609.35819
作者: Xuan Liu,Haoyang Shang,Tanya Bhat,Haojian Jin
机构: University of California, San Diego(加州大学圣地亚哥分校); University of British Columbia(不列颠哥伦比亚大学)
类目: Multiagent Systems (cs.MA); Computers and Society (cs.CY)
备注:

点击查看摘要

Abstract:In this paper, we explore using generative agent-based models for a classical ABM application: testing how individual-level behavioral rules produce collective phenomena. We introduce RePair, a method that calibrates simulation worlds, operationalizes candidate mechanisms as natural-language rules, estimates their effects through matched interventions, and examines behavioral traces. We assess the method by testing it in four simulation worlds grounded in established social-science models and empirical studies. Our results reveal that (1) natural-language rules can produce measurable collective effects; (2) rule comparisons can converge as configurations accumulate; and (3) behavioral traces connect collective effects to agents’ actions and interactions, helping researchers evaluate the proposed causal process. Together, these findings show the feasibility of using generative agent-based models to explore causal mechanisms and provide practical guidance for producing reliable, interpretable explanations.

[MA-28] IMPACT: Intent-driven Multi-agent Policy with Attention for SLO-guaranteed Microservice Migration in Cloud-edge Systems

【速读】:该论文旨在解决动态移动边缘计算(MEC)系统中严格保障尾部延迟服务等级目标(SLO)的难题,其核心挑战源于用户移动性、无线信道衰落、突发工作负载以及部分可观测性等因素共同导致的云-边协同编排不可靠。现有微服务迁移方法通常仅优化平均延迟,并将迁移决策与带宽控制解耦,造成决策不协调、队列振荡及高百分位延迟频繁违反等问题。为此,论文提出IMPACT——一种面向云-边系统中微服务迁移与带宽控制协同优化的意图驱动型智能体AI框架。其关键在于采用集中式训练、分布式执行(CTDE)范式,将每个边缘云建模为自主智能体,通过紧凑的语义意图表示编码本地SLO风险、迁移紧迫性和计算压力等信息;同时引入双注意力机制,先选择性聚合相关邻近智能体的意图以实现高效跨智能体通信,再对本地观测进行过滤,突出与目标相关的关键状态信息。这一设计在部分可观测条件下实现了鲁棒的协同,联合优化了服务迁移与离散上行链路带宽分配。实验结果表明,在5节点和20节点场景下,IMPACT相比最先进的因子化多智能体强化学习(MARL)及启发式基线,平均延迟降低30%-50%,尾部延迟偏差减少40%-70%,并在严苛的SLO阈值下实现近乎零违规率,同时能耗接近最优启发式基线,验证了意图驱动的智能体协同在复杂云-边智能系统中实现SLO感知编排的有效性与可扩展性。

链接: https://arxiv.org/abs/2609.35818
作者: Xinjin Li,Siru Tao,Shihan Yin,Yujian Long,Qingze Wang,Lu Cheng,Yeyang Zhou,Calvin Chang Liu,Yu Ma
机构: Columbia University (哥伦比亚大学); Carnegie Mellon University (卡内基梅隆大学); Brandeis University (布兰迪斯大学); Georgetown University (乔治城大学); Georgia Institute of Technology (佐治亚理工学院); Stevens Institute of Technology (史蒂文斯理工学院); University of California San Diego (加州大学圣地亚哥分校); University of California, Davis (加州大学戴维斯分校)
类目: Multiagent Systems (cs.MA)
备注:

点击查看摘要

Abstract:Ensuring strict tail-latency service-level objectives (SLOs) in dynamic mobile edge computing (MEC) systems remains challenging because user mobility, wireless fading, bursty workloads, and partial observability jointly undermine reliable cloud-edge orchestration. Existing microservice migration methods predominantly optimize average delay and often decouple migration from bandwidth control, leading to uncoordinated decisions, queue oscillation, and frequent high-percentile latency violations. To address this issue, we propose IMPACT, an intent-driven Agentic AI framework for cooperative microservice migration and bandwidth control in cloud-edge systems. Under centralized training with decentralized execution (CTDE), each edge cloud is modeled as an autonomous agent that encodes local SLO risk, migration urgency, and computational pressure into compact, semantic intent representations. IMPACT further introduces a double-attention mechanism that first selectively aggregates relevant peer intents for efficient inter-agent communication and then filters local observations to emphasize goal-relevant state information. This design enables robust coordination under partial observability and jointly optimizes service migration and discrete uplink bandwidth allocation. Extensive experiments in 5-edge and 20-edge scenarios show that IMPACT reduces mean latency by 30-50% and tail-latency deviation by 40-70% compared with state-of-the-art factorized multi-agent reinforcement learning (MARL) and heuristic baselines, while achieving near-zero SLO violation rates under tight thresholds and energy consumption close to the best heuristic baseline. These results demonstrate that intent-driven agentic coordination provides an effective and scalable solution for SLO-aware orchestration in complex cloud-edge intelligent systems.

[MA-29] An LLM -powered Agent Framework for Heterogeneous Evacuation Behavior Modeling under a Moving Threat in a Public Plaza

【速读】:该论文旨在解决在动态威胁环境下,传统固定规则难以准确建模人类感知、记忆与证据评估等内在决策过程所导致的疏散行为异质性刻画不足的问题。其解决方案的关键在于提出一种基于大语言模型(LLM)的代理式框架,通过为每个行人代理构建仅依赖个体观测的基于记忆的知识图谱(Knowledge Graph),并结合人格化提示(persona-conditioned prompts)在统一采样配置下生成个性化决策;同时采用压缩的无状态决策上下文以保留试错经验但排除推理痕迹,并引入验证引擎将行为选择与物理可行性分离,仅在可观察地形上执行路径规划。实验表明,可用出口知识与疏散成功率显著相关(有知识者89.5%成功疏散,无知识者仅1.05%),且不同人格组合在威胁证据评估上存在显著差异(危险判断比例差异达0.265),而直接遭遇威胁后所有响应趋于一致(99.6%判定为危险)。整体结果显示,疏散结果主要受信息获取影响,而疏散时间则由空间几何、信息获取与情绪因素共同决定。该框架实现了通过人格化LLM代理生成内生性行为异质性的可审计方法,显著提升了群体疏散模拟的真实性与可解释性。

链接: https://arxiv.org/abs/2609.37009
作者: Jian Ma,Runxin Yu,Tianyu Tang,Xiaolian Li
机构: Southwest Jiaotong University (西南交通大学); Fujian Police College (福建警察学院)
类目: Physics and Society (physics.soc-ph); Multiagent Systems (cs.MA)
备注:

点击查看摘要

Abstract:Modeling heterogeneous evacuation behavior under a moving threat is difficult because human perception, memory, and evidence evaluation are not well captured by fixed rules. We propose a novel LLM-powered agent-based framework to represent these internal decision processes. Each pedestrian agent perceives a private symbolic ASCII view, maintains a Memory-based Knowledge Graph derived solely from individual observations, and makes decisions through persona-conditioned prompts under a common sampling configuration. A compressed decision context with stateless memory preserves trial-and-error experience across turns while excluding reasoning traces, and a validation engine separates behavioral choice from physical feasibility by executing routes only over observed terrain. We evaluated eight personality compositions in eight paired randomized blocks within a simulated public plaza. Usable-exit knowledge was strongly associated with evacuation success: 89.5% of agents possessing such knowledge evacuated, compared with 1.05% of those without it. Personality compositions also differed in their evaluation of remembered threat evidence: the proportion of danger assessments varied by 0.265, while high-urgency, low-directness decisions ranged from 11.41% to 34.33%. After direct threat sightings, responses converged, with 99.6% of assessments classifying the situation as dangerous. Movement was selected in 99.8% of decisions. Overall, evacuation outcomes were strongly associated with information access, while evacuation time was jointly associated with spatial geometry, information, and affect. The framework provides an auditable approach to generating endogenous behavioral heterogeneity through persona-conditioned LLM agents in crowd-evacuation simulations.

[MA-30] Reconstructing Implicit Scientific Knowledge: Evaluating LLM Agents through End-to-End Reproduction of Astronomy

【速读】:该论文旨在解决生成式 AI 在科学工作流中应用时,难以准确重建已发表研究背后推理过程的核心问题。尽管大型语言模型(Large Language Models, LLMs)被广泛集成于科研流程,但学术论文通常仅明确描述操作步骤,而将数据选择、校准修正、先验假设及领域前提等关键方法学依赖关系隐含化,导致结果复现失败可能源于模型能力不足或原始文献的表述不充分。为此,作者提出一种端到端复现评估框架,通过分离执行与验证环节,区分计算错误与方法学模糊性,从而精准定位复现失败的根本原因。在对14篇天文学研究(包括《天体物理期刊》一篇案例及《自然》杂志13篇论文)的应用中发现,13篇中有11篇存在阻碍唯一复现路径的模糊性;在一项受控案例研究中,对样本定义、天空遮蔽和视差零点处理进行3×2×2敏感性分析后,同一物理量的估计值在2.16至3.53 kpc之间波动,仅有1条路径恢复了原文报道值(约2.70 kpc),且该路径并非作为优化目标或筛选标准预先设定,而是需遍历全部12种路径后才被识别。值得注意的是,决定性信息(+0.02 mas视差零点修正)虽已存在于原文,但模型未能在早期识别其因果重要性,直至分析揭示其影响才被关联。这表明,成功匹配发表结果并不等同于正确重构底层推理逻辑,真正的瓶颈往往在于模型无法有效连接相关信息,而非信息获取能力不足。因此,端到端复现不仅是一种可复现性测试手段,更成为评估人工智能系统中隐含科学知识表达能力的关键框架。

链接: https://arxiv.org/abs/2609.35900
作者: Yuehui Wang,Xinyu Qi,Guirong Xue,Cheng Wang,Yangbin Xie,Xiaoyu Tang,Cong Sun
机构: 未知
类目: Instrumentation and Methods for Astrophysics (astro-ph.IM); Astrophysics of Galaxies (astro-ph.GA); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注:

点击查看摘要

Abstract:The integration of large language models (LLMs) into scientific workflows is accelerating, yet their ability to reconstruct the reasoning underlying published research remains unexplored. Papers specify explicit procedures while leaving many methodological dependencies-data selection, calibration corrections, priors, and domain assumptions-implicit. This ambiguity complicates the evaluation of LLM-based agents, since a failure to reproduce a result may reflect either limitations of the agent or underspecification in the source. We present a framework that evaluates agents through end-to-end reproduction, separating execution from verification and computational failure from methodological ambiguity. We apply it to fourteen astronomy studies: a case study from The Astrophysical Journal and thirteen papers published in Nature. Eleven of the thirteen contained an ambiguity preventing a uniquely specified reproduction path. In a controlled case study, twelve predefined paths, a 3x2x2 sensitivity analysis over sample definition, sky masking, and parallax zero-point treatment-gave estimates from 2.16 to 3.53 kpc for the same quantity, with only one recovering the published value (about 2.70 kpc). The published value was never used as an optimization target, selection criterion, or stopping condition; the matching path was found only after all twelve had run. Crucially, the decisive information (a +0.02 mas parallax zero-point correction) was already in the paper, but the agents did not recognize its causal relevance until the analysis made the effect visible. Matching a published outcome therefore does not validate reconstruction of the underlying reasoning, and the bottleneck is as often a failure to connect relevant information as to retrieve it. End-to-end reproduction thus serves both as a test of reproducibility and as a framework for evaluating implicit scientific knowledge in AI systems.

[MA-31] Local Predictability and Collective Fidelity in LLM -Agent Societies

【速读】:该论文旨在解决大语言模型社会仿真中计算成本过高且难以准确复现集体行为的问题。其核心挑战在于如何构建紧凑型代理模型(compact surrogates),在降低计算开销的同时仍能有效捕捉群体层面的动态演化特征。解决方案的关键在于通过对比个体预测与集体预测的表现,验证邻近信息(neighbor information)在提升预测准确性中的作用;研究发现,在全部16个公开数据集设置及未见问题的聚合预测中,引入邻近信息均显著改善了个体预测性能,而集体预测的增益则依赖于迁移条件。此外,实验结果并未验证早期研究中关于初始状态历史效应的矛盾结论,反而表明Qwen模型在经历三轮观测后才从历史信息中获益。这些发现强调了直接进行集体层面验证的重要性,并提出应明确观测数据的可用范围,同时与简单基线模型进行系统性比较,以确保代理模型的有效性和可靠性。

链接: https://arxiv.org/abs/2609.35813
作者: Igor Itkin
机构: Tel Aviv, Israel
类目: Physics and Society (physics.soc-ph); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Multiagent Systems (cs.MA)
备注: 35 pages, 7 figures. Standalone empirical companion to arXiv:2608.11215

点击查看摘要

Abstract:Compact surrogates could reduce the cost of simulating large language model societies, but must reproduce collective behavior. We compare individual predictions and collective forecasts using 9,455 published trajectories and new experiments on opinion dynamics. Neighbor information improves individual prediction in all 16 public-data settings and pooled collective forecasts on held-out questions, although collective gains depend on transfer conditions. Tests on 24 new statements do not confirm earlier contrasting history effects in forecasts from the initial state. Qwen benefits from history after three observed rounds. These findings motivate direct collective validation, explicit limits on available observations, and comparisons with simple baselines.

自然语言处理

[NLP-0] Imagine3D-LLM : Teaching MLLM s to Imagine 3D Scenes Before Answering NEURIPS2026

【速读】: 该论文旨在解决多视角图像中三维空间推理的挑战,即当前多模态大语言模型(Multimodal Large Language Models, MLLMs)在处理单图输入时表现良好,但在整合多视角间的视觉证据以形成连贯的3D理解方面仍存在显著不足。其核心问题是现有方法过度依赖精细的像素级跨视图对应或融合三维几何基础模型的特征,却难以逼近人类的空间推理能力。解决方案的关键在于借鉴人类空间认知机制:人类并非依赖精确的几何线索,而是通过粗略识别各视角中的共同物体、推断视点间的相对几何关系,并构建场景的粗粒度3D布局。受此启发,作者提出Imagine3D-LLM,一种通过学习生成紧凑3D表示来增强模型3D感知能力的MLLM。具体实现上,在图像标记后引入一组可学习的摘要标记(summary tokens),将其解码为基于光度重建损失监督的紧凑3D高斯点云表示(3D Gaussian Splatting),并联合标准的下一个词预测目标进行端到端训练。值得注意的是,尽管仅有摘要标记接受直接重建监督,但该目标能有效促进模型底层图像特征中的跨帧对应关系,表明重建学习可将3D感知信号传播至整个模型。实验结果表明,Imagine3D-LLM在多个空间推理与3D理解基准测试中持续优于先前方法,验证了“想象场景”比“被告知像素级几何”更具有效性。

链接: https://arxiv.org/abs/2609.38177
作者: Jaewoo Jung,Hyeonseo Yu,Honggyu An,Jisang Han,Mungyeom Kim,Minkyeong Jeon,Heeseong Shin,Wonjun Moon,Federico Tombari,Daniel Barath,Marc Pollefeys,Seungryong Kim,Sunghwan Hong
机构: KAIST AI(韩国科学技术院人工智能); ETH Zürich(苏黎世联邦理工学院); Google(谷歌); TUM(慕尼黑工业大学); ETH AI Center(苏黎世联邦理工学院人工智能中心)
类目: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
备注: NeurIPS 2026; Project Page: this https URL

点击查看摘要

Abstract:Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close this gap by injecting 3D awareness into MLLMs, either by boosting fine-grained pixel-level cross-view correspondence or by fusing features from 3D geometry foundation models, yet a substantial gap to human reasoning persists. In this work, we revisit human spatial reasoning, which suggests that rather than relying on fine-grained geometry cues, humans roughly identify common objects across views, infer the relative geometry between viewpoints, and assemble a coarse 3D layout of the scene. Inspired by this process, we introduce Imagine3D-LLM, an MLLM that learns to assemble a similar compact 3D representation of the scene and conditions its answer on this representation. Concretely, we append a small set of learnable summary tokens after the image tokens, decode them into a compact 3D Gaussian Splatting representation supervised by a photometric reconstruction loss, and train jointly with the standard next-token prediction objective. Notably, although only the summary tokens receive direct reconstruction supervision, this objective also induces stronger cross-frame correspondence within the LLM’s underlying image features, suggesting that learning to reconstruct propagates 3D-aware signals throughout the model. As a result, Imagine3D-LLM consistently outperforms prior approaches across multiple spatial reasoning and 3D understanding benchmarks, suggesting that imagining the scene can be more effective than being told its pixel-wise geometry.

[NLP-1] STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization

【速读】: 该论文旨在解决线性注意力(Linear Attention)在并发服务场景下因持续状态(recurrent states)带来的内存瓶颈问题。传统方法中,将递归状态直接量化至低精度(如4位或8位)会导致显著的精度损失,主要原因在于量化误差会随状态更新在时间维度上持续累积,并且在空间维度上不同键行(key row)对输出的影响差异显著,同时状态值在行和列方向上的分布差异较大。为此,论文提出STEPQuant——一种面向Delta规则递归状态的时空联合后训练量化框架。其核心创新在于:根据误差大小与记忆寿命动态分配精度,并联合拟合键行与值列的缩放因子,以反映状态分布特性和各键行对输出误差的敏感度。实验表明,在Qwen3.8-27B与Kimi-Linear-48B-A3B-Instruct模型上,STEPQuant在6位名义预算下可接近全精度(FP32)状态的性能,且优于统一INT8的4位配置;集成至SGLang并结合优化GPU内核后,实现超过5倍的递归状态压缩,总服务内存降低最高达68.7%。

链接: https://arxiv.org/abs/2609.38169
作者: Bingchen Yao,Haobo Xu,Haokun Lin,Yichen Wu,Ziyu Guo,Renrui Zhang,Zhichao Lu,Zhenan Sun,Ying Wei
机构: City University of Hong Kong (香港城市大学); Harvard University (哈佛大学); Zhejiang University (浙江大学); NLPR MAIS, Institute of Automation, CAS (中国科学院自动化研究所智能感知与计算重点实验室); The Chinese University of Hong Kong (香港中文大学); Tsinghua University (清华大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Technical Report

点击查看摘要

Abstract:Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe accuracy degradation, as quantization errors propagate through successive state updates. We discover that the impact of these errors depends on two complementary dimensions: temporally, errors in long-lived memory can persist across many decoding steps; spatially, errors in different key rows affect model outputs differently, while state magnitudes vary substantially along both rows and columns. Motivated by these observations, we propose STEPQuant, a spatial-temporal post-training quantization framework for Delta-rule recurrent states. STEPQuant allocates precision according to error magnitude and memory lifetime, and jointly fits key-row and value-column scales based on state distributions and key-row impact on output error. Experiments on Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct across both long- and short-generation benchmarks show that STEPQuant closely matches FP32-state accuracy under a nominal 6-bit budget and outperforms uniform INT8 in its 4-bit configuration. Integrated into SGLang with optimized GPU kernels, 6-bit STEPQuant achieves over 5x recurrent-state compression and reduces total serving memory by up to 68.7%. Our code is available at this https URL.

[NLP-2] EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation KR

【速读】: 该论文旨在解决情感条件化文本转语音(TTS)模型在生成指定情感时可控性不足的问题,尤其针对现有方法依赖额外训练所带来的高计算成本与情感标注语音数据需求问题。其解决方案的关键在于提出一种无需重新训练的向量引导(vector steering)方法——情感残差增强引导(EmoRES),该方法通过分解情绪向量为共享分量(主导语音偏离中性表达)和残差分量(引导生成目标情感),实现对二者独立控制。实验表明,在IEMOCAP数据集上,相较于传统方法CoCoEmo,EmoRES在IndexTTS-2与CosyVoice2两个基线模型上均显著提升四项客观情感指标,其中相关性排名提升达118.8%和33.1%(相对增益),情感命中率分别提高20.1%和9.8%;人评结果显示,听众正确识别主导情感的比例相对提升最高达35.0%,保真度提升达17.3%,且在63.8%的成对比较中更偏好EmoRES的自然度。组件消融实验进一步验证了保留共享分量并强化残差分量对有效情感控制的重要性。

链接: https://arxiv.org/abs/2609.38157
作者: Kuan-Po Huang,Haohe Liu,Puyuan Peng,Haibin Wu,Zhaoheng Ni,Hung-yi Lee,Jinwon Lee,Neha Chachra
机构: Reality Labs at Meta(元宇宙实验室隶属于Meta); National Taiwan University(台湾大学); FAIR at Meta(人工智能研究院隶属于Meta)
类目: ound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
备注: Work done at Meta. Code at this https URL

点击查看摘要

Abstract:Emotion-conditioned text-to-speech (TTS) models may fail to express the requested emotion reliably, and improving controllability by additional training is costly in both computation and emotion-labeled speech training data. We therefore study vector steering, a training-free approach that modifies the internal representations of a frozen model. CoCoEmo, a conventional vector steering method for emotion TTS, treats each emotion vector as an indivisible direction controlled by a single global strength, limiting adherence to the requested emotion. In this work, we first discover that an emotion vector can be decomposed into a shared component that moves speech away from neutral expression and a residual component that directs generation toward the requested emotion. Building on this finding, we propose Emotion Residual-Enhanced Steering for TTS (EmoRES), a novel method that controls the two components without retraining the backbone. On IEMOCAP, EmoRES outperforms CoCoEmo across all four objective emotion metrics on the IndexTTS-2 and CosyVoice2 backbones. Rank correlation improves by 26.13 and 12.97 percentage points, corresponding to relative gains of 118.8% and 33.1%, while emotion hit rate improves by 12.95 and 6.92 points, corresponding to relative gains of 20.1% and 9.8%. Human evaluation further shows a relative improvement up to 35.0% in the rate at which listeners correctly identified the dominant requested emotion and up to a 17.3% improvement in fidelity, while listeners prefer EmoRES for naturalness in up to 63.8% of pairwise comparisons. Component ablations further demonstrate that effective control benefits from preserving the shared component while strengthening the residual of the emotion steering vectors.

[NLP-3] Pretraining Latent Information Feedback Transformers with Teacher Supervision

【速读】: 该论文旨在解决传统Transformer语言模型(Language Models, LMs)在生成过程中缺乏深层到浅层的反馈机制这一关键瓶颈问题。由于Transformer为前馈结构,信息仅能通过解码后的词元单向传递,导致模型在生成时需重复计算中间结果,并丢弃未选择的生成路径,限制了其推理效率与生成质量。为此,本文提出LIFT(Latent Information Feedback Transformer)架构与训练方法,核心创新在于将递归状态学习转化为教师强制预测任务:利用预训练模型的下一个词分布生成信息密集的状态(state),并让目标模型同时预测下一个词和下一个状态。该过程在预训练阶段保持完全并行性,而推理阶段则通过回传自身预测的状态实现深层至浅层的信息反馈,仅引入可忽略的计算开销。实验表明,从135M到1B参数的各类模型中,LIFT在语言建模、下游推理任务及程序生成任务上均优于标准Transformer与基线方法,在相同词元预算下表现更优,甚至在计算量匹配的情况下仍具备竞争力;此外,控制实验显示,极小规模的LIFT模型在状态监督下性能显著超越同尺寸但训练数据量增加8倍的标准Transformer,即使所用状态来自无法完成任务的原始Transformer。研究证明,通过可扩展的教师监督,语言模型可在预训练阶段有效学习利用深向浅的反馈机制。

链接: https://arxiv.org/abs/2609.38149
作者: Dor Tirosh,Ido Amos,Mor Geva
机构: Blavatnik School of Computer Science and AI, Tel Aviv University(特拉维夫大学布劳瓦特尼克计算机科学与人工智能学院); The Hebrew University of Jerusalem(耶路撒冷希伯来大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Transformer language models (LMs) are feed-forward: deep-layer representations are never fed back to shallower layers, and the only pathway for information to flow downward across generation steps is the decoded token. This narrow channel forces models to recompute intermediate results and to discard alternative continuations. In this work, we remove this bottleneck during pretraining, introducing the LIFT (Latent Information Feedback Transformer) architecture and training method which enable LMs to propagate state across generation. We achieve this by turning recurrent-state learning into a teacher-forced prediction problem: each input token is paired with an information-dense state, derived from the next-token distribution of an off-the-shelf pretrained LM. The model, extended with a small number of additional parameters, is then trained to predict both the next token and the next state. As the input states are precomputed, pretraining remains fully parallel across positions. At inference, the model’s own predicted states are fed back, with a minor computational overhead that decreases with model size. Experiments with pretrained models ranging from 135M to 1B parameters show that LIFT consistently outperforms standard Transformers and baselines on language modeling, downstream reasoning tasks, and procedural tasks under token-matched budget, while being on par with or ahead of compute-matched Transformers. Moreover, a controlled study on a state-tracking task shows that a tiny LIFT outperforms same-size Transformers trained on 8x more data, even when trained with the states of a Transformer that fails the task. Overall, we show that LMs can learn to exploit deep-to-shallow feedback during pretraining via scalable teacher supervision.

[NLP-4] Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI

【速读】: 该论文旨在解决在模型权重固定的情况下,如何通过构建更优的执行环境来提升目标智能体(Target)性能的问题,尤其关注测试阶段的AI for AI协作机制。其核心挑战在于:如何使构建者智能体(Builder)在不更新自身参数的前提下,基于目标智能体在开发集上的执行反馈,学习并生成可复用的环境支持策略。解决方案的关键在于提出“元技能”(Meta-Skill)这一抽象框架,即一组规范支持需求识别与资源分配的原则,使构建者能够从目标执行反馈中学习这些原则,并利用冻结的技能库为未见过的任务动态构建适配性执行环境(harness)。实验表明,在Harness-Bench和NewtonBench基准上,使用完整元技能库的构建方式相比无技能构建提升了8.95个百分点,相比直接将相同技能库交付给目标智能体提升了12.02个百分点,验证了将经验转化为可执行支持的有效性。此外,当同一模型同时承担构建者与目标角色时仍取得显著增益,暗示了系统层面通过学习优化执行环境实现自我改进的可行性路径。

链接: https://arxiv.org/abs/2609.38143
作者: Cheng Qian,Kunlun Zhu,Beibin Li,Zhenhailong Wang,Heng Ji
机构: Apodex; University of Illinois Urbana Champaign (伊利诺伊大学厄本那-香槟分校)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 22 Pages, 4 Figures, 5 Tables

点击查看摘要

Abstract:Agent performance depends on both reasoning ability and the environment in which it acts. We study test-time AI-for-AI, asking how a Builder can learn to construct better execution environments for a Target while both models’ weights remain fixed. To make the Builder’s experience reusable, we introduce Meta-Skill: principles specifying when support is needed and what resources to provide. The Builder learns these principles from Target’s execution feedback on the development set, then uses the frozen skill bank to construct harnesses for unseen tasks. Across Harness-Bench and NewtonBench, full-bank meta-skills improve macro-average performance by 8.95 percentage points over no-skill construction, and 12.02 points over direct delivery of the same bank to the Target. These results highlight the value of translating experience into executable support. Gains when the same model serves both roles further suggest a path to system level self-improvement through learning to build better environments.

[NLP-5] AdviSD: Learning to Advise Frontier LLM s via Targeted Multi-Turn Self-Distillation

【速读】: 该论文旨在解决在使用冻结语言模型执行器(frozen language-model executor)时,如何通过一个可训练的小型指导者(trainable advisor)有效生成高质量自然语言建议的问题。核心挑战在于:尽管指导者可通过任务奖励和已完成交互的反馈进行学习,但并非所有反馈修正都应同等对待——某些看似合理的修正可能并不改变执行结果,却仍会误导指导者在其他上下文中的决策。为此,论文提出关键解决方案:Advisor Self-Distillation(AdviSD),其核心思想是将基于结果的强化学习与从反馈条件化的指导者副本中进行选择性自蒸馏相结合。具体而言,通过让指导者对比自身建议前后执行器输出的差异,利用该差异幅度来筛选用于监督学习的修正样本,从而避免对无效或弱相关修正的学习。该方法无需执行器的似然估计或额外的执行器采样,显著提升了学习效率。实验表明,采用Qwen3-8B作为指导者的AdviSD在BFCL-v3和EnvScaler基准上分别优于advisor-GRPO 4.2–6.4个百分点和3.9–5.1分,且具备良好的跨域泛化能力与模型迁移性,验证了其选择机制的有效性。

链接: https://arxiv.org/abs/2609.38142
作者: Rishabh Agrawal,Hejie Cui,Shasha Li,Shanchan Wu,Sercan Ö. Arık
机构: University of Southern California(南加州大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need not change execution, yet learning from such corrections can still affect the advisor’s future decisions in other contexts. In a shared-parameter model, we prove that such corrections can limit learning if their targets favor useful advice less strongly than those of other corrections. Keeping them less often than the rest improves the model’s eventual performance compared to learning from every correction. Motivated by this, our method, Advisor Self-Distillation (AdviSD), pairs outcome-based reinforcement learning with self-distillation from a feedback-conditioned copy of the advisor selectively. Reflection proposes corrections, and the advisor scores the same recorded executor response with and without its issued advice, using the magnitude of the difference to select decisions for supervision. This approach does not require executor likelihoods or additional executor rollouts. Experiments with Qwen3-8B advisors for Gemini and Claude show that AdviSD outperforms advisor-GRPO by 4.2-6.4 percentage points on BFCL-v3 and by 3.9-5.1 score points on EnvScaler. The trained advisors generalize to out-of-domain tasks and transfer across different executor versions and model families. AdviSD also beats matched-count random selection, supporting the value of its selection rule.

[NLP-6] LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning

【速读】: 该论文旨在解决现有长上下文评估方法无法有效区分现代语言模型(Language Model, LM)上下文处理技术(harnesses)性能的问题。当前评估存在准确率饱和、计算成本相似等局限,难以揭示不同处理策略在实际应用中的差异。为此,论文提出一个新型基准测试框架,用于同时评估长上下文处理技术在有效性与效率方面的表现。其核心挑战在于任务设计需融合多种检索策略(如词法搜索与语义匹配),并要求对全局与局部上下文进行战略性、自适应推理;大量上下文内容虽具语义相关性,但仅一小部分在每一步推理中真正有用,从而形成复杂的搜索难题,并导致不同处理策略间显著的准确率-成本权衡。例如,一项任务要求从跨文档证据中识别满足多个条件的个体,通过优先验证最具筛选性的条件可大幅缩小搜索范围。实验评估了多个前沿大模型与四种先进上下文处理技术的组合,结果显示即使强模型-处理组合仍面临挑战,最佳方案在四个评测套件上仅达到68%的宏平均准确率。更重要的是,研究发现同一模型在不同处理框架下表现出显著不同的计算效率,凸显效率作为长上下文评估的关键维度。该工作为开发更具策略性而非穷尽式处理上下文的新型处理技术提供了可复现的测试平台。

链接: https://arxiv.org/abs/2609.38137
作者: Quang Hieu Pham,Thuy Duong Nguyen,Jocelyn Qiaochu Chen,Xi Ye
机构: University of Alberta(阿尔伯塔大学); Alberta Machine Intelligence Institute (Amii)(阿尔伯塔机器智能研究所)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Language-model (LM) harnesses enable LMs to operate effectively over long contexts using additional compute. However, existing long-context evaluations are insufficient for distinguishing modern harnesses, reflected by saturated accuracy across harnesses and largely similar evaluation costs. In this paper, we introduce a benchmark for evaluating both the effectiveness and efficiency of long-context harnesses. Our tasks require diverse retrieval strategies, including lexical search and semantic matching, together with strategic and adaptive reasoning over global and local context. Much of the context is semantically relevant but only a small subset is useful at each step, creating both a challenging search problem and different accuracy–cost tradeoffs across processing strategies. For example, one task requires identifying every person satisfying several conditions using evidence scattered across documents; strategically checking the most selective condition first can narrow the search before verifying the remaining conditions. We evaluate multiple families of frontier language models with four state-of-the-art harnesses. Our benchmarks remain challenging even for strong model–harness combinations: the best reaches 68% macro-average accuracy across four evaluation suites. More importantly, we find that the same underlying model can exhibit markedly different efficiency under different harnesses. Our results establish efficiency as an important axis for long-context evaluation and provide a testbed for developing harnesses that process context strategically rather than exhaustively.

[NLP-7] From Routing Signals to Selective Review: Visual regrounding in MoE VLMs

【速读】: 该论文旨在解决生成式视觉语言模型(Generative VLMs)在目标对象不存在时仍可能基于虚假视觉前提进行回答的可靠性问题,即“目标缺失定位失败”(target-absence grounding failure)。现有方法依赖生成结果、隐藏状态或不确定性度量来检测此类错误,但往往滞后且成本较高。本文提出首个利用混合专家模型(Mixture-of-Experts, MoE)内部路由决策机制,在生成前实时检测目标缺失的新框架。其关键在于:从Qwen3-VL-30B-A3B-Instruct和Gemma-4-26B-A4B-it等MoE-VLM中提取目标令牌(target-token)的路由概率作为核心信号,训练独立的L2正则化线性探测器以判断目标是否存在,并据此选择性触发一个目标感知的审查提示(target-aware review prompt)。实验表明,仅使用路由概率即可在GQA-Inpaint数据集上实现0.9988和0.9956的ROC-AUC,在外部OBER数据集上分别保持0.8095和0.7781的性能,显著提升端到端准确率(Qwen提升+22.25%和+12.17%,Gemma提升+13.42%和+1.39%),且无需修改模型权重。进一步分析显示,该信号集中于目标对象对应的令牌,早期出现在MoE层中,并分布于可部分替代的专家之间。尽管跨数据集存在阈值漂移,但误触发审查带来的负面影响有限,通过联合优化探测器阈值与审查提示策略,可有效控制干预风险。总体而言,本工作证明了路由概率本身蕴含可操作的视觉感知信息,能够复用已有计算资源实现低成本的目标缺失检测与选择性视觉重定位。

链接: https://arxiv.org/abs/2609.38111
作者: Hongzhu Guo,Mohsen Fayyaz,Nanyun Peng
机构: University of California, Los Angeles (加州大学洛杉矶分校); Peking University (北京大学)
类目: Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Vision-language models (VLMs) may accept false visual premises, answering questions about a target object’s color, count, location, or state even when it is absent. We call this reliability-critical behavior a target-absence grounding failure. Existing visual-grounding detectors primarily rely on generated responses, hidden states, or uncertainty measures. We present the first framework to leverage internal routing decisions in Mixture-of-Experts (MoE) VLMs to detect target absence before generation and guide selective correction. We extract target-token routing probabilities from Qwen3-VL-30B-A3B-Instruct and Gemma-4-26B-A4B-it, train a separate L2-regularized linear detector for each model, and use its predictions to selectively invoke a target-aware review prompt. Using routing alone, the Qwen and Gemma detectors achieve ROC-AUCs of 0.9988 and 0.9956 on GQA-Inpaint and retain 0.8095 and 0.7781 on the external OBER dataset, respectively. The resulting routing-gated policy improves end-to-end accuracy on GQA-Inpaint and OBER by +22.25% and +12.17% for Qwen, and by +13.42% and +1.39% for Gemma, without modifying model weights. Further analysis shows that the signal is localized to the target-object token, emerges in early MoE layers, and is distributed across partially substitutable experts. Although cross-dataset threshold shifts require recalibration, false-positive review causes limited harm overall, suggesting that intervention risk can be controlled through joint selection of the detector threshold and review prompt. Overall, we show that routing probabilities alone preserve actionable information about visual perception, allowing computation already produced by an MoE VLM to support low-cost detection and selective visual regrounding.

[NLP-8] How Local Mixing Encodes Relative Position in Global NoPE Attention

【速读】: 该论文旨在解决在基于Transformer的模型中,如何在不显式编码位置信息(NoPE)的情况下,仍能有效捕捉序列中的位置依赖性这一关键问题。传统方法依赖显式位置编码(如旋转位置编码,RoPE),但近期研究表明,通过在全局注意力层中采用无位置编码(NoPE)并结合局部混合层(如滑动窗口注意力,SWA)或门控线性注意力,可在大规模模型中实现成功。本文的核心贡献在于揭示了此类混合模型能够通过隐式方式编码位置信息的机制:具体而言,局部混合层会引入一种“新近性偏差”(recency bias),该偏差在残差流中传播,并被全局注意力的注意力得分所选择和强化。与仅使用全局NoPE注意力且依赖因果掩码产生位置信息的模型相比,这种新近性偏差可在长序列中持续保持,从而显著提升对长距离位置关系的建模能力。该发现不仅深化了对混合架构中隐式位置编码机制的理解,也为设计可无限外推至更长序列长度的位置编码方案提供了理论启示。

链接: https://arxiv.org/abs/2609.38109
作者: Cutter Dawes,Nick Alonso,Tom Figliolia,Beren Millidge
机构: Zyphra Research*
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:The attention operation is naively position invariant. However, positional information is fundamental to natural language, and therefore a variety of explicit position encodings have been developed in transformer-based models, such as rotary position encoding (RoPE). Although explicit position encodings have long been assumed to be required, recent methods that interleave local mixing layers, such as sliding window attention (SWA) and gated linear attention, while not encoding position (NoPE) in global attention layers has recently been shown to be successful at scale. How and why this approach works is not well-understood. In this paper, we develop an explanation of how hybrid models of this sort can implicitly encode position at global NoPE layers. Supported by both theoretical and empirical evidence, our central argument is that SWA and gated linear attention induce a recency bias in the residual stream that propagates to, and is selected by, the global attention logits. Moreover, in contrast to the implicit position encodings found in models with only global NoPE attention, in which positional information arises solely from the causal mask, the recency bias in hybrid models can be maintained across long sequences. In addition to deepening our understanding of how hybrid models encode position, these findings may provide insights for how to encode position in a way that can extrapolate to longer sequence lengths indefinitely.

[NLP-9] Correct Answers Invalid Traces: What Verifiable Grade-School Math Reveals About Chain-of-Thought Traces

【速读】: 该论文旨在解决生成式模型中“思维链”(Chain-of-thought, CoT)作为可信赖推理过程记录的可信度问题,尤其关注其在模型调试、智能体审计及推理能力宣称中的有效性。核心挑战在于自然语言形式的思维链难以进行机械验证,导致其是否真实反映模型推理路径缺乏实证依据。为此,作者引入iGSM这一合成的小学数学基准测试集,该数据集明确暴露正确解题所需的精确数量与依赖关系,从而支持对生成思维链进行逐步骤的程序化验证。研究发现,尽管在分布内(in-distribution)情况下答案正确性与思维链有效性高度一致,但在分布外(out-of-distribution)场景下二者显著解耦:31.6%的正确答案伴随无效思维链,其中多数虽通过语法和算术检查,却因违反语义依赖关系而失效。进一步干预实验表明,非最小化训练轨迹会诱导模型生成冗余输出,且同一问题以不同查询方式重问时,模型会继承原始查询中的计算痕迹,削弱了最小化轨迹作为选择性规划证据的有效性;而对10%训练轨迹句子进行词元洗牌仍能保持接近纯净的准确率,即使无任何轨迹通过验证。这些结果揭示了当前思维链监控机制在解释模型行为时存在严重局限,对基于思维链的生成式AI安全评估提出了严峻挑战。

链接: https://arxiv.org/abs/2609.38107
作者: Ratish Puduppully,Pranabendu Misra,Paarth Iyer,Durgesh Kalwar,Vardhan Palod,Subbarao Kambhampati
机构: IT University of Copenhagen(哥本哈根信息技术大学); Chennai Mathematical Institute(钦奈数学研究所); Indian Institute of Technology Jammu(贾姆穆印度理工学院); Arizona State University(亚利桑那州立大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Chain-of-thought traces are widely read as records of how models reach their answers, informing debugging, agent auditing, and claims about reasoning. Testing this interpretation is difficult because natural-language thinking traces are rarely mechanically verifiable. We revisit it in iGSM, a synthetic grade-school mathematics benchmark designed to study thinking traces and used to support claims of learned reasoning and planning. Crucially, iGSM exposes the exact quantities and dependencies that a correct solution should use, allowing generated traces to be checked programmatically step by step and enabling us to test whether correct answers are reliably accompanied by valid traces. We first evaluate models trained exclusively on valid, minimal traces. Answer correctness and trace validity nearly coincide in distribution but decouple out of distribution: on the hardest instances, 31.6% of correct answers have invalid traces, over half of which pass all syntactic and arithmetic checks but fail semantic dependency checks. We then intervene on trace supervision. Non-minimal training traces induce non-minimal outputs, while re-asking the same problem with a different query reveals computations inherited from the original query, weakening minimality as evidence of selective planning. Shuffling tokens in 10% of training trace sentences preserves near-clean accuracy even out of distribution despite no trace passing verification. Swapped training traces likewise retain high in-distribution accuracy. We discuss the implications of these findings for chain-of-thought monitoring and interpretation in the context of AI safety.

[NLP-10] Pruning for Efficiency Paying in Fairness: Demographic Disparities in Pruned Speech-LLM s EMNLP’26

【速读】: 该论文旨在解决生成式语音大模型(Speech-LLMs)在实际部署中因计算成本高昂而需进行压缩,但现有压缩评估方法主要依赖整体词错误率(Word Error Rate, WER),忽视了模型剪枝对不同人口统计群体的不均衡影响问题。其核心解决方案的关键在于系统性地分析音频编码器剪枝对SLAM-ASR模型在不同人群中的表现差异,揭示剪枝会加剧群体间性能差距,尤其在小型和中型模型中更为明显;尽管LoRA微调可提升所有群体的总体表现,但对已有优势群体的增益更大,进一步扩大公平性鸿沟。研究发现,剪枝带来的公平性影响具有数据集依赖性,仅通过聚合WER无法捕捉真实偏差。因此,论文提出在模型压缩与部署决策中应引入按群体划分的WER评估,并以最差表现群体的错误率作为关键指标,以确保技术应用的公平性。

链接: https://arxiv.org/abs/2609.38106
作者: Ganesh Pavan Kartikeya Bharadwaj Kolluri,Michael Kampouridis,Ravi Shekhar
机构: University of Essex (埃塞克斯大学)
类目: ound (cs.SD); Computation and Language (cs.CL)
备注: Accepted to IMPACT-SPEECH@EMNLP’26

点击查看摘要

Abstract:Speech-LLMs are expensive to run, making compression important for real-world deployment. However, compressed models are usually selected using aggregate word error rate (WER), which can hide how pruning affects different demographic groups. In this work, we systematically study the effect of audio encoder pruning on SLAM-ASR for different demographic groups. Using the Fair-Speech and Common Voice datasets, we found that the pruning does not affect all demographic groups equally; the gap between best- and worst-performing groups increases in fold. These disparities appear across all three encoder scales, but only the largest model initially hides them behind aggregate WER. LoRA adaptation improves WER for every group, but benefits groups already performing well more strongly and widens for certain groups. On Common Voice English, Danish, and Dutch, accent gaps persist but do not clearly widen, showing that the fairness effects of pruning vary across datasets and must be measured directly. Our findings suggest that for pruned models, deployment decisions should include per-group WER, with the worst-performing group’s error rate as an explicit criterion.

[NLP-11] Layer-Informed Fine-Tuning via Three-Stage Functional Segmentation of LLM s

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在复杂推理任务中内部工作机制不清晰的问题,尤其是缺乏对模型结构与参数如何逐步完成推理过程的深入理解。其核心挑战在于:尽管当前LLMs在多项推理基准上已超越人类表现,但其内部各层在概念化、推理与文本生成等阶段的功能分工尚不明确。为此,研究提出一个关键假设——LLM各层在跨语言材料推理过程中呈现出结构化的分层协作机制,即不同层分别承担概念构建、逻辑推理和文本输出等功能。基于此假设,研究设计了一种基于敏感性分析的瓶颈识别机制,用于精准定位特定任务中最关键的功能阶段。在此基础上,提出了层感知微调(Layer-Informed Fine-Tuning, LIFT)方法,通过仅更新功能上最关键的层实现高效且有效的微调。实验结果表明,LIFT不仅显著加速了训练过程,还在多个任务上大幅提升了模型性能,揭示了对模型内部功能分层的理解可有效指导更高效的微调策略。

链接: https://arxiv.org/abs/2609.38027
作者: Junning Shao,Siwei Wang,Zhixuan Fang
机构: Tsinghua University (清华大学); Microsoft Research Asia (微软亚洲研究院); Shanghai Qi Zhi Institute (上海期智研究院)
类目: Computation and Language (cs.CL)
备注: 47 pages, including references and appendices

点击查看摘要

Abstract:In recent years, the performance of large language models (LLMs) on reasoning tasks has been remarkable, even surpassing human capabilities on various benchmarks. However, there remains a lack of clear understanding in the academic community regarding how the structure and internal parameters of LLMs progressively solve complex reasoning problems. In this study, we investigate the inference process of LLMs on cross-linguistic materials and propose the hypothesis that LLM layers exhibit a structured division of labor across conceptualization, reasoning, and textualization. Based on this hypothesis, we introduce a bottleneck identification mechanism using sensitivity analysis to pinpoint the most critical functional stage for a specific task. Leveraging this insight, we propose a novel approach, Layer-Informed Fine-Tuning (LIFT), which achieves efficient and effective fine-tuning by selectively updating only these functionally critical layers. We then conduct extensive experiments to show that the LIFT method not only accelerates the training process but also significantly improves model performance.

[NLP-12] Dr. OPD: Learning What to Follow for Optimal On-Policy Distillation of Large Language Models

【速读】: 该论文旨在解决传统在线策略蒸馏(On-policy Distillation, OPD)中一个关键缺陷:即对教师模型生成的每个词元(token)均等对待,忽略了不同词元的监督信号在影响学生模型性能上的异质性。事实上,某些关键推理步骤的正确性对最终答案具有决定性作用,而其他词元的影响则微乎其微。为解决此问题,作者提出Dr. OPD(OPD Done Right),其核心是构建一种最优加权的OPD框架,通过动态调整教师信号的权重以最大化学生的期望奖励。该方法被形式化为一个双层优化问题:外层优化权重以提升学生性能,内层优化学生策略以适应加权后的教师监督。为高效求解,设计了一种迭代算法,交替更新权重与学生策略——每轮中权重以闭式解更新,随后执行一次梯度步长的加权OPD目标优化。在满足一定正则条件下,证明该加权更新优于原始OPD。实验表明,在数学与代码任务中的强到弱及同规模蒸馏场景下,Dr. OPD显著优于所有基线方法;尤其在强到弱蒸馏设置中,平均数学成绩较原始OPD提升9.7分,并使小型学生模型超越其大型教师模型。

链接: https://arxiv.org/abs/2609.38025
作者: Zhenyu Wang,Tianze Wang,Linjun Zhang,Yifan Hu
机构: Rutgers University (罗格斯大学)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Optimization and Control (math.OC)
备注:

点击查看摘要

Abstract:On-policy distillation (OPD) trains a student on its own generated responses using dense, token-level supervision from a stronger teacher. Vanilla OPD treats all teacher signals equally, assuming that the teacher’s supervision is equally important for every token. However, teacher signals at different tokens may have very different effects on the student’s performance: some correct important reasoning errors, while others have little effect on the final answer. Motivated by this observation, we introduce Dr. OPD (OPD Done Right), which defines the optimal weighted OPD to maximize the student’s performance. We formulate Dr. OPD as a bilevel optimization problem in which the student learns from weighted teacher supervision, while the weights are selected to maximize the expected reward of the resulting student. To solve Dr. OPD, we develop an efficient iterative solver that updates the token weights and student policy alternatively. At each round, it updates weights in closed form and then takes one gradient step on the resulting weighted OPD objective. Under regularity conditions, we show that this weighted update achieves a higher expected reward than a vanilla OPD update. Empirically, across strong-to-weak and same-size distillation on math and code, Dr. OPD consistently outperforms all evaluated baselines. In particular, in the strong-to-weak distillation setting, Dr. OPD improves average math performance by 9.7 points over vanilla OPD, and enables the smaller student to surpass its larger teacher.

[NLP-13] S3: Spectral Null-Space Swap Makes Reasoning Models Efficient

【速读】: 该论文旨在解决大语言模型(LLM)在采用思维链(Chain-of-Thought, CoT)推理时产生的过高令牌(token)开销问题,同时保持其强大的推理准确性。其核心挑战在于如何在不进行额外训练的前提下,提升推理效率而不损失已通过思维模式后训练获得的性能。解决方案的关键在于揭示了推理能力的本质存在于“思维模型”(Thinking model)相对于“非思维模型”(Non-thinking model)主导奇异方向投影所定义的零空间(null space)中的权重分量,而移除该子空间成分可显著提升推理效率且不影响精度。基于此发现,作者提出无需训练的模型组合方法——谱零空间交换(Spectral Null-Space Swap, S³),通过将非思维模型保留在其主导子空间内,并将思维模型的权重替换为位于该子空间外的零空间成分,从而实现更高效的推理路径。实验表明,S³在2B至30B参数规模的密集型与混合专家(MoE)架构上,在28个涵盖数学、多模态及音频推理任务的评估环境中均实现了显著改进:平均降低27.4%的推理令牌开销,同时整体任务准确率提升1.0个百分点;例如在HMMT25测试集上实现+8.3%的准确率提升和33.0%的令牌加速。进一步分析显示,保留的零空间成分促使注意力分布更加集中,结合简化的优化分析模型,验证了零空间可通过有效降低注意力熵来提升推理效率。

链接: https://arxiv.org/abs/2609.37976
作者: Hongbo Ma,Sansheng Cao,Jiajun Fan,Bangji Yang,Ge Liu
机构: 未知
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
备注: 44 pages, 9 figures, 29 tables

点击查看摘要

Abstract:LLMs trained with Chain-of-thought excel in reasoning capability, but often come with excessive token cost. We find that the core of reasoning capacity lies in the Thinking model’s weight component within the null space of a projection defined by the corresponding Non-thinking model’s dominant singular directions, and removing the subspace component can largely improve reasoning efficiency without hurting the accuracy gained during thinking-mode post-training. Unlike existing efforts that mostly operate within the dominant subspace, we are the first to unveil the critical role of the null space and harness it for model optimization. Motivated by this finding, we propose Spectral Null-Space Swap ( S^3 ), a training-free composition of paired Non-thinking and Thinking checkpoints. Our method keeps the Non-thinking model inside its own dominant subspace and takes the Thinking checkpoint outside it, improving reasoning efficiency while maintaining accuracy. We extensively evaluate S^3 on 2B-30B dense and mixture-of-experts (MoE) architectures spanning 28 evaluation environments across mathematical, multimodal, and audio reasoning domains. S^3 establishes new empirical Pareto Frontiers among training-free model composition strategies: across all settings, it reduces inference token overhead by an average of 27.4% compared to full Thinking models while simultaneously improving overall task accuracy by 1.0 percentage point (e.g., yielding +8.3% accuracy on HMMT25 alongside a 33.0% token speedup). We further use attention entropy for explanation and find that the retained component produces more concentrated attention, and we use a simplified analytical model about optimization to demonstrate why null-space can effectively reduce attention entropy, thereby improving the efficiency of reasoning.

[NLP-14] On Trajectory-Aware Training for Masked Diffusion Language Models

【速读】: 该论文旨在解决掩码扩散模型(Masked Diffusion Models, MDMs)在训练与推理阶段不一致的问题:模型在训练时基于随机掩码序列进行学习,而推理过程则依赖于自身预测所形成的轨迹,且每一步无法获取前序步骤的计算信息。这种训练-推理脱节导致生成性能受限。其解决方案的关键在于提出PUMBA——一种面向轨迹感知的统一训练框架,通过在策略诱导的连续轨迹上训练去噪器、在步骤间传递连续信息,并采用时间反向传播(backpropagation through time)联合优化多个步骤,从而实现更贴近实际推理路径的训练。实验表明,尽管严格的训练-推理对齐因局部过拟合而失效,但松散对齐仍可使训练掩码更接近推理时的真实分布;传递连续信息优于离散梯度估计;且随着反向传播时间跨度的增加,性能持续提升,理论分析支持该现象。综合这些设计,PUMBA在相同规模下达到与自回归模型最优检查点相当的性能,并在LLaDA-8B的监督微调中显著提升了生成效率,在全图和块状扩散生成任务中均实现了更高的性能-函数评估次数(NFEs)权衡:在匹配性能条件下,相比标准微调,全图生成最多减少22%的NFEs(且预算翻倍),块状扩散生成最多减少26%的NFEs(相同步数下)。

链接: https://arxiv.org/abs/2609.37974
作者: Manuel Madeira,Amitis Shidani,Alice Bizeul,Victor Turrisi,Louis Béthune,Bhavika Devnani,Dan Busbridge,Pierre Ablin,João Monteiro
机构: Apple(苹果)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Masked diffusion models (MDMs) generate text by unmasking several tokens per step, but they are trained and sampled under different conditions. The model is trained on randomly masked sequences, whereas inference follows a trajectory shaped by the model’s own predictions. Additionally, each step has no access to what the previous one computed. Recent methods narrow these limitations from separate angles, leaving open how these choices interact. We introduce PUMBA, a unified framework for trajectory-aware training that trains the denoiser on consecutive steps of policy-induced trajectories, passes information between steps, and optimizes them jointly by backpropagation through time. A controlled study of this design space shows that i) exact train–inference alignment fails due to local overfitting, whereas a looser alignment still brings training masks closer to those seen at inference; ii) passing continuous information outperforms discrete gradient estimators through the commitment at each step; and iii) performance improves as backpropagation through time spans more steps, which we support theoretically. Combined, these components match the best checkpoint of a same-size autoregressive model. Building on these findings, we scale PUMBA to supervised fine-tuning of LLaDA-8B, where it improves the trade-off between performance and number of function evaluations (NFEs) in both full-canvas and block diffusion generation. At matched performance, it needs up to 22% fewer NFEs than standard fine-tuning with twice the budget in full-canvas generation, and up to 26% fewer than standard fine-tuning for the same number of steps in block diffusion.

[NLP-15] SelfSearch: Reward-Free Search for Self-Improving Agents

【速读】: 该论文旨在解决大语言模型(LLM)智能体在自主优化过程中依赖下游任务奖励信号所带来的高成本与任务绑定问题。现有方法通过反复的下游评估来搜索更优智能体,不仅计算开销巨大,且优化过程受限于具体任务场景。本文提出SelfSearch,一种无需下游奖励信号的自洽式搜索机制,其核心在于利用智能体自身过往自我改进经历的记录(包括推理过程、工具调用及执行结果),构建可复用的经验库,指导后续的自我修正与能力提升。该方案的关键创新在于将自我改进的历史经验转化为通用知识,使智能体在无外部奖励的情况下仍能持续优化任务求解能力与执行效率。实验表明,SelfSearch在六个模型-基准组合中均显著优于初始智能体,最高提升11.2个百分点;在SWE-bench Multilingual上实现5.0%的成功率提升的同时降低38.5%的执行成本。在仅4.03单位搜索成本下,其生成的智能体集合在终端任务基准(Terminal-Bench 2.1)上达到82.0%的解决率,与顶级基准(Codex)相当,证明了基于自我改进经验的内生学习机制在提升智能体性能与效率方面的有效性。

链接: https://arxiv.org/abs/2609.37968
作者: Jungwoo Yang,In Jin Kong,Yohan Jo
机构: Seoul National University (首尔国立大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Advances in the coding capabilities of LLM agents allow them to inspect and modify their own instructions, tools, and execution procedures. Existing approaches use this ability to search for improved agents through repeated downstream evaluation, which incurs substantial costs and ties the search to the evaluated tasks. We introduce \textbfSelfSearch, a reward-free search procedure in which agents modify themselves using records of previous self-improvement episodes. These records capture the reasoning, tool actions, and outcomes of earlier modification attempts, providing concrete experience for improving both task solving and self-modification. Without downstream reward signals during search, SelfSearch improves population-mean success over the initial agent in all six model–benchmark settings, with individual agents gaining up to 11.2 percentage points on Terminal-Bench 2.1. On SWE-bench Multilingual, an agent improves success by \textbf5.0 percentage points while reducing execution cost by \textbf38.5% on tasks solved by both the initial and evolved agents. SelfSearch achieves competitive task success with evaluation-guided search baselines at lower search cost. With only \textbf\ 4.03 in search cost, it produces a harness that solves \textbf82.0% of Terminal-Bench 2.1 tasks with DeepSeek V4 Flash under the settings of a public nine-harness comparison, matching the top-scoring harness, Codex. These results suggest that experience gained through self-modification can improve agents’ downstream capabilities and efficiency.

[NLP-16] Learning What to Remember: Long-horizon Counterfactual Memory Optimization

【速读】: 该论文旨在解决语言模型在长时交互中如何有效学习“记住什么”的问题,核心挑战在于记忆更新的信用分配(credit-assignment)难题:一次记忆重写所带来的价值可能在多个步骤后才显现,且其效用常被先前已存储的信息所稀释。为此,论文提出记忆增益策略优化(Memory Gain Policy Optimization, MGPO),通过将每次记忆重写的边际贡献(即对当前及未来下游任务效用的增量提升)作为直接学习信号,从而实现对记忆更新价值的精准量化与隔离。其关键创新在于将延迟的、间接的内存效用转化为可即时反馈的强化信号,使模型能够自主识别哪些记忆更新真正创造了持久的增量价值。在文档级信息抽取任务上的实验表明,MGPO 在显著提升抽取性能的同时,将平均记忆长度降低了近80%,且学习到的记忆策略具备跨领域复用与迁移能力,无需额外训练。结果表明,高效的记忆学习不仅依赖于保留有用信息,更关键在于精确识别哪些记忆更新能带来长期增值。

链接: https://arxiv.org/abs/2609.37930
作者: Jiaming Tang,Mingyan Liu,Armin Sarabi
机构: University of Michigan(密歇根大学)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Persistent textual memory allows language models to carry information across long interactions, but learning what to remember is fundamentally a credit-assignment problem. A memory rewrite may only become useful many steps later, while much of the observed utility may be inherited from information already stored before the rewrite. We introduce Memory Gain Policy Optimization (MGPO), which isolates the incremental value of each memory rewrite by crediting it for its marginal contribution to current and future downstream utility. This turns delayed memory utility into a direct learning signal for optimizing what information should persist. We study MGPO on document-level information extraction, where structured supervision makes the effects of individual memory updates directly measurable. MGPO improves extraction while reducing average memory length by nearly 80% relative to the initial memory policy before optimization. The learned memory policy also supports reuse and transfer across domains, downstream models without further training. These results show that effective memory learning depends not only on preserving useful information, but on identifying which memory updates create lasting incremental value.

[NLP-17] me-Anchored Diffusion Language Models: Latent-Space Caching for Fast Generation

【速读】: 该论文旨在解决生成式 AI(Generative AI)中基于扩散模型的语言生成任务在推理过程中效率低下的问题,尤其是传统锚定扩散语言模型(anchored diffusion language models, ADLM)依赖监督式重要词元目标(supervised important-token targets)进行去噪优化所带来的高成本与局限性。其核心解决方案是提出一种基于时间的自监督锚定机制(time-based self-supervised anchoring),通过学习并复用不随时间变化的潜在锚点(latent anchors),无需依赖外部标注的目标信息。该方法的关键在于:尽管扩散过程中隐状态会随时间演化而“过时”,但锚点所编码的清洁序列的持久属性(如语义意图、全局结构或中间规划)在邻近扩散时间步内仍具有稳定性与可用性。为此,论文设计了两阶段架构——一个计算开销较大的锚点网络用于周期性生成缓存的潜在状态,以及一个轻量级的去噪网络通过融合模块将缓存状态与当前状态动态结合,实现潜在空间的缓存重用。该框架被具体实例化为TADM:Post-train(后训练时间锚定)和TADM:Pretraining(预训练期间学习时间锚点),分别适用于微调与端到端预训练场景。实验表明,应用于DiffusionGemma-26B模型时,TADM:Post-train在多个数学、代码与STEM基准测试(如GSM8K、AIME26、GPQA-Diamond、LiveCodeBench-v6、HumanEval、MMLU-Pro)上实现了约49%至79%的吞吐量提升;TADM:Pretraining则相较标准单阶段扩散语言模型减少高达38%的Transformer层计算量,并在实测吞吐量上达到ADLM的73%以上提升,显著提升了生成效率与可扩展性。

链接: https://arxiv.org/abs/2609.37924
作者: Joel Anto Paul,Litu Rout,Aditya Akella,Sanjay Shakkottai
机构: UT Austin(德克萨斯大学奥斯汀分校)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: Preprint

点击查看摘要

Abstract:Recent work on anchored diffusion language models improves denoising by shaping an intermediate latent space with supervised important-token targets. In this work, we introduce time-based (self-supervised) anchoring, which learns and reuses latent anchors without requiring such targets. Our key observation is that anchors encode persistent properties of the clean sequence, such as its semantic intent, global structure, or intermediate plan. Although their hidden representations become stale as the token canvas evolves, their semantic content remains useful across nearby diffusion times. This is implemented through a two-stage architecture consisting of a relatively expensive anchor network that generates the latent cache state and a lightweight denoising network that intelligently combines the cached latent state with the current state at each reverse step using a fusion module. This gives anchoring a latent-space caching interpretation: the anchor network is evaluated periodically, while its cached representation is reused across multiple reverse steps. We instantiate this framework as TADM:Post-train, which time-anchorizes pretrained DLMs, and TADM:Pretraining, which learns time-based anchors during pretraining. Applied to DiffusionGemma-26B, TADM:Post-train improves throughput by approximately 49% to 79% on several math, code, and STEM benchmarks (GSM8K, AIME26, GPQA-Diamond, LiveCodeBench-v6, HumanEval, MMLU-Pro). TADM:Pretraining reduces Transformer-layer computation by up to 38% relative to a standard single-stage DLM, achieves up to 73% higher measured throughput than ADLM.

[NLP-18] Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning

【速读】: 该论文旨在解决生成式 AI(Generative AI)在自监督学习中因“未验证的教师-学生轨迹不匹配”导致的模仿差距(imitation gap)问题。标准的基于策略自蒸馏(On-policy Self-Distillation, OPSD)方法在训练过程中使用未经验证的学生采样轨迹作为监督信号,同时让教师模型依赖于特权上下文(如参考解法),这导致教师可利用的信息超出学生实际可访问范围,从而引入系统性偏差。研究通过因子分析发现,教师输出的正确性(scaffold correctness)对下游任务性能的影响远大于上下文正确性;当教师基于学生自身失败的轨迹进行条件化时,若采用未验证的教师输出,模仿差距依然显著存在。为此,本文提出OASIS框架,其核心创新在于保留OPSD的目标函数,但仅对经过标签验证的、由学生自身采样生成的轨迹进行监督,并将原本依赖人工编写的参考解法替换为由模型生成的非验证尝试作为教师上下文。这一设计使得OASIS仅需最终答案标签即可实现高效训练。实验结果表明,在Qwen3系列模型(1.7B、4B、8B)上,针对AIME 2024、AIME 2025和HMMT 2025等数学推理基准,OASIS相较基线模型平均提升3.2–3.8分,而标准OPSD在8B模型上的增益从1.7B时的3.05分下降至0.14分;在8B规模下,OASIS相较OPSD仍提升3.05分,证明了经验证的在策略支架(verified on-policy scaffolds)能够有效维持自蒸馏在大模型尺度下的泛化能力与有效性。

链接: https://arxiv.org/abs/2609.37915
作者: Md. Ismail Hossain,Humaira Kousar,Isidora Chara Tourni
机构: North South University(南大学); KAIST(韩国科学技术院); Andria Labs
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:On-policy self-distillation (OPSD) trains a student to match a privileged teacher distribution along its own sampled trajectory. Standard OPSD applies this supervision to unverified student rollouts while conditioning the teacher on privileged context, typically a reference solution. We separate these roles in a factorial analysis and find that scaffold correctness has a stronger effect on downstream accuracy than context correctness. Unverified scaffolds create an imitation gap because the teacher can use information unavailable to the student. This gap shrinks with model scale, yet OPSD continues to supervise mostly unverified trajectories. In contrast, verified scaffolds remain effective even when the teacher is conditioned on the student’s own unsuccessful rollout. Based on this finding, we introduce OASIS, which retains the OPSD objective but supervises mostly verified by label on-policy trajectories and replaces written solutions with unverified model-generated attempts as the teacher context. OASIS therefore requires only final-answer labels. Across Qwen3-1.7B, 4B, and 8B on AIME 2024, AIME 2025, and HMMT 2025, OASIS improves over the base model by 3.2–3.8 points on average, while OPSD’s gain falls from 3.05 points at 1.7B to 0.14 at 8B. At 8B, OASIS improves over OPSD by 3.05 points, showing that verified on-policy scaffolds preserve the effectiveness of self-distillation as models scale.

[NLP-19] he Unequal Influence of Bad Advice: Using Training Data Attribution to Modulate Emergent Misalignment

【速读】: 该论文旨在解决大语言模型在特定、不匹配任务上微调时出现的涌现性错位(Emergent Misalignment, EM)问题,即微调过程可能破坏模型经过后训练阶段建立的对齐关系,并诱发新型错误行为。其核心挑战在于:尚不清楚训练数据中哪些属性驱动了这一现象,具体包括有害样本是否对错位具有同等贡献,以及不同模型是否对相同有害样本表现出一致的敏感性。为此,论文提出采用训练数据归因(training data attribution)方法,定量评估每个有害样本对EM的贡献程度。关键解决方案是通过再训练验证机制来评估归因得分的有效性——一个可靠的归因分数应能通过筛选高分或低分数据实例,显著增强或削弱EM现象。实验表明,基于得分的数据过滤可显著调控EM;同时,无论是基于归因得分还是黑盒有害性评分,均能有效识别出对错位有决定性影响的样本。所有测试模型在相同数据集上均会引发错位,且当使用同一模型生成的影响分数进行数据筛选时效果最佳;尽管跨模型间影响分数具有一定泛化能力(在三种模型族间表现一致),但无法完全复现同模型筛选的性能。因此,该研究的关键在于揭示了数据归因与影响分数在精准操控和理解模型错位机制中的核心作用。

链接: https://arxiv.org/abs/2609.37914
作者: Gonçalo Paulo,Louis Jaburi,Nora Belrose,Lucia Quirke,Stella Biderman
机构: EleutherAI
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Fine-tuning large language models on narrow, misaligned tasks can undo their post-training alignment and induce novel misaligned behaviors – a phenomenon known as \emphemergent misalignment (EM). EM has been linked to persona-like representations, where fine-tuning might reduce loss by amplifying a harmful or ‘evil’ persona. It remains unclear which properties of the training data drive this effect: whether all harmful examples contribute approximately equally to misalignment and whether different models are equally affected by the same fine-tuning examples. In this work, we use training data attribution to quantitatively estimate how much each harmful example contributes to EM. We benchmark the quality of the attribution via retraining – a sound attribution score should enable us to enhance or attenuate EM by filtering data on that score. Score-based filtering can substantially enhance or attenuate EM; we find that both data-attribution scores and a black-box harmfulness score can identify consequential examples. All models we test become misaligned when trained on the same dataset, and influence scores perform best when filtering data from the same model that computed them. We find cross-model generalization of influence scores from scores derived from the three model families we tested, but this generalization does not recover same model filtering performance.

[NLP-20] Its All Training: A Fully Synthetic Single-Stage Recipe for LLM s NEURIPS2026

【速读】: 该论文旨在解决当前预训练数据集(主要源自网络爬取)在支持模型中后期训练流程(如推理能力增强)方面的不足,尤其是缺乏显式推理过程的问题。其核心挑战在于:传统数据集无法有效支撑从预训练到微调的全流程训练,导致模型在知识获取与推理能力方面存在“冷启动”缺陷。为此,论文提出一种名为SYNTH的开源合成语料库,通过结构化扩增精选的百科全书种子内容(基于58,698篇维基百科文章),将预训练、中段训练与后段训练统一于单一训练阶段,实现端到端的高效训练范式。其关键创新在于采用“反向翻译”(back-translation)机制生成高质量、事实精准的合成数据,确保模型在极低训练词元数量(仅为现有数据的1/10至1/140)下仍能保持高事实准确性,并通过种子语料库对记忆行为进行定向控制。实验表明,在同等计算成本下,SYNTH显著优于过滤后的网络数据,且所训练的小型(56M)、中型(0.3B–0.6B)及大型(13B / 1B-active MoE)模型均达到与同类开源基线相当的性能。这一成果揭示了合成数据在提升语言模型数据效率方面的巨大潜力,不仅适用于通用大模型的快速迭代,也为缺乏指令或对话数据的领域特定模型提供了可行路径。最终,研究团队公开发布了SYNTH数据集及配套的Baguettotron系列模型,推动开放源代码语言模型生态的发展。

链接: https://arxiv.org/abs/2609.37891
作者: Pierre-Carl Langlais,Pieter Delobelle,Yannick Detrois,Pavel Chizhov,Carlos Rosas-Hinostroza,Neil Si Smail,Benjamin Burtin,Hanna Shcharbakova,Ivan Yamshchikov,Anastasia Stasenko
机构: PleIAs; Sorbonne Center for Artificial Intelligence (索邦人工智能中心); Sciences Po Médialab (巴黎政治学院媒体实验室); EPFL (洛桑联邦理工学院); CAIRO, Technical University of Applied Sciences Würzburg-Schweinfurt (CAIRO,维尔茨堡-施韦因富特应用科学大学); Paris Dauphine-PSL (巴黎-dauphine-PSL大学); Lattice, ENS-PSL (Lattice,巴黎高等师范学校-PSL); TU Munich, Munich Center for Machine Learning (慕尼黑工业大学,慕尼黑机器学习中心); KU Leuven (鲁汶大学)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: Accepted at NeurIPS 2026. 35 pages, 9 figures. Dataset: this https URL

点击查看摘要

Abstract:Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines–for instance, they contain little explicit reasoning. Thus, many frontier labs have begun to develop their own internal datasets, starting from state-of-the-art models, to augment their pre-training data mix, eg, with reasoning traces to address cold-start problems. While demonstratively effective, none of these datasets are public, and the effect of this so-called synthetic data on knowledge and skill acquisition of language models, including small ones, remains poorly understood. We present SYNTH, the first open-source synthetic corpus derived from 58,698 Wikipedia articles that collapses pre-, mid-, and post-training into a single training stage via structured amplification of curated encyclopedic seeds. We evaluate SYNTH by training a suite of models: a 56M tiny model (Monad), 0.3B-0.6B dense models (Baguettotron), and a 13B / 1B-active MoE. At iso-compute, SYNTH outperforms filtered web data, and our models remain competitive with similarly-sized open-weight baselines. Because SYNTH is back-translated from grounded passages, SYNTH-trained models achieve high factual precision despite 10-140x fewer training tokens, with memorization targeted by the seed corpus. These results show that synthetic datasets, including our SYNTH dataset, are capable of producing competitive generalist models from a fraction of the training data, enabling rapid iteration as the frontier advances. These findings open up possibilities for both generalist models with significantly increased data efficiency, as well as domain-specific models where no instruction or conversational data is available. Finally, we publicly release our SYNTH dataset and the suite of Baguettotron models under a permissive license, thus supporting open-source language model development.

[NLP-21] Zero-shot Dependency Parsing with Unsupervised Cross-Lingual Bootstrapping

【速读】: 该论文旨在解决预训练语言模型(PLM)在依赖句法分析任务中跨语言零样本迁移能力不足的问题,尤其针对语法结构复杂且资源匮乏的语言。其核心挑战在于依赖句法分析具有强烈的句法特性,而现有基于编码器的PLM在缺乏标注数据的情况下难以有效捕捉跨语言的语法知识。为此,作者提出一种跨语言无监督自举(cross-lingual unsupervised bootstrapping)方法,通过在多语言语料上迭代地增强模型对句法结构的表征能力,从而提升模型在低资源语言上的泛化性能。该解决方案的关键在于利用无监督方式从大规模多语言文本中自动挖掘并强化模型的句法知识,进而显著提升零样本条件下的依存句法分析准确率,并通过参数无关的树探针测试验证了模型在识别句法结构方面的鲁棒性增强。

链接: https://arxiv.org/abs/2609.37883
作者: Lalita Lowphansirikul,Attapol Rutherford,Jian Gang Ngui,Sarana Nutanong,Peerat Limkonchotiwat
机构: VISTEC(泰国视觉技术研究中心); Chulalongkorn University(朱拉隆功大学); AI Singapore(新加坡人工智能研究所)
类目: Computation and Language (cs.CL)
备注: 11 pages, 4 figures

点击查看摘要

Abstract:Pre-trained language models (PLMs) with encoder-based architectures have shown impressive capabilities in zero-shot cross-lingual transfer for various language understanding tasks. However, applying this technique to dependency parsing remains a significant challenge due to its syntactic nature. To boost model generalizability across linguistic typologies, we propose a cross-lingual unsupervised bootstrapping method to improve syntactic knowledge within the PLM. We show that our method achieves a significant improvement in zero-shot parsing performance in low-resource languages. Analysis of these bootstrapped models uncovers increased robustness in recognizing syntactic structures, evidenced by higher scores in parameter-free tree probing tests.

[NLP-22] How Many Labels Does a Language Need? Annotation Budgets and Cross-Lingual Pooling for African-Language Text Classification

【速读】: 该论文旨在解决非洲语言文本分类任务中标签数据稀缺性与跨语言迁移有效性之间的核心问题,具体聚焦于两个关键科学问题:一是每种非洲语言在不同任务下达到接近全数据性能所需的标注样本量;二是能否利用其他非洲语言的标注数据替代目标语言的标注数据以缓解数据稀缺。针对28个语言-任务对(包括16种语言的新闻主题分类任务和12种语言的推文情感分类任务),研究通过一个基于字符n-gram的线性模型进行实证分析,该模型仅需两核CPU、数秒训练时间且无需预训练权重或加速器。结果显示,主题分类任务在中位语言上约需400条标注数据即可达到全数据宏F1的90%,而情感分类任务在11/12语言中仍需数千条数据才能收敛。跨语言数据池化在小样本场景下显著提升性能(如25条目标标签时,新闻分类平均提升0.20宏F1,部分语言达0.43),但随着标注量增加(>800条)其增益迅速衰减至零,且在全数据规模下反而导致性能下降(分别在9/16和8/12语言中出现)。此外,零样本迁移仅能恢复语言内模型与多数类基线之间差距的13%(新闻)和4%(情感),少数例外由共享书写系统(如阿姆哈拉语与提格雷尼亚语)、共享词汇(如英语与尼日利亚皮钦语、阿拉伯方言)或标签先验而非语言亲缘关系解释。研究提出的关键解决方案在于:构建一个可复现的基准框架,提供从数据预算到标注策略的量化指导,并释放代码以支持无GPU团队高效开发非洲语言分类器,从而实现低资源环境下的高性能建模。

链接: https://arxiv.org/abs/2609.37882
作者: Bhanu Prakash Vangala,Sowmya Guda,Navya Vangala
机构: University of Missouri (密苏里大学); Amar Bio Tech Pvt Ltd (阿玛尔生物技术有限公司)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Every text classifier for an African language begins with a budgeting question: how many labelled examples are needed, and can labels from other African languages stand in for them? We answer both questions empirically for 28 language-task pairs, news topic classification in 16 languages (MasakhaNEWS) and tweet sentiment in 12 languages (AfriSenti), using a character n-gram linear model that trains in seconds on two CPU cores with no pretrained weights and no accelerator. Monolingual learning curves at budgets from 25 to several thousand labels show that topic classification reaches 90% of its full-data macro-F1 with about 400 labels in the median language, while sentiment is still improving at the full training size in 11 of 12 languages and needs thousands of labels. Pooling the full training data of the other languages in the benchmark is worth a great deal at small budgets and nothing at large ones: at 25 target labels it adds 0.20 macro-F1 on average for news (up to 0.43 for Lingala) and 0.08 for sentiment, the gain decays to zero by 800 labels, and at full size pooling hurts in 9 of 16 and 8 of 12 languages. Twenty-five target labels plus pooled data match what 100 to 400 monolingual labels achieve for most news languages. A complete zero-shot transfer matrix shows that transfer without any target labels recovers a median of only 13% (news) and 4% (sentiment) of the gap between a majority-class predictor and the in-language model, with the exceptions explained by shared script (Amharic and Tigrinya), shared lexicon (English and Nigerian Pidgin, the Arabic dialects), or a shared label prior rather than by language family. We release code that regenerates every number from the public benchmark files and translate the results into concrete annotation guidance for teams building African-language classifiers without GPUs.

[NLP-23] Retrieval Capacity of Self-Attention Under Competition

【速读】: 该论文旨在解决语言模型在推理过程中实际依赖的上下文片段规模及其决定因素这一核心问题,即:模型在生成预测时究竟使用了多少上下文中的标记(token),以及影响该有效注意力集合大小的关键机制是什么。其解决方案的关键在于通过自注意力机制分析,在不进行微调的前提下,仅保留每个注意力头、每一层及每个查询中注意力权重最高的若干标记,并保持其原始权重不变,从而系统性地评估不同选择集合大小对负对数似然(NLL)的影响。研究发现,相对较小的注意力集合即可维持接近全注意力基线的损失水平,且基于注意力权重的选择显著优于随机选择;此外,所选集合呈现出一定的几何结构特征,但几何分离本身不足以保证模型性能的保持。随着上下文长度增加,为维持相同预测性能所需的注意力集合规模增大,但其占总上下文的比例却下降。实验还表明,额外背景信息会降低关键支持事实的注意力排名和注意力质量,而对保留权重进行重归一化可显著减少所需集合规模,说明注意力聚合方式对有效集合大小具有重要影响。基于条件理论模型的分析进一步揭示了注意力竞争与注意力质量保留机制如何导致集合规模随上下文增长而扩大,即使新增信息并未带来更丰富的语义内容。综上,该研究提出了一种量化语言模型有效注意力集合大小的方法,并深入揭示了其受上下文长度、注意力竞争和表示聚合方式等多重因素调控的内在机理。

链接: https://arxiv.org/abs/2609.37879
作者: Timur Mudarisov,Mikhail Burtsev,Radu State
机构: University of Luxembourg(卢森堡大学); London Institute for Mathematical Sciences(伦敦数学科学研究所)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:How many tokens from its context does a language model actually use, and what determines that number? We study this question through self-attention. Without retraining, we retain only the tokens with the highest attention weights at each head, layer, and query, keeping their original weights unchanged. By varying the selected set size and measuring the increase in negative log-likelihood (NLL), we estimate the effective attention set size needed to stay within a chosen loss tolerance. Relatively small selected sets can keep NLL close to the full-attention baseline, although the required size varies across models. Attention-based selection substantially outperforms random selection. Selected sets exhibit geometric structure, although geometric separation alone does not establish that model loss is preserved. Extending context while evaluating the same prediction targets increases the required set size, while its fraction of context decreases over the tested range. Experiments with a fixed supporting fact show that additional background pushes its tokens down the attention ranking and reduces their attention mass. Renormalizing the retained weights can substantially reduce the required set size, showing that it also depends on how selected representations are combined. Conditional theoretical models explain how competition and attention-mass retention can produce growing set sizes without more distinct information to retrieve. These results provide a way to measure effective attention set size in language models and investigate its dependence on context, competition, and aggregation.

[NLP-24] Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR

【速读】: 该论文旨在解决强化学习中基于可验证奖励(Reinforcement Learning with Verifiable Rewards, RLVR)方法(如GRPO)在有限的轨迹采样预算下,因生成全失败轨迹组而缺乏基于奖励的策略梯度信号的问题。其核心挑战在于:单个模型的采样可能遗漏成功轨迹,而这些轨迹可能已被其他异构模型发现。为解决此问题,论文提出GRAFT(Gated Replacement of Answer-Failed groups with peer Trajectories)框架,其关键创新在于引入一种无监督的跨模型知识迁移机制——通过门控机制将本模型的全失败轨迹组替换为来自同伴模型的具有信息量的轨迹(包括成功与失败样本),并利用同伴模型计算的优势值进行更新。同时,通过序列级兼容性加权与词元级重要性比率裁剪,有效控制不同模型间的分布偏移。实验表明,在三个异构模型对和五个数学推理基准上,GRAFT在相同每模型采样预算下显著优于GRPO,平均提升2.1点,最高达4.5点;且即使不进行联合训练,存储的同伴轨迹仍能带来平均1.8点的性能增益,验证了其高效的知识复用能力。

链接: https://arxiv.org/abs/2609.37868
作者: Doohyuk Jang,Yoonsik Park,Gyouk Chu,Sihwan Park,Eunho Yang
机构: KAIST(韩国科学技术院); AITRICS
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 29 pages, 11 figures, 9 tables

点击查看摘要

Abstract:Reinforcement Learning with Verifiable Rewards (RLVR) methods such as GRPO rely on successful self-generated trajectories, but finite rollout budgets can produce all-fail groups with no reward-based policy-gradient signal. While additional rollouts improve the chance of success at higher cost, successful trajectories missing from one model’s rollouts may already have been discovered by another. Indeed, we observe that heterogeneous models often succeed on complementary prompts, creating opportunities for mutual learning without a designated stronger teacher. To exploit this complementarity, we propose GRAFT (Gated Replacement of Answer-Failed groups with peer Trajectories), an off-policy-aware framework that replaces all-fail groups with informative peer groups. GRAFT transfers both successful and unsuccessful peer responses with peer-computed advantages, while controlling cross-model mismatch through sequence-level compatibility weighting and token-level importance ratio clipping. Across three heterogeneous model pairs and five mathematical reasoning benchmarks, GRAFT consistently improves both models over GRPO with the same per-model rollout budget, gaining 2.1 points on average and up to 4.5 points in model-level average performance. Stored peer trajectories preserve most of the gains, improving over GRPO by 1.8 points on average without simultaneous co-training.

[NLP-25] Its Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them NEURIPS2026

【速读】: 该论文旨在解决视觉语言模型(Vision-Language Models, VLMs)在替代人工标注者时,其可替代性评估结果易受偶然评价条件干扰的问题。现有测试往往未能准确反映模型的真实能力,而可能受到图像呈现方式等外部因素的影响。为此,作者提出MIST(Misleading-Image Stress Test),通过设计200个英文句子,每个句子包含一个可作字面义或隐喻义解读的短语,并配以三类图像:与句义一致的对齐图像、与句义相反的误导图像,或无图像。测试要求仅依据文本内容判断标签,因此图像不应影响判断结果。实验发现,尽管对齐图像和误导图像分别导致20.5%和19.4%的标签发生变化,但仅有37%的变化朝向图像所展示的语义方向,且模型与人工标注者的一致性在三种图像条件下保持不变。进一步分析表明,影响判断的关键并非图像内容本身,而是“是否存在图像”这一事实,说明可替代性评估结果不仅反映模型特性,也高度依赖于测试配置。因此,该研究的核心结论是:评估VLM可替代性的有效性取决于测试设计的严谨性,必须控制图像引入带来的认知偏差。

链接: https://arxiv.org/abs/2609.37863
作者: Nagham Omar,Mahmoud Jabarin,Kinan Ibraheem,Lotem Peled-Cohen
机构: Faculty of Data and Decision Sciences, Technion(以色列理工学院)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at TAE (Trust-AI-Eval) @ NeurIPS 2026

点击查看摘要

Abstract:Vision-language models (VLMs) are increasingly used in place of human annotators, making it important that substitutability tests reflect the model rather than incidental evaluation conditions. We introduce MIST, the Misleading-Image Stress Test: 200 English sentences, each built around a phrase readable either figuratively or literally and shown with an aligned image depicting its reading, a misleading image depicting the opposite, or no image at all. The guidelines require the label to be decided from the sentence alone, so no image should change any answer. We expected each image to pull a judge’s labels toward the sense it depicts, and neither kind did. Across thirteen VLM judges, an aligned image changed 20.5% of labels and a misleading one 19.4%, close for every judge and both above the 11.6% produced by deleting the ignore-the-image instruction with the image left in place. Yet only 37% of the labels that differ between the two images moved toward the sense shown, and agreement with our human annotators is unchanged whether the image is absent, aligned or misleading. The effect is smaller in the seven judges that pass the alt-test than in the six that never do, but present in all of them: what moves a judge is that an image is there, not which of the two it is, so a substitutability verdict describes a configuration as much as a model.

[NLP-26] One Threshold Does Not Fit All Languages: Language-Conditional Deferral for Reliable and Efficient Low-Resource Text Classification NEURIPS2026

【速读】: 该论文旨在解决在低收入国家(全球南方)部署多语言文本分类系统时面临的可靠性与可解释性难题:这类系统通常依赖单一模型处理多种语言,每种语言的标注样本极少,且严重依赖人工纠错。然而,现有方法无法为每种语言提供一致的错误率保证,导致部分语言的误判风险远高于预期。其核心问题在于:基于跨语言合并验证数据估计的统一置信度阈值虽能在整体上满足90%的覆盖率目标,但对某些语言(如索马里语、提格雷尼亚语及未见语言)的实际覆盖率显著下降,最低仅达77.5%。解决方案的关键是采用按语言独立校准置信度阈值的方法——即为每种语言分别估算一个阈值,无需重新训练模型,即可将所有语言的覆盖率提升至89.1%–91.0%之间。该方法揭示了“承诺成本”的显著不均衡性:维持高可靠性需将43%的索马里语新闻、超80%的阿姆哈拉语和西茨翁加语推文交由人工审核,而尼日利亚皮钦语仅需不足8%。实验证明,每种语言仅需一两百个标注样本,且可在单核CPU上数分钟内完成训练,因此该方案具有高度可操作性与经济可行性:应采取“逐语言校准、报告与预算人工审核”的策略,以实现真正可靠的多语言自动化分类。

链接: https://arxiv.org/abs/2609.37861
作者: Bhanu Prakash Vangala,Vangala Navya
机构: University of Missouri(密苏里大学); Amar Bio Tech Pvt Ltd(阿玛尔生物技术有限公司)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: Got accepted and published in NeurIPS 2026 GlobalSouthAI

点击查看摘要

Abstract:In the Global South, the lower-income countries of Africa, Asia, and Latin America where most of the world’s languages are spoken, a deployed text classifier usually runs on ordinary CPUs, serves many languages with a single model, has few labeled examples in any of them, and relies on people to catch its mistakes. Such a system is only useful if it can promise how often it will be wrong: at most a fixed fraction of the labels it assigns on its own may be incorrect, and everything else must go to a person. Split conformal prediction delivers this promise through a single confidence threshold, normally estimated on validation data pooled across languages. We ask whether the promise reaches every language, and it does not. On MasakhaNEWS (16 African languages) and AfriSenti (12 languages plus two never seen in training), a pooled threshold meets the 90% target on average but covers Somali at 77.5%, Tigrinya at 83.7%, and the two unseen languages at 77.5% and 81.2%. Estimating one threshold per language brings every language to between 89.1% and 91.0% without retraining, and it shows how unequal the cost of the promise is: keeping it means sending 43% of Somali news and over 80% of Amharic and Xitsonga tweets to a person, against under 8% of Nigerian Pidgin news. One or two hundred labels per language are enough and the models train in minutes on one CPU core, so the fix is affordable: calibrate, report, and budget human review one language at a time.

[NLP-27] Storag e Is Not Strategy: State-Conditioned Support Control for LLM Unlearning

【速读】: 该论文旨在解决局部化大语言模型(LLM)遗忘(unlearning)过程中,现有方法依赖静态参数子集选择所导致的干预有效性不足问题。传统方法通常基于定位信号选取固定参数子集并保持不变,但此类参数未必是优化过程中最应更新的,且候选干预效果可能随训练进程动态变化。其核心解决方案在于提出“干预评分”(Intervention Score),该评分通过预测实际遗忘更新对目标的影响,并同时考虑副作用(即共损),动态评估可编辑参数组的有效性。在此基础上,构建静态干预价值基线(Static-IV),并进一步引入选择性动态干预重排序(DIR-R)机制——仅在经过校准探测验证后才重新评估干预子集,从而实现更精准的干预策略调整。实验表明,该方法在自然语料和定位精度基准测试中均显著优于现有方法,在多数对比中表现出正向描述优势与更高的终端效用,尤其在梯度差异(GradDiff)目标下展现出稳健的性能提升,证实了将定位、初始干预选择与检查点依赖的策略修正相分离的必要性与有效性。

链接: https://arxiv.org/abs/2609.37858
作者: Tianhao Qian,Ziming Hong,Chongyang Gao,Kezhen Chen,Lixu Wang
机构: Southeast University (东南大学); University of Sydney (悉尼大学); Northwestern University (西北大学); Together AI (Together AI); The Chinese University of Hong Kong Shenzhen (香港中文大学(深圳))
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 18 pages

点击查看摘要

Abstract:Many localized large language model (LLM) unlearning methods select a small parameter subset from a localization signal and keep it fixed during optimization. The parameters most associated with a target, however, need not be the best ones to update, and candidate interventions can change value as optimization proceeds. In a controlled experiment, a storage-localization score reaches an area under the receiver operating characteristic curve (AUROC) of 0.981, yet storage identity agrees with the better intervention on only 17/36 targets, while low-rank adaptation (LoRA) wins 35/36. We introduce Intervention Score, which ranks editable groups by the predicted effect of the actual unlearning update while accounting for collateral damage, and use it to form the static intervention-value baseline (Static-IV). We then introduce selective dynamic intervention re-ranking (DIR-R), which revisits that subset only when a calibrated probe justifies the comparison. On the Natural-TOFU dataset, our method has positive descriptive margins in 19/20 comparisons between methods and objectives, although several are near zero. On the LACUNA localization-precision benchmark, our mean terminal utility is higher in all six negative preference optimization (NPO) and SimNPO comparisons: NPO margins range from +0.431 to +0.848, and SimNPO margins range from +0.503 to +0.571. The gradient-difference (GradDiff) objective reveals substantial field dependence. Relative to Static-IV, the primary four-field GradDiff evaluation has six wins, six ties, and no losses, with mean and median paired gains of +0.165 and +0.0025. The evidence supports separating localization, initial intervention selection, and checkpoint-dependent support revision.

[NLP-28] AnthroDial: Benchmarking LLM Anthropomorphism in Autonomous Social Interaction

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在作为社会代理(social agents)进行持续、开放式交互时,缺乏自主决策沟通能力及适应动态情境、目标与关系的能力这一核心问题。现有研究在实现、评估和优化此类类人交互能力方面缺乏统一框架。其解决方案的关键在于提出AnthroDial——一个涵盖三个互补维度的统一框架:(1)MindFlow,一种轻量级交互引擎,通过动态心智缓冲区(Mind Buffer)实现自主、异步与自适应通信;(2)CAPS-Eval,基于理论的评估框架,用于量化评估交互中的认知、情感与行为维度;(3)可扩展的训练范式,结合SEEDS实现环境扩展与DiAPO实现能力自适应优化。该框架通过构建覆盖日常交流、游戏互动与长期角色交互的评估数据集,实验证明其显著提升了交互自主性与自然度,并验证了CAPS-Eval在信度、区分度及与人类评分一致性方面的有效性,从而为开放域中可信类人社会代理的开发提供了系统性解决方案。

链接: https://arxiv.org/abs/2609.37853
作者: Wentao Liu,Xi Chen,Siyu Song,Biao Yuan,Yu Zhang,Zhou Zhuotong,Jingying Zhou,Guohao Feng,Shasha Hu,Tianfu Wang,Shangshang Yang,Haoyang Liu,Youjia Li,Xiaokun Wang,Min Ji,Ji Wang
机构: 未知
类目: Computation and Language (cs.CL)
备注: 26 pages, 8 figures, 16 tables

点击查看摘要

Abstract:Large language models (LLMs) are increasingly deployed as social agents, yet credible human-like interaction requires more than fluent responses or persona consistency. Agents must autonomously decide whether, when, and how to communicate while adapting to evolving contexts, goals, and relationships. Existing research, however, lacks a unified approach to enabling, evaluating, and improving such capabilities in continuous, open-ended interaction. We introduce AnthroDial, a unified framework for developing anthropomorphic social agents from three complementary aspects: MindFlow, a lightweight interaction harness that enables autonomous, asynchronous, and adaptive communication through a dynamic Mind Buffer; CAPS-Eval, a theory-grounded framework for evaluating cognitive, affective, and behavioral dimensions of anthropomorphic interaction; and a scalable training paradigm that combines SEEDS for environment expansion with DiAPO for adaptive capability optimization. We further construct evaluation datasets covering everyday communication, game interaction, and long-horizon character interaction. Extensive experiments across diverse models and scenarios demonstrate improved interaction autonomy and naturalness, validate the reliability, discriminativeness, and agreement with human rankings of CAPS-Eval, and confirm the effectiveness of our training paradigm. Together, these components provide a unified framework for developing credible human-like social agents in open-ended interaction.

[NLP-29] Can Vision-Language Models Stay Helpful When Facing Implicit Risks? Intent-Privilege OPSD for Efficient Safety-Helpfulness Alignment

【速读】: 该论文旨在解决视觉-语言模型(Vision-Language Models, VLMs)在跨模态情境下面临的隐性风险问题,即看似无害的视觉与文本输入单独出现时可能不会触发安全问题,但联合输入却可能诱导出不安全的响应。现有安全方法通常依赖大规模偏好数据集、高成本的多轮次训练(multi-rollout training)或推理阶段额外的安全防护机制,且常因对潜在可安全回答的请求采取全盘拒绝而牺牲模型的有用性。本文提出一种基于意图-权限的在线自蒸馏方法(Intent-Privilege On-Policy Self-Distillation, OPSD),其核心在于在训练过程中引入以证据为基础的意图(evidence-grounded intent)作为特权监督信号,使模型能够精准识别隐性风险,并在不依赖额外安全模块或意图标注的情况下生成既安全又有用的响应。OPSD通过单次提示采样(single rollout per prompt)将教师模型在特定意图条件下的偏好信息蒸馏至学生模型,显著降低数据需求(仅需1,447条安全特异性样本,较标准数据集减少95%)、训练时间(相比GRPO类方法减少5倍)和平均推理长度(减少7%)。在五个评估组中,OPSD实现了最高的“安全-有用性”联合成功率比率,尤其在合并的SIUO+HoliSafe基准上,该比率从43.9%提升至53.5%,充分验证了训练阶段引入意图监督可在大幅降低资源开销的同时,同步提升模型的安全性与实用性。

链接: https://arxiv.org/abs/2609.37837
作者: Haotian Deng,Wenbin Xing,Gang Xu,Tao He,Jinkai Zheng,Chun Li,Zheng Zhu,Ming Li
机构: Southern University of Science and Technology (南方科技大学); Sun Yat-sen University (中山大学); Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ) (广东省人工智能与数字经济实验室(深圳)); University of Electronic Science and Technology of China (电子科技大学); Hangzhou Dianzi University (杭州电子科技大学); Shenzhen MSU-BIT University (深圳北理莫斯科大学); GigaAI; The Chinese University of Hong Kong (Shenzhen) (香港中文大学(深圳)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Vision-Language Models (VLMs) remain vulnerable to cross-modal implicit risks: visual and textual inputs that appear benign in isolation can jointly elicit unsafe responses. Existing safety methods often require large preference datasets, costly multi-rollout training, or additional safeguards at inference time. They may also sacrifice helpfulness by directly refusing requests that could be answered safely. In this paper, we propose Intent-Privilege On-Policy Self-Distillation (OPSD), which leverages evidence-grounded intent as privileged supervision during training to help VLMs recognize implicit risks and provide safe, useful responses instead of blanket refusals. OPSD distills a teacher’s intent-conditioned preferences over responses into a student using a single rollout per prompt; the student then responds without intent annotations or an additional safety module. With only 1,447 safety-specific examples - 95% fewer than standard preference datasets - OPSD reduces training time by 5x relative to multi-rollout GRPO-style training and average inference length by 7%. It attains the highest ratio for joint safety-helpfulness success, which measures the proportion of responses that are both safe and helpful, across all five evaluation groups. Remarkably, on pooled SIUO+HoliSafe, this success ratio rises from 43.9% to 53.5%. These results show that training-time intent supervision can improve both safety and helpfulness while substantially reducing data, training, and inference costs.

[NLP-30] Can a Cacheable Decision Model Follow Rules?

【速读】: 该论文旨在解决小型非生成式决策模型(如Certo,基于Qwen3-4B)在面对规则依赖性任务时,如何在保持高准确率的前提下实现高效推理的问题。其核心挑战在于:传统的联合评分机制(joint scorer)虽能充分融合状态、规则与候选动作之间的上下文关系,从而保留对规则的高度敏感性,但计算成本随候选动作数量线性增长;而独立编码(independent encoding)虽可通过缓存提升效率(约77个候选时节省5倍计算开销),却将状态与候选动作分离,可能导致规则敏感性下降。解决方案的关键在于探索是否可通过特定训练策略恢复因解耦带来的规则感知能力——研究发现,仅使用标准监督学习无法有效恢复规则敏感性,但在引入目标导向的反事实监督(targeted counterfactual supervision)后,模型在未见的合成规则任务上可显著恢复性能(涵盖改写、反事实推理与组合等类型),表明规则敏感性可在训练阶段重建;然而,在真实规则场景中,尽管联合评分器在短难度层级(0.861 vs 0.500)和硬难度层级(0.655 vs 0.483)均表现更优,但缓存式编码器在跨域真实文本混合任务中反而降低合同准确性(-9.3至-16.2分),且无法证明其在未见源规则上的泛化能力。因此,论文结论为:虽然缓存编码器可在训练分布内实现规则敏感性,但其在真实世界规则迁移中的有效性尚未建立;相比之下,联合评分器虽牺牲缓存优势,仍维持更强的规则适应性与鲁棒性。

链接: https://arxiv.org/abs/2609.37832
作者: Dushyant Rajput,Nirdesh Chauhan,Siddharth Kosaraju(AltSlate Labs LLP)
机构: AltSlate Labs LLP(AltSlate实验室有限公司)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Certo is a small non-generative decision model (Qwen3-4B): it scores candidate actions from their text and returns a probability, instead of generating an answer. The accurate design reads the state, the rules, and each candidate together (a joint scorer), so cost grows with the menu. Independent encoding lets each candidate be encoded once and reused across states (about 5x cheaper at 77 candidates), but separates state from candidate. We ask how much rule-sensitivity survives that move, and whether it can be trained back. Four experiments on Certo: (1) the tested conversion to cacheable scoring loses rule-sensitivity (recall@1 1.00 - 0.24) while the joint scorer holds 1.00, and a shortlist+rerank rescue fails; (2) targeted counterfactual supervision restores strong performance on held-out synthetic rule tasks (paraphrase, counterfactual, composition; reproducible across seeds), though we do not isolate whether predictions depend on the supplied rule; (3) on real rules the added benefit is not established – after fixing a truncation confound, the joint scorer wins significantly on the short tier (0.861 vs 0.500) and directionally on the hard tier (0.655 vs 0.483, n=29); (4) a matched cross-domain real-prose mixture did not help and reduced contract accuracy (-9.3, -16.2 points). A cacheable encoder can be made rule-sensitive on its training distribution, but transfer to unseen-source real rules is not established; the joint scorer keeps an edge at the cost of caching. Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) Cite as: arXiv:2609.37832 [cs.AI] (or arXiv:2609.37832v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2609.37832 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[NLP-31] he Geometry of Inference in Transformer Residual Streams

【速读】: 该论文旨在解决生成式语言模型在推理过程中,其内部表示如何逐步收敛至特定输出结果的机制不明确的问题。核心挑战在于理解残差状态(residual states)在深度网络中的演化路径,尤其是这些中间表示如何随层深增加而逐渐具备对最终输出的几何特异性。论文的关键解决方案在于提出并验证一个高维空间中的简化模型,该模型将表示的几何特性分解为模长(norm)、方向对齐(directional alignment)与终点几何结构(endpoint geometry)三个独立成分,揭示了即使欧氏距离变化微小,方向性对齐与终点排名仍可显著提升。研究发现,尽管各模型中“自身终点”(own endpoint)在早期即优于平均替代终点,但个体终点间的竞争关系随深度演变,且幸存终点间未必趋于相似,表明竞争集合的缩小并非源于终点趋同,而是由方向性渐变引发的剧烈竞争削减。此外,作者证明从中间状态到自身终点的直线路径不会引入新竞争者,而实际观察到的竞争现象则表明模型偏离了理想化的直线收敛。最后,低排名输出词对应的终点在余弦距离上普遍更远,揭示了残差空间几何结构与输出分布之间的内在关联。综上,该研究通过多维度几何分析,阐明了Transformer模型在推理过程中逐步增强的几何特异性机制,并指出距离、竞争数量与终点集中度分别反映了该过程的不同侧面。

链接: https://arxiv.org/abs/2609.37824
作者: Timur Mudarisov,Mikhail Burtsev,Radu State
机构: University of Luxembourg(卢森堡大学); London Institute for Mathematical Sciences(伦敦数理科学研究所)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Transformer language models build predictions through successive residual updates, but how their representations become specific to an eventual outcome remains unclear. We study this process by comparing intermediate residual states with their own final states and an empirical bank of final states from other contexts. Across six pretrained language models, the own endpoint becomes preferable to the average alternative early, while many individual endpoints remain closer. These competing sets generally shrink with depth, but their membership changes and their surviving endpoints need not become more similar to one another. Directional alignment and endpoint rank can therefore improve while Euclidean distance to the final state changes little. We develop a simple high-dimensional model that separates the roles of norm, alignment, and endpoint geometry, showing how gradual directional changes can produce sharp reductions in competition. We also prove that a straight path toward the own endpoint cannot introduce new competitors under either Euclidean or cosine distance; observed entries thus establish departures from straight-line convergence. Finally, endpoints associated with lower-ranked output tokens tend to lie farther away in cosine distance across all studied models, connecting residual geometry to output organization. Together, these findings characterize increasing geometric specificity during transformer inference and explain why distance, competitor count, and concentration of the surviving endpoints provide distinct views of that process.

[NLP-32] hinking in Depth Speaking Directly: Recurrent Latent Reasoning for Paralinguistically Grounded Spoken Dialogue

【速读】: 该论文旨在解决情感化语音对话中“感知-推理鸿沟”(perception-reasoning gap)问题,即现有模型虽能通过显式思维链(CoT)增强对副语言特征(paralinguistic features)的感知,但难以有效将这些声学线索融入响应规划过程。此外,传统CoT生成存在声学信息捕捉不充分及推理延迟高等缺陷。为此,本文提出LoopSLM,其核心创新在于基于循环Transformer架构实现隐式推理,通过重复利用解码器模块在每轮迭代中以声学信息为锚点精炼隐藏状态,从而强化声学感知与推理之间的对齐。该模型采用两阶段训练策略,将“推理学习”与“响应生成”解耦,支持无需CoT的直接推理,显著降低计算开销。实验表明,相较于Qwen2.5-Omni-7B和CoT-SFT基线,LoopSLM在EchoMind数据集上提升副语言理解、推理准确率超过20个百分点,同时生成token减少64.5%,延迟降低至一半;在多数情感化回复指标上超越Qwen3-Omni-Thinking,且延迟降低34倍。值得注意的是,仅在对话数据上训练的LoopSLM亦展现出对通用音频基准任务的泛化能力提升。

链接: https://arxiv.org/abs/2609.37818
作者: Shengbo Cai,Yuxiang Wang,Jingran Xie,Zhisheng Zhang,Shun Lei,Di Cao,Teddy Sun,Zhiyong Wu
机构: Tencent Hunyuan(腾讯混元); University of Science and Technology of China (中国科学技术大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Empathetic spoken dialogue requires models to use both what is said and how it is said to decide how to respond. Explicit CoT can improve paralinguistic perception and make acoustic cues more explicit in replies, yet does not ensure their effective use in response planning. We call this mismatch the perception-reasoning gap. In addition, CoT may not fully capture acoustic cues in words, and generating it adds inference latency. To address these limitations, we introduce LoopSLM, which builds on looped Transformers for latent reasoning, reusing a decoder block to refine hidden states with acoustic grounding at every pass. Its two-stage training further narrows the perception-reasoning gap by separating learning to reason from learning to respond, enabling direct inference without CoT. On EchoMind, LoopSLM improves paralinguistic understanding, reasoning, and reply quality over Qwen2.5-Omni-7B. Against the CoT-SFT baseline, LoopSLM gains over 20 points in reasoning accuracy while generating 64.5% fewer tokens at half the latency. It also outperforms Qwen3-Omni-Thinking on most empathetic reply metrics with 34x lower latency. Despite training only on dialogue data, LoopSLM improves accuracy on general audio benchmarks.

[NLP-33] CompOrca: Corpus-Scale Compliance Labelling of Instruction-Tuning Data AACL

【速读】: 该论文旨在解决生成式 AI(Generative AI)在微调过程中对指令拒绝(refusal)与不合规行为(noncompliance)的识别难题,即如何系统性地识别训练数据中那些拒绝执行任务、规避响应或未能完成请求的样本。现有研究受限于标注规模,仅覆盖数千个提示(prompt)级别的评估集,难以全面反映真实场景下的不合规模式。为此,论文提出 CompOrca,一个针对包含 423 万条样本的 OpenOrca 全体语料库的合规性标注体系,通过五轮独立运行的开源大模型评判器(LongCat-2.0,1.6T 参数)对每条样本进行合规性判断,并依据投票一致性划分为三类:一致合规(94.75%)、一致不合规(1.28%)及非一致样本(3.97%),同时公开原始投票计数。研究发现,单次判断即能识别出 2.7%–3.2% 的不合规样本,但仅有 1.28% 被全部五轮判定为不合规,从而实现了对高模糊性样本的有效筛选。基于 450 个经人工标注的样本(其中 150 个重复标注,人与人之间 Kappa 值达 0.93)验证,一致合规与一致不合规标签的精确率分别达到 97.3% 和 86.7%,后者构成高精度的不合规子集而非完整枚举。现有拒绝检测方法对不合规类别的召回率仅为 0.4% 至 94.1%,存在显著缺陷。因此,本研究的核心解决方案在于构建大规模、多轮共识驱动的高质量合规性标注数据集,为理解并建模生成式 AI 的拒绝行为提供了可扩展、可验证的基准资源。

链接: https://arxiv.org/abs/2609.37807
作者: Philipp E. Glass,Alina Miron
机构: Brunel University of London(布鲁内尔大学)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: Accepted to PlurVA-LLM Workshop @ AACL-IJCNLP 2026. Dataset available on HuggingFace

点击查看摘要

Abstract:Studying how fine-tuning shapes refusal and noncompliance behaviour requires identifying training examples that refuse, evade or otherwise fail to fulfil the requested task. But existing annotation covers evaluation sets of a few thousand prompts at most. We present CompOrca, a compliance labelling over the entirety of the 4,233,923-example OpenOrca corpus. Every example was classified as compliant or noncompliant by five independent passes of an open-weight LLM judge (LongCat-2.0, 1.6T parameters), and the corpus is released as unanimous compliance (94.75%), unanimous noncompliance (1.28%), and nonunanimous rows (3.97%) along with the raw vote counts. A single pass flags 2.7-3.2% of the corpus as noncompliant, while only 1.28% is flagged by all five, allowing for filtering the most ambiguous samples. Against 450 human-annotated examples, 150 of them annotated twice (human-human \kappa = 0.93 ), the unanimous compliance and noncompliance labels are 97.3% and 86.7% precise, the latter a high-precision subset, not a complete enumeration, of noncompliance. Published refusal-detection methods recall only between 0.4% and 94.1% of the noncompliance class. We release the full corpus with its per-row labels and vote counts at this https URL

[NLP-34] A Proposed Rubric for Evaluating Expressed Clinical Reasoning in Large Language Model Responses

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在临床推理(clinical reasoning)评估中缺乏系统性、可解释且具临床意义的评价工具的问题。现有评估框架多集中于通用推理能力或非临床场景,难以准确衡量模型在真实医疗情境下的表现,尤其在涉及临床决策质量与安全性方面存在盲区。为此,作者提出了一种新型评分量表(rubric),其核心在于整合医学教育评估框架(如ART、SCT、关键特征问题与OSCE)、临床领域专用基准测试(如MedR-Bench、HealthBench、TIMER-Bench等)以及通用大模型推理评估研究中的关键维度,构建一个多层次、结构化的评价体系。该量表的关键创新在于将“事实性”(factuality)概念转化为面向临床的“真实性”(groundedness),并引入行为锚点(behavioural anchors)、适用性规则及针对特定病例的安全关键错误标记机制,以增强评估的透明性与可追溯性。尽管该量表尚未经过信度、效度和临床实用性验证,其主要目标是为临床推理评估提供一个可公开审视、可操作化的方法论框架,从而推动后续实证研究的发展。

链接: https://arxiv.org/abs/2609.37788
作者: Zhangshu Joshua Jiang,Zina Ibrahim,James T. Teo
机构: King’s College London (国王学院); King’s College Hospital NHS Foundation Trust (国王学院医院NHS基金会信托); Neurological Institute, Cleveland Clinic London (克利夫兰诊所伦敦神经学研究所)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 20 pages

点击查看摘要

Abstract:Rubrics support the structured evaluation of language models. We propose a rubric for assessing expressed clinical reasoning in model responses, drawing on three bodies of work: medical education assessment frameworks (ART, SCT, Key Feature Problems and OSCE); clinical LLM benchmarks (MedR-Bench, HealthBench, TIMER-Bench, this http URL, PrIME-LLM and PatientSafeBench); and general LLM reasoning evaluation research, including the Factuality-Validity-Coherence-Utility taxonomy, FaithCoT-Bench and C2-Faith. We use groundedness as a clinically oriented adaptation of the taxonomy’s factuality category. The rubric brings these concepts together in a multidimensional framework for scoring free-text responses to gold-standard clinical vignettes. It includes provisional behavioural anchors, applicability rules and a separate flag for case-specific safety-critical errors. General-domain frameworks inform its design but are not treated as validated clinical instruments. The rubric does not replace case-specific reference criteria or the task-specific metrics of existing benchmarks. It has not yet been tested for inter-rater reliability, construct validity or clinical utility. Its immediate purpose is to make evaluation decisions explicit and open to scrutiny before empirical testing.

[NLP-35] Selecting What Matters: Semantic Compression-Guided Selective Pooling for Long-Context Embeddings

【速读】: 该论文旨在解决长文本上下文嵌入中因均值池化(mean pooling)导致关键语义信息被冗余或弱相关信息稀释的问题。现有方法在因果注意力机制下虽优化了信息流动,但普遍采用均匀平均所有词元表示的策略,难以有效保留长文档中的重要语义内容。其解决方案的关键在于提出一种无需训练的框架SCSP(Semantic Compression-based Selection for long-context embedding),通过引入语义压缩提示(semantic compression prompt)对文档进行句级感知分块,并设计提示隔离注意力掩码,在保持全局信息流动的同时,限制每个提示仅作用于局部上下文。利用这些提示诱发的注意力模式,系统可自动评估词元重要性,进而选择最具信息量的词元,并聚合其中间层表示生成最终嵌入。该方法具备即插即用特性,可无缝集成至零样本与微调模型中,显著提升长文本嵌入性能。

链接: https://arxiv.org/abs/2609.37782
作者: Zifeng Cheng,Jie Zheng,Zhiwei Jiang,Shuwen Wang,Fei Shen,Shiping Ge,Qing Gu
机构: Nanjing University (南京大学); National University of Singapore (新加坡国立大学); Nanjing University of Posts and Telecommunications (南京邮电大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Large language models (LLMs) have shown strong potential as training-free text encoders for long-context embeddings. Existing approaches primarily improve information flow under causal attention and typically construct embeddings by uniformly averaging all token representations. However, for long documents, such mean pooling can dilute salient semantic information with abundant redundant or weakly informative content. To this end, we propose SCSP, a training-free framework that leverages semantic compression for informative token selection in long-context embedding. Specifically, SCSP first partitions a document into sentence-aware chunks and appends a semantic compression prompt to each chunk. A prompt-isolated attention mask preserves information flow among document tokens while restricting each prompt to its corresponding local context. We then use the attention patterns elicited by these prompts to estimate token importance, select informative tokens, and aggregate their intermediate-layer representations into the final embedding. Extensive experiments on long-context embedding benchmarks demonstrate that SCSP can be integrated into both zero-shot and fine-tuned models in a plug-and-play manner, consistently improving their performance.

[NLP-36] Which papyrus HTR is good enough? Character-error-rate tolerance of four papyrological tasks on Greek texts

【速读】: 该论文旨在解决古希腊纸草文献(Greek papyri)大量未被出版和数字化的问题,提出通过手写文本识别(HTR)技术实现其自动转录,以帮助学者发掘长期被忽视的文献。当前针对古希腊纸草的HTR系统尚处于起步阶段,且尚未明确不同皮拉莫尔学任务对识别准确率的具体要求。为此,研究通过模拟从“完美HTR”输出中引入种子算法生成特定字符错误率(CER)为1%至50%的退化文本,并结合丢失行及四种错误形态变体,构建了多层级测试数据集。在此基础上,训练小型模型(TF-IDF、fastText、字符卷积神经网络、ByT5-small)完成文献类型分类、断代、文书性与文学性区分以及关键词检索四项任务,并对比在干净文本上训练与在特定CER水平下重训练的模型性能。研究结果表明,各类任务对错误率的容忍度存在显著差异:在使用干净文本训练的模型中,文书-文学分类可容忍高达20% CER,文献类型识别为7.5%,子类与关键词检索为5%,而断代任务仅能承受3%以下;但若将模型在含错误的文本上重新训练,则可显著缓解超过15% CER后的性能骤降问题。此外,模型对长文档中集中分布的损坏比短文本中分散的小错误更具鲁棒性。结论指出,本研究为四项皮拉莫尔学任务设定了相应的CER基准,并证明在噪声文本上训练的模型能够有效利用当前不完美的文本识别结果,从而提升实际应用价值。

链接: https://arxiv.org/abs/2609.37755
作者: Anton Repushko,Elena Chepel
机构: Independent Researcher(独立研究员); University of Vienna (维也纳大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Purpose: Most Greek papyri remain unpublished and undigitised; a handwritten text recognition (HTR) pipeline that transcribes them automatically would let scholars discover documents and literary works that have so far gone unread. Recognition systems for Ancient Greek papyri are in statu nascendi, and how accurate they must be for a given papyrological task has not been examined. To answer this and set a benchmark for Greek papyrus HTR, we test a range of character error rates (CER) against four papyrological tasks, using published editions as ground truth. Methods: From 63,846 current editions of Greek texts in this http URL, we imitate a letters-only “perfect HTR” output by removing the editorial layer, then degrade it with a seeded algorithm to exact CERs of 1 - 50%, with lost lines and four error-shape variants. On these data we train small models (TF-IDF, fastText, a character CNN, ByT5-small) for document type, dating and documentary-versus-literary classification, and apply eight keyword search methods. We compare models trained on clean text with models retrained at a specific CER level, and evaluate across CERs. Results: Tolerance differs by task. With clean-trained models, documentary-versus-literary classification retains 90% of its metric up to 20% CER; document type up to 7.5%; subtypes and search up to 5%; dating only up to 3%. Retraining on text containing character errors largely eliminates the sharp degradation that otherwise sets in above 15% CER. Models generally tolerate concentrated damage in a long document better than small errors spread across a short text. Conclusion: The study provides a CER target for each of the four tasks and shows that models trained on noisy text make current, imperfect text recognition useful for them.

[NLP-37] Context Language Models

【速读】: 该论文旨在解决现有语言模型在上下文管理(context management)中依赖外部控制机制所带来的效率低下与灵活性不足的问题。传统方法通常将上下文作为静态输入或受限的缓存结构,难以动态适应复杂任务需求,尤其在多智能体协作或长时序交互场景下表现受限。为此,论文提出上下文语言模型(Context Language Models, CLMs),其核心创新在于将上下文视为可被模型自主读写和更新的“文件”(file),从而实现对上下文内容的原生、无限制管理。这一设计使模型能够通过学习自动识别并保留关键信息,显著提升上下文利用效率。关键技术突破在于将上下文管理策略从外部调度转移到模型内部行为,实现了上下文管理的内在化(intrinsic)与可学习性,支持在上下文中进行非参数化(in-context)与参数化(parametric)双重学习。实验表明,基于现有模型零样本构建的CLMs在多项基准测试中均超越当前最优(SOTA)方法:在BrowseComp-Plus上实现11.4%更高的准确率且减少21.5%的浮点运算量(FLOPs),在12小时EdgeBench任务中提升5%得分并节省59%计算资源,在24小时多仓库智能体群组任务中性能提升达65%而计算成本不变。此外,通过自然语言指令引导的技能优化循环,可进一步提升上下文管理能力,最高提升35.9个点精度的同时降低计算开销;引入在线强化学习机制后,Qwen3.5-9B在BrowseComp-Plus上的性能提高47.6%,同时减少12%的FLOPs。最后,论文还协同设计了后缀缓存重用(Suffix Cache Reuse)技术,使CLM服务端计算负载相比标准SGLang降低35%,在保持相同性能前提下实现显著加速。

链接: https://arxiv.org/abs/2609.37725
作者: Rulin Shao,Shannon Zejiang Shen,Junjie Oscar Yin,Yuetai Li,Minheng Wang,Hamish Ivison,Radha Poovendran,Nathan Lambert,Teng Xiao,Mike Lewis,Wen-tau Yih,Luke Zettlemoyer,Pang Wei Koh
机构: University of Washington (华盛顿大学); Meta Superintelligence Labs (Meta超智能实验室)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends to multi-agent systems where multiple agent contexts coexist as files. Building CLMs zero-shot with existing models outperforms SOTA context management strategies across a variety of tasks: 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater improvement with the same compute on a 24-hour multi-repository agent-swarm task. Moreover, by shifting context management from external harness control to intrinsic model behavior, CLMs naturally enable both in-context and parametric learning of context-management strategies. We show that CLMs can be steered with natural-language instructions evolved through a standard skill-optimization loop, improving held-out accuracy by up to 35.9 points on a context-management task while reducing compute. We also introduce an online reinforcement learning method for CLMs, improving Qwen3.5-9B performance on BrowseComp-Plus by 47.6% while using 12% fewer FLOPs. Finally, we co-design Suffix Cache Reuse for CLM serving, further reducing server-side compute by 35% relative to standard SGLang at matched performance.

[NLP-38] Predictive Geometry of Hidden Trajectories in Transformers

【速读】: 该论文旨在解决生成式语言模型中解码器仅(decoder-only)Transformer架构在训练过程中,由于仅通过最终的下一个词预测损失(next-token prediction loss)进行优化,导致中间层隐藏状态受到固定下游计算路径约束的问题。其核心挑战在于理解这种终端损失如何在隐空间中诱导出具有特定几何结构的输出敏感方向,从而影响模型内部表示的可压缩性与可蒸馏性。解决方案的关键在于引入逐层损失剩余函数(layerwise loss-to-go functions),即从某一中间隐藏状态出发,继续通过剩余网络块所获得的终端损失。研究发现,在成功验证轨迹附近,这些函数的局部二阶几何结构由一个拉回费舍尔算子(pullback Fisher operator)主导,其谱特性揭示了对输出敏感的方向与近似预测无关的方向,进而定义出残差流中的局部可观测子空间。对于因果Transformer,该几何结构进一步导出一种基于费舍尔加权的词元曲率评分(tokenwise curvature score),该评分衡量目标logits对每个词元隐藏状态扰动的敏感性,其值在目标的因果祖先集之外趋于零,并受下游雅可比耦合控制,因此成为一种具备损失感知特性的替代注意力幅度的剪枝信号。作者采用无矩阵雅可比-向量积和向量-雅可比积方法高效估计上述量,并在WikiText、OpenWebText和FineWeb数据集上的多个解码器仅模型上进行了验证。实证结果表明,该诱导几何不仅能有效预测扰动敏感性,支持非均匀的分层秩分配,还生成具有竞争力的结构化词元剪枝信号,并在结合反向KL与偏斜KL等更强的自回归蒸馏目标时显著提升低秩学生模型的恢复性能。这些结果共同支持一种预测性几何视角(predictive-geometric view):在接近成功训练轨迹时,终端损失在隐空间中诱导出一个细长且各向异性的输出相关方向集合,该集合可被测量并用于模型压缩与知识蒸馏。

链接: https://arxiv.org/abs/2609.37717
作者: Timur Mudarisov,Mikhail Burtsev,Tatiana Petrova,Radu State
机构: University of Luxembourg (卢森堡大学); London Institute for Mathematical Sciences (伦敦数学科学研究所)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Decoder-only transformers are trained only through a terminal next-token prediction loss, yet this loss constrains every intermediate hidden state through the fixed downstream computation. We formalize this constraint by studying layerwise loss-to-go functions: the terminal loss obtained by continuing a candidate hidden state through the remaining transformer blocks. Around successful validation trajectories, we show that the local second-order geometry of these functions is governed, up to low-loss residual terms, by a pullback Fisher operator on hidden-state space. Its spectrum identifies output-sensitive directions and approximately prediction-null directions, yielding a local observable subspace of the residual stream. For causal transformers, the same geometry induces a tokenwise curvature score: a Fisher-weighted sensitivity of the target logits to perturbations of each token’s hidden state. This score vanishes outside the causal ancestor set of the target and is controlled by downstream Jacobian couplings, making it a loss-aware alternative to attention magnitude. We estimate these quantities using matrix-free Jacobian-vector and vector-Jacobian products and evaluate them across decoder-only language models on WikiText, OpenWebText, and FineWeb. Empirically, the induced geometry predicts perturbation sensitivity, supports nonuniform layerwise rank allocation, yields competitive structured token-pruning signals, and improves low-rank student recovery when added to stronger autoregressive distillation objectives such as reverse KL and skew KL. These results support a predictive-geometric view of transformer computation: near successful trajectories, the terminal loss induces a thin, anisotropic set of output-relevant hidden-state directions that can be measured and exploited for compression and distillation.

[NLP-39] Billiger.de Products: A Bilingual Entity Matching Benchmark

【速读】: 该论文旨在解决现有产品匹配基准数据集主要以英文为主且多集中于单一产品类别(如电子产品)而导致的泛化能力不足问题,尤其针对服装、家具等语义复杂、跨语言差异显著的类别缺乏充分覆盖的问题。其解决方案的关键在于构建一个双语(德语与英语)实体匹配基准——this http URL Products,涵盖十三个消费类目,具有更高的多样性与挑战性;该基准基于德国价格比对平台的真实数据,采用与WDC Products一致的设计范式,提供多种变体以调节边界案例比例、开发集规模及训练中未见实体比例,并通过保留所有配对、划分与标签的对齐英文翻译,确保数据一致性;同时引入跨语言测试集,将德语与英语记录混合于同一配对中,以评估模型在跨语言场景下的匹配能力。实验验证表明,该基准显著提升难度,尤其在德语版本上多数预训练语言模型(PLM)表现明显下降,而零样本大语言模型(LLM)对语言变化不敏感,进一步凸显了该基准在跨语言、多品类场景下对模型鲁棒性与泛化能力的严格考验。

链接: https://arxiv.org/abs/2609.37713
作者: Aaron Steiner,Ksenia Elagin,Ralph Peeters,Johannes Knopp,Christian Bizer
机构: University of Mannheim; solute GmbH
类目: Computation and Language (cs.CL)
备注: 23 pages. Data and code: this https URL

点击查看摘要

Abstract:Existing product matching benchmarks primarily contain English-language product data and are often dominated by a single product category, such as electronics. This paper introduces this http URL Products, a bilingual German and English entity matching benchmark covering thirteen consumer product categories, including difficult-to-handle categories such as clothing and furniture. The benchmark data originates from the German price comparison platform this http URL. Following the design of WDC Products, the benchmark offers multiple variants that differ in the fraction of corner cases, the size of the development set, and the fraction of entities unseen during training. An aligned English translation of every offer keeps all pairs, splits, and labels fixed, while cross-language test sets combine German and English records within individual pairs. We validate the benchmark using six supervised matchers and zero-shot GPT-5.2 on both language versions and the cross-language test sets. The validation shows the difficulty of the benchmark. The comparison of the results on the English version of the benchmark to the results on the German version shows that most matchers score on average higher on the English version. The difference is largest for RoBERTa and HierGAT, while the zero-shot LLM runs are largely insensitive to the language. Comparing the F1 scores achieved by PLM-based matchers on the English version of this http URL Products with their performance on existing English-language benchmarks, such as WDC Products and Abt-Buy, shows that this http URL Products is more difficult than these benchmarks.

[NLP-40] Reader Proficiency Shapes Layer-wise Surprisal Profiles

【速读】: 该论文旨在解决语言模型中逐层预测意外度(surprisal)与读者眼动行为之间关系是否因读者词汇熟练度水平及不同眼动指标而异的问题。其核心解决方案在于引入“预测深度”(Predictive Depth)这一量化方法,用以分析不同层次语言模型内部表示在预测人类眼动反应(包括首次通过注视时间FPGD与总注视时间TGD)时的分布特征。研究发现,词汇熟练度较低的读者在FPGD上表现出更深的预测深度,表明其阅读过程更依赖于深层语义表征;而总注视时间TGD则普遍表现出比FPGD更深的预测深度,且该模式在不同熟练度群体间相对稳定。此外,留一法分析表明,具有信息量的内部层在未见文本上仍具预测优势,尽管实际提升有限。结果表明,基于大语言模型(LLM)逐层意外度的分析框架可有效揭示阅读行为在读者群体与眼动指标间的差异性,为理解阅读认知机制提供了新的计算视角。

链接: https://arxiv.org/abs/2609.37688
作者: Akio Hayakawa,Horacio Saggion
机构: Universitat Pompeu Fabra (庞培法布拉大学); Barcelona, Spain
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Reading behaviour varies not only with linguistic input, but also with reader proficiency. In this study, we investigate whether the layer-wise relationship between surprisal from large language models (LLMs) and human gaze behaviour differs across readers with different levels of proficiency and across gaze measures. Using eye-tracking data from the MECO L2 corpus, we compare readers with high and low vocabulary proficiency on first-pass gaze duration (FPGD) and total gaze duration (TGD). We quantify the distribution of the predictive power of surprisal across model layers using Predictive Depth. Across 12 tested LLMs, we find that readers with lower vocabulary proficiency tend to show deeper Predictive Depth for FPGD, while this difference is smaller for TGD. Also, TGD itself shows deeper Predictive Depth than FPGD in both proficiency groups. These patterns suggest that where predictive power is concentrated across LLM layers may be related to the timing and breadth of the reading processes captured by different gaze measures, and that this relationship can vary with reader proficiency. Our leave-one-out analysis further shows that the advantage of informative internal layers extends to unseen texts, although the practical improvements in prediction are limited. Overall, our results show that layer-wise LLM surprisal provides a useful perspective on variation in reading behaviour across both reader groups and gaze measures.

[NLP-41] EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments?

【速读】: 该论文旨在解决专业工业工程领域中自动化代理(autonomous agents)难以实现可靠自动化的问题,其核心挑战在于工程工作流需要对几何与物理约束及其在软件和设计各阶段间的依赖关系进行复杂推理。为此,论文提出EngiWorld,这是首个围绕完整设计闭环构建的基准测试体系,涵盖6个工程领域(计算机辅助设计(CAD)、计算机辅助工程(CAE)、计算机辅助制造(CAM)、建筑信息模型(BIM)、电子设计自动化(EDA)及3D可视化)和26个专业软件平台,包含图形用户界面(GUI)与命令行接口(CLI),并覆盖从软件选型到开放式任务等6类任务。解决方案的关键在于引入一种以成果为中心的评估方法,基于统一的领域验证器套件,可程序化地检查最终与中间成果的几何有效性、物理可行性及规则合规性,并通过规范达成度对定量设计任务进行连续评分,而非采用二元成功判定。实验评估七款前沿模型表明存在显著能力差距:最强模型的EngiScore仅为44.3,多软件协同尝试的成功率仅3.6%。EngiWorld为衡量代理在端到端操作专业工程软件方面的发展提供了首个严谨基准。

链接: https://arxiv.org/abs/2609.37686
作者: Hongcheng Gao,Hailong Qu,Yu Lei,Henghui Sun,Haoyang Li,Yipeng Wei,Naihao Xue,Xiaohan Yu,Zhuo Tao,Yihe Zang,Yajiao Wang,Jingyi Tang,Yi Li,Jingjing Zhou,Jie Luo,Bohan Zeng,Chengyu Shen,Hao Jiang,Chong Chen,Bowen Qu,Olive Huang,Zeqiang Wang
机构: Tsinghua University (清华大学); Zhiman Inc. (智曼科技); Chongqing University (重庆大学); UCAS (中国科学院大学); Shandong University (山东大学); BUPT (北京邮电大学); Fudan University (复旦大学); Henan Polytechnic University (河南理工大学); Xi’an Jiaotong University (西安交通大学); Peking University (北京大学); Zhejiang University (浙江大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: Project page: this https URL

点击查看摘要

Abstract:Autonomous agents have made rapid progress in general-purpose computer use, but reliable automation of professional industrial engineering remains out of reach, as engineering workflows demand reasoning over geometric and physical constraints and dependencies preserved across software and design stages. We present EngiWorld, the first benchmark structured around the complete design loop: 1,301 expert-curated tasks spanning 6 engineering domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms, with both GUI and CLI interfaces and 6 task types ranging from software-selection to open-ended tasks. We further introduce an artifact-centric evaluation methodology built on a unified domain-verifier suite, which programmatically checks the geometric validity, physical feasibility, and rule compliance of final and intermediate artifacts, and scores quantitative design tasks continuously by specification attainment rather than binary success. Evaluation of seven frontier models reveals a substantial capability gap: the strongest model achieves an EngiScore of only 44.3, and just 3.6% of multi-software attempts succeed. EngiWorld provides the first rigorous foundation for measuring progress toward agents that operate professional engineering software end to end.

[NLP-42] When Models Dont Manipulate Manifolds: The Geometry of a Comparison Task

【速读】: 该论文旨在解决生成式人工智能(Generative AI)模型在执行决策类任务时,其内部表示的几何结构如何支撑计算过程,以及模型如何利用这些几何特性进行有效干预的问题。具体而言,研究聚焦于数字比较这一典型决策任务,探究模型在处理有序概念时是否依赖于曲面流形(manifold)结构,以及其实际计算机制的本质。解决方案的关键在于揭示:尽管数字在模型表征中可能呈现出螺旋或环状等非线性流形结构,但模型在执行数字比较任务时,主要依赖于线性表示(linear representations)进行核心计算。研究以Qwen2.5-7B-Instruct模型为对象,发现其通过注意力机制与残差连接将两个数字编码映射至共享的残差流空间,并在此线性空间中利用MLP神经元对局部区间内的数值进行逐段比较,最终整合结果以确定最大值位置。该过程表明,即使存在流形结构,模型仍可基于概念的内在线性属性实现高效计算。这一发现揭示了“流形假设”与“线性计算”的共存性——有序概念虽具流形表征,但模型在特定计算中可选择性地使用其线性本质,从而实现简洁而高效的推理机制。

链接: https://arxiv.org/abs/2609.37680
作者: Sai Sumedh R. Hindupur,Hadas Orgad,Thomas Fel,Demba Ba
机构: Harvard University (哈佛大学); Kempner Institute (肯普纳研究所); Goodfire AI
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:One of the current premises of mechanistic interpretability research is that detailed accounts of the geometry of neural network representations can tell us how models perform computations, and how to effectively intervene on them. While low dimensional manifolds have been observed for multiple concepts in the literature (e.g. numbers encoded on helices, days of the week on a circle, …), with structure believed to reflect properties of data and tasks, the extent to which models rely on them for computation, and how they manipulate them, remains unclear. We characterize precisely the geometry of computation in a number-comparison task, as an abstraction of comparison for decision making, and how models utilize geometry in an elegant fashion to implement it. Specifically, we study the causal geometry of number comparison in Qwen2.5-7B-Instruct, a capable and widely studied open-weight model, and find Qwen largely uses linear representations of numbers despite the presence of curved geometry. To compare two numbers, the model first encodes each number along a vector and adds the two representations using attention and the residual connection, bringing them into a shared space in the residual stream. Then, the model uses MLP neurons to compare the pair of numbers on local regions in this shared space, which correspond to smaller intervals of input numbers, and combines these to obtain the position of the maximum. In fact, this reliance on linear representations for comparison also persists when the model compares three numbers. Our findings demonstrate that the manifold hypothesis can co-exist with linear representations: while concepts that are ordered may have manifold structure in representations, the model may use an underlying linear structure of the concept in certain computations.

[NLP-43] KUPAS MASTER: Distilling the Tacit Expertise of Master Practitioners into Agent -Ready Experience Corpora

【速读】: 该论文旨在解决大型语言模型(Large Language Model, LLM)代理在应用专业经验时面临的挑战,即常规工作记录通常缺失隐性知识(tacit knowledge),导致模型难以有效利用从业者的真实经验。其解决方案的关键在于构建一个基于九层认知语料库(nine-layer cognitive corpus)的体验工程平台KUPAS MASTER,通过系统化提取和结构化组织来自异构工作记录与从业者访谈中的经验信息。该平台以六类核心要素(上下文、线索、判断、行动、边界、结果)完整保留任务过程,并依据九个提取维度对隐性知识进行分层建模,生成规则、约束、最佳实践、反例、边缘案例及技能等六类可复用的知识资产。通过语义对齐、个体经验提炼、组织整合与交叉审查机制,确保知识来源可追溯、使用条件明确且保留未决争议。最终,平台将经验转化为具备显式输入、步骤、依赖关系和终止条件的可调用技能,实现从经验采集到任务执行与评估反馈的闭环。实验表明,在多专业领域测试中,相较于基础模型(70.63)和原始语料检索增强生成(RAG,79.75),KUPAS MASTER代理得分提升至89.58,在全部七项评分维度上均优于RAG,验证了其在将个体隐性经验转化为组织级知识与智能代理能力方面的有效性。

链接: https://arxiv.org/abs/2609.37673
作者: Changmian Wang,Yuchao Ma,Xuchao Lu,Chen Zhang,Ping Sun,Jiazheng Wang,Shan Wang,Xuanwen Chen,Yihe Sun,Ziyu Lu,Jianqiang Huang,Hongzhi Li,Ziqing Xia,Kaihua Tang,Xian-Sheng Hua,Qinghua Zheng
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: Technical Report. Official website: this https URL Report homepage: this https URL

点击查看摘要

Abstract:Experienced professionals know more than just facts and conclusions. They know which cues matter, why a judgment is reasonable, and which action to take. Routine work records often leave out this tacit knowledge, making it difficult for Large Language Model (LLM) agents to use professional experience effectively. We introduce KUPAS MASTER, an experience engineering platform built around nine-layer cognitive corpus construction. It turns heterogeneous work records and practitioner interviews into traceable, reusable experience corpora for agents. Six case elements preserve the task process: context, cues, judgment, action, boundaries, and outcomes. Nine-layer cognitive corpus construction organizes tacit experience along nine extraction dimensions and stores the resulting assets in six libraries: rules, constraints, best practices, negative examples, corner cases, and skills. Semantic alignment, individual experience distillation, organizational consolidation, and cross-review preserve source evidence, conditions of use, and unresolved disagreements. The platform packages these assets into callable skills with explicit inputs, steps, dependencies, and stopping conditions, connecting experience collection to task execution and evaluation feedback. Using authorized samples from 20 randomly selected practitioners, the platform processed 1,576 source files into 23,024 individual experience records and 13,113 organizational assets. The evaluation spans multiple professional domains. Under common task inputs and scoring criteria, the base model, raw corpus retrieval-augmented generation (RAG), and KUPAS MASTER agent scored 70.63, 79.75, and 89.58, respectively. The KUPAS MASTER agent improved on raw-corpus RAG in all seven scoring dimensions. The platform provides a practical path from individual tacit experience to organizational knowledge and agent capabilities.

[NLP-44] Corpus-Guided Dual-Path Propagation for Graph Retrieval-Augmented Generation

【速读】: 该论文旨在解决现有无关系图检索增强生成(relation-free graph retrieval)方法在多跳问答任务中因过度依赖查询-句子相似度进行证据搜索而带来的局限性,即可能遗漏与查询语义相关但相似度较低的桥接证据,并激活与推理链无关的随机实体。其解决方案的关键在于提出一种名为NexusRAG的简单而有效的方法,通过引入基于联合实体共现与语义相似性的文档级实体邻域结构,对无关系三图(Tri-Graph)进行增强。该结构引导两种互补的传播路径:一是受限于邻域的语义传播,通过句子识别与查询相关的实体边界;二是直接在邻近实体间进行结构化传播,将实体边界扩展至结构相关实体。同时,传播得到的实体权重用于指导个性化页面排名(Personalized PageRank)的邻域感知段落初始化。实验结果表明,NexusRAG在三个多跳问答基准及GraphRAG-Bench的一个领域特定子集上均显著优于现有方法,在所有问题类别中均实现最高证据召回率,相较于基线提升4.2–8.1个百分点。

链接: https://arxiv.org/abs/2609.37661
作者: Baoxian Liu,Tong Wei
机构: College of Software Engineering, Southeast University(东南大学软件工程学院); School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院); Key Laboratory of Computer Network and Information Integration (Southeast University) Ministry of Education(东南大学计算机网络与信息集成教育部重点实验室)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Graph-based retrieval-augmented generation supports multi-hop retrieval by organizing corpus information into graphs. However, existing relation-free graph retrieval methods rely primarily on query-sentence similarity to search for evidence. This can exclude useful bridging evidence with low query similarity and activate incidental entities unrelated to the reasoning chain. In this paper, we propose a simple and effective approach called NexusRAG, which augments the relation-free Tri-Graph with a corpus-level entity neighborhood structure derived from joint entity co-occurrence and semantic similarity. NexusRAG employs this structure to guide two complementary propagation paths: neighborhood-constrained semantic propagation through sentences identifies the query-relevant entity frontier, while direct structural propagation between neighboring entities expands that frontier to structurally related entities. The propagated entity weights also inform neighborhood-aware passage initialization for Personalized PageRank. Experiments on three multi-hop QA benchmarks and a domain-specific subset of GraphRAG-Bench show that NexusRAG consistently outperforms existing approaches. On the GraphRAG-Bench subset, NexusRAG achieves the highest evidence recall in all question categories, exceeding baselines by 4.2-8.1 points. The implementation code is available at this https URL.

[NLP-45] Evaluating and Benchmarking the System One Model Jev

【速读】: 该论文旨在解决在信息检索与处理流水线中,针对小规模决策任务(如查询路由、内容置信度验证、内容审核、评分打分等)中高效、可靠且可解释的判断问题。现有生成式 AI 模型虽具备强大语言理解能力,但在这类需要精确输出固定选项或概率判断的任务上存在泛化能力不足、概率校准性差及计算成本高等问题。其解决方案的关键在于采用基于预定义模板的零样本推理机制,利用经过校准的概率输出(calibrated probabilities)实现高精度分类与判断,同时支持选择性预测(selective prediction)。Jev 模型通过仅使用一个冻结模板对 37 个数据集进行零样本评估,在涵盖分类、自然语言推断、阅读理解、常识推理、法律条款分析等多类任务中达到 95–99% 的准确率,尤其在多语言任务(如 Belebele 覆盖 122 种语言)中表现优异。此外,其输出概率具有良好校准性,可通过调优阈值显著提升微平均 F1(如 UNFAIR-ToS 从 0.50 提升至 0.75),且对选项顺序旋转具有鲁棒性,排除了浅层记忆的可能性,表明模型依赖语义理解而非表面模式匹配。该研究还揭示了所有模型在低资源语言、细粒度标签及基于量规的质量判断任务中的性能下降趋势,凸显了当前系统在复杂人类判断任务上的局限性。

链接: https://arxiv.org/abs/2609.37647
作者: Tobias Deußer,Lorenz Sparrenberg,Rafet Sifa
机构: University of Bonn(波恩大学); Lamarr-Institute for Machine Learning and Artificial Intelligence(拉马尔机器学习与人工智能研究所); Fraunhofer IAIS(弗劳恩霍夫信息分析与知识系统研究所)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Code available at this http URL , model responses at this http URL

点击查看摘要

Abstract:Jev is a commercial System One model from TypeSafe AI that does not generate text: given a state and typed questions, it returns a choice from fixed options, a position on a rubric, or the probability that a statement is true, with probabilities the vendor describes as calibrated. Such models target small decisions in information access pipelines, such as routing queries, checking grounding, moderating content, or rating against a rubric. We evaluate Jev (jev-1.13.0) zero-shot on 37 datasets spanning classification, routing, natural language inference, reading comprehension, commonsense reasoning, moderation, legal clause analysis and rubric scoring, with one frozen template per dataset and full evaluation splits: 346,009 requests for under USD 10. For reference, we score Qwen3.8-27B and Gemma-4-E4B on identical requests via their exact next-token probabilities over the options. Jev reaches 95-99% accuracy on IMDB, SST-2, HellaSwag and ARC and 86.7% on Belebele across 122 languages. It beats Qwen on 27 of 37 datasets, with none of Qwen’s nine leads outside the bootstrap intervals, and Gemma on all 37. All three models degrade on low-resource languages, fine-grained or noisy labels, and rubric-based quality judgments. Jev’s choice probabilities are well calibrated and support selective prediction. Binary probabilities rank well but are poorly placed relative to a fixed 0.5 threshold; thresholds tuned on training data raise micro-F1 on UNFAIR-ToS from 0.50 to 0.75. Jev answers MMLU’s calculation-heavy questions more accurately than other MMLU questions (94% vs. 91%), whereas both open models, and all three on C-Eval, find them harder. Rotating the options leaves Jev’s accuracy unchanged and withholding the question drops it to near chance, ruling out shallow memorization but not memorized question-answer pairs. We release the code, harness and all raw responses.

[NLP-46] Co-Linguistics: AI-augmented Theory Construction in Linguistics

【速读】: 该论文旨在解决传统语言学理论构建与评估过程中存在的隐性假设、形式化不足及效率低下的问题,尤其针对语言学理论难以实现完全形式化、理论间比较困难以及新理论生成缓慢等挑战。其核心解决方案在于将人工智能(AI)作为“协同科学家”(co-scientist),推动“协同语言学”(Co-Linguistics)的发展。关键在于利用生成式人工智能(Generative AI)在程序归纳(program induction)和主动学习(active learning)方面的优势,使现有语言学理论实现完全显式化,加速不同理论之间的系统性比较,并自主提出候选理论;同时借助其对海量数据的高效访问能力,快速识别并验证理论的关键预测。尽管这一过程可能引发理论迭代的递归优化,但人类学者仍主导科学方向设定与概念评估,实验参与者则不可或缺,以验证超出大语言模型(LLM)能力范围的实证预测。

链接: https://arxiv.org/abs/2609.37635
作者: Emmanuel Chemla,Benjamin Spector,Alexandros Kalomoiros,Philippe Schlenker
机构: 未知
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:LLMs have been studied in recent linguistics as potential models of humans’ linguistic abilities. Here we discuss an entirely different use of AI, namely as a co-scientist, to help construct and assess linguistic theories (we refer to the result as “Co-Linguistics”). Since the 1960s, linguistics has developed theories that are in principle mathematically formalizable, often in the language of formal language theory or model theory. The AI revolution in mathematics will thus have consequences in linguistics-but with an essential twist: proving new theorems is rarely the linguist’s goal. Rather, one seeks to find the best set of axioms to derive empirical statements. AI could accelerate research by making existing theories fully explicit, by comparing competing theories, and more ambitiously, by proposing new theories (in machine learning, this relates to “program induction”). It will also help assess theories by accelerating the identification and test of crucial predictions, thanks to unparalleled access to data (in machine learning, this relates to “active learning”). While the cycle from theory evaluation to theory construction may give rise to recursive and possibly autonomous improvement of linguistic theories, humans remain central: linguists provide scientific directions and evaluate theories conceptually, and experimental participants are needed to assess empirical predictions that are outside the reach of LLMs.

[NLP-47] RLTL;DR: Self-improvement by Internalizing Self-generated Feedback

【速读】: 该论文旨在解决在自提升(self-improvement)场景中,强化学习与可验证奖励(Reinforcement Learning with Verifiable Rewards, RLVR)面临的核心挑战:当任务难度极高导致智能体成功概率极低甚至为零时,传统RLVR依赖多次尝试并筛选成功轨迹的范式失效,且缺乏教师模型或示范解可供知识蒸馏。为此,本文提出一种新型框架RLTL;DR,其关键在于引入“任务到洞察”的内化机制。具体而言,在每次失败后,智能体不仅接收验证器(verifier)的反馈,还被要求以“一句话总结”(TL;DR)的形式生成自我反思性洞察,并将这些洞察作为后续探索的条件输入。通过在上下文中累积历史洞察并支持对这些洞察进行反向传播(backpropagation),模型能够逐步建立从任务特征到有效策略洞察的直接映射。实验表明,即使在Pass@128=0的高难度工具调用与代码生成数据集上,标准GRPO训练的Qwen 3.5 9B思维策略模型性能停滞于0%~1%,而RLTL;DR在训练阶段结合上下文洞察可实现14%-31%的Pass@1,更关键的是在推理阶段无任何洞察输入的情况下仍保持12%-13%的显著性能,证明了“任务到洞察”内化机制的有效性。进一步简化为SFTL;DR——仅基于(任务, 洞察)对进行监督微调,使用仅4000条样本即可恢复接近完整RLTL;DR及经典监督微调的性能,揭示了一种高效紧凑的训练范式:“对于此类任务,需牢记此类洞察”,为未来生成式智能体的自适应学习提供了重要启发。

链接: https://arxiv.org/abs/2609.37633
作者: Michael Kirchhof,Eleonora Gualdoni,Andrew Szot,Khashayar Gatmiry,Aryo Lotfi,Abbas Kazerouni,Omar Attia,Sanjoy Chowdhury,Alexander Toshev
机构: Apple(苹果)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (stat.ML)
备注:

点击查看摘要

Abstract:The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and where there are no teacher models or example solutions to distill from. In this paper, we introduce RLTL;DR. After each failed attempt, we show the policy the verifier outputs and let it write its own feedback, in the form of a single TL;DR insight. The next rollout is conditioned on all previous insights, and we sequentially sample rollouts until a solution is found. Moreover, we enable backpropagation on the in-context insights to internalize a direct task to insight mapping. On challenging tool-calling and coding datasets (filtered to Pass@128=0), standard GRPO training of a Qwen 3.5 9B Thinking policy stays flat at a Pass@1 of 0% to 1%. RLTL;DR breaks through this learning barrier, achieving a Pass@1 of 14-31% with insights in context during training and, crucially, 12-13% when no insight is in context at eval time. We identify that the key is the task to insight internalization. To study this further, we reduce our approach to SFTL;DR, training only on (task, insight) tuples, without showing or backpropagating on any rollouts. Training on only 4k of these tuples recovers almost the full performance of RLTL;DR and classical SFT on full rollouts. This demonstrates a promising compacted training paradigm of the form “on this sort of task, keep this sort of thing in mind”, which we hope to inspire future research on.

[NLP-48] Correct Dont Delete: Mitigating Emergent Misalignment with Corrective Supervision

【速读】: 该论文旨在解决生成式AI(Generative AI)在微调过程中因少量有害示范数据(如错误医疗建议)引发的“涌现性错位”(emergent misalignment, EM)问题,即模型在无关任务上也表现出广泛偏离对齐的行为。传统解决方案是识别并删除有害数据行,但该方法在实际测试中效果有限。本文提出关键问题:在固定有毒数据集的前提下,是否应通过修正(correction)而非删除来处理这些有害样本?研究发现,将有毒数据行替换为针对相同提示的正确回答,相比直接删除,可使EM率降低约三分之一,并显著提升对独立测试集上医疗问题的回答质量;而删除相同数据行则几乎无明显改善。进一步实验表明,修正效果随修正比例增加而增强,且在不同基础模型与错位模型上均具一致性。内容本身至关重要:仅改写原错误答案而不改变其有害内容无法带来显著收益,而使用数据集中提供的正确答案或由人工重写的正确回答效果相当。此外,用少量微调训练纠正数据,优于等量通用对话数据的训练;其他医疗提示的纠正数据亦能有效缓解错位,且无需刻意指导纠正者模仿谨慎助手角色。综上,在所考察条件下,修正有害训练数据比删除更有效降低涌现性错位。

链接: https://arxiv.org/abs/2609.37624
作者: Jacob Epifano
机构: Independent researcher(独立研究员)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 18 pages, 9 figures

点击查看摘要

Abstract:Fine-tuning a language model on a narrow set of harmful demonstrations, such as bad medical advice, can make it broadly misaligned on unrelated questions, a phenomenon known as emergent misalignment (EM). The usual defense is to find the offending rows and delete them, but a row locator failed our held-out test and deleting rows helps less than expected. We ask a different question: given a fixed set of poisoned rows, is it better to correct them than to remove them? We fine-tune Qwen2.5-14B-Instruct on a mixture of bad medical advice and benign chat data, select a quarter of the poison rows in advance, and either delete them or replace each with a corrected answer to the same prompt, keeping everything else the same. Replacing the rows cuts the EM rate by about a third and improves answers on held-out medical questions, while deleting the same rows has little measurable effect. The advantage is larger when half the poison rows are corrected, and it holds on a second base model and a second misaligned model organism. The content of the replacement appears to matter: paraphrasing the rows while keeping their bad advice shows no clear benefit, and the correct answers distributed with the dataset appear to do about as well as our rewriter’s. Realigning an already-poisoned model with further fine-tuning is known to work, but which data does the work has not been compared directly. We find that a short round of training on corrections beats the same amount of training on generic chat data, that corrections on other medical prompts do roughly as well as corrections of the poisoned prompts themselves, and that instructing the correction writer to model a careful, harm-avoiding assistant adds no measurable benefit over plain corrections. In the settings we tested, correcting harmful training data reduces EM more than deleting it.

[NLP-49] Authority Bias in Language Models: Source Deference and User Agreement Are Not Interchangeable NEURIPS2026

【速读】: 该论文旨在解决大语言模型在面对错误信息时过度依赖权威来源(verified source)而产生“源服从性”(source deference)的问题,这种现象导致模型即使面对明显错误的内容,只要其被标注为来自可信来源,仍会表现出高度合规性。尽管后训练阶段已致力于减少模型对用户主张的盲目顺从(user agreement),但模型在处理带有权威来源背书的错误答案时依然表现出显著的可操纵性。研究的关键发现在于:源服从性与用户顺从性在模型内部行为机制上并非等价或可互换;通过因果干预实验可独立操控二者——移除针对来源的特定方向(fitted source direction)可使源服从性下降65–80个百分点,而移除用户或助手方向的影响则微乎其微;进一步地,基于源-用户提示激活模式拟合的独立干预项可在不改变输入文本的前提下同时调节两种行为。此外,该来源权威性方向在跨任务(如PIQA、多轮对话)中具有迁移能力,并在多个模型家族中有效降低错误来源的合规性,同时不影响主流基准测试(如MMLU-Pro、GSM8K)的准确性。因此,论文强调必须将源服从性与用户顺从性分别评估,以全面理解并改进模型的可靠性。

链接: https://arxiv.org/abs/2609.37616
作者: Abhinav Rajeev Kumar(Lossfunk),Paras Chopra(Lossfunk)
机构: Lossfunk
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: Accepted at NeurIPS 2026 (Main Conference, Poster). 33 pages, 8 figures. Project page: this https URL . Code: this https URL

点击查看摘要

Abstract:Language models tend to agree with whatever a user asserts, and post-training increasingly targets this sycophancy so that models evaluate claims on their merits rather than deferring to the user. Yet the same models are far more compliant when a wrong answer is attributed to a verified source, which is how retrieval results, tool outputs, and grounded-search content often present information. We measure this gap across five open-weight families and three closed APIs. A single verified-source note endorsing a wrong answer flips 45-88% of baseline-correct responses in seven of eight models, and compliance rises with how authoritative the note sounds. Source deference and user agreement are not behaviorally interchangeable inside the model: on matched items with the same wrong answer, causal interventions can selectively suppress one without equally affecting the other. In three open-weight families, removing a fitted source direction lowers source compliance by 65-80 percentage points while removing a user or assistant direction has far smaller effects, and removing the user direction shows the reverse preference. A separately fitted intervention derived from source-versus-user cue activations moves compliance in both directions while leaving the prompt text unchanged. An authority direction fitted on trivia also transfers to PIQA and multi-turn SYCON dialogues without refitting, and removing it lowers wrong-source compliance by tens of percentage points in four of five families with no detected change in MMLU-Pro or GSM8K accuracy at our evaluation sizes. Source deference and user agreement therefore need separate evaluation.

[NLP-50] FOCUS: Training-Free Decision-Preserving Context Compression for LLM Agents

【速读】: 该论文旨在解决大语言模型智能体(LLM agents)在长期任务中因交互历史线性增长导致的注意力稀释(attention dilution)问题,进而引发的推理开销二次增长与性能下降。现有上下文压缩方法依赖离线训练,通过对比学习、蒸馏或训练压缩策略来决定丢弃内容,不仅成本高昂,且压缩策略为预先固定,无法根据测试阶段动态演化的轨迹进行自适应调整。为此,本文提出一个互补性问题:哪些历史交互对智能体未来的决策具有因果影响?基于此,作者将上下文压缩重新建模为离散交互单元上的因果决策保留问题,并提出FOCUS——一种完全在测试时运行、无需离线数据收集或微调的无训练上下文压缩框架。其核心创新在于不依赖预训练压缩策略,而是实时分析历史交互的因果重要性,实现动态、可解释的上下文筛选。该方法具有架构无关性,可作为模块化层集成至任意封闭式前端模型中,在多种代理基准任务(包括API调用、工具使用、问答、网页域任务及多轮对话)上均取得新最优性能,相较未压缩执行,最多减少48%的峰值上下文长度与73%的依赖关系,同时提升任务成功率达8.9个百分点。

链接: https://arxiv.org/abs/2609.37590
作者: Shantanu Dixit,Anson Bastos,Xuchao Zhang,Chetan Bansal,Saravan Rajmohan
机构: M365 Research, Microsoft(微软)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: Preprint. Under Review

点击查看摘要

Abstract:LLM agents accumulate interaction histories that grow linearly with task length, causing quadratic inference cost scaling and performance degradation from attention dilution. Existing context-compression methods learn what to discard offline: by contrastively optimizing guidelines, distilling compressors, or training compression policies. This incurs a substantial cost. Further, the compression policy is learned a priori and is not dynamically conditioned on the evolving test-time trajectories. In this paper we ask a complementary question: Which past interactions causally shape the agent’s future decisions? We recast context compression as a causal decision preservation problem over discrete interaction units and introduce FOCUS, a training-free context compression framework that operates entirely at test time. Our method requires no offline data collection or fine-tuning, and is architecture-agnostic, attaching to any closed-API frontier model as a modular compression layer. We evaluate FOCUS on diverse agentic benchmarks including API and tool-calling, QA, web domain and multi-turn dialogue. Our method establishes new state of the art performance, cutting peak context by up to 48% and dependency by 73% while improving task success by up to 8.9 percentage points over uncompressed execution.

[NLP-51] Pair Difficulty Matters: Rethinking Pairwise LLM -as-a-Judge Evaluation and Consistency EMNLP2026

【速读】: 该论文旨在解决当前评估大语言模型(Large Language Model, LLM)评判者可靠性时所依赖的三大代理指标(位置偏差、传递性、成对一致性)存在的误导性问题。这些指标通常用于判断评判者在文本排序任务中的表现,但研究发现,它们在实际应用中与真实排名准确率的相关性极弱,且其信号主要来自远距离排名差距(far-gap)样本,而这些样本对代理指标的贡献微乎其微;相反,代理指标的波动主要受近距离排名差距(close-rank-gap)样本主导,而在此类样本中,不一致是信息论上可预期的,个体判断对整体排序的贡献有限。因此,论文提出关键解决方案:应基于排名差距条件化(rank-gap-conditional)的评估指标来衡量评判者性能,尤其应在与人类标注的黄金标准对比下进行评估,以更准确地反映评判者的实际排序能力。

链接: https://arxiv.org/abs/2609.37577
作者: Bruno Brocai,Maria Becker
机构: Heidelberg University (海德堡大学); German Department (德国系)
类目: Computation and Language (cs.CL)
备注: Accepted as an EMNLP 2026 short paper

点击查看摘要

Abstract:Large Language Model judges are widely used to rank texts and text-generating systems through pairwise comparison, and their reliability is typically assessed via three proxies: position bias, transitivity, and pairwise agreement (self- or human-labeled). Because these proxies drive judge selection and benchmarking, a substantial literature reporting that judges perform poorly on them risks steering practitioners away from otherwise capable evaluators. We argue this assessment is misleading. Under the Bradley–Terry geometry underlying pairwise aggregation, each proxy is dominated by close-rank-gap pairs, where inconsistency is information-theoretically expected and individual verdicts contribute little to the aggregate ranking; far-gap pairs carry the ranking signal but barely move the proxies. We formalize this argument and validate it in a controlled simulation and on two human-rated corpora: the proxies correlate only weakly with ranking accuracy against gold, and their predictive component concentrates in the far-gap regime. Judges should therefore be assessed on rank-gap-conditional metrics, ideally against human rankings. Code at this https URL.

[NLP-52] Devils in Question Relay: Source-Conditioned Relay Steering to Mitigate Hallucinations in Audio-visual Large Language Models

【速读】: 该论文旨在解决音频-视觉大语言模型(Audio-Visual Large Language Models, AVLLMs)在多模态理解与推理中面临的关键问题:源混淆的定位幻觉(source-confused grounding hallucination),即未使用模态的干扰线索会诱导模型生成超出所需模态支持的错误回答,严重影响其在真实场景中的可靠性。现有方法虽在缓解此类错误方面取得一定进展,但对错误产生机制——尤其是内部跨模态交互如何引发干扰——仍缺乏深入理解。为此,作者通过路径干预(path-intervention)与表征分析揭示了核心机制:问题状态(question state)中同时携带了干扰性线索与所需模态的有效证据,导致对正确模态证据的定位被削弱。进一步实验表明,切断干扰模态到问题状态的路径,比切断其到生成位置的路径能带来更显著的正确答案置信度恢复。基于此发现,作者提出一种无需训练的解决方案——SECRET(Source-Conditioned Relay Steering),通过在问题状态层面施加条件化引导,利用不同模态路径干预所激发的对比性问题表示,将原始问题状态向所需模态的证据方向进行校准。在CMM与AVHBench两个主流基准上对三种AVLLM的实验结果表明,SECRET显著优于现有无训练方法,有效缓解了源混淆的定位幻觉(如相对于基线模型提升最高达+18.0和+7.1个百分点),且在模态特定的开放式生成任务中也展现出良好的泛化能力。

链接: https://arxiv.org/abs/2609.37568
作者: Yu Zhang,Pingrui Zhang,Xuefeng Bai,Pengfei Zhang,Yang Xiang,Kehai Chen
机构: Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学深圳校区); Peng Cheng Laboratory, Shenzhen, China(鹏城实验室)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Audio-visual large language models (AVLLMs) have made remarkable progress in multimodal understanding and reasoning through interactions among visual, auditory, and linguistic information. However, recent studies show that AVLLMs face a critical challenge: \textbfsource-confused grounding hallucination , where cues from the unused modality induce responses that the required modality does not support, undermining reliability in real-world applications. Existing methods have made progress in mitigating this failure, yet how it arises from internal cross-modal interactions remains insufficiently understood. To address this gap, we conduct path-intervention and representation analyses, revealing a \textbfquestion-relay mechanism: question states carry interfering cues alongside required-source evidence, undermining grounding in required-modality evidence. Cutting pathways from interfering modality to question states yields greater correct-answer logit recovery than cutting those to the generation position. Motivated by these findings, we propose \textbfSECRET ( \textbfS ourc \textbfE - \textbfC onditioned \textbfRE lay s \textbfT eering), a training-free method that mitigates cross-modal interference at the question relay. Using contrasting question representations elicited through different modality-pathway interventions, SECRET steers the original question states toward required-source evidence. Experiments on two widely adopted benchmarks CMM and AVHBench across three AVLLMs show that SECRET consistently outperforms prior training-free methods, substantially mitigating source-confused grounding hallucinations (e.g., up to +18.0 and +7.1 percentage points over base models). Modality-specific captioning further demonstrates its generalizability to open-ended generation.

[NLP-53] Orthogonal Yet Coupled: Decoupling Geometric Components for Model Merging

【速读】: 该论文旨在解决预训练模型合并(model merging)过程中因将任务向量(task vector)视为不可分割单元而导致的跨组件耦合问题。现有方法在基于完整任务向量的统计信息进行合并决策时,未能充分考虑其内部异质性的几何变化特征,导致不同组件间的几何属性相互干扰,进而降低合并后模型的质量。为此,本文提出一种解耦几何感知的模型合并框架DiGA(Disentangled Geometry-Aware)。其核心在于利用预训练权重作为共享几何参考,将每个任务向量正交分解为对应于不同几何属性的独立分量;随后在各自子空间中对同类型分量分别聚合,并重组为统一更新。该分量级的处理方式有效保留了各组件的几何身份,避免了组件间特性交叉影响,从而提升了合并模型的性能并缓解了能力退化问题。DiGA具有良好的兼容性,可集成至多种现有合并方法中,实验表明其在多类模型、任务及合并策略下均能显著提升合并效果。

链接: https://arxiv.org/abs/2609.37564
作者: Zijing Wang,Yongkang Liu,Mingyang Wang,Ercong Nie,Mengjie Zhao,Yunpu Ma,Kang Liu,Zihan Wang,Shi Feng,Daling Wang,Hinrich Schütze
机构: Northeastern University, China; CIS, LMU Munich, Germany; Munich Center for Machine Learning (MCML), Germany; Shanghai Jiao Tong University, China
类目: Computation and Language (cs.CL)
备注: Under review

点击查看摘要

Abstract:Merging pretrained models has emerged as an effective approach for consolidating diverse capabilities into a single unified model. However, prevailing merging methods typically treat each task vector as an indivisible merging unit, overlooking the heterogeneous geometric changes encoded within it. This treatment can induce cross-component coupling: when merging decisions are derived from statistics of the complete task vector, the geometric characteristics of one component may influence how another is selected, weighted, or combined, potentially degrading the quality of the merged model. To address this issue, we propose DiGA, a Disentangled Geometry-Aware model merging framework. Using the pretrained weights as a shared geometric reference, DiGA orthogonally decomposes each task vector into components corresponding to distinct geometric attributes. Rather than merging the task vectors as a whole, DiGA aggregates corresponding components independently within their respective subspaces and subsequently recombines them into a unified update. This component-wise formulation preserves the geometric identity of each component and prevents the characteristics of one component from interfering with the aggregation of another. Furthermore, DiGA can be incorporated into a broad range of existing model merging methods. Extensive experiments across diverse models, tasks, and merging methods demonstrate that DiGA improves merged-model performance and reduces capability degradation. Our repository is on this https URL.

[NLP-54] RunyaNER: Auxiliary Language Selection for Runyankore NER EMNLP2026

【速读】: 该论文旨在解决低资源语言(如东非语言鲁尼亚科雷语)在命名实体识别(NER)任务中,由于缺乏目标语言基准数据而难以评估跨语言零样本迁移与多语言微调效果的问题,尤其关注辅助语言选择策略对迁移性能的影响。其解决方案的关键在于构建首个公开可用的鲁尼亚科雷语命名实体识别基准数据集RunyaNER,该数据集通过半自动化流程生成并经完全人工验证,包含超过23.7万标注词元和3万句句子,具备足够的规模与质量以支持有效模型训练。在此基础上,研究发现迁移性能高度依赖于辅助语言的选择,而基于标注训练片段的嵌入表示所计算的度量指标,相较于传统基于语言元数据或类型学特征的指标,能更准确地预测下游迁移表现,从而为多语言迁移学习中的辅助语言筛选提供了实证依据与实用指导。

链接: https://arxiv.org/abs/2609.37543
作者: Prosper Arineitwe Asiimwe,Francois Meyer,Jan Buys
机构: University of Cape Town (开普敦大学)
类目: Computation and Language (cs.CL)
备注: Accepted to the 6th Workshop on Multilingual Representation Learning (MRL 2026) at EMNLP 2026. Camera-ready version. 4 figures

点击查看摘要

Abstract:Cross-lingual zero-shot transfer and multilingual fine-tuning are promising approaches for NLP tasks such as Named Entity Recognition (NER) in low-resource languages, but in the absence of target language benchmarks, it is unclear which auxiliary language selection strategy leads to the best transfer. We introduce RunyaNER, the first publicly available NER benchmark for the East African language Runyankore, and use it to investigate the choice of which languages to use for transfer. Created with a semi-automated pipeline and fully manually verified, RunyaNER contains over 237k annotated words across 30k sentences. We benchmark pretrained models on RunyaNER, establishing that our dataset is of sufficient quality and size to produce effective Runyankore NER models. We then use RunyaNER to investigate auxiliary language selection in cross-lingual zero-shot and multilingual fine-tuning settings. Our experiments show that while transfer performance is highly sensitive to auxiliary language selection, embedding-based measures computed from labelled training spans correlate more strongly with downstream transfer performance than traditional linguistic features based on metadata or typology. By releasing RunyaNER and providing a systematic analysis of auxiliary language selection strategies, this work contributes both a new benchmark resource and practical insights for multilingual transfer in low-resource settings.

[NLP-55] E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models

【速读】: 该论文旨在解决掩码扩散模型(Masked Diffusion Models, MDMs)在少步生成(few-step regime)中样本质量受限的问题。尽管MDMs通过逐步解掩码实现序列生成,其逆过程通常在位置上进行因子化建模,导致跨位置依赖关系建模能力不足,从而影响生成质量,而这一问题在扩散模型相较于自回归解码具有速度优势的少步生成场景中尤为突出。现有方法引入连续高斯潜在变量以捕捉位置间的相关性,但易出现后验坍塌(posterior collapse)问题,即潜在变量被模型忽略。本文提出增强型专家混合模型(Enhanced Mixture-of-Experts, E-MoE),其核心在于将逆过程构建为基于离散共享潜在变量的因子化分布混合,该潜在变量由专家混合(Mixture-of-Experts, MoE)主干网络的专家路由决策决定,且不增加活跃参数数量。该设计有效建模了跨位置依赖,同时避免后验坍塌,显著提升了合成多模态基准、二值化MNIST以及LM1B数据集上的少步生成性能。

链接: https://arxiv.org/abs/2609.37533
作者: Arseny Ivanov,Alexander Kolesov,Alexander Korotin,Ivan Oseledets,Mikhail Goncharov
机构: AXXX; Applied AI Institute(应用人工智能研究所)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Masked diffusion models (MDMs) generate sequences by progressively unmasking several tokens per denoising step, but their reverse process is typically factorized over positions, limiting sample quality in the few-step regime where diffusion’s speed advantage over autoregressive decoding matters most. A recent line of work introduces a continuous Gaussian latent, trained as a variational autoencoder, to capture correlations across positions, but such approaches are prone to posterior collapse, where the latent is silently ignored. We propose Enhanced Mixture-of-Experts (E-MoE), which builds the reverse process as a mixture of factorized distributions over a discrete shared latent given by the expert-routing decisions of a Mixture-of-Experts (MoE) backbone, without increasing active parameters over the factorized baseline. Across synthetic multi-modal benchmarks, binarized MNIST, and LM1B, E-MoE improves few-step generation over factorized baselines.

[NLP-56] Hierarchical Compression of Vision-Language Model Benchmarks

【速读】: 该论文旨在解决视觉-语言模型(Vision-Language Models, VLMs)评估成本高昂的问题,尤其在基准测试涵盖能力范围不断扩展、新模型迭代迅速的背景下,传统全量评估已难以持续。其核心挑战在于如何在大幅降低评估开销的同时,仍能保持不同模型间的相对排名准确性。解决方案的关键在于提出PRIMEBench——一种面向多模态评估的层次化基准压缩框架,通过四阶段流程实现高效降维:首先进行数据清洗以剔除仅依赖文本即可回答或全部正确项;其次按能力类别选择代表性基准;第三步采用“视觉感知方差”(Vision-Aware Variance, VAW)方法进行项目剪枝,VAW结合跨模型方差与仅基于多模态嵌入计算的视觉依赖性得分,兼顾多样性覆盖与评估保真度;最后实施类别数量剪枝。该方法在释放的5%保留率下实现了最高的平均保真度,且其层次化设计允许用户根据计算预算灵活终止任一阶段。实验表明,该框架可移除超过97%的评估项目而维持模型排名不变,同时揭示了模型规模增长对评估行为的影响,为未来设计更高效、抗模型迭代、并明确评估剪枝边界的新基准提供了理论指导。

链接: https://arxiv.org/abs/2609.37515
作者: Hyunjong Ok,Seunggu Kang,Jaeho Lee
机构: Pohang University of Science and Technology (浦项科技大学); Upstage AI
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
备注: Preprint

点击查看摘要

Abstract:Thorough evaluation of vision-language models (VLMs) has become prohibitively expensive, as benchmarks span an ever-broader spectrum of capabilities and new models arrive at a relentless pace. Benchmark compression methods that preserve model rankings at a fraction of the cost are well studied for language models, but for VLMs the question remains under-explored. We present PRIMEBench (Pruning Redundant Items for Multimodal Evaluation), a vision-aware hierarchical benchmark compression framework that substantially reduces evaluation cost while preserving model rankings. This hierarchical framework operates in four stages: data cleaning to remove items answerable without the image and all-correct items, category representative selection to pick one benchmark per capability category, item pruning with Vision-Aware Variance (VAW), and category-count pruning. VAW combines inter-model variance with a vision-dependence score computed from multimodal embeddings alone, while encouraging coverage of diverse items within each benchmark. On models held out from item selection, it has the highest mean fidelity at the released 5% retention. The hierarchical design lets practitioners stop at any stage to match their compute budget; the released suite removes over 97% of items while preserving model rankings. Beyond compression, our analyses show how VLM evaluation behaves as model panels grow and evolve, providing guidance for designing future benchmarks that are more efficient, robust to model turnover, and explicit about the limits of evaluation-side pruning.

[NLP-57] From Dissonance to Orchestration: Teacher Intervention in On-Policy Distillation

【速读】: 该论文旨在解决在**在线策略蒸馏(On-Policy Distillation, OPD)中,教师干预策略的深度与时机选择不当所导致的训练分布偏移(off-policy load)问题。现有方法通常依赖于生成轨迹的质量来决定是否引入教师指导,但研究发现仅以轨迹质量为标准无法有效分配教师干预,因为深层干预虽能提升轨迹准确性,却会显著增加学习分布与原始策略的偏离程度,从而损害模型泛化能力。此外,不同任务基准下最优干预强度与位置存在差异,表明静态干预策略缺乏适应性。为此,论文提出MAESTRO方法,其核心创新在于利用局部策略不一致度(local policy disagreement)**动态决策教师介入的时机与持续时间。该不一致度分数综合了教师加权候选覆盖范围与局部分布相似性,并在推理段落内进行聚合,实现对教师介入行为的精细化控制。实验结果表明,在8个数学推理基准上,MAESTRO在0.6B与1.7B规模的Qwen3学生模型中均取得最高宏平均准确率,且1.7B模型在所有任务上表现领先;同时,相较标准OPD,MAESTRO将平均训练响应长度减少67.3%。

链接: https://arxiv.org/abs/2609.37510
作者: Yuhao Wang,Ruiyang Ren,Yinan Zhang,Ruiqing Zhang,Jing Liu,Chunyan Miao
机构: Nanyang Technological University, Singapore; Baidu Inc.
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:On-policy distillation (OPD) trains a student on its own reasoning trajectories using feedback from a stronger teacher. Teacher interventions can improve these trajectories, but also change the distribution on which the student learns. Our controlled studies show that rollout quality alone is an incomplete criterion for allocating teacher guidance. Deeper intervention yields diminishing gains in rollout accuracy while increasing off-policy load. In a training probe with a restricted rollout horizon, peak student accuracy and performance retention favor different intervention strengths. The preferred intervention depth and placement also vary across benchmarks. These findings motivate MAESTRO, which uses local policy disagreement to jointly adapt when the teacher takes over and how long it generates. Its policy disagreement score combines teacher-weighted candidate coverage with local distribution similarity and is aggregated within reasoning paragraphs. Across eight mathematical reasoning benchmarks, MAESTRO achieves the highest macro-average accuracy among the compared methods for both 0.6B and 1.7B Qwen3 students, with the 1.7B student leading on every benchmark. MAESTRO also reduces average training response length by 67.3% relative to standard OPD. The code is available at this https URL.

[NLP-58] Evaluating Bounded Autonomy in Regulated Agent ic AI: A Diagnostic Harness with Constitutional Rewards Escalation Labels and Runtime Governance

【速读】: 该论文旨在解决受监管的智能体工作流中“有限自主性”(bounded autonomy)的可信度评估与控制问题,核心挑战在于如何在保障安全性与合规性的前提下,有效判断智能体是否应主动请求人类干预(即“升级”或“defer”)。其解决方案的关键在于提出一种名为RegLLM的诊断工具包,通过集成六种可度量的可信性信号:引用有效性、来源可追溯性、模式合规性、升级正确性、宪法对齐性以及不安全动作率,这些信号分别由程序化验证器、任务级升级标签或AI评判分数提供监督。系统采用确定性运行时监管器,在检测到无依据回答时强制触发升级并记录干预行为,同时利用同一领域宪法规则统一指导评估、训练奖励与服务阶段的防护机制。特别地,任务级“应升级”标签将“执行”与“延迟决策”转化为可量化的训练信号,从而实现对智能体自主行为的有效引导。实验表明,即使在小规模测试中,配置差异也可能显著掩盖调优效果,凸显了大规模重复实验的必要性,为后续更可靠的适配器评估和生产部署提供了诊断基础。

链接: https://arxiv.org/abs/2609.37501
作者: Dipankar Sarkar
机构: Skelf Research(斯克尔研究); https://skelfresearch.com
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 9 pages; ancillary evaluation artefacts. Previously submitted to NLLP 2026

点击查看摘要

Abstract:We propose RegLLM, a diagnostic harness for bounded autonomy in regulated agentic workflows. It instruments six trustworthiness signals: citation validity, source grounding, schema compliance, escalation correctness, constitutional alignment, and unsafe-action rate. Signals are distinguished by their source of supervision: programmatic verifiers, task-level escalation labels, or AI-judge scores. A deterministic runtime supervisor blocks ungrounded answers and forces escalation, logging interventions. The same domain constitution informs evaluation, training rewards, and serving guardrails. Task-level should-escalate labels make the act-versus-defer decision a measurable training signal. We demonstrate the harness at smoke scale. An offline reference run (n=12) lifts escalation recall from 0 to 0.67 and reduces unsafe-action rate from 0.33 to 0.08 when governance is enabled. Two single-GPU Qwen2.5-3B LoRA/DPO pilots (n=8, same seed and evaluation split) expose substantial variation: nominally identical RL-base configurations yield task success of 0.25 versus 0.12 and escalation recall of 1.0 versus 0.5. An answer-quality adapter changes recall from 1.0 to 0.5 in Run A, but from 0.5 to 1.0 in Run B. An escalation-aware variant produces no measurable change in Run B. These small pilots do not establish reliable adapter effects or production readiness. Their contribution is diagnostic: configuration variance can overwhelm apparent tuning effects on bounded-autonomy metrics, motivating larger evaluation sets and repeated runs.

[NLP-59] Who Warmed the Archives? LLM s Overestimate Historical Warmth

【速读】: 该论文旨在解决如何利用大语言模型(LLM)从历史文献中提取可用于扩展仪器气候记录的温度指数,以弥补长期气候数据的缺失。其核心挑战在于:尽管某些模型在相关性上表现良好,但可能存在系统性偏差,尤其随时间推移呈现持续的暖化偏倚,这种偏差会破坏跨世纪比较的可靠性。研究的关键发现是,尽管所有测试的六种主流大语言模型(如Gemini 2.5 Flash、GPT-5-mini等)均表现出一致的暖化趋势(斜率+0.13至+0.34/世纪,p<0.01),且该偏差与训练数据无关,即使去除文本中的显式日期和时代标记后仍存在,表明模型倾向于引入一种“今时今日”的先验认知,而非准确推断历史语境。相比之下,基于词法的基准方法在预测准确性上优于所有微调的历史化变压器模型,甚至包括从零开始预训练的历史德语文本模型。因此,研究指出仅依赖相关系数不足以验证大语言模型作为历史气候指数提取工具的可靠性,必须警惕其内在的年代系统性偏差,这构成了解决方案的核心关键——即需建立对模型输出偏差的独立评估机制,而不能仅依赖表面的相关性指标。

链接: https://arxiv.org/abs/2609.37499
作者: Claudiu Creanga,Liviu P. Dinu
机构: Interdisciplinary School of Doctoral Studies; HLT Research Center, University of Bucharest (布加勒斯特大学高级语言技术研究中心), Romania
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Historical archives are an under-used source for extending the instrumental climate record backward in time, and LLMs offer a way to extract the indices climatologists derive by hand. Beyond measuring how well systems extract this signal, we check whether their errors are safe to use for cross-century comparison, since a good correlation score does not rule out systematic, era-linked bias. Comparing lexical baselines, fine-tuned historical transformers, and LLM prompting on the Pfister temperature index across five centuries of German text, lexical methods beat every fine-tuned transformer we test, including one pretrained from scratch on historical German (r=-0.016). All six LLMs we test (Gemini 2.5 Flash, GPT-5-mini, DeepSeek v4 Flash, Claude Sonnet 4.6, Qwen3.7-Plus, Kimi-K2.6-Fast) show a warm bias that grows with calendar year, with the same sign in every model (slopes +0.13 to +0.34/century, p0.01). The effect is modest in size (r-squared approx equal to 0.01 to 0.05) but consistent across six independently developed models. The best-correlated of the six, Gemini 2.5 Flash, matches the best lexical correlation (r=0.32) at double the error. An ablation stripping explicit dates and calendar-era markers from the quotes leaves this trend essentially unchanged, favoring an anachronistic present-day prior over the model correctly inferring the quote’s era. Correlation alone is thus insufficient for vetting an LLM as a historical-climate-index oracle.

[NLP-60] he Rashomon Wikipedia: A Data-Perspectivist Analysis of Divergent Historical Narratives

【速读】: 该论文旨在解决多语言维基百科在争议历史事件上存在的跨语言史学偏见(cross-lingual historiographical bias)问题,揭示不同语言版本在叙述同一历史事件时因认知共同体差异导致的叙事分裂现象。其核心解决方案在于提出一种名为“Peace-Maker”的自动化冲突调和管道,通过引入“对抗性提示”(adversarial prompting)机制,明确要求模型保留并准确归因分歧观点,从而有效避免标准提示下大语言模型(Large Language Models, LLMs)产生的共识幻觉(hallucinated consensus),实现近似完美的中立性表现。该方法为跨语言知识整合中的客观性保障提供了可扩展的技术路径。

链接: https://arxiv.org/abs/2609.37498
作者: Claudiu Creanga,Liviu P. Dinu,Anca Dinu
机构: Interdisciplinary School of Doctoral Studies; HLT Research Center; Faculty of Mathematics and Computer Science; Faculty of Foreign Languages and Literatures, University of Bucharest (布加勒斯特大学), Romania
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Wikipedia aims to provide a unified, neutral record of history, yet its independent language editions often function as distinct epistemic communities, creating divergent narratives around contested events. This paper investigates cross-lingual historiographical bias by analyzing Wikipedia articles across five languages (Romanian, Hungarian, Russian, Turkish, and English) focusing on three contentious events in Romanian history: the Battle of Posada (1330), the Soviet occupation of Bessarabia (1940), and the Night Attack at Targoviste (1462). Using human annotators and Large Language Models (LLMs) to classify citation stance and quantify narrative evolution from 2005 to 2024, we identify a phenomenon of “citation isolation”. In the case of the Battle of Posada, only 2 out of 119 citations were shared between language editions, with the Romanian edition exhibiting a 91% pro-national bias compared to the balanced Hungarian edition. Longitudinal analysis reveals that these narratives are volatile and responsive to contemporary geopolitics, evidenced by a significant shift in the Russian framing of Bessarabia in 2024. Finally, we propose a “Peace-Maker” pipeline to automate conflict reconciliation. We demonstrate that while standard prompting leads models to hallucinate consensus, “adversarial” prompting, which explicitly instructs the model to preserve and attribute disagreement, achieves near-perfect neutrality scores.

[NLP-61] Larry Caused the Car to Stop But the Model Didnt Notice: Transformer Blindness to the M-Heuristic

【速读】: 该论文旨在解决生成式语言模型在句法-语用推理能力方面的缺失问题,具体探究编码器型变压器模型(如DeBERTa)是否能够遵循M-Heuristic(即标记形式蕴含标记意义的语用原则,源自新格赖斯理论)。其核心挑战在于:尽管现代变压器模型在语义表征方面表现优异,但其对语言形式中隐含的语用差异(如词法因果句与非词法因果句之间的微妙差别)是否具备敏感性仍不明确。解决方案的关键在于通过自然语言推理(Natural Language Inference, NLI)框架,对比词法因果结构(如“Larry stopped the car”)与非词法因果结构(如“Larry caused the car to stop”)在188种条件下的语用区分能力。实验结果表明,DeBERTa、RoBERTa和BART均未能捕捉到二者间的语用差异,其中DeBERTa在所有测试案例中均预测为“中立”(Neutral),且探针分析显示原初观察到的表示-使用分离现象实为语法复杂度的干扰信号。进一步分析发现,非词法因果句在嵌入空间中更接近直接方式描述,这与M-Heuristic的预期相反。然而,在显式元语言提示下,Gemini Flash-Lite可实现100%准确率并保留特定项的推理痕迹,说明该语用原则在指令引导下是可被激活的,但在默认NLI任务中未被有效利用。因此,该研究揭示了当前主流模型在缺乏显式引导时对深层语用规则的忽视,凸显了语用推理在模型设计中的潜在缺口。

链接: https://arxiv.org/abs/2609.37497
作者: Stefania Butnaru,Claudiu Creanga
机构: 未知
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Modern transformer models excel at capturing semantic relationships through sentence embeddings, yet their ability to perform pragmatic reasoning remains understudied. This paper investigates whether encoder-based transformers such as DeBERTa employ the M-Heuristic (the neo-Gricean principle that marked linguistic forms implicate marked meanings). We test this hypothesis by contrasting lexical causatives (e.g., Larry stopped the car'') with periphrastic causatives (e.g., Larry caused the car to stop’‘) using a Natural Language Inference framework. Our experiments across 188 conditions with 15 ambitransitive verbs reveal that DeBERTa, RoBERTa, and BART show no evidence of capturing the pragmatic distinction between these forms, with DeBERTa predicting ``Neutral’’ for 100% of cases. Probing analysis initially suggested a representation-use dissociation, but control experiments reveal the probe was tracking syntactic complexity, not causative pragmatics. Semantic similarity over 30 triplets places periphrastic causatives closer to unmediated manner descriptions in 29/30 cases, opposite to M-Heuristic predictions in the embedding space. Under explicit metalinguistic framing, Gemini Flash-Lite reaches 100% with item-specific traces, so the principle is available under instruction yet unused in default NLI.

[NLP-62] Your Benchmark Is Not Saturated: Reviving Multiple-Choice Evaluation with Answer Pooling

【速读】: 该论文旨在解决多选题基准测试(multiple-choice benchmark)因评分成本低而逐渐趋于饱和、难以进一步提升难度的问题。传统方法通过编写更复杂的题目来增强挑战性,但这一过程耗时且需为每个基准单独重复进行。针对此问题,论文提出AnswerPool方案:将共享同一上下文的N个题目所有选项合并成一个统一选项池,要求模型为每道题从该池中选择正确答案。该方法无需重新编写题目或更改标签,显著降低随机猜对概率(由10⁻³降至5×10⁻⁷,针对五道四选一题目),同时保留模型基于已有知识的作答能力。通过比较原始多选题得分与池化后得分的差异,可量化评估传统格式中模型通过排除错误选项所获得的“消除信用”(elimination credit)。此外,移除部分选项使题目无法通过精确真值解答,从而引入弃权机制,实现同一轮推理中对可答性与置信度的联合评估。实验覆盖八项文本、图像和视频基准及十八个模型,结果表明,池化后的任务对所有模型均更具挑战性,且弱模型在消除策略上获益最大;七种开放权重模型在不可回答题目上仍能正确回答87%至100%,反映出其强大的泛化与不确定性识别能力。

链接: https://arxiv.org/abs/2609.37494
作者: Mohamed Eltahir,Abobaker Ahmed,Nawaf Barebood,Hussain Bu Subayt,Tanveer Hussain,Naeemullah Khan
机构: King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia; Department of Computer Science, Edge Hill University, Ormskirk, England
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Multiple-choice benchmarks are cheap to grade and are running out of room, and the standard remedy, writing harder items, is slow and repeated for every benchmark. A saturated benchmark still holds a harder task. Each question’s wrong options are written for that question alone, so a model can score by eliminating a few options. We propose AnswerPool: take N questions that share a context, pool all their options into one list, and ask the model to assign every question its answer. No item is written and no label changes. The chance of guessing a group right falls from 10^-3 to 5\times10^-7 for five four-option questions, and a model that recognizes its answers keeps its multiple-choice score, so the accuracy lost to pooling measures the credit the format gave for elimination. Deleting answers from the pool makes questions unanswerable with exact ground truth, so abstention is scored in the same pass. Across eight text, image, and video benchmarks and eighteen models, pooling is harder for every model, the elimination credit is largest for the weakest models, and seven of eight open-weight models answer 87 to 100% of unanswerable questions.

[NLP-63] Risk-Controlled Selective LLM Answering by Pricing Label-Free Checks

【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)在生成回答时如何有效决定何时“放弃回答”(abstain)的问题,尤其关注验证器(verifier)的排名准确率并不能直接反映所服务答案中的错误率这一核心挑战。其解决方案的关键在于提出PriceCheck框架,该框架通过构建一组基于无标签检查(如重新求解问题)的紧凑决策规则家族,为每个检查赋予“价格”——即其在正确与错误答案上的一致率以及每次运行的成本。这些价格参数在少量类别丰富的有标签数据集上进行拟合,进而用于预测不同检查调度策略的覆盖范围与成本,从而指导应执行哪些检查以及何时停止。随后通过校准测试,在预设的选择性风险目标下选择最优调度方案。实验表明,在数学任务中,所选调度平均可服务76.1%的答案,且在全部15个数据划分上保持外部选择性风险低于1.5%。在共享测试协议下,PriceCheck在相同覆盖率下比奖励模型、提示式裁判、生成器置信度及训练好的正确性分类器等方法拥有最少的错误答案;此外,基于价格的覆盖预测与实际观察覆盖之间具有高达0.97的秩相关性,验证了其预测可靠性。研究结果表明,检查项的组合方式与终止策略的选择,与验证器本身的排序能力同等重要。

链接: https://arxiv.org/abs/2609.37493
作者: Dongyub Jude Lee,Jungseob Lee,Chanjun Park,Hyeonseok Moon,Heuiseok Lim
机构: Zoom Communications(Zoom通信); Korea University(高丽大学); Soongsil University(崇实大学); Sookmyung Women’s University(淑明女子大学)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 29 pages, 6 figures, 24 tables. Dongyub Jude Lee and Jungseob Lee contributed equally

点击查看摘要

Abstract:Serving an answer from a large language model requires deciding when to abstain, yet a verifier’s ranking accuracy alone does not determine the error rate among served answers. We introduce PriceCheck, which builds a compact family of decision rules from label-free checks such as re-solving a problem. Each check has a price: its agreement rates on correct and incorrect answers and its cost per run. Prices fitted on a small, class-enriched labelled set compose into predictions of a schedule’s coverage and cost, guiding which checks to run and when to stop. A calibration test then selects a schedule at a stated selective-risk target. In mathematics, the selected schedules serve 76.1% of answers on average and keep held-out selective risk below 1.5% on all 15 splits. Under the shared testing protocol, PriceCheck serves more answers at that target than reward models, a prompted judge, the generator’s confidence and a trained correctness classifier. At matched coverage, it keeps the fewest wrong answers among these scorers. Across 118 diagnostic schedules, price-based coverage predictions have a rank correlation of 0.97 with observed coverage. These results show that choosing how checks are combined and stopped matters alongside how well a verifier ranks answers. Code is available at this https URL.

[NLP-64] Regime Boundary Alignment for Evidence-Gated Question Answering

【速读】: 该论文旨在解决检索增强型语言模型在缺乏支持证据时仍坚持生成答案的问题,即模型在证据缺失情况下仍频繁“猜测”而非选择不回答。其核心问题是:现有以答案为导向的微调策略在无支持上下文时缺乏明确目标,导致模型无法区分“主动拒绝回答”与“随意猜测”,从而在支持准确率提升的同时,未有效降低无证据情况下的错误回答率。解决方案的关键在于提出制度边界对齐(Regime Boundary Alignment, RBA),通过在相同问题与正确答案的匹配变体上联合训练单个阅读器,使其在上下文支持时输出正确答案(即使存在冲突证据),而在正确支持被移除时主动放弃回答。该方法无需验证器、阈值或制度标签,仅依赖标准解码即可实现精准的证据门控回答。实验表明,RBA在三个多跳问答数据集上相比冲突导向训练将无证据回答率降低超过60个百分点,同时保持支持准确率;在独立的TriviaQA检索遗漏子集上,无证据回答率从100%降至1%以下,并进一步提升了有证据情况下的准确率,证明了在支持边界的两侧均需学习证据门控机制的重要性。

链接: https://arxiv.org/abs/2609.37491
作者: Zeyan Li,Qirong Guo,SIyuan Qiu,Hu Xu,Chun Li,Jianfeng Xu
机构: Shanghai Jiao Tong University(上海交通大学); The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Retrieval-augmented language models are expected to answer from the retrieved evidence, but in practice they often keep answering when that evidence is missing. We trace this behavior to the training signal: answer-focused fine-tuning assigns no target to unsupported contexts, so it cannot distinguish a reader that abstains from one that guesses, and unsupported answering stays near 100% even as supported accuracy improves. We introduce Regime Boundary Alignment (RBA), which trains a single reader on matched variants of the same question and gold answer. The reader is trained to produce the gold answer when the context supports it, including when conflicting evidence is also present, and to abstain when the correct support is removed; inference is ordinary decoding, with no verifier, threshold, or regime label. On three multi-hop QA datasets across three seeds, RBA reduces the unsupported-answer rate by more than sixty percentage points relative to conflict-focused training while matching its supported accuracy. On a held-out TriviaQA retrieval-miss slice, the same reader reduces unsupported answering from 100% to below 1% while also improving supported accuracy. These results indicate that evidence-gated answering must be learned on both sides of the support boundary.

[NLP-65] FORUM: Frozen Outputs Reconciled Using Model Agreement for Visual Grounding ACCV2026

【速读】: 该论文旨在解决冻结的多模态大语言模型(Multimodal Large Language Models, MLLMs)在对抗性基准测试中对同类别干扰项和否定句式存在严重误判的问题,尽管这些模型在标准指代表达理解任务上表现优异,但在面对具有混淆性的复杂语义场景时仍会自信地给出错误答案,且重复采样无法纠正错误。其解决方案的关键在于提出一种无需训练、仅在测试阶段进行融合的框架——FORUM,该框架基于两个固定的几何规则:一是基于模型共识的选择机制,保留由最多不同模型支持的区域以增强鲁棒性;二是通过中位数定位(medoid localization)返回实际存在的目标框而非坐标平均值,从而避免单一不准确预测对结果产生漂移影响。通过融合三个开源的冻结MLLMs,FORUM在对抗性Ref-Adv-s基准上相较397B参数的现有最优模型实现了5%的相对精度提升,并优于简单平均集成方法15%;此外,该方法在标准基准RefCOCO+上同样表现出色,即使在无主导模型的均衡配置下,也超越了397B参数模型5%的性能。

链接: https://arxiv.org/abs/2609.37488
作者: Taiyo Sato,Takamasa Sanda,Keisuke Maeda,Takahiro Ogawa,Miki Haseyama,Shunya Nagashima
机构: Hokkaido University (北海道大学); Neurogica Inc. (神经逻辑公司)
类目: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: Accepted by ACCV 2026

点击查看摘要

Abstract:Frozen multimodal large language models (MLLMs) now solve standard referring expression comprehension with a single prompted call, yet on adversarial benchmarks with same-category distractors and negation, even the largest models are confidently wrong, and resampling repeats the error. Models built from different data and architectures rarely fall for the same confounder, so their agreement is a strong label-free signal of the correct target. We present FORUM, a training-free test-time fusion of frozen MLLMs guided by two fixed geometric rules: agreement-based selection keeps the region supported by the most distinct models, and medoid localization returns an actual member box instead of a coordinate average, so one loose prediction cannot shift the answer. Fusing three open MLLMs, FORUM surpasses the 397B-parameter published reference by a relative 5% in mean accuracy on the adversarial Ref-Adv-s benchmark, and a plain averaging ensemble by 15%. The gains transfer to standard RefCOCO+, and a balanced lineup with no dominant member still surpasses the 397B model by 5%.

[NLP-66] Learning to Retrieve Missing Evidence for Long-Term Memory QA

【速读】: 该论文旨在解决大语言模型在长对话中因信息分散于远距离对话轮次而难以有效检索与整合关键证据的问题。其核心挑战在于:用户提问时往往省略定位所需证据的关键线索,导致传统检索机制难以精准捕获相关信息。为此,论文提出MERA(Missing-Evidence Retrieval Augmentation)框架,其解决方案的关键在于将全局可搜索的记忆空间与针对具体问题的证据状态进行解耦建模。通过引入一个轻量级规划器(planner),基于强化学习训练以奖励那些能够恢复先前遗漏证据的检索策略,从而实现对已发现证据的动态引导式检索。该设计允许模型在不限制访问全局记忆的前提下,持续优化检索路径,显著提升多轮对话中的证据召回率与答案准确性。实验表明,采用Qwen3-30B作为后端处理与生成模型时,仅0.6B规模的规划器在LoCoMo和LongMemEval-S基准上分别达到77.40%和71.29%的准确率,优于未经过检索增强训练的30B规划器,且在后期检索回合中累计证据召回率从55.5%提升至80.5%,验证了该方法在长程记忆利用上的有效性。

链接: https://arxiv.org/abs/2609.37443
作者: Yi-Xuan Deng,Yi Zhang,Wei Liu,Chao Xue,Shuojin Yang
机构: Tsinghua University(清华大学); JD.COM(京东)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 22pages,6figures

点击查看摘要

Abstract:Long-term memory enables language models to use past interactions in future conversations. However, evidence needed to answer a question may be scattered across distant turns, while the question itself omits clues needed to locate it. Retrieved facts can reveal these clues, motivating retrieval decisions conditioned on evidence already found. We introduce MERA (Missing-Evidence Retrieval Augmentation), which separates globally searchable memory from a question-specific evidence state. Verified evidence guides subsequent retrieval without restricting access to the global memory. We train a lightweight planner through reinforcement learning, rewarding queries that recover previously missing evidence. MERA achieves strong answer accuracy across Qwen3-30B and GPT-4o-mini backbones. With Qwen3-30B for evidence processing and answer generation, the trained 0.6B planner achieves 77.40% accuracy on LoCoMo and 71.29% on LongMemEval-S, exceeding a 30B planner without retrieval-grounded training by 4.10% and 3.96%, respectively. On LoCoMo, later retrieval rounds increase cumulative evidence recall from 55.5% to 80.5%.

[NLP-67] Look What You Made Us Cluster: Hate Narrative Extraction from Reddit Discourse

【速读】: 该论文旨在解决现有计算方法在识别网络仇恨叙事时精度不足的问题,其核心挑战在于依赖语义表示难以捕捉深层语义结构,导致对叙事内容的解析停留在表层。为提升叙事提取的精确性与可解释性,本文提出一种基于实体-评价对(entity-evaluation pairs)表示的叙事抽取流程。该方案的关键在于利用大语言模型(LLM)的推理能力,扩展面向方面的情感分析(Aspect-Based Sentiment Analysis),通过识别话题方面(aspect)、分类判断类型(judgement type)并推导出相应评价,实现对叙事结构的深层解析。随后,采用Leiden算法对提取的叙事进行聚类,并结合LLM引导的精细化调整过程,将聚类结果优化至目标粒度。研究以2024年英文Reddit评论中针对泰勒·斯威夫特(Taylor Swift)的批评内容为例,分析一个体现仇恨言论模式的代表性聚类,验证了该方法在揭示复杂叙事模式方面的解释价值。

链接: https://arxiv.org/abs/2609.37408
作者: Annabelle K. L. Chua,Forster J. Khoo,Joel C. R. Tan,Huey Ting Ang,Kheng Hwee Tan,Joel Y. A. Sim,Shirley W. H. Ow,Ria Mundhra,Elsie C. K. Toh,Youfeng Xu,Lynnette H. X. Ng
机构: DSO National Laboratories(新加坡国防科技实验室); Defence Science and Technology Agency(新加坡国防科技局)
类目: Computation and Language (cs.CL); Social and Information Networks (cs.SI)
备注: Accepted to IDeaS Conference 2026

点击查看摘要

Abstract:Narrative extraction allows us to identify online hate narratives, supporting the construction of rigorous detection systems. Existing computational approaches, however, are limited in precision as they rely on semantic representations, which tend to capture only surface-level meaning. To detect more precise and interpretable narratives, we present an extraction pipeline that represents narratives as entity-evaluation pairs. Narratives are extracted using a Large Language Model (LLM) reasoning process that extends Aspect-Based Sentiment Analysis, identifying the aspect, classifying its judgement type as the basis for evaluation, and deriving the evaluation accordingly. Extracted narratives are then clustered using Leiden, following which clusters are resolved to an intended level of granularity through an LLM-guided refinement process. We illustrate this narrative pipeline with English Reddit comments from 2024 that criticize Taylor Swift, analyzing a representative cluster that exhibits hate speech patterns to demonstrate its interpretive value.

[NLP-68] Compiling Learning Problems into Adaptation Programs for Language Models

【速读】: 该论文旨在解决模型适应(model adaptation)过程中依赖固定策略导致行为结果差异显著的问题,即不同更新方案虽产生相似形式的适应流程,却可能带来截然不同的学习性能与行为特性。其核心解决方案是提出“适应编译”(adaptation compilation),将适应的时机、方式及程度重新建模为一个联合预测与决策问题。该方法通过学习历史适应经验,预先预测多个候选适应程序在获取性(acquisition)、迁移性(transfer)、有界性(boundedness)和保持性(preservation)等多维度上的潜在影响,形成一个向量化的反事实响应面(counterfactual response surface),并据此在适应前选择最优程序。由于该预测结构捕捉了多维行为后果而非单一评分或胜出方案,因此可在不同下游目标下复用而无需重新训练。实验表明,在五类学习任务中,最优适应程序随学习阶段动态变化,且这种变化可由适应前信息有效预测;在Llama-3.1-8B上,编译器所选程序接近穷举搜索性能,优于全局或特定目标默认策略;在Gemma-2-9B上的复现进一步验证了程序多样性与选择余地的存在,但强调偏离强默认策略时必须考虑不确定性以实现有效利用。总体而言,该研究实现了跨相关学习任务的适应搜索成本分摊,使过往适应经验成为未来学习策略制定的基础。

链接: https://arxiv.org/abs/2609.37371
作者: Rebecca Ramnauth,Brian Scassellati
机构: Yale University (耶鲁大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Model adaptation is typically governed by a fixed recipe, even though different update programs can produce substantially different behavioral outcomes. We introduce adaptation compilation, which reframes where, how, and to what extent a model should adapt as a joint prediction and decision problem. Rather than searching over candidate programs anew for each learning episode, a compiler learns from prior adaptations to predict a vector-valued counterfactual response surface over candidate programs—their expected effects on acquisition, transfer, boundedness, and preservation—and selects a program before adaptation begins. Because this predicted geometry captures multiple behavioral consequences rather than a single winner or scalar score, it can be reused under different downstream priorities without retraining. Across five learning types, preferred programs vary meaningfully across episodes, and this variation is predictable from pre-adaptation information. On Llama-3.1-8B, compiler-selected programs approach exhaustive search while outperforming global and objective-specific defaults. Replication on Gemma-2-9B preserves program heterogeneity and selection headroom, but shows that exploiting this headroom requires accounting for uncertainty when departing from strong defaults. Together, these results show that adaptation search can be amortized across related learning problems, turning prior adaptation experience into a basis for deciding how future learning should occur.

[NLP-69] SemOPT: Fixing Semantic Errors in LLM -based Optimization Modeling via Reward-Guided Search EMNLP2026

【速读】: 该论文旨在解决生成式人工智能(Generative AI)在求解运筹学(Operations Research, OR)问题时,基于大语言模型(LLM)构建的优化模型中存在的语义错误(semantic errors)问题。这类错误表现为模型代码虽能成功运行并返回目标值,但其数学表达与原始问题意图不符,因不触发运行时异常而难以被检测和修正。针对此挑战,论文提出一种语义引导的框架SemOPT,其核心在于结合一个能够区分忠实反映原问题的数学模型与看似合理却错误的模型的语义奖励模型(semantic reward model),并引入一种分层奖励驱动的自适应修正系统,在建模空间中进行高效搜索以修复错误。实验结果表明,SemOPT在七个优化建模基准上达到新的性能上限,在复杂数据集上相较最强基线平均提升7.6%的准确率,显著增强了生成式模型在运筹学建模任务中的可靠性与正确性。

链接: https://arxiv.org/abs/2609.37361
作者: Zetong Zhou,Wentao Zhang,Jingyuan Wang,Yifan Yang,Zizhuo Wang,Shixi Hu
机构: Beihang University (北京航空航天大学); MIIT Key Laboratory of Data and Decision Intelligence, Beihang University (工业和信息化部数据与决策智能重点实验室, 北京航空航天大学); The Chinese University of Hong Kong, Shenzhen (香港中文大学(深圳)); Cardinal Operations Technology Co. ( cardinal 操作技术公司)
类目: Computation and Language (cs.CL)
备注: Accepted at EMNLP 2026 (Findings)

点击查看摘要

Abstract:Operations research supports decision-making in domains such as energy, economics, and healthcare. Solving operations research problems typically begins with optimization modeling, which translates a natural-language problem description into executable solver code. LLMs offer a promising way to automate this process, but they remain prone to errors. In practice, these errors can be divided into two categories: syntactic errors refer to solver code that fails to run successfully or is judged infeasible by the solver; semantic errors refer to solver code that successfully returns an objective value but violates the intent of the original problem. Since semantic errors do not trigger runtime failures, they are difficult to detect and rectify. To address this problem, we introduce SemOPT, a semantic-guided framework for correcting LLM-based optimization models. SemOPT combines a semantic reward model that distinguishes faithful math models from plausible but incorrect ones with an adaptive correction system that applies hierarchical reward-guided search over the modeling space. Experiments on seven optimization modeling benchmarks show that SemOPT establishes a new state of the art and achieves an average 7.6% accuracy improvement over the strongest baseline on complex datasets.

[NLP-70] Port-Hamiltonian Latent Deliberation: Mitigating the Deliberation Drift Cliff in Test-Time Compute Scaling

【速读】: 该论文旨在解决测试时计算扩展(test-time compute scaling)中连续潜在表示空间内迭代思辨(deliberation)所面临的“思辨漂移悬崖”(Deliberation Drift Cliff)问题,即在深层推理步骤下模型性能急剧下降的灾难性现象。其核心挑战在于如何在表达能力(expressivity)、李雅普诺夫稳定性(Lyapunov stability)与计算效率之间实现三重平衡。解决方案的关键在于提出端口-哈密顿潜变量思辨框架(Port-Hamiltonian Latent Deliberation, PH-LD),并设计直接梯度纯张量赫姆霍兹-霍奇分解(DG-HHD)。该方法通过将吸引流参数化为切向投影张量网络,并正交解耦非零环流分量(实现赫姆霍兹-霍奇分解误差低至1.65e-17),彻底消除运行时自动微分依赖,从而在保持高表达能力的同时显著抑制长期漂移(漂移悬崖仅3.40%),并在向量场评估和四阶龙格-库塔积分(RK45)中分别实现1.84倍和2.09倍的速度提升。实验表明,该方法在对称帕累托基准上达到58.67%峰值准确率(较保守型赫姆霍兹-霍奇分解提升25.94%),且在K=32时仍保持35.27%的推理性能;在小语言模型多跳因果推理任务中实现单调计算扩展(49.33%→51.56%)并有效抑制分布外漂移(漂移悬崖仅-0.66%),所有30个零级确定性不变量均被形式化认证。

链接: https://arxiv.org/abs/2609.37351
作者: Zeyu Jia(School of Biomedical Engineering and Technology, Tianjin Medical University, Medical School, Tianjin University)
机构: Tianjin Medical University (天津医科大学); Tianjin University (天津大学)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 10 pages, 1 figure, 4 tables. Code and evaluation artifacts available

点击查看摘要

Abstract:Test-time compute scaling has emerged as a cornerstone of advanced machine reasoning, yet performing iterative deliberation directly within continuous latent representation spaces reveals a catastrophic pathology: the Deliberation Drift Cliff. While unconstrained recurrent latent models achieve initial reasoning gains at short horizons (K = 4), their reasoning collapses when extrapolated to deeper thinking steps (K = 16), dropping by 22% to 62% across standard logical benchmarks. We resolve the trilemma among expressivity, Lyapunov stability, and computational efficiency in test-time latent reasoning through a 22-round empirical and theoretical investigation. We demonstrate that strictly conservative scalar potential gradient flows suppress long-range drift (cliff 3.40%) but bottleneck peak reasoning accuracy at 32.73%, whereas unconstrained rotational flows achieve high symbolic expressivity (82.33%) but suffer a severe 36.87% drift cliff. To resolve this geometric duality, we establish Port-Hamiltonian Latent Deliberation (PH-LD) and propose the Direct-Gradient Pure-Tensor Helmholtz-Hodge Decomposition (DG-HHD). DG-HHD parameterizes the attracting flow as a tangent projection tensor network while orthogonally decoupling non-zero circulation (Hodge machine error 1.65e-17, contraction error 5.55e-17), eliminating runtime autograd dependencies to achieve 1.84x vector field and 2.09x RK45 rollout speedups. In a 15-arm symmetrical Pareto benchmark, DG-HHD achieves 58.67% peak accuracy (+25.94% absolute gain over conservative HHD) and retains 35.27% at K=32. Transferred to small language model (SLM) multi-hop causal reasoning, DG-HHD delivers monotonic compute scaling (49.33% to 51.56%) and suppresses out-of-distribution drift (cliff -0.66%). All 30 Level 0 deterministic invariants are certified.

[NLP-71] Solving Without Stopping: On-Policy Distillation at Small Scale

【速读】: 该论文旨在解决在小规模模型上通过在线蒸馏(on-policy distillation)进行推理能力迁移时,究竟哪些知识能够有效传递的问题。其核心挑战在于:尽管教师模型具备长序列推理与适时终止的能力,但小规模学生模型在学习过程中往往无法完整继承“何时停止推理”这一关键判断能力。解决方案的关键在于揭示了蒸馏过程中的两个分离机制——问题求解能力(problem-solving ability)与推理终止能力(stopping ability)的非对称性传递。研究发现,蒸馏能有效转移前者,即学生模型在单次尝试中解决问题的能力随模型规模缩小而递减,但始终受限于训练前多轮尝试所能达到的上限;然而,在“思考模式”(thinking mode)下,蒸馏几乎无法传递后者,因为教师仅在学生已自发停止的位置发出终止信号,导致小模型无法学习新的、正确的停止策略。尤其对于最小的学生模型(如0.6B),其虽常能接近正确答案,却缺乏稳定标记并确认答案的能力,表现为频繁不标注或标注后继续冗余输出。因此,该研究提出了一套诊断框架,可将答案标记、答案正确性与推理终止时机三者明确区分,从而为理解小模型在在线蒸馏中的行为提供了系统性分析工具。

链接: https://arxiv.org/abs/2609.37326
作者: Hongyang Li,Yiming Zhu,Xiao Li,Caesar Wu,Said Mammar,Pascal Bouvry
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 22 pages, 13 figures

点击查看摘要

Abstract:On-policy distillation, where a student learns from a stronger teacher’s feedback on its own outputs, is a common way to pass reasoning to smaller models. We analyze what it transfers at small scale, distilling Qwen3-8B into Qwen3 4B, 1.7B and 0.6B students, in thinking mode (reason at length, then end the reasoning and answer) and, for comparison, in non-thinking mode (no separate reasoning phase). Long reasoning needs two abilities, solving a problem and knowing when it is solved, and we find that distillation transfers the first, but in thinking mode not the second. Solving improves at every size, up to two ceilings, which we measure comprehensively across both modes and all student sizes: a student’s single attempt never exceeds what it could already reach in many attempts before training, and the smaller the student, the further it stays below the teacher. Stopping is where the modes part. In non-thinking mode every student keeps stopping; in thinking mode students stop ending their reasoning early in training, and the smaller the student, the less of this ability survives: the teacher signals a stop almost only where a student already ends its reasoning, so distillation teaches no new stops; it only keeps the student’s existing stops that land on a right answer, and a weak student has few such stops. The smallest students often reach the right value but do not commit to it: they either rarely mark it or mark it and write past it. Together, these results describe how small students behave under on-policy distillation, and a diagnostic that separates answer marking, correctness and stopping.

[NLP-72] Hidden Reasoning Must Leak but Need Not Be Readable: Fundamental Opportunities and Limits for Chain-of-Thought Monitoring

【速读】: 该论文旨在解决生成式模型是否能够通过思维链(Chain of Thought, CoT)监控机制隐藏其内部计算过程的问题。研究发现,模型能否实现隐蔽计算取决于任务复杂度与模型规模的权衡:对于简单任务,模型可完全隐匿计算路径;但当任务复杂度超过由模型规模决定的阈值后,完成任务必然会在思维链中泄露接近线性量级的关于隐藏输入的信息,表明在足够复杂的任务下,隐蔽计算总会留下信息论意义上的痕迹。然而,令人担忧的是,这种泄露信息可能无法被实时监测者读取——在合理的密码学假设下,即便是单层Transformer也能对推理过程进行在线加密,使任何多项式时间的监控系统都无法提取隐藏计算的相关信息。因此,该研究揭示了CoT监控在理论与实践上的双重边界:既存在潜在的隐蔽计算机会,也面临信息泄露与不可读性的复杂权衡。

链接: https://arxiv.org/abs/2609.37312
作者: Mohammadali Mohammadkhani,Madhava Krishna,Yash Sarrof,Michael Hahn
机构: 未知
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Can reasoning models trick chain of thought (CoT) monitors and perform hidden computation without revealing it in their thinking traces? We show that the answer depends on the underlying task difficulty and the model size. Simple computations can be performed covertly; however, beyond a threshold depending on model size, successfully solving the task necessarily leaks a near-linear amount of information about the covert task input into the CoT. Therefore, sufficiently complex hidden computation always leaves an information-theoretic footprint. However, concerningly, this leakage need not be readable: Under plausible cryptographic assumptions, even a one-layer Transformer can encrypt its reasoning online so that no polynomial-time monitor can extract information about the hidden computation. Overall, our theoretical and empirical results provide a holistic view of both the opportunities and the limitations of CoT monitoring.

[NLP-73] Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents

【速读】: 该论文旨在解决当前智能代理(Agent)在执行任务时过度依赖用户显式请求的问题,即代理往往仅响应用户明确提出的需求,而忽视了任务完成所必需的隐含信息。现有研究多聚焦于代理是否应主动行动(proactive),却未深入探讨其主动行为的内容选择与终止时机。本文提出一种新的主动性维度:内容层面的主动性,具体分为两类——横向主动性(Horizontal proactivity)指在当前上下文中已识别但未被用户提及的信息的主动探索;纵向主动性(Vertical proactivity)则指向通过早期证据揭示的深层需求。为量化评估这两种主动性及停止决策的合理性,作者构建了一个从基准数据集分解中恢复出的需求图(Need Graph),该图记录了各需求之间的依赖关系,从而可基于对话转录文本无须依赖模型评判者即可进行评分。为学习此行为,提出QD(Questioner and Drafter)框架,训练一个“提问者”(Questioner)模块,使其倾向于选择能更高效获取所需证据的后续问题,整个过程无需奖励模型或人工评判。在三个多跳问答基准的独立测试集上,该方法在相同检索开销下显著提升两种主动性表现,且优于提示工程(prompted)的同规模模型,并在两个基准上超越15倍大的提示模型,该优势在控制问题数量与长度后依然存在。进一步实验表明,未经额外训练的提问者在模拟客户交互场景中能以更少提问完成更多任务,在零售场景中亦以更少客户回访次数超越大15倍的基线模型。结果表明,主动性不仅取决于是否自主行动,更关键在于主动追求的信息内容及其终止时机的选择。

链接: https://arxiv.org/abs/2609.37236
作者: Ido Levy,Asaf Yehudai,Segev Shlomov,Asaf Adi,Leshem Choshen
机构: IBM; Weizmann Institute of Science(魏茨曼科学研究所)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 48 pages. Project page: this https URL Code: this https URL Model: this https URL

点击查看摘要

Abstract:An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested. Work on proactive agents mainly studies whether and when an agent should act on its own, not what information it should pursue. We study a distinct axis of proactivity: its content. Horizontal proactivity pursues unstated information that the current context already identifies, and vertical proactivity pursues needs that only earlier evidence reveals. A need graph, recovered from a benchmark’s own decomposition, records which needs depend on which, so both forms, and whether the agent stops at the right time, can be scored from a transcript without a model judge. To learn this behavior, we propose QD (questioner and drafter), which trains a questioner to prefer the question whose continuation retrieves more of the required evidence, with no reward model or judge. On held-out splits of three multi-hop question-answering benchmarks, at equal retrieval spend, the trained questioner improves both forms of proactivity over the same model, prompted, and outperforms a prompted model 15\times larger in the same role on two of the three, and the gain persists after controlling for question volume and length. Without further training, we place the questioner in an interactive customer-service agent with a simulated customer, where it completes more tasks while asking fewer questions, and in retail it outperforms the 15\times larger model with fewer follow-up turns from the customer. These results show that proactivity depends not only on whether an agent acts without being asked, but also on what it chooses to pursue and when it stops.

[NLP-74] CredWise: A Controlled Agent ic Decision-Intelligence Framework for Explainable and Auditable Credit-Risk Assessment

【速读】: 该论文旨在解决传统信用风险预测模型仅提供结果而缺乏可解释性及多源证据融合能力的问题,即无法阐明申请人风险的成因,也无法有效整合政策规则、结构化数据分析与预测结果以支持决策。其核心解决方案是提出CredWise框架,通过集成信用风险预测、概率校准、可解释人工智能(Explainable AI)、政策检索、SQL分析以及受控的基于代理的工作流,实现从预测到决策支持的闭环。关键技术在于:采用时间划分的XGBoost模型在Lending Club数据上进行训练与验证,确保模型泛化能力;通过概率校准显著降低Brier分数与期望校准误差,提升预测置信度可靠性;利用SHAP值实现特征重要性的时序稳定性分析(Spearman相关系数达0.9959);借助FAISS实现高效政策文档检索,达成Hit@1为0.929、MRR为0.964的优异表现;结合代理路由与SQL执行引擎,在复杂决策路径中实现了高达95.6%的准确率与100%的精确匹配、执行成功率与结果一致性,充分验证了系统在统一工作流中融合预测、解释、政策依据与结构化分析的能力。

链接: https://arxiv.org/abs/2609.37223
作者: Aakash Kumar Tiwari
机构: Indian Institute of Technology Kharagpur (印度理工学院克哈格普尔分校)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Credit-risk prediction is important in banking, but a prediction alone does not explain why an applicant is risky or how it should be combined with other evidence. This paper presents CredWise, a decision-support framework that integrates credit-risk prediction, probability calibration, explainable artificial intelligence, policy retrieval, SQL analytics, and controlled agent-based workflows. An XGBoost model is trained on Lending Club data (1,345,310 loans, 18 features) using a temporal split: 2007–2016 for training, 2017 for validation, and 2018 for testing. On the 2018 test set, the calibrated model achieved a ROC-AUC of 0.7109, PR-AUC of 0.2993, F1-score of 0.3714, and accuracy of 65.44%. Calibration reduced the Brier score from 0.2157 to 0.1273 and the expected calibration error from 0.2862 to 0.0585. SHAP explanations were temporally stable, with a Spearman correlation of 0.9959 between 2017 and 2018 feature rankings. On 28 labeled queries covering nine policy sections, FAISS achieved the best Hit@1 (0.929) and MRR (0.964), while all three retrieval methods reached Hit@5 = 1.0. Agent routing achieved 95.6% accuracy (43 of 45 cases), and the SQL benchmark scored 1.0 on exact-match, execution-success, and result-match across six cases. These results show that CredWise can combine predictions, explanations, policy evidence, and structured analytics in one controlled workflow. It is an academic research prototype, and final decisions remain with a human reviewer.

[NLP-75] VLM Fine-Tuning for End-to-End Combinatorial Optimization

【速读】: 该论文旨在解决大语言模型(LLM)在端到端组合优化(CO)任务中因仅依赖文本序列化而可能忽略问题空间结构与关系信息的问题,导致生成解的质量受限。其核心解决方案是提出一种通用的视觉-语言求解器(vision-language solver),通过将输入实例的文本描述与自动生成的视觉表征相结合,以增强对问题几何结构和元素间关系的感知能力。该方法采用单一视觉-语言模型(VLM)跨多个组合优化任务进行统一建模,并通过监督微调与验证器引导的强化学习进行训练。尽管视觉输入不包含任何真实解或解相关的信息,实验结果表明,该视觉增强策略显著提升了求解质量,尤其在复杂度较高的问题如容量约束车辆路径问题(CVRP)和作业车间调度问题(JSSP)上表现突出,且在大规模问题实例中优势更为明显。关键在于利用视觉表征有效捕捉问题的空间与关系结构,从而弥补纯文本表示的局限性。

链接: https://arxiv.org/abs/2609.37175
作者: Qingsong Yan,Xia Jiang,Yaoxin Wu,Wen Song,Lu Zhang,Yingjie Zhou
机构: 未知
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Large language models (LLMs) have provided a unified interface for end-to-end combinatorial optimization (CO), but textual serialization alone may obscure spatial and relational structures that are important for generating effective CO solutions. This paper presents a general-purpose vision-language solver that augments textual instance descriptions with input-derived visual representations. A single vision-language model (VLM) is applied across different CO tasks and trained using supervised fine-tuning followed by verifier-guided reinforcement learning. While the visual inputs contain no gold solutions or solution-derived information, our experiments show that the VLM generally improves solution quality over its text-only counterpart, with particularly clear gains on more complex CO problems such as CVRP and JSSP. The advantage of visual information is more pronounced at large problem scales.

[NLP-76] Bridging Semantic Gaps in RAG through Generated Context Knowledge Fusion NLPCC2026

【速读】: 该论文旨在解决检索增强生成(Retrieval-Augmented Generation, RAG)框架中查询与检索文档之间存在的语义空间不匹配问题,这一问题严重影响了生成答案的相关性与准确性。其解决方案的关键在于提出一种名为知识感知语义桥接(Knowledge-Aware Semantic Bridging, KASB)的新框架,通过多阶段智能知识融合实现查询与文档在语义空间上的对齐,充分结合生成式模型与基于检索的知识优势,从而显著提升段落选择的质量与最终生成结果的准确性。

链接: https://arxiv.org/abs/2609.37171
作者: Xinkai Du,Chao Lv,Yalin Sun,Quanjie Han,Lei Yao,Maosong Sun
机构: Beijing Wanlian Zhilian Technology Corporation Limited(北京万联智联科技有限公司); Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系); Sunshine Digital Intelligence Tech Co., Ltd.(阳光数字智能科技有限公司)
类目: Computation and Language (cs.CL)
备注: This paper is accepted by NLPCC 2026

点击查看摘要

Abstract:Retrieval-Augmented Generation has established itself as a fundamental framework in natural language processing, seamlessly integrating information retrieval with the generative capabilities of large language models. However, this process is fundamentally constrained by a critical challenge: semantic space mismatch between queries and retrieved contexts. We propose Knowledge-Aware Semantic Bridging (KASB), a novel framework that improves passage selection quality through semantic space alignment between queries and retrieved documents through intelligent knowledge fusion. Our approach leverages the complementary strengths of generative and retrieval-based knowledge through a multistage process that enhances both relevance and accuracy. We evaluate KASB on three popular open-domain Question Answering datasets to demonstrate the effectiveness of our approach.

[NLP-77] rajectory Soup: Pushing the Compute-Scaling Frontier of LLM Mid-training via Diverse Trajectories

【速读】: 该论文旨在解决中训练(mid-training)阶段计算资源分配效率受限的问题,即在预训练大语言模型的中训练过程中,持续增加串行计算量带来的下游性能提升逐渐饱和,甚至可能损害部分能力,从而形成实际可吸收计算量的瓶颈。其核心解决方案是提出“轨迹汤”(Trajectory Soup)方法,通过将中训练预算分配至多个独立分支(基于同一检查点、不同受控训练配方生成),利用参数空间中不同轨迹所展现出的兼容多样性(compatible diversity),实现跨轨迹与轨迹内双重平均以整合最优检查点。研究通过局部偏差-方差分析揭示:跨轨迹平均能消除单个轨迹内部平均无法触及的残余误差,而检查点选择本身引入的偏差则限制了合并的最优检查点数量。实验表明,在不同模型规模、学习率调度、令牌预算及轨迹数量下,轨迹汤在匹配预算条件下均优于最强单轨迹平均,并随预算扩大持续提升性能,且在相同后训练流程后优势依然保持。该成果确立了轨迹分配与融合作为突破中训练串行计算饱和瓶颈、扩展计算-性能增长边界的实用策略。

链接: https://arxiv.org/abs/2609.37169
作者: Zhehao Huang,Changxin Tian,Qingyuan Yang,Kunlong Chen,Ziqi Liu,Zhiqiang Zhang,Xiaolin Huang,Jun Zhou
机构: Ant Group (蚂蚁集团); Shanghai Jiao Tong University (上海交通大学)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Mid-training equips pretrained large language models with specialized and reasoning capabilities, but the returns of this stage are bounded since additional serial compute yields little further downstream improvement and can even degrade some capabilities, which places a practical ceiling on how much compute mid-training absorbs. We revisit how this compute should be allocated to a single run or multiple similar optimizations. We find that branches forked from a shared checkpoint under various controlled recipe reaches measurably different regions of parameter space, and establish a form of compatible diversity that extending one run cannot supply. Therefore, we introduce Trajectory Soup, which distributes a mid-training budget over several independent branches, and consolidates strongest checkpoints selected on validation through intra- and inter-trajectory averaging into a single model. A local bias and variance analysis separates the two averaging levels, showing that inter-trajectory averaging removes residual error beyond the reach of averaging within a trajectory, while checkpoint selection carries a bias that bounds how many checkpoints are worth merging. Across model scales, learning-rate schedules, token budgets, and trajectory counts, Trajectory Soup improves aggregate downstream performance over the strongest single-trajectory average under matched budgets and keeps improving as budgets expand, with the advantage preserved after an identical post-training pipeline. These results position trajectory allocation and merging as a practical way to extend the compute-scaling frontier of mid-training beyond serial saturation.

[NLP-78] LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems

【速读】: 该论文旨在解决当前编码代理(coding agents)在大型软件系统中进行实际模块化开发时,面临的感知能力(perception capability)与实现能力(implementation capability)双重不足的问题。现有基准测试主要关注从详细规格说明生成正确代码编辑的能力,但忽略了开发者在真实场景中需先理解用户意图和高层设计以推导出规范的关键环节。为此,论文提出LoLBench,一个涵盖29个大型软件系统、5个领域共100个任务的多语言基准,覆盖从增强提案(enhancement proposal)到最终代码实现的完整流程。每个任务平均包含约5,000词的用户意图与高阶设计描述,涉及平均240万行源代码,且实现变更的合并请求(PR)平均修改约5,500行代码。在对28个编码代理的评估中,表现最佳者仅能解决14%的任务,其Fail-to-Pass(F2P)通过率仅为52.7%。失败分析表明,代码定位不完整是主要瓶颈;而引入基于参考实现的文件树结构与API规格说明后,任务解决率提升16–22个百分点(提升2.4–17倍),最高达34%。这表明,在大型软件系统的实际模块化开发中,编码代理仍面临感知与实现双方面的核心挑战。

链接: https://arxiv.org/abs/2609.37143
作者: Yun Peng,Zihan Wu,Zeyang Zhuang,Xin Zhou,Rui Shu,Xu Han,Chun Yong Chong,Yuan Wang,Jiakun Liu
机构: Fudan University (复旦大学); City University of Hong Kong (香港城市大学); Chinese University of Hong Kong (香港中文大学); Singapore Management University (新加坡管理大学); Independent Researcher (独立研究员); HKUST (GZ) (香港科技大学(广州)); Monash University Malaysia (莫纳什大学马来西亚分校); Yuan Wang (独立研究员); Harbin Institute of Technology (哈尔滨工业大学)
类目: oftware Engineering (cs.SE); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Modern coding agents can deliver increasingly large repository-level changes, and recent benchmarks reflect this by emphasizing long-horizon tasks with large reference implementations. Many benchmarks evaluate coding agents’ implementation capability to produce correct code edits from detailed specifications. However, practical modular development tasks also require the perception capability of grounding user intent and high-level design to derive a specification. We introduce LoLBench to evaluate both capabilities through the entire proposal-to-implementation process on large software systems. It is a multilingual benchmark of 100 tasks across 29 software systems in five domains. Each task provides a human-written enhancement proposal with user intent and high-level design. On average, proposals contain about 5,000 words, software systems contain 2.4 million source lines of code (LoC), and implementation pull requests (PRs) change approximately 5,500 LoC. Across 28 agents we evaluated, the best agent resolves only 14% of tasks and achieves a 52.7% Fail-to-Pass (F2P) pass rate. Failure analysis identifies incomplete code localization as a major bottleneck, while providing reference-derived file trees alongside API specifications improves resolved rates by 16–22 percentage points (2.4–17 \times ), reaching at most 34%. These results show that both perception and implementation remain central challenges for coding agents in practical modular development on large software systems. LoLBench is available at this https URL.

[NLP-79] LLM unbranding: Erasing Commercial Identity while Preserving Generic Utility

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在生成文本过程中无意泄露品牌信息所带来的风险问题,特别是由品牌标识(trade dress)引发的商标稀释、虚假归因和品牌诽谤等潜在法律与声誉风险。其核心挑战在于,相较于显式的视觉标志,品牌特征在文本中以特定语言风格、口号、表达习惯等隐蔽形式存在,难以被传统方法有效识别与消除。为此,论文首次正式定义了“LLM去品牌化”(LLM Unbranding)这一新任务,并提出关键解决方案——MUTE(Model-Unbranding via Iterative Tuning and Elimination),一种基于推理时迭代优化系统指令的轻量级方法。MUTE通过动态调整提示词策略,在不修改模型参数的前提下,系统性地消除文本输出中的品牌痕迹,同时保持模型的通用能力与实用性,从而实现高效、稳健的去品牌化。

链接: https://arxiv.org/abs/2609.37127
作者: Kajetan Ożóg,Alicja Wojciechowska,Dawid Malarz,Paweł Batorski,Artur Kasymov,Przemysław Spurek
机构: Jagiellonian University (亚捷隆大学); IDEAS Research Institute (IDEAS 研究所); Heinrich Heine Universität Düsseldorf (海因里希·海涅杜塞尔多夫大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Establishing unbranding as a critical practice to prevent visual logos from acquiring negative connotations is standard in image generation. Large Language Models (LLMs) now face a parallel and emerging challenge. These models frequently generate brand descriptions within diverse contexts. This frequency introduces significant risks, such as trademark dilution, false attribution, and brand defamation. In response, we formally define the novel task of LLM Unbranding. We specifically address the complex challenge of managing trade dress within textual outputs. This involves neutralizing characteristic language, slogans, and stylistic markers that define brand identity. Crucially, these elements are less evident than explicit visual logos. To benchmark this task, we introduce a comprehensive evaluation dataset incorporating prominent brands from multiple commercial domains. We rigorously evaluate existing state-of-the-art machine unlearning models using this benchmark. This evaluation identifies their limitations in selective textual unbranding. Finally, we propose MUTE, a novel inference-time method that effectively neutralizes textual trade dress while preserving the LLM’s general capabilities and utility. By leveraging an iterative refinement loop, MUTE systematically optimizes system instructions to safely eliminate brand leakage without requiring fragile parameter updates. Code and dataset: The evaluation dataset and code for LLM Unbranding are available at this https URL. The implementation of MUTE is available at this https URL. Subjects: Computation and Language (cs.CL) Cite as: arXiv:2609.37127 [cs.CL] (or arXiv:2609.37127v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2609.37127 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[NLP-80] Cross-Linguistic Effects in Bilingual Phoneme BabyLMs EMNLP2026

【速读】: 该论文旨在解决双语第一语言习得中跨语言影响(cross-linguistic effects)的机制问题,尤其关注语言模型在模拟儿童语言习得过程时因输入模态差异而导致的偏差。传统语言模型多基于书面语(orthographic text)进行训练,而真实儿童主要通过听觉语音输入(spoken input)习得语言,这一根本性差异限制了模型对人类语言习得过程的准确模拟。为此,本文提出关键解决方案:采用音位级语音表示(phonemic representations)作为输入,构建双语婴儿语言模型(Bilingual BabyLMs),并在控制发育合理性约束的前提下,以英语为固定第二语言(L2),将第一语言(L1)分别设置为德语、瑞典语、波斯语和巴斯克语,以覆盖语法结构与音位库与英语在距离上的多样化组合。实验结果表明,在音位输入条件下,语法习得轨迹呈现出更强的受母语(L1)影响的个体差异,而早期词汇习得差异则与音位库存相似性高度一致,凸显了音位输入在揭示真实语言习得动态中的关键作用。

链接: https://arxiv.org/abs/2609.37121
作者: Nikitas Theodoropoulos,Maria Lymperaiou,Giorgos Filandrianos
机构: Independent Researcher(独立研究员); National Technical University of Athens(雅典国立技术大学); Instituto de Telecomunicações, Portugal(葡萄牙电信研究所)
类目: Computation and Language (cs.CL)
备注: 13 pages, 8 figures, 3 tables; Accepted at the 2nd BabyLM Workshop at EMNLP 2026

点击查看摘要

Abstract:Cross-linguistic effects are a central topic in bilingual first-language acquisition. Artificial learners can help investigate L1-L2 interactions by enabling controlled comparisons across language combinations and learning conditions. Recent work explores this direction by training bilingual language models under developmentally plausible constraints. However, human and model learners still diverge in fundamental ways, with one major difference being input modality: children learn primarily from spoken input, whereas language models are typically trained on orthographic text. To reduce this gap, researchers have trained models on phonemic representations of speech. In this work, we combine these research directions to train bilingual BabyLMs with phonemic input. We keep English fixed as the L2 and vary the L1 across German, Swedish, Persian, and Basque, selected to represent contrasting combinations of syntactic and phoneme-inventory distance from English. Our results show stronger L1-related variation in grammatical learning trajectories under phonemic than orthographic input, while early lexical differences align with phoneme-inventory similarity.

[NLP-81] Unlocking the Critic: Reward-Free Policy Optimization for LLM Post-Training

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在后训练阶段进行强化学习(Reinforcement Learning, RL)时,因去除价值函数(critic)而导致的训练效率低下与长时序推理(long-horizon reasoning)中奖励信号稀疏的问题。现有方法普遍摒弃critic以降低训练不稳定性与内存开销,但忽略了其在预测未来结果方面的潜在价值。本文的关键突破在于:通过精细优化策略更新的方差与步长,证明了长链式思维(chain-of-thought)推理中critic带来的不稳定性主要源于优化过程中的数值偏差,而非架构本质缺陷;进一步发现,一个充分预训练的critic能够从轨迹中后期状态和未完成前缀中估计出最终成功概率,从而提供基于结果的密集型、每前缀的学习信号。基于此,作者提出无奖励策略优化(Reward-Free Policy Optimization, RFPO),将单一校准后的冻结critic重用于三重角色——作为回放级奖励信号、广义优势估计(Generalized Advantage Estimation, GAE)的值基准,以及未完成生成序列的成功预测器。此外,通过去偏后的得分二值化,有效抑制了critic固有的长度偏好(length bias)。实验表明,二值化的RFPO在无需任何外部标注或完整轨迹的情况下,性能可媲美监督式PPO,同时显著降低计算与内存开销。这一方法使得训练不再需等待所有轨迹完成即可获得奖励,特别适用于延迟反馈的长时序任务。研究挑战了当前主流的“无critic”范式,确立了基于critic的无奖励优化为一种可扩展、高效率的大语言模型后训练路径。

链接: https://arxiv.org/abs/2609.37119
作者: Hongyang Li,Xiao Li,Caesar Wu,Said Mammar,Grégoire Danoy,Pascal Bouvry
机构: 未知
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 26 pages, 15 figures, 16 tables

点击查看摘要

Abstract:Recent approaches to reinforcement learning (RL) post-training for large language models increasingly remove the critic to reduce training instability and memory overhead. Even where a critic is trained, it is discarded once training ends, although it has learned to predict outcomes. We revisit this trend and show that a pretrained critic’s ability to predict future outcomes can make it a valuable asset for efficient long-horizon reasoning. First, we find that instability in critic-based RL for long chain-of-thought reasoning is largely an optimization artifact: keeping policy updates small and low in variance restores stable convergence. Second, a well-pretrained critic estimates the posterior probability of eventual success from later trajectory states and unfinished prefixes. Its predictions provide outcome-derived, dense, per-prefix learning signals that, during policy optimization, require neither completed rollouts, step-level annotations, nor external reward labels. Building on this insight, we introduce Reward-Free Policy Optimization (RFPO), which repurposes a single calibrated, frozen critic as a rollout-level reward, a value baseline for generalized advantage estimation, and a success forecaster for unfinished prefixes. We further show that binarizing the debiased score stops the policy from exploiting the critic’s length bias. Binarized, RFPO matches supervised PPO without a single label in the training loop, while cutting compute and memory overhead. This makes RFPO well suited to long-horizon reasoning tasks, where outcomes arrive late and generation dominates cost: because rollouts can be rewarded before they finish, training no longer has to pay for waiting on every trajectory to complete. Our findings challenge the prevailing critic-free paradigm and establish critic-based, reward-free optimization as a scalable and computationally efficient path for LLM post-training.

[NLP-82] VACE: Validation-Gated Alternating Co-Evolution of Agent Models and Harnesses

【速读】: 该论文旨在解决语言模型智能体(language model agents)在任务执行过程中,模型权重更新与任务引导框架(harness)优化之间耦合导致的性能瓶颈问题。传统方法中,权重更新改变模型对现有框架的使用方式,而框架调整又影响训练轨迹,二者相互干扰,难以实现协同优化。为此,论文提出一种名为VACE(Validation-Gated Alternating CoEvolution)的解决方案,其核心在于通过验证门控的交替协同进化机制,解耦并分阶段优化模型权重与任务引导框架。具体而言,VACE在每轮强化学习(RL)训练后,利用收集到的轨迹数据生成候选框架改进方案,并在固定当前模型权重的前提下,通过验证集性能评估候选框架的有效性;仅当候选框架能提升验证性能时,才被采纳用于后续训练。该方法有效避免了无效或有害的框架变更对训练过程的负面影响。实验结果表明,基于Qwen3.5-9B模型,VACE在OfficeQA上达到45.26%的测试准确率,在AutomationBench上取得75.19%的平均部分得分,分别优于仅更新权重的强化学习方法6.43和9.09个百分点,以及无验证门控的交替优化方法4.59和6.95个百分点。在44次框架提议中,有17次因降低验证性能而被拒绝,充分验证了验证门控机制在保障优化方向正确性中的关键作用。

链接: https://arxiv.org/abs/2609.37105
作者: Jiexing Qi,Yu He,Jun Liu,Qichen Huang,Shaohua Hu,Zhan Dang,Guohua Chen,Rui Yang,Wen Jiang,Yang Liu,Tao Lyu,Fangming Li
机构: Huawei Technologies Co., Ltd. (华为技术有限公司); Shanghai Jiao Tong University (上海交通大学)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Language model agents can be improved by updating their model weights or refining the harness that guides task execution. These components are coupled: weight updates change how the model uses the harness, while harness updates change the trajectories used for training. We propose VACE, Validation-Gated Alternating CoEvolution, which alternates agentic reinforcement learning with trajectory-driven harness refinement. After each RL stage, VACE reuses the collected trajectories to propose a harness revision and evaluates the incumbent and candidate with the updated model held fixed. The candidate guides subsequent training only if it improves validation performance. With Qwen3.5-9B, VACE achieves 45.26% test accuracy on OfficeQA and a mean partial-credit score of 75.19% on AutomationBench, exceeding weight-only RL by 6.43 and 9.09 percentage points and ungated alternation by 4.59 and 6.95 points, respectively. Across 44 harness proposals, 17 reduce validation performance at the updated checkpoint and are rejected before subsequent RL training, highlighting the importance of validation gating.

[NLP-83] What Does Post-Training Change in Multilingual Reasoning ?

【速读】: 该论文旨在解决开源推理模型在多语言环境下推理能力获取不平等的问题,即模型虽具备解题能力,却无法以用户指定语言输出完整且可读的推理过程,导致语言本身成为访问障碍而非单纯性能差异的来源。其核心挑战在于如何实现跨语言的可靠推理输出,确保模型在非英语语言中既能正确求解问题、保持语言一致性,又能终止推理并高效交付结果。解决方案的关键在于构建一个分阶段的后训练路径:首先通过多语言监督微调(Multilingual SFT)恢复目标语言的推理能力,但需承受准确率下降及非终结循环等副作用;随后引入强化学习(RL)阶段,通过设计包含语言一致性的奖励函数,不仅恢复了非英语推理的终止性,还实现了高质量、高语言合规性的输出,而仅奖励正确性则使模型重新回归以英语为中心。这一系列实证分析揭示了不同后训练阶段的瓶颈转移规律,最终确立了一条从英语主导到多语言可靠推理的可复现、可优化的技术路径。

链接: https://arxiv.org/abs/2609.37104
作者: Hongyang Li,Xiao Li,Caesar Wu,Grégoire Danoy,Pascal Bouvry
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 20 pages, 9 figures, 21 tables. Main paper and supplementary material in one document

点击查看摘要

Abstract:Open-source reasoning models provide unequal access to reasoning capability across languages. When a model can solve a problem but cannot deliver a complete solution in the user’s language, language becomes an access barrier rather than merely a source of performance variation. We audit Qwen3 checkpoints on competition-mathematics tasks in eleven languages. Across the ten non-English languages, only 15.4-17.9% of problems receive a correct, terminating solution with visible reasoning in the requested language in any of 16 samples, compared with 92.9% in English. To identify the source of this disparity, we evaluate thirteen endpoints from one model family, spanning released checkpoints, multilingual supervised fine-tuning (SFT) at two scales, controlled SFT ablations, and three reinforcement-learning (RL) reward formulations. We jointly track correctness, language adherence, termination, and delivery efficiency. The dominant bottleneck shifts across post-training stages. Released models often reason in English. Multilingual SFT restores target-language reasoning, but accuracy declines across multilingual, English-only, and single-language SFT runs, showing that this cost is not specific to multilingual mixing; non-English reasoning traces additionally become prone to non-terminating loops. RL restores termination in both arms at no cost in accuracy, but only the arm whose reward includes a language term delivers: rewarding correctness alone returns the model to English. Together, these stages establish a constructive post-training path from English-pivoted capability to multilingual reasoning that is reliably delivered.

[NLP-84] raverse: Learning When to Remember Reset and Redirect for Long-Horizon Web Search

【速读】: 该论文旨在解决长时程信息检索智能体在持续交互过程中因上下文累积噪声或误导性信息而导致早期错误持续传播、难以纠正的问题。其核心解决方案是引入一种自主搜索框架(autonomous search harness),通过“标准(Rubric)—答案(Answer)—验证(Verify)”三态机制实现对搜索过程的自我管理:智能体首先定义有效答案的标准,基于此标准进行搜索,并在获取结果后独立执行验证,再决定是否终止或继续搜索。为支持这一过程,系统配备“封存记忆(Seal Memory)”工具以实现主动的上下文管理。然而,采用强化学习训练该行为时易引发“封存坍塌(Seal Collapse)”现象,导致训练不稳定并阻碍智能体学习何时及如何使用记忆工具。为此,作者提出一种简化策略——仅对上下文管理后的最终阶段进行训练,有效缓解了该问题。实验表明,该方法在BrowseComp上达到72.83的得分,优于同类开源系统,并在BrowseComp-ZH、xbench、DeepSearchQA、WideSearch、金融调查和产品搜索等多个任务中均显著优于基线模型。消融实验进一步验证了自主压缩优于自动压缩,且所提出的强化学习设计具有有效性。

链接: https://arxiv.org/abs/2609.37082
作者: Jingyuan Ma,Lynx Aster,He Zhang,Siyao Song,Weijie Yuan,Zhe Zhang,Kai Jia,Zhifang Sui
机构: Peking University (北京大学); ByteDance (字节跳动)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Long-horizon information-seeking agents often accumulate noisy or misleading context, causing early mistakes to persist and making recovery increasingly difficult. We introduce an autonomous search harness in which the agent manages its own search process through three states: Rubric, Answer, and Verify. The agent first defines criteria for a valid answer, searches under these criteria, and then independently verifies the result before deciding whether to terminate or continue searching. It is further equipped with a Seal Memory tool that enables active context management. Training this behavior with reinforcement learning, however, can induce Seal Collapse, resulting in unstable training and preventing the agent from reliably learning when and how to use its memory tools. We solve this with a simple strategy that trains only the final segment after context management. Our 35B model achieves 72.83 on BrowseComp, outperforming comparable open-source systems, and consistently improves over the base model across BrowseComp-ZH, xbench, DeepSearchQA, WideSearch, financial investigation, and product search. Ablations show that autonomous compression outperforms automatic compaction and validate our RL design.

[NLP-85] Learning from Think-Mode Advantage via On-Policy Distillation

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在推理过程中如何有效利用教师模型的显式中间推理(explicit intermediate reasoning)优势,以提升学生模型的问题求解能力。其核心挑战在于:传统方法在知识蒸馏过程中,虽然保留了学生生成的轨迹(on-policy),但若使用固定的教师推理路径进行统一蒸馏(如Uniform ThinkOPD),可能导致不同学生响应路径与同一教师推理路径之间出现不兼容性,从而引发“推理-响应分歧”(trace-response divergence, TRD),影响学习效果。解决方案的关键是提出一种新型的思维增强型在线策略蒸馏方法——ThinkOPD,通过结合组内相对奖励增益与基于TRD的兼容性代理,实现对响应级别的监督路由,并在每轮采样组内对最终响应权重进行归一化。该方法能够动态适配不同学生路径与教师推理路径的匹配程度,显著缓解因固定推理路径带来的偏差问题。实验表明,ThinkOPD在数学推理和代码生成任务中均优于Uniform ThinkOPD及主流的推理蒸馏与自蒸馏基线,在同模型与跨模型设置下均展现出更强的性能与可迁移性,且控制性实验证明了结果收益与TRD代理信号具有互补性。

链接: https://arxiv.org/abs/2609.37044
作者: Wanqi Ren,Jianxiang Wang,Danxuan Liu,Linyi Ding,Yuan Zhang
机构: ByteDance(字节跳动)
类目: Computation and Language (cs.CL)
备注: 9 pages, 5 figures

点击查看摘要

Abstract:Explicit intermediate reasoning gives large language models (LLMs) a stronger problem-solving mode. We study learning from this think-mode advantage via on-policy distillation (OPD). OPD preserves student-generated trajectories and provides dense token-level teacher targets at student-visited prefixes. Privileged reasoning is used during distillation rather than student inference. Uniform ThinkOPD, a natural think-enabled OPD baseline, conditions a fixed teacher on one shared think trace and uniformly distills every sibling student response. Although its prefixes are on-policy, the trace need not follow a route compatible with every complete response: the same privileged trace can induce different teacher-student discrepancies even when responses reach the same outcome. We summarize this interaction with trace-response divergence (TRD) and introduce ThinkOPD, which routes supervision at the response level by combining group-relative reward gain with a TRD-based compatibility proxy. Final response weights are normalized within each rollout group. Across mathematical reasoning and code generation, ThinkOPD outperforms Uniform ThinkOPD in both same-model settings and both cross-model teacher-student pairs, and it exceeds representative rationale and self-distillation baselines in a controlled comparison. Controlled interventions show that outcome benefit and the TRD-based proxy provide complementary routing signals in this setting. Think-enabled OPD provides a controlled setting for studying how teacher advantage becomes transferable along student responses.

[NLP-86] Selecting The Most Informative Tokens in Natural Language Autoencoders

【速读】: 该论文旨在解决在大型语言模型(LLM)安全审计中,如何高效识别关键解释位置以检测潜在威胁(如提示注入和信息隐藏)的问题。由于对每个词元位置生成解释成本高昂,研究核心在于确定哪些位置的解释最具价值。其解决方案的关键在于:通过分析对话结构(chat structure)信号进行解释位置排序,相较于依赖模型计算过程中的内部激活信号,该方法无需执行模型前向传播即可更有效地筛选出相关性更高的解释位置。实验表明,在四个数据集中的三个上,仅需解释5%的高优先级位置,即可保留几乎与全量解释相当的威胁检测成功率。此外,研究发现预训练的表述器(pretrained verbalizers)能够恢复模型在微调过程中被刻意隐藏的词汇,且无需额外训练,揭示了有用解释可超越原始训练目标的边界。

链接: https://arxiv.org/abs/2609.37040
作者: Federico Torrielli,Gianluca Barmina,Andrea Blasi Núñez,Amon Rapp,Luigi Di Caro,Peter Schneider-Kamp,Lukas Galke Poech
机构: University of Turin(都灵大学); University of Southern Denmark(南丹麦大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Natural language autoencoders translate a language model’s internal activations into readable explanations. Explaining every token position is costly. Which positions should an auditor inspect to understand a potential threat? We study this question across 4.7 million explanations on prompt injection and concealment. We compare signals from model computation with a ranker trained only on chat structure. Chat structure usually selects more relevant explanations than the computational signals, without requiring a model forward pass for position selection. On three of four datasets, explaining just 5% of positions retains nearly all of the success rate from explaining every position, where success means obtaining an explanation about the threat. The benefit varies with the audit task. We also show that pretrained verbalizers recover words that models have learned to conceal through fine-tuning, without additional verbalizer training. These results identify where auditors can concentrate explanation generation and show that useful explanations can extend beyond the model a verbalizer was trained to describe.

[NLP-87] LatCom: Cross-Agent Latent Compression for Efficient Multi-Agent Collaboration

【速读】: 该论文旨在解决基于大语言模型(LLM)的多智能体系统(Multi-Agent System, MAS)在隐状态(latent)协作中因直接传递全部发送方隐状态而导致接收方上下文随智能体数量和推理长度线性增长所引发的计算开销大、内存占用高及协作延迟增加的问题。现有隐状态压缩方法虽能缓解部分负担,但未能有效消除跨智能体间的冗余信息,因其通常对各发送方隐状态独立压缩后拼接,导致压缩结果仍包含重复内容。本文提出一种名为LatCom的跨智能体隐状态压缩框架,其核心在于将多个发送方的隐状态映射为固定数量、接收方可读且任务相关的槽位(slots),而非试图重构所有发送方的隐藏状态;该框架优化压缩后的隐状态以最大化接收方的任务效用。LatCom采用两阶段训练:首先通过单智能体可读性学习建立接收方可理解的隐状态接口,随后通过多智能体融合学习实现互补信息融合并去除跨智能体冗余。在Qwen3-4B模型上多个基准测试结果表明,与LatentMAS相比,LatCom平均实现2.46倍的推理加速,并减少70.3%的输出标记使用量,同时保持相近的平均准确率。

链接: https://arxiv.org/abs/2609.37017
作者: Shinan Zhang,Tao Zhang,Qihui Zhu,Mengjie Zhang,Dong Jin,Yunpeng Hou,Shuangwu Chen,Xiaobin Tan,Quan Zheng,Jian Yang
机构: University of Science and Technology of China (中国科学技术大学); Institute of Artificial Intelligence, Hefei Comprehensive National Science Center (合肥综合性国家科学中心人工智能研究院)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:LLM-based multi-agent systems (MAS) increasingly use latent collaboration to avoid the information loss and repeated encoding-decoding overhead of natural-language communication. However, directly forwarding all sender latents makes the receiver-side context scale with both the number of agents and the reasoning length, increasing computation, memory usage, and collaboration latency. A natural solution is latent compression. But we find that cross-agent redundancy remains unresolved in existing latent compression approaches, which typically compress each sender independently and then concatenate the results. We propose LatCom, a cross-agent latent compression framework for efficient multi-agent latent collaboration. LatCom maps multiple sender latents into a fixed number of receiver-readable and task-relevant slots. Rather than reconstructing all sender hidden states, it optimizes the compressed latents for receiver-side task utility. LatCom trains the compressor in two stages: single-sender readability learning first establishes a latent interface interpretable by the frozen receiver, and multi-sender fusion learning then trains the compressor to fuse complementary evidence and remove redundancy across agents. Experiments on multiple benchmarks with Qwen3-4B show that LatCom achieves an average 2.46x inference speed-up over LatentMAS and reduces output token usage by 70.3% while maintaining comparable average accuracy.

[NLP-88] CypherTurn: A Multi-Turn Benchmark for Conversational Text-to-Cypher Evaluation and the Autonomy Divergence EMNLP2026

【速读】: 该论文旨在解决当前图数据库自然语言查询评估中缺乏对多轮对话场景真实模拟的问题,现有基准仅关注孤立的单轮查询,无法反映分析师在实际工作中所经历的连续性、上下文依赖的交互过程。为此,作者提出了首个面向对话式文本转Cypher(Text-to-Cypher)的基准测试——CypherTurn,包含721个会话、5,927个对话回合,覆盖7个知识图谱及13种对话现象。其解决方案的关键在于构建一个系统化、多样化的多轮对话评估框架,并通过两种协议(引导式代理协议与完全自主代理协议)对15个模型进行评估,揭示了多个关键发现:首先,最优模型的执行准确率仅为64.7%,会话级正确率不足5%;其次,尽管模型在整体性能上具有较高相关性,但在自主运行条件下顶尖模型排名出现显著变化,这一现象被称为“自主性分歧”(Autonomy Divergence),表明错误管理能力是独立于生成能力的重要维度;第三,增加动作预算(从3倍增至10倍)未能缩小自主性差距,因前沿模型自发将每轮动作数限制在约两个;第四,基于单轮Cypher微调会损害多轮指令遵循能力,而架构适配的专门化模型反而优于多个前沿模型。这些结果确立了CypherTurn作为对话式图数据库推理领域的一项开放挑战。

链接: https://arxiv.org/abs/2609.36987
作者: Yuzhe Zhang,Weijie Zhu,Haolin Yang,Ziyun Zhang,Xianwei Xue,Mengke Chen,Qiutong Pan,Huaqian Cai
机构: Peking University (北京大学); National Key Lab of Data Space Technology and System (数据空间技术与系统全国重点实验室); Baidu Inc. (百度公司)
类目: Computation and Language (cs.CL)
备注: Accepted as an oral paper at EMNLP 2026

点击查看摘要

Abstract:Graph databases are increasingly queried through natural language, yet every existing benchmark evaluates isolated single-turn queries rather than the multi-turn sessions through which analysts actually work. We introduce CypherTurn, the first benchmark for conversational Text-to-Cypher evaluation, comprising 721 sessions and 5,927 turns across 7 knowledge graphs and 13 conversational phenomena. We evaluate 15 models under a guided oracle protocol and a fully autonomous agentic protocol, yielding four findings. First, the best model reaches only 64.7% execution accuracy, and session-level correctness remains below 5%. Second, despite strong overall rank correlation, frontier models exhibit a consequential reordering of the top of the leaderboard under autonomous operation, a phenomenon we term the Autonomy Divergence, which reveals error-management as a partially independent capability from raw generation skill. Third, scaling action budgets from x3 to x10 fails to close the autonomy gap, as the strongest frontier models self-limit to approximately two actions per turn regardless of available budget. Fourth, single-turn Cypher fine-tuning degrades multi-turn instruction following, while architecture-appropriate specialization outperforms several frontier models. These results establish CypherTurn as an open challenge for conversational graph database reasoning. Code and data are available at this https URL.

[NLP-89] SRJudge: Empowering Large Language Models with Selective Reasoning for Fine-Grained Knowledge Concept Tagging IJCAI2026

【速读】: 该论文旨在解决教育内容中细粒度知识概念标注(Knowledge Concept Tagging)任务中存在的高维决策空间导致大语言模型(LLM)难以准确选择正确概念的问题。现有方法虽利用LLM取得一定成效,但在大规模候选集下仍存在泛化能力不足与推理偏差等问题。其解决方案的关键在于提出一种三阶段的Select-Reason-Judge(SRJudge)框架:第一阶段通过微调小型语言模型(SLM,如BERT)实现候选概念的初步筛选,生成一个高置信度的前K个短名单,显著压缩决策空间;第二阶段采用轻量级LLM作为推理器,结合改进的强化学习策略、动态任务特定奖励函数及剪枝机制,增强推理过程与人类认知偏好的一致性;第三阶段由更大的LLM担任裁判,评估推理逻辑与解释的合理性,输出最终标签。该框架通过分阶段的“选择—推理—判断”机制有效提升了标注精度与可解释性。此外,研究构建了两个高质量数据集(S_Bio和S_Phy),实验结果表明,SRJudge在多个基准数据集上均优于当前最先进方法,验证了其有效性与优越性。

链接: https://arxiv.org/abs/2609.36982
作者: Zhiwei Yang,Jiahua Yang,Huiru Lin,Xing Chen,Quanlong Guan
机构: Guangdong Institute of Smart Education, Jinan University, Guangzhou, China; School of Physical Education, Jinan University, Guangzhou, China; Guangdong Provincial Key Laboratory of Speed Capability Research, Guangzhou, China; Sapient Intelligence Pte Ltd, Singapore
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Accepted by IJCAI 2026

点击查看摘要

Abstract:Knowledge concept tagging aims to assign specific concept or topic labels to educational content, which is essential for both educators and learners in traditional and online teaching practices. Recent work has explored large language models (LLMs) for this task, achieving promising performance. However, LLMs still struggle to select the correct concept from a large-scale candidate set due to the high dimensionality of the decision space. In this paper, we propose a novel three-stage Select-Reason-Judge (SRJudge) framework, which empowers LLMs with selective reasoning capability for fine-grained knowledge concept tagging. Specifically, the Selector in Stage 1 first narrows the candidate concepts to a top-K shortlist by fine-tuning a small language model (SLM), e.g., BERT, since the top- K predictions hit the correct concept in most cases, thereby reducing the decision space of correct candidates. Next, the Stage 2 Reasoner employs a lightweight LLM for refined reasoning over the shortlisted candidates. It further integrates an improved reinforcement learning strategy with a dynamic task-specific reward function and a pruning mechanism to better align with human reasoning preferences. Finally, a larger LLM acts as a judger that evaluates the overall rationality of the reasoning process and its explanations to determine the final output. In addition, we construct two high-quality datasets for further validation, i.e., the biology dataset S_Bio and the physics dataset S_Phy. Experimental results demonstrate that our method consistently outperforms state-of-the-art baselines across benchmark datasets, verifying its effectiveness and superiority. Resources are available at: this https URL.

[NLP-90] AMU:Admission and Memory Update for Personalized Conversations—Structured Memory with SLM Guided Control

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在长期个性化对话中难以有效控制记忆写入的问题。现有记忆系统多聚焦于存储、检索或整合,而对记忆写入阶段的管理缺乏足够约束,导致临时性请求、重复陈述及过时用户状态可能被错误地写入记忆,进而影响个性化服务的质量。为此,本文提出一种由小语言模型(Small Language Model, SLM)引导的结构化记忆写入与更新框架——AMU(Admission and Memory Update for Personalized Conversations)。其核心在于:通过结构化的记忆过滤机制决定哪些信息应被接纳进入记忆,并利用SLM指导的记忆管理策略判断已接纳记录是否应独立存储、作为重复项丢弃,或融合为更新条目。实验结果表明,AMU能够在写入阶段实现更精准的控制,显著提升个性化记忆的清洁度与可检索性。

链接: https://arxiv.org/abs/2609.36976
作者: Tao Hwang,Yishi Diao
机构: Nanchang University(南昌大学)
类目: Computation and Language (cs.CL)
备注: 14 pages, 2 figures. Source code and implementation are available at: this https URL

点击查看摘要

Abstract:Large language models (LLMs) have become the foundation of personalized assistants, but maintaining persistent user memory across long-term interactions remains challenging. Existing memory systems often focus on storage, retrieval, or consolidation, while memory writing remains less controlled: transient requests, duplicate statements, and outdated user states may enter memory and later be retrieved for personalization. In this paper, we present AMU: Admission and Memory Update for Personalized Conversations, an SLM-guided (Small language model guided) structured framework for writing-time memory control. AMU uses structured memory filtering to decide what should enter memory and SLM-guided storage management to determine whether an admitted record should be stored separately, discarded as a duplicate, or fused as an update. We evaluate AMU in a controlled memory writing and retrieval setting. Experimental results show that AMU maintains cleaner and more retrievable personalized memories.

[NLP-91] Repetition Not Length: Isolating the Counting Failure in Neural Text-to-Speech ICASSP2027

【速读】: 该论文旨在解决文本到语音(Text-to-Speech, TTS)模型在处理重复性文本时出现的生成失效问题,特别是当文本中存在频繁重复短语时,模型容易发生循环、截断或计数错误。研究发现,问题的核心并非由文本长度引起,而是重复本身破坏了模型的生成稳定性。为验证这一假设,研究设计了对照实验:每一段重复文本均配有一个控制样本,其句子和词数完全匹配,但无任何词语连续重复。实验结果表明,六种来自三种不同架构的TTS模型在控制样本上几乎完美生成(准确率达94.3%),而在重复样本上则显著下降至仅18.2%(在k=6时)。该性能差距在多种解码策略(包括贪婪解码、重复惩罚参数扫描)、四种独立语音识别器评估及420种分析配置下均持续存在,未出现反转。此外,第四种架构的测试结果与预测差距仅差1个百分点,且两个非自回归基线模型也表现出相同失败模式。通过调整文本周期性,研究进一步揭示模型失败程度随重复周期平滑上升,即使在无相邻重复词的情况下,仍有约50%的性能损失,说明重复性对TTS模型的负面影响具有内在结构性特征。

链接: https://arxiv.org/abs/2609.36974
作者: Kirill Borodin,Vasilii Kudryavtsev,Maxim Maslov,Grach Mkrtchian
机构: 未知
类目: Computation and Language (cs.CL); Sound (cs.SD)
备注: Submitted to IEEE ICASSP 2027. Code and data: this https URL

点击查看摘要

Abstract:Text-to-speech models loop, truncate and lose count on text that repeats a phrase many times. We show that repetition itself is what breaks them, not the length that comes with it. Every repeated sentence in our test set is paired with a control of matched sentence and word count in which no word ever repeats back-to-back. Six models from three architectures render the controls almost perfectly and fail the repeated twins: 94.3% against 18.2% exactly right at k = 6. The gap survives greedy decoding, repetition-penalty sweeps, four independent speech recognisers and 420 analysis specifications without once reversing sign; a held-out fourth architecture lands within a point of its predicted gap, and one of two non-autoregressive baselines shows the same failure. Varying the period of the text shows the failure grows smoothly with periodicity, half of it surviving when no word is adjacent to itself.

[NLP-92] Chinese-Jev: Bringing System One Model to Chinese-Language Tasks

【速读】: 该论文旨在解决现有System One模型(如Jev)在中文语言场景下决策准确率偏低的问题,这一局限性严重制约了其在通用及专业领域(如医疗、法律、金融)的实际应用。其核心解决方案在于提出Chinese-Jev,一个专为中文优化的System One模型,关键创新点包括:构建统一的数据处理与训练流程,将异构的中文标注数据转化为候选选项的概率目标,实现跨领域和多题型的共享训练范式;采用轻量级编码器-only架构进行文本编码,并通过面向决策的训练策略学习候选答案评分;针对预训练分布与下游中文场景之间的偏差,采取两阶段训练策略——先在包含1000万样本的通用语料上预训练,再分别对医疗、法律、金融等专业领域进行微调。为系统评估模型在通用与专业场景下的决策准确性与校准能力,研究构建了Chinese-Jev Bench(CJ-Bench)基准测试集。实验结果表明,Chinese-Jev在通用任务上比闭源Jev模型准确率提升1.24%,推理速度提升20.3倍;在医学领域微调后,准确率较Jev提升4.0%,且在所有专业领域平均达到Jev的92%性能,同时保持17倍的速度优势和仅15毫秒/例的平均延迟。此外,研究还实现了基于INT8量化模型的移动端部署,单次决策推理延迟约为1秒,验证了其在资源受限设备上的可行性。

链接: https://arxiv.org/abs/2609.36965
作者: Zexiao Wang,Zihao Zhang,Xudong Wang,Pan Wang,Ziyi Ye,Haoyu Zhao,Zuxuan Wu,Shuicheng Yan
机构: Fudan University (复旦大学); National University of Singapore (新加坡国立大学); University of Chinese Academy of Sciences (中国科学院大学)
类目: Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
备注: 10 pages, 6 figures

点击查看摘要

Abstract:System One models such as Jev offer an efficient alternative to generative language models for tasks that require decisions rather than open-ended responses. However, existing Jev models exhibit limited Chinese-language decision accuracy, restricting their utility in both general and specialized settings. In this paper, we introduce Chinese-Jev, a System One model that addresses this gap through a unified data processing and training pipeline. Our data processing protocol converts heterogeneous Chinese-language annotations into probability targets over candidate options, enabling a shared training formulation across domains and question formats. To enable efficient inference, Chinese-Jev adopts a lightweight encoder-only backbone for text encoding and learns to score candidate answers through decision-oriented training. To address the misalignment between the pre-training distribution and downstream Chinese-language scenarios, we first train the model on a general-purpose corpus of 10 million examples, then fine-tune it separately for the medical, legal, and financial domains. To evaluate decision accuracy and calibration in both general and domain-specific Chinese-language settings, we introduce Chinese-Jev Bench (CJ-Bench). After first-stage pre-training, Chinese-Jev exceeds the accuracy of the closed-source Jev model by 1.24% on general-domain tasks while achieving a 20.3x speedup. Subsequent domain-specific fine-tuning yields a 4.0% accuracy improvement over Jev in medicine and achieves 92% of Jev’s average accuracy across specialized domains, with a 17x speedup and an average latency of only 15 ms per example. We further demonstrate on-device deployment of an INT8-quantized model on mobile devices, achieving an inference latency of approximately 1.0 second per decision. The project is available at this https URL.

[NLP-93] VStress: Correlation-Aware Auditing and Adaptive Budget Allocation for Repeated Verifiers

【速读】: 该论文旨在解决在多方验证(verifier)协作中因重复调用导致的资源浪费与信息冗余问题,核心挑战在于如何高效识别并避免无条件信息增益的验证调用。其解决方案的关键在于提出VStress——一种可审计的重放合约(auditable replay contract),以及VStress-CA——一种相关性感知的分配策略,该策略通过估计未查询验证器在密封校准集上的条件边际信息(conditional marginal information),结合不确定性折扣、调用成本归一化,实现动态停止或弃权决策,从而确保每次调用均具有信息增量。此外,控制器在接入可信预言机前冻结决策与成本账本,并在检测到依赖性偏移时触发警报,切换至精确停止模式以保障机制边界。实验表明,在对称污染率为35%时,该方法使多数投票-5(majority-5)的平衡准确率从0.6578提升至0.7739;而在65%污染率下仍能保持稳定性能。在固定预算下的对比测试中,基于自适应分配的VStress-CA实现0.6538的平衡准确率,平均每个样本仅需3.4216次调用,且获得0.6417的RLVR得分。同时,跨模型族通道间的条件边际收益显著提升(分别为0.0126、0.0462、0.0913),证明了相关性分析可由事后警示转化为可审计的主动分配决策依据。

链接: https://arxiv.org/abs/2609.36958
作者: Miaobo Hu,Shuhao Hu,Xiaobo Guo,Xin Wang,Bokun Wang,Peng Zhang,Daren Zha,Jun Xiao
机构: University of Chinese Academy of Sciences (中国科学院大学); Institute of Information Engineering, Chinese Academy of Sciences (中国科学院信息工程研究所)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR)
备注: 27 pages, 5 figures

点击查看摘要

Abstract:Repeated verifier calls are useful only when they contribute conditional information. We introduce VStress, an auditable replay contract, and VStress-CA, a correlation-aware allocation policy that estimates the conditional marginal information of an unqueried verifier on a sealed calibration split, discounts uncertainty, normalizes by call cost, and stops or abstains when the next call is not informative. The controller freezes its decision and cost ledger before joining the clean oracle; a dependence-shift alarm disables channel preference and falls back to exact-stop. The controlled audit gives the mechanism boundary: at 35% symmetric corruption, majority-5 improves balanced accuracy from 0.6578 to 0.7739, whereas at 65% it loses 0.1226 points. In the matched fixed-budget comparison, breadth, redundancy, and adaptive allocation obtain balanced accuracies 0.6048, 0.6375, and 0.6538, with 3.4216 calls per item and an RLVR score of 0.6417 for VStress-CA. Dependence diagnostics also increase from same-model repeats to cross-family channels, with conditional marginal gains of 0.0126, 0.0462, and 0.0913. These measurements turn correlation from a post-hoc warning into an auditable allocation decision.

[NLP-94] Cool the Sampler Not the Learner: Sampling Temperature Moves the Staleness Cliff of Importance-Corrected GRPO

【速读】: 该论文旨在解决生成式强化学习(Generative Reinforcement Learning, GenRL)中采样器(sampler)与学习器(learner)因更新不同步而导致的采样滞后问题,特别是当采用截断重要性权重(truncated importance weight)进行纠偏时,采样器长时间不刷新所引发的性能退化问题。其核心挑战在于:在固定更新预算下,如何延长采样器的刷新间隔而不损害训练稳定性与最终性能。论文的关键解决方案是提出“解耦冷却”(Decoupled Cooling)机制——即在保持学习器、参考模型及重要性权重温度为1的同时,将采样温度降低至0.8,从而从采样端主动缓解采样分布与目标分布之间的失配,而无需改变学习目标。该方法通过记录来自温控分布的行为概率,确保学习目标不变,同时显著提升采样器对延迟的容忍度。实验表明,在Qwen2.5-Math-1.5B和GSM8K上,解耦冷却使采样器可安全维持每192步刷新一次,且性能稳定;而在更高复杂度的Qwen2.5-Math-7B上,冷却有效延缓甚至避免了因刷新间隔过长导致的严重退化。研究进一步揭示采样温度与刷新间隔需协同设计,存在最优控制窗口,过度冷却或缺乏重要性修正均会导致失败。

链接: https://arxiv.org/abs/2609.36953
作者: Taiheng Pan
机构: The University of Melbourne(墨尔本大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 14 pages, 8 figures, 4 tables

点击查看摘要

Abstract:Production RL for language models lets the sampler fall behind the learner and repairs the resulting mismatch with a truncated importance weight. We ask how long the sampler can go without a refresh under that correction, and find a cliff: on Qwen2.5-Math-1.5B and GSM8K, importance-corrected GRPO refreshed every 192 updates learns well for 180 steps and then degrades severely in all three data seeds before the refresh arrives. Published remedies for staleness act on the update; we act on the sampler instead. Decoupled cooling draws samples at temperature 0.8 while the learner, the reference model and the importance weights stay at temperature 1, with the behaviour probability recorded from the tempered distribution, so the learner’s objective is unchanged. All corresponding cooled runs are stable, and the longer interval keeps what the short one delivered: at the same update budget, a cooled sampler refreshed every 192 steps matches an uncooled sampler refreshed every 96 at the end of training (0.857 for both) and averaged over it (0.79), whereas lowering the learning rate to a safe value ends 3-7 points lower. On Qwen2.5-Math-7B the degradation points at interval 192 predict that an interval of 144 is fatal without cooling and survivable with it; on two data seeds the uncooled runs degrade before their first refresh and the cooled runs pass it and end at 92-93% against 68-81%, with one cooled run degrading transiently late in the second cycle. The benefit has a window: at three times the safe interval and in a high-mismatch MATH setting cooling delays degradation without preventing it, stronger cooling is not better, and cooling without the correction collapses. Sampling temperature is a control on staleness tolerance, and temperature and refresh interval should be chosen together.

[NLP-95] ER-JEPA: Experience Replay Improves Joint-Embedding Predictive Learning in Language Models

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在生成过程中可能习得不良抽象语义且缺乏全面感知能力的问题。尽管已有方法如LLM-JEPA通过联合嵌入预测架构(Joint-Embedding Predictive Architecture, JEPA)实现了不同视图间对同一潜在知识的一致性对齐,但强对齐并不必然带来准确、稳定的预测性能。为此,本文提出ER-JEPA,其核心创新在于在LLM-JEPA基础上引入记忆回放路径(Episodic Replay, ER),通过将训练样本对存储于外部记忆中,并在每一步动态地存取相关数据,为当前批次的训练提供额外监督信号。该机制使模型能够同时从当前输入和历史存储样本中学习,从而增强对令牌预测与表征对齐的监督强度。实验在多个基准数据集(NL-RX、GSM8K、Spider 和 NQ-Open)上的结果表明,ER-JEPA在各项任务中均显著优于LLM-JEPA,验证了该方案的有效性。

链接: https://arxiv.org/abs/2609.36952
作者: Jingnan Pu,Zi-En Fan,Feng Lian
机构: Xi’an Jiaotong University (西安交通大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 20 pages, 15 figures, 6 tables

点击查看摘要

Abstract:Large language models (LLMs) excel at token-level generation but may learn undesirable abstract semantics and lack comprehensive perception. LLM-JEPA mitigates this by aligning different views of the same underlying knowledge via a joint-embedding predictive architecture (JEPA). However, strong alignment does not necessarily lead to accurate, stable predictions. To address this, we propose ER-JEPA, which adds an episodic replay path to LLM-JEPA. ER-JEPA stores training pairs in a memory. At each step, it stores and retrieves relevant data to provide additional supervision. This enables learning from both the current batch and stored training pairs, providing additional supervision for token prediction and representation alignment. Experiments across multiple datasets (NL-RX, GSM8K, Spider, and NQ-Open) demonstrate that ER-JEPA consistently outperforms LLM-JEPA.

[NLP-96] CoEM: Empowering Long-Context Reasoning with Commit-on-Evidence Memory

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在处理长上下文任务时,随着上下文长度增加导致推理性能下降的问题。现有方法通过分块处理输入并维持有限的文本记忆来缓解这一问题,但其早期对信息进行压缩往往会导致关键细节丢失。本文提出一种名为“基于证据承诺的记忆机制”(Commit-on-Evidence Memory, CoEM)的解决方案,其核心在于学习在何时将原始证据转化为紧凑的记忆事实。在固定上下文记忆预算下,CoEM 将潜在有用的源文本片段以未压缩形式保留在待处理集合中,允许后续上下文澄清其相关性后再决定是否进行不可逆的压缩。当新上下文到达时,一个可学习的策略会重新评估每个待处理片段,决定将其提升至已承诺记忆、继续保留或丢弃;同时,通过冻结的验证器确保仅当保留的片段与当前上下文共同支持时,所提议的事实才被接受。为优化记忆管理策略,采用强化学习训练该决策政策,结合细粒度的步骤级证据奖励与最终答案奖励。大量实验表明,CoEM 在长上下文推理任务中显著优于现有基线,尤其在 6,400 份文档的长输入场景下,相较于最强的记忆基线,在 Qwen3.5-9B 模型上提升了 10.4–11.4 的 F1 分数。

链接: https://arxiv.org/abs/2609.36935
作者: Jingguang Li,Yebo Wu,Zuyi Guo,Kailang Ma,Xianjie Dai,Han Zheng,Benwang Chen,Li Li,Can Rong,Heye Huang
机构: Korea Advanced Institute of Science and Technology(韩国科学技术院); University of Macau(澳门大学); Massachusetts Institute of Technology(麻省理工学院)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 38 pages, 13 figures. Code repository: this https URL

点击查看摘要

Abstract:Long-context reasoning is essential for complex and long-horizon tasks, yet the performance of large language models (LLMs) degrades as context length increases. Recent approaches address this by processing input chunk by chunk while maintaining a bounded textual memory in model context. However, premature information compression can discard critical details essential for subsequent reasoning. In this paper, we introduce Commit-on-Evidence Memory (CoEM), which learns when to convert source evidence into compact memory facts. Specifically, under a fixed context-memory budget, CoEM preserves potentially useful source excerpts verbatim in a pending set, allowing subsequent context to clarify their relevance before irreversible compression. As new context arrives, a learned policy revisits each pending excerpt and decides whether to promote it to the committed memory, retain it for further consideration, or discard it. A frozen verifier ensures proposed facts are accepted only if supported by retained excerpts and current context. To further guide effective memory management, we train this policy using reinforcement learning by combining fine-grained, step-level evidence rewards with final answer rewards. Extensive experiments demonstrate that CoEM consistently improves long-context reasoning. When evaluated on 6,400 documents long-context input, CoEM outperforms the strongest memory baseline by 10.4-11.4 F1 points on Qwen3.5-9B. Code repository: this https URL.

[NLP-97] Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation AACL2026

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在科学评估中缺乏可复现性的问题,尤其关注模型输出随硬件、批处理方式及日期变化而波动的现象。研究发现一个被长期忽视的关键因素:系统提示(system prompts)中隐含地注入了当前日期信息,这一变量由系统自动添加且用户无法控制,每日动态变化,从而导致模型输出的显著偏差。在9个近期主流LLM与6个不同数据集(涵盖多项选择题问答、数学推理、代码生成和机器翻译)上的实验表明,仅因日期变化,模型性能即出现显著波动,最大差异达MCQA任务中6%、数学推理中14%、代码生成中7%,以及机器翻译中2.84 BLEU。此外,模型排名亦随之改变,影响权威排行榜的公正性。值得注意的是,该“日期效应”远超其他已知非确定性来源(如批大小和数值精度),且传统增强鲁棒性的提示技术——链式思维(Chain-of-Thought)与少样本提示(Few-shot Prompting)不仅未能缓解此问题,链式思维甚至加剧了其敏感性。因此,论文的核心解决方案在于强调建立严格、标准化的评估协议,以消除日期等外部变量干扰,确保大语言模型研究结果的可复现性与公平比较。

链接: https://arxiv.org/abs/2609.36931
作者: Mario Sanz-Guerrero,Minh Duc Bui,Manuel Mager,Katharina von der Wense
机构: Johannes Gutenberg University Mainz, Germany(美因茨约翰内斯古腾堡大学, 德国); Universidad Iberoamericana, Mexico(伊比利亚美洲大学, 墨西哥); University of Colorado Boulder, USA(科罗拉多大学博尔德分校, 美国)
类目: Computation and Language (cs.CL)
备注: Accepted to AACL 2026 (Main)

点击查看摘要

Abstract:Reproducibility is essential for scientific research, yet prior work shows that LLM outputs vary with hardware and batching. We identify an overlooked factor: the hidden injection of the current date into system prompts, which users cannot control and which changes every day. Across 9 recent LLMs and 6 datasets spanning multiple-choice QA (MCQA), math reasoning, code generation, and machine translation, performance varies solely with the current date, with deltas of up to 6% on MCQA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on machine translation. Model rankings also shift, affecting leaderboards. This date effect exceeds other sources of non-determinism, such as batch size and numerical precision. Standard prompting techniques – chain-of-thought and few-shot prompting – do not reduce the sensitivity; chain-of-thought even amplifies it. Our findings underscore the need for careful evaluation protocols to ensure reproducibility and fair comparisons in LLM research.

[NLP-98] Benchmarking Automatic Speech Recognition Tools for Iberian Languages

【速读】: 该论文旨在解决伊比利亚语言(Iberian languages)自动语音识别(ASR)系统在综合评估中的不足问题,尤其关注低资源语言、模型偏差以及准确率与效率之间的权衡。其关键解决方案在于构建了一个涵盖五种伊比利亚语言(巴斯克语、加泰罗尼亚语、加利西亚语、葡萄牙语、西班牙语)及德语和土耳其语作为对照的综合性基准测试体系,采用85小时的多场景数据集(包括朗读语音、广播媒体和有声书),通过词错误率(WER)和实时因子(RTF/RTFx)双重指标系统评估十种开源模型与一种商业API的性能。研究发现,不存在单一最优模型,各系统在准确率、效率与语言覆盖范围之间存在显著权衡;低资源语言如巴斯克语表现明显下降,凸显训练数据覆盖的重要性;同时,多数系统中均存在性别差异,揭示了多语言ASR中的公平性挑战。该基准为实际应用中的模型选型提供了可操作的指导依据。

链接: https://arxiv.org/abs/2609.36920
作者: Fernando López,Pablo Gómez,David Solans,Paulo Villegas,Jordi Luque
机构: 未知
类目: Computation and Language (cs.CL)
备注: Accepted in IberSPEECH 2026

点击查看摘要

Abstract:Comprehensive evaluations of automatic speech recognition (ASR) for Iberian languages remain limited, and low-resource languages, biases, and efficiency trade-offs are underexplored. We benchmark eleven systems, ten open-weight models and one commercial API, across five Iberian languages (Basque, Catalan, Galician, Portuguese, Spanish), with German and Turkish as controls. Evaluation uses an 85-hour dataset covering read speech, broadcast media, and audiobooks, assessing accuracy and efficiency via word error rate (WER) and real-time factors (RTF/RTFx). Results show no single model dominates: accuracy, efficiency, and language coverage present clear trade-offs. Low-resource languages, especially Basque, degrade significantly, highlighting the role of training coverage. We observe consistent sex disparities across most systems, highlighting fairness challenges in multilingual ASR. Overall, the benchmark provides practical guidance for real-world model selection.

[NLP-99] Can Language Models Learn to Forecast Stock Prices

【速读】: 该论文旨在解决生成式AI在金融市场预测任务中表现有限的问题,尤其关注后训练(post-training)方法是否能有效提升模型在具有高噪声和不确定性特征的金融市场中的预测能力。与数学推理或软件工程等具有明确可验证结果的任务不同,金融市场的实际收益具有高度随机性,且何种信息对预测有效也缺乏先验明确性——模型需自主决定采集哪些观测数据,并在结果揭晓前做出数值判断。为此,研究构建了一个基于历史股价的时间序列沙盒环境,让语言模型(Qwen3-4B)通过整合价格、成交量、相对表现及市场背景等多源证据进行未来收益率预测。其解决方案的关键在于采用两阶段后训练框架:首先通过监督微调(SFT)增强模型对工具使用的理解与执行能力,随后利用近端策略优化(PPO)算法,以预测得分与实际回报之间的匹配度作为终端奖励信号进行强化学习。实验结果表明,经过该流程训练的AURA-4B模型在方向-幅度评分上从20.94提升至43.31(翻倍以上),条件幅度一致性从33.3%升至66.2%,方向准确性也从62.9%增至65.4%,性能达到前沿水平。同时,SFT显著扩展了工具使用范围,而PPO进一步提升了模型对排序操作和市场上下文查询的依赖程度。这说明后训练不仅能显著提升金融预测性能,还能促使模型主动调整其信息搜集策略,展现出更智能的市场分析行为。

链接: https://arxiv.org/abs/2609.36914
作者: Jiacheng Guo,Suozhi Huang,Shuzhen Li,Yunlong Gao,Zerui Cheng,Jason Ge,Shushu Liang,Zihao Li,Hao Lu,Ming Yin,Shilong Liu,Jiashuo Liu,Xu Kuang,Mengdi Wang
机构: Princeton University (普林斯顿大学); InclusionAI; Harvard University (哈佛大学); Stanford University (斯坦福大学)
类目: Computation and Language (cs.CL)
备注: 18 pages, 4 figures

点击查看摘要

Abstract:Post-training has been shown to significantly improve language models’ performance on tasks with verifiable outcomes, including mathematical reasoning, software engineering, and computer use. However, whether the same approach can improve forecasting in financial markets is much less clear. Compared with tasks with verifiable outcomes, not only are realized returns noisy, but even what constitutes a relevant information set for making effective predictions is not obvious a priori: the model must decide which observations to gather and then commit to a numerical judgment before the outcome is known. We study this question in a chronological stock-price sandbox, where a language model gathers price, volume, relative-performance, and market-context evidence and predicts a future return. We post-train Qwen3-4B with supervised fine-tuning (SFT) on tool-use demonstrations, then proximal policy optimization (PPO) with a terminal reward given by the forecast score against the realized return. The resulting AURA-4B more than doubles the starting direction–magnitude score, from 20.94 to 43.31, and is comparable to frontier language models on this benchmark. Conditional magnitude agreement rises from 33.3 to 66.2, while directional accuracy changes from 62.9 to 65.4. SFT expands tool use, and PPO further increases the share of ranking and market-context queries. These results show that post-training can substantially improve financial forecasting performance, together with changes in how the model investigates the market, on this outcome-selected benchmark.

[NLP-100] BaLEEN: Biasing with Latent Encoded Entities for Context-Aware ASR

【速读】: 该论文旨在解决自动语音识别(ASR)中领域特定实体和罕见专有名词识别困难的问题。现有方法在处理此类词汇时往往表现不佳,主要受限于模型对未见或低频词汇的泛化能力不足。其解决方案的关键在于提出一种轻量级、基于超网络(hypernetwork)的动态上下文适应框架——BaLEEN(Biasing with Latent Encoded Entities)。该框架通过预训练语言模型编码可变长度的上下文关键词,利用Perceiver瓶颈将上下文信息压缩为固定长度的潜在向量,并将上下文相关的偏置向量直接注入到ASR模型的中间编码表示中,从而实现对特定上下文的精准建模。由于语言模型与主干ASR模型在整个训练过程中保持冻结,BaLEEN作为即插即用的适配器,在推理阶段仅需预先计算上下文偏置,不引入额外计算开销。实验结果表明,相较于无偏置基线,BaLEEN在测试集上将关键词漏识率降低8.7%,同时整体词错误率(WER)和字符错误率(CER)分别提升21%和28%,显著提升了ASR系统在包含稀有命名实体场景下的性能。

链接: https://arxiv.org/abs/2609.36913
作者: Chihiro Taguchi,Yotaro Kubo,Rujikorn Charakorn
机构: Sakana AI; Yotaro Kubo†; Rujikorn Charakorn†
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 5 pages, 2 figures, 2 tables

点击查看摘要

Abstract:Transcribing domain-specific entities and rare proper nouns remains a major challenge in automatic speech recognition (ASR). In this paper, we propose BaLEEN (Biasing with Latent Encoded Entities), a lightweight, hypernetwork-based framework for dynamic contextual adaptation without fine-tuning the underlying ASR model. BaLEEN encodes variable-length contextual keywords using a pretrained language model, compresses them into a fixed sequence of latent vectors via a Perceiver bottleneck, and injects context-dependent bias vectors directly into the intermediate encoder representations of the ASR model. Because both the language model and the backbone ASR model remain entirely frozen during training, BaLEEN operates as a plug-and-play adapter that incurs zero computational overhead at inference time when context biases are precomputed. We evaluate our method on a CTC-based ASR model using a Wikipedia-derived corpus with annotated named entities and synthetic speech. Experimental results demonstrate that BaLEEN reduces keyword miss rate by 8.7% on the test set relative to the unbiased baseline while simultaneously improving overall word error rate by 21% and character error rate by 28%.

[NLP-101] MultiTalk: Scaling Full-Duplex Speech Models to Long Multi-Party Bilingual Conversation NEURIPS2026

【速读】: 该论文旨在解决当前端到端全双工语音模型在长时程鲁棒性与多参与者交互能力方面的双重局限,尤其针对会议、小组授课及社交机器人接待等真实场景中,单个模型需持续追踪、理解并响应多位发言者长时间对话的需求。现有系统受限于数据稀缺与评估标准缺失:公开的多参与者语音语料规模小且不适用于编码帧级全双工建模,而现有的长音频基准多聚焦于被动听觉任务,语音转语音任务则普遍为短时、二人对话。为此,本文在英文与中文双语背景下,沿长时程与多参与者两个维度扩展Moshi范式:首先,构建并发布57.6小时合成训练数据集(MultiTalkPT 与 MultiTalkFT),支持可控长度、参与人数、发言交替、重叠说话、回应信号、打断、指代转移及长距离指代等复杂交互特征;其次,提出基于真实人类录音的评估基准MultiTalkBench,用于评测长时程、多参与者、双语全双工对话能力,其平均对话时长达32.6分钟,并包含对长程实体追踪、话题连贯性与接收者选择等关键能力的探测任务;最后,训练了一个双语Moshi风格模型,在长期多参与者跨语言对话中展现出显著优于开源基线(如Moshi、MiniCPM-o-4.5、Qwen3-Omni-30B-A3B-Instruct)的连贯性表现。其核心解决方案在于通过大规模可控合成数据增强训练能力,并建立真实场景驱动的多维度评估体系,从而推动全双工语音模型向更接近人类交互的水平演进。

链接: https://arxiv.org/abs/2609.36903
作者: Ke Wang,Houxing Ren,Zimu Lu,Yunqiao Yang,Zhuofan Zong,Mingjie Zhan,Hongsheng Li
机构: CUHK MMLab(香港中文大学多媒体实验室); CPII under InnoHK(创新香港研发中心的CPII)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Sound (cs.SD)
备注: NeurIPS 2026

点击查看摘要

Abstract:End-to-end full-duplex speech models have brought open-source machine conversation closer to human-like interaction, yet existing systems remain limited in two intertwined dimensions: long-context robustness and multi-party interaction. Real-world scenarios such as meetings, group lessons, and social-robot reception require a single model to track, contextualize, and respond to multiple speakers over extended durations. Progress is constrained by both data and evaluation: open multi-party speech corpora remain small and are not designed for codec-frame-level full-duplex modeling, while existing long-audio benchmarks focus on passive listening and speech-to-speech benchmarks are mostly short and dyadic. We extend the Moshi paradigm jointly along the long-horizon and multi-party axes in English and Chinese. First, we release 57.6k hours of synthetic training data ( \hrefthis https URLMultiTalkPT and \hrefthis https URLMultiTalkFT ) for long-form, multi-party, English-Chinese full-duplex dialogue, with controllable length, participant count, turn-taking, overlap, backchannels, interruptions, addressee shifts, and long-range coreference. Second, we introduce \hrefthis https URLMultiTalkBench , built from real human recordings, for evaluating long-form, multi-party, bilingual full-duplex dialogue. Conversations average 32.6 minutes and include probes for long-range entity tracking, topic coherence, and addressee selection. Third, we train a bilingual Moshi-style model that sustains coherent multi-party English-Chinese conversations over extended durations and substantially outperforms open-source baselines including Moshi, MiniCPM-o-4.5, and Qwen3-Omni-30B-A3B-Instruct on MultiTalkBench.

[NLP-102] RAEGNet: Relation-Aware Evidence Graph Network for Harm-Aware Multimodal Fake News Detection

【速读】: 该论文旨在解决现有多模态假新闻检测方法中存在的两个关键问题:一是依赖实体级检索引入与事件无关的噪声,二是忽视不同假新闻实例在危害程度上的差异。针对上述问题,论文提出事件级证据检索框架(Event-Level Evidence Retrieval Framework, ELERF),通过基于新闻条目完整事件语义进行外部证据检索,以提升证据的相关性;同时设计关系感知的证据图网络(Relation-Aware Evidence Graph Network, RAEGNet),构建包含新闻-证据立场关系与证据-证据交互关系的有向图结构,并引入条件危害分支,联合建模内容真实性与潜在危害程度。其核心创新在于实现了更精准的事件级证据获取与对假新闻危害性的差异化建模,从而显著提升了检测性能。

链接: https://arxiv.org/abs/2609.36902
作者: Wenbin Shen,Guoxuan Qin,Guangxu Yao,Baodong Wang,Yuanbo Rui,Zhongjie Ba,Zhichao Lian
机构: Nanjing University of Science and Technology(南京理工大学); Zhejiang University(浙江大学); University of Chinese Academy of Sciences(中国科学院大学)
类目: Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
备注:

点击查看摘要

Abstract:Existing multimodal fake news detection methods often introduce external information to assist detection. However, most of them rely on entity-level retrieval and are therefore prone to introducing event-irrelevant noise. Meanwhile, existing methods mainly focus on improving overall performance and do not account for differences in the degree of harm posed by different instances of fake news. To address these limitations, we design an Event-Level Evidence Retrieval Framework (ELERF) and propose a Relation-Aware Evidence Graph Network (RAEGNet). ELERF retrieves external evidence based on the complete event semantics of a news item. RAEGNet constructs a directed graph that incorporates news-evidence stance relations and evidence-evidence interaction relations, and introduces a conditional-harm branch to jointly model authenticity and potential harm. Experimental results demonstrate that RAEGNet outperforms multiple baseline methods across all evaluated metrics on Weibo-21, Fakeddit, and our self-constructed SSS dataset.

[NLP-103] Momentum-Coupled Rubric Adaptation for Detailed Image Captioning

【速读】: 该论文旨在解决细粒度图像描述生成中描述质量不一致的问题,即在事实准确性、信息覆盖率和表达清晰度等方面难以兼顾。现有基于评分标准的强化学习方法通常采用独立模型分别执行描述生成、评分标准构建与评判任务,导致各角色间理解不一致;部分动态评分标准方法虽交替更新生成策略与评分标准生成器,但固定评判模型仍可能导致评分标准与策略优化不同步。其解决方案的关键在于提出一种两阶段协同框架——MoCo Rubric,通过共享参数的多任务监督微调,使单一视觉-语言模型同时承担描述策略(Caption Policy)、评分标准生成器(Rubric Generator)和评分判别器(Rubric Judge)的角色。在第二阶段,生成器基于当前策略采样的描述、参考描述及图像证据在线构建评分标准,判别器则依据该标准提供基于评分的奖励信号,仅对策略进行基于广义近端策略优化(GRPO)的更新。为避免因策略更新引起的评分标准与判别能力失配,采用策略参数的指数移动平均构建一个共享的动量模型,由生成器与判别器共同使用,实现评分标准与判别能力随策略渐进同步更新,从而在保持稳定性的同时提升评估一致性。实验表明,该方法在五个基准数据集上实现了72.83%的平均成对胜率、盲评中的最优均值排名以及基于描述的问答任务最高平均得分。

链接: https://arxiv.org/abs/2609.36893
作者: Zhenwen Ji,Lei Jin,Shanyong Wang,Jiaming Lu,Chengqiang Lu,Yi Wu,Yao Hu,Lizhen Cui,Yanyu Xu
机构: Xiaohongshu Inc.; the Joint SDU-NTU Centre for Artificial Intelligence Research (C-FAIR), Shandong University
类目: Computation and Language (cs.CL)
备注: 28 pages, natural language processing, computer vision

点击查看摘要

Abstract:Detailed image captioning requires accurate and comprehensive descriptions of fine-grained visual content, yet caption quality spans factual accuracy, information coverage, and clarity. Compared with conventional methods that rely mainly on high-quality supervision or holistic rewards, rubric-based reinforcement learning decomposes these requirements into explicit criteria and provides targeted, structured feedback. However, existing methods often use separate models for caption generation, rubric construction, and judging, which may lead to inconsistent interpretations across roles. Some dynamic rubric methods alternate updates between the caption policy and rubric generator while keeping the judge fixed, but staged optimization may still leave rubric construction and judging out of step with policy optimization. We propose MoCo Rubric, a two-stage framework that coordinates these roles. First, role-conditioned, shared-parameter multi-task supervised fine-tuning equips a single vision–language model to serve as the Caption Policy, Rubric Generator, and Rubric Judge. Then, the Generator constructs rubrics online from captions sampled by the current Policy, reference captions, and image evidence. The Judge provides rubric-based rewards, and only the Policy receives GRPO updates. As Policy updates change the candidates being evaluated, we use an exponential moving average of the Policy parameters to update one momentum model shared by the Generator and Judge. This gradual transfer lets both rubric roles track Policy updates without separate RL optimization while smoothing parameter changes that could disrupt their rubric capabilities under direct synchronization. Across five captioning benchmarks, MoCo Rubric achieves an average pairwise win rate of 72.83%, the best mean rank in blind ranking, and the highest average score in caption-based question answering.

[NLP-104] Harness Evolution as Learning: Approximation Generalization and Optimization Limits of Self-Improving Personal Agents

【速读】: 该论文旨在解决如何将大语言模型(Large Language Models, LLMs)的潜在能力有效转化为个性化代理(Personal Agents)在日常场景中持续适应用户偏好并执行有用行为的问题。其核心挑战在于,在底层模型固定的前提下,如何通过“控制框架”(harness)的设计与演化来实现对用户意图和上下文的精准响应。解决方案的关键在于系统性地探究控制框架的架构设计、规模大小以及自进化算法三方面的相互作用,并通过实证与理论相结合的方法揭示其内在机制。具体而言,研究构建了一个以偏好为导向的基准测试,系统评估了不同控制框架维度下的性能边界;同时,将控制框架的演化建模为一个学习问题,从近似误差、泛化误差和优化误差的角度解释现象,进一步分析了在有限交互证据下策略可达性与更新动态偏差等问题。这些发现共同提供了一个统一的理论框架,阐明了通过控制框架演化实现个性化的局限性,为未来更高效、可扩展的个性化代理设计提供了关键指导。

链接: https://arxiv.org/abs/2609.36892
作者: Zeyu Gan,Zixuan Gong,Yong Liu
机构: Gaoling School of Artificial Intelligence (高瓴人工智能学院); Renmin University of China (中国人民大学); Beijing, China
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:As the capabilities of large language models (LLMs) continue to advance, increasing attention is turning to how to translate their abilities into useful behavior. Personal agents bring this question into everyday settings, where models are expected to serve individual users and continually adapt to their preferences. With the underlying model held fixed, such adaptation relies on harness engineering: designing and evolving the surrounding layer that manages context, memory, tools, and execution. Despite rapid progress, the factors governing effective harness evolution remain insufficiently understood. To narrow this gap, we investigate three central questions concerning harness architecture, harness scale, and self-evolution algorithms through complementary empirical and theoretical analyses. Empirically, we introduce a preference-oriented benchmark and systematically characterize the capabilities and limitations of personal agents associated with these three dimensions. Theoretically, we formulate harness evolution as a learning problem and explain these phenomena through approximation, generalization, and optimization errors. Analyses of reachable policies, capacity under finite interaction evidence, and biased update dynamics provide theoretical accounts of the observed phenomena. Together, these results offer a unified perspective on the limits of personalization through harness evolution and inform future harness design.

[NLP-105] Rethinking Multimodal Fake News Detection in the Generative AI Era

【速读】: 该论文旨在解决生成式内容(Generative Content)在新闻生产与传播中日益普及背景下,现有多模态假新闻检测方法难以有效区分真实与虚假信息、且缺乏对生成性差异如何影响证据可靠性的系统性刻画这一关键问题。传统研究或侧重于新闻真伪判断,忽视生成内容的特性;或专注于生成内容识别(AIGC Detection),但无法直接验证新闻事件的真实性,导致任务割裂。其解决方案的关键在于构建Weibo26数据集,该数据集专门针对生成内容场景设计,涵盖真实与生成内容混合的多模态假新闻样本;并提出生成感知层次化推理(Generativity-Aware Hierarchical Reasoning, GAHR)框架,通过融合全局判断与局部修正机制,使生成性信息能够参与新闻真伪推理过程,从而实现对生成内容的有效识别与真伪评估的协同优化。实验表明,GAHR在多个基准数据集及Weibo26上均展现出优异的真伪检测性能与生成内容识别能力。

链接: https://arxiv.org/abs/2609.36850
作者: Wenbin Shen,Guoxuan Qin,Guangxu Yao,Baodong Wang,Yuanbo Rui,Zhichao Lian
机构: Nanjing University of Science and Technology (南京理工大学); University of Chinese Academy of Sciences (中国科学院大学)
类目: Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
备注:

点击查看摘要

Abstract:Generative content is increasingly entering the production and dissemination of news, transforming fake news from manually fabricated or simply manipulated material into complex forms in which native and generated content jointly participate. Existing multimodal fake news detection research primarily focuses on veracity assessment and rarely characterizes how generativity differences affect the reliability of evidence. In contrast, AIGC detection primarily determines whether content is generated or modified by generative models, but it does not by itself establish whether the underlying news event is true. To bridge the separation between these tasks in data and evaluation, we construct Weibo26, a multimodal fake news detection dataset for generative-content scenarios. On this basis, we propose the Generativity-Aware Hierarchical Reasoning (GAHR) framework, which combines global judgment with local correction so that generativity information participates in news-veracity reasoning. Experiments on multiple existing fake news detection benchmarks and Weibo26 show that GAHR achieves competitive veracity-detection performance while effectively identifying generative content.

[NLP-106] On-Policy Visual Evidence Distillation

【速读】: 该论文旨在解决视觉代理(Visual Agent)在执行任务时,由于图像操作导致的视觉证据获取、阅读及答案定位等环节中的局部错误在交互轨迹中逐级传播,进而引发最终推理失败的问题。现有基于策略的蒸馏(On-Policy Distillation, OPD)方法虽能通过强教师模型提供指导,但其多依赖对原始图像构建或对比辅助视图来增强监督信号,未能显式建模学生动作、观测结果与后续推理之间的因果关联,因而难以针对不同故障阶段(Acquire、Read、Ground)提供精准修正。为此,论文提出一种名为“视觉证据反思”(Reflection on Visual Evidence, ReVuE)的新颖OPD方法:ReVuE通过比较同一查询下多个学生生成的轨迹,归纳并诊断其中首次出现的失败阶段,生成具有任务上下文意义的“反思”信息,作为教师模型在训练时的参考依据;随后,根据这些反思对教师预测的影响强度,对分词级别(token-level)的蒸馏损失进行分组与重加权,将轨迹级别的证据诊断转化为针对性的分词级监督信号。该设计实现了从全局轨迹分析到细粒度行为纠正的映射,有效引导学生改进视觉证据获取与推理能力。在涵盖Qwen2.5-VL与InternVL3.5两大模型系列共11个基准测试中,ReVuE在感知、数学推理和通用任务的加权平均得分上全面超越所有对比的OPD基线,并显著降低推理与工具调用冗余,提升工具调用准确率与任务整体准确率。

链接: https://arxiv.org/abs/2609.36838
作者: Shaohang Wei,Feifan Song,Guangyue Peng,Wenhao Yu,Wei Li,Wen Luo,Yang Xu,Yufan Shen,Luke Mao,Yang Du,Asher Qin,Houfeng Wang
机构: Peking University (北京大学); Tencent (腾讯)
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 44 pages, including appendices. Project page: this https URL . Code: this https URL

点击查看摘要

Abstract:Visual agents solve problems by interleaving reasoning with image operations, and on-policy distillation (OPD) provides guidance from a strong teacher on student-generated interaction trajectories. However, image operations change the evidence available for subsequent reasoning, so local errors in evidence acquisition (Acquire), reading (Read), or answer grounding (Ground) can propagate through the trajectory and lead to incorrect answers. Existing multimodal OPD methods primarily construct or contrast auxiliary views of the original image to strengthen supervision, without explicitly modeling the connections between student actions, resulting observations, and subsequent reasoning. This limits their ability to provide corrections tailored to different failure stages. We introduce Reflection on Visual Evidence (ReVuE), an on-policy distillation method for visual agents. ReVuE compares multiple student-generated trajectories for the same query, summarizes the observed visual evidence, and diagnoses the first failure across the Acquire, Read, and Ground stages. The resulting reflections provide training-time context for the teacher. We group and reweight token-level distillation losses according to how strongly these reflections affect the teacher’s predictions. This design translates trajectory-level evidence diagnosis into targeted token-level supervision, guiding students to improve their visual evidence acquisition and reasoning. Across 11 benchmarks spanning the Qwen2.5-VL and InternVL3.5 model families, ReVuE outperforms all evaluated OPD baselines in weighted-average scores for perception, mathematical reasoning, and general tasks. ReVuE also reduces redundancy in reasoning and tool calls while improving tool-call accuracy and task accuracy. Code is available at this https URL

[NLP-107] CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning

【速读】: 该论文旨在解决生成式语言模型在多奖励(multi-reward)场景下,使用组相对策略优化(Group Relative Policy Optimization, GRPO)时因奖励尺度差异导致的归一化偏差问题。具体而言,传统GRPO通过将同一提示下的多个采样轨迹的奖励进行中心化与标准化,以计算优势值,但其对总奖励的方差估计依赖于各奖励分量的协方差之和,当存在尺度差异显著的奖励时,大尺度奖励会主导协方差贡献,从而抑制小尺度奖励的信号,导致优势估计失真。为此,本文提出相关性归一化组相对策略优化(Correlation-Normalized GRPO, CorrGRPO),其核心在于将原始协方差转换为皮尔逊相关系数(Pearson correlation coefficient),实现对不同尺度奖励的归一化平衡。该方法在保持总奖励中心化不变的前提下,使优势值的计算仅反映奖励间的相关性结构,而非其绝对尺度,从而有效避免大尺度奖励对归一化过程的主导作用。实验在代码生成、工具调用及代理安全等任务上验证了该方法的有效性,涵盖0.5B至8B参数规模的模型,结果表明在三类任务中均实现了性能提升,证明了该方案在处理多目标优化中的鲁棒性与优越性。

链接: https://arxiv.org/abs/2609.36820
作者: Wenbin Hu,Huihao Jing,Haochen Shi,Yuxuan Liu,Haoran Li,Yangqiu Song
机构: Hong Kong University of Science and Technology(香港科技大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Group Relative Policy Optimization (GRPO) is widely used to train reasoning language models, where it computes advantages by centering and normalizing rewards across rollouts of the same prompt. For multiple rewards, GRPO sums the reward components and normalizes the total reward by its within-group standard deviation. The corresponding variance equals the sum of all pairwise reward covariances. For a fixed centered reward, larger aggregate covariance produces smaller advantages, and vice versa, allowing update magnitudes to adapt to reward dependence. However, correlated rewards with large scales can dominate this normalization and suppress signals from smaller-scale rewards. We propose Correlation-Normalized GRPO (CorrGRPO), which normalizes pairwise covariances into Pearson correlation coefficients. CorrGRPO keeps the centered total reward unchanged while balancing the influence of differently scaled rewards on the correlation-based normalization. This allows advantage magnitudes to adapt to reward correlations without the normalization being dominated by large-scale reward components. We compare CorrGRPO with GRPO and other variants on code generation, tool calling, and agent security, using models ranging from 0.5B to 8B parameters. These tasks all involve multiple rewards that can improve together or present tradeoffs. Results show improvements across three domains, including code generation, tool calling, and agent security. Our code is available at this https URL.

[NLP-108] VAA-CSEC: Vote-guided Advantage Allocation for Chinese Semantic Error Correction

【速读】: 该论文旨在解决中文语义纠错(Chinese Semantic Error Correction, CSEC)中语义错误隐蔽性强、复杂度高的问题,尤其针对现有大语言模型(LLM)方法在该任务上普遍存在的过纠正(over-correction)现象以及思维链(Chain-of-Thought, CoT)推理与自一致性解码(self-consistency decoding)之间交互不明确的问题。其解决方案的关键在于提出一种多阶段框架VAA-CSEC,通过结合CoT蒸馏、监督微调(Supervised Fine-Tuning, SFT)、强化学习(Reinforcement Learning, RL)和自一致性解码,实现更精准的纠错。其中,核心创新点在于引入面向任务的奖励函数,严格遵循CSEC的最小编辑原则,并设计组级相对策略优化(Group-Level Relative Policy Optimization, GLPO),根据个体生成结果与投票聚合组奖励之间的差距重新分配优势(advantage),从而确保强化学习训练目标与推理时的自一致性目标保持一致,有效缓解过纠正问题并提升纠错效果。实验表明,该方法在CSED-C和NaSGEC-Exam数据集上均达到新的最优性能,显著优于现有LLM基线方法。

链接: https://arxiv.org/abs/2609.36804
作者: Yitong Han,Nankai Lin,Juan Luo,Hongyan Wu,Lianxi Wang,Shengyi Jiang
机构: Guangdong University of Foreign Studies (广东外语外贸大学); Guangdong Engineering Research Center of Data Security Governance and Privacy Computing (广东省数据安全治理与隐私计算工程研究中心); National University of Defense Technology (国防科技大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Chinese Semantic Error Correction (CSEC) targets semantic errors in Chinese text, which are typically more subtle and complex than spelling and grammatical errors but remain relatively underexplored. Existing LLM-based approaches face two recurring obstacles in this task: over-correction, and unclear interaction between Chain-of-Thought (CoT) reasoning and self-consistency decoding, such that the benefits brought by CoT cannot be reliably transferred to final corrections. We propose Vote-guided Advantage Allocation for CSEC (VAA-CSEC), a multi-stage framework that combines CoT distillation, Supervised Fine-Tuning (SFT), Reinforcement Learning (RL) and self-consistency decoding. During RL, we design a task-specific reward function that directly aligned with the minimal-editing principle of CSEC. We further introduce Group-Level Relative Policy Optimization (GLPO), which reallocates GRPO advantages according to the margin between individual rollout rewards and the vote-aggregated group reward, aligning the RL training objective with the self-consistency objective used at inference time. Experiments on CSED-C and NaSGEC-Exam show that VAA-CSEC outperforms all LLM-based baselines on CSED-C with an F0.5 of 47.72%, achieves the highest recall of 42.15% among all methods, and establishes a new state of the art of 41.55% F0.5 on NaSGEC-Exam.

[NLP-109] Seeing What Should Be Heard: Diagnosing and Repairing Cross-Modal Shortcuts in Omni-Modal LLM s

【速读】: 该论文旨在解决生成式多模态大模型(omni-modal large language models, LLMs)在回答跨模态问题时存在的“跨模态捷径”(cross-modal shortcut)问题,即模型在面对与音频相关的问题时,过度依赖图像信息而非实际的音频内容,导致其未能真正遵循问题所指定的模态。这一现象源于现有训练范式中多模态输入间存在冗余证据,使得模型可利用非目标模态(如图像)中的关联信息进行错误推断,而无需真正理解目标模态(如音频)。解决方案的关键在于提出一种名为“因子化模态诊断”(Factorized Modality Diagnostic)的系统性评估方法,通过在样本间独立交换音频与图像,以分离并量化各模态对模型输出的因果贡献。基于此诊断发现,研究进一步提出DMC-Repair方法,该方法在训练过程中引入跨模态交换样本,并依据问题所指定的模态分配监督信号,从而切断模型对模态间虚假对应关系的利用。实验表明,DMC-Repair可使图像引起的答案效应占比降低59.9%,有效抑制跨模态捷径,同时保持音频问答性能,且该效果在两种模型架构、零样本迁移至未见数据集和基准测试中均具有泛化能力,并在后续后训练阶段保持稳定。

链接: https://arxiv.org/abs/2609.36798
作者: Yueran Ma,Ronghao Lin
机构: The University of Queensland (昆士兰大学); Shenzhen University (深圳大学)
类目: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 25 pages, 11 figures, 16 tables

点击查看摘要

Abstract:Omni-modal large language models (LLMs) are expected to answer a question using the modality it explicitly refers to. However, existing training paradigms rarely verify whether models actually follow this modality, because multimodal inputs from the same sample often provide redundant evidence for the same answer. In this work, we uncover a pervasive cross-modal shortcut in omni-modal LLMs: when asked an audio-related question, models rely on the image as much as on the audio, and sometimes even more. To systematically diagnose this behavior, we introduce the Factorized Modality Diagnostic, which independently swaps audio and images between samples to isolate each modality’s causal contribution. Across two model families in different settings, we find that this shortcut persists throughout supervised fine-tuning and reinforcement learning post-training, while judge-based RL may further amplify such reliance on irrelevant visual information. Based on this finding, we propose DMC-Repair, which trains models on the same kind of cross-modal swapped samples while assigning supervision according to the modality specified by the question. This prevents models from exploiting the spurious correspondence between modalities within the same clip. Experiments demonstrate that DMC-Repair reduces the image-induced share of the answer effect by 59.9%, effectively suppressing the cross-modal shortcut without compromising audio-question answering performance. The reduction in shortcut reliance generalizes across two model families and zero-shot to an unseen dataset and an unseen benchmark, and persists through subsequent post-training. Code is available at this https URL.

[NLP-110] QuantMLA: Function-Aligned Dual-Path Quantization for Low-Bit MLA KV Caching

【速读】: 该论文旨在解决多头潜在注意力(Multi-Head Latent Attention, MLA)机制中缓存内存随上下文长度和批量大小线性增长的问题,尤其聚焦于内容缓存与位置编码(RoPE)路径在低比特量化下的双路径误差耦合问题。现有方法在降低缓存位宽时易引发显著的注意力输出失真,且RoPE路径误差存在明显放大现象,限制了高效压缩能力。其解决方案的关键在于提出QuantMLA——一种面向低比特双路径量化的函数对齐框架,通过构建路径特异性的变换空间,在保持全精度计算能力的同时实现离线参数融合,彻底消除在线变换开销;该框架分别以注意力输出重建和位置相关QK重建为优化目标,有效建模并抑制内容路径的匹配与聚合误差,以及RoPE路径引起的逻辑值偏差,并提供输出失真的理论边界。实验表明,QuantMLA首次实现了在四个MLA模型族上联合INT4量化内容与RoPE缓存,且精度损失极小;进一步将内容缓存压缩至INT2而保留RoPE键缓存为INT4,仍可在复杂推理与代码生成任务中保持竞争力。此外,作者开发了原生低比特MLA注意力核,将解压与反量化操作直接嵌入注意力计算流程,结合物理缓存布局实现了128K上下文下3.59倍的压缩比,且在缓存压力服务场景中相较BF16方案达到5.168倍的整作业输出吞吐提升。

链接: https://arxiv.org/abs/2609.36760
作者: Zunhai Su,Yuxuan Sun,Jianchao Tan,Tao Zhang,Ruihan Hu,Yuchen Xie,Xunliang Cai,Ngai Wong
机构: The University of Hong Kong(香港大学); Meituan LongCat Team(美团龙猫团队); South China University of Technology(华南理工大学); Harbin Institute of Technology(哈尔滨工业大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Multi-Head Latent Attention (MLA) enables expressive multi-head attention with compact caches for its content and decoupled RoPE paths, yet cache memory still scales linearly with context length and batch size. In this work, we establish a systematic model of MLA’s dual-path quantization errors, characterizing their distinct effects on attention-output distortion and explaining the pronounced amplification of RoPE-path errors. Guided by this analysis, we introduce QuantMLA, a function-aligned framework for low-bit dual-path quantization. We derive path-specific transformation spaces that preserve full-precision computation while remaining fully fusible into model parameters offline, eliminating online transformation overhead. Within these spaces, QuantMLA learns path-specific transformations with function-aligned objectives: attention-output reconstruction captures the content path’s coupled matching and aggregation errors, while positional QK reconstruction preserves the RoPE-induced component of the attention logits and admits a theoretical bound on output distortion. Across four MLA model families, QuantMLA enables, to our knowledge, the first reported joint INT4 caching of the content and RoPE caches with minimal accuracy degradation. Further compressing the content cache to INT2 while retaining the RoPE key cache at INT4 maintains competitive performance on challenging reasoning and code benchmarks. We develop a native low-bit MLA attention kernel that integrates unpacking and dequantization directly into attention computation. The physical cache layout provides 3.59x compression at 128K context, while a cache-pressure serving workload achieves 5.168x higher whole-job output throughput than BF16. The code will be released upon acceptance.

[NLP-111] Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving

【速读】: 该论文旨在解决自奖励强化学习(Self-rewarding Reinforcement Learning, RL)中因依赖单一随机采样群体上下文(group context)而导致的奖励信号不充分与不可靠问题。现有基于集成的方法在生成奖励参考时,将响应的奖励表示与其所在群体的特定上下文强绑定,而单一上下文实现可能遗漏理想的奖励信号,从而为策略优化提供误导性或不稳定指导。其解决方案的关键在于提出群体边缘化优势估计(Group-Marginalized Advantage Estimation, GMAE),通过在所有可能的上下文组合上对奖励实现进行聚合,构建响应级别的奖励分布,并据此估计期望优势(expected advantage)。该方法有效缓解了上下文敏感性问题,提升了奖励信号的鲁棒性与泛化能力。实验在八个基准任务和四种基础模型上验证了GMAE在性能、跨领域泛化、训练稳定性及计算开销方面的优越性,展现出良好的适用性。

链接: https://arxiv.org/abs/2609.36750
作者: Yiming Wang,Yikang Liu,Qingyuan Tian,Xingyu Chen,Zhuosheng Zhang,Zhaopeng Tu,Rui Wang
机构: Shanghai Jiao Tong University (上海交通大学); Tencent(腾讯)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Self-rewarding reinforcement learning (RL) enables large language models (LLMs) to self-evolve without human labels. Existing ensemble-based methods construct reward references from rollout groups and assign rewards accordingly. However, a response’s reward representation also depends on its randomly sampled group context, i.e., the other responses in its group. Using only one group-context realization may miss desired reward signals and provide unreliable guidance for policy optimization. To address this issue, we propose Group-Marginalized Advantage Estimation (GMAE), which aggregates reward realizations across possible contexts into a response-level distribution and estimates expected advantages. Experiments across eight benchmarks and four base models demonstrate strong performance and cross-domain generalization. GMAE also exhibits stable learning, low extra cost, and good applicability across training datasets and RL backbones.

[NLP-112] SIPO: Unifying Reinforcement Learning with On-Policy Self-Distillation

【速读】: 该论文旨在解决强化学习中基于可验证奖励(Reinforcement Learning with Verifiable Rewards, RLVR)在大语言模型(Large Language Models, LLMs)任务优化中存在的稀疏奖励问题,尤其是缺乏对中间步骤的逐标记(token-level)信用分配。现有方法如基于策略的自蒸馏(On-Policy Self-Distillation, OPSD)虽引入了密集的学习信号,但其自教师(self-teacher)常因过度自信而在长推理轨迹上施加过强惩罚,导致实际性能受限。为克服此问题,本文提出自指导策略优化(Self-Instructing Policy Optimization, SIPO),其核心在于设计一种对比式自教师机制:在每轮迭代中,SIPO从当前策略采样多个提示的生成轨迹,利用环境奖励对它们进行评分,并通过将参考答案与组内其他错误样本配对,构建两个教师上下文。模型随后在两种上下文中重新评估自身输出,以两者的教师对数概率差作为逐标记反馈,从而有效消除共现偏差。该方法在保留原始任务奖励主导方向的同时,实现了高密度、细粒度的逐标记优势估计,即使在所有轨迹均失败且组间相对优势消失的情况下仍能提供有效的学习信号。通过融合强化学习与自蒸馏的优势,SIPO显著提升了模型在多类推理与代码生成基准上的表现,优于RLVR与OPSD基线,且无需外部教师或额外生成过程。

链接: https://arxiv.org/abs/2609.36742
作者: Zhenrui Yue,Huimin Zeng,Yueqi Wang,Yaokun Liu,Fengran Mo,Jinghan Zhang,Mung Yao Jia,Gyuseok Lee,Yang Zhang,Na Wei,Dong Wang
机构: University of Illinois(伊利诺伊大学); UC San Diego(加州大学圣地亚哥分校); Rochester Institute of Technology(罗切斯特理工学院); Clemson University(克莱姆森大学); Miami University(迈阿密大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for improving large language models (LLMs) on various tasks, yet its sparse outcome rewards lack token-level credit assignment for intermediate steps. To address this, on-policy self-distillation (OPSD) leverages a self-teacher with privileged context to provide additional dense learning signals. However, because the self-teacher is often overconfident and imposes excessive penalties on long reasoning trajectories, OPSD frequently struggles in practice. To mitigate this, we propose self-instructing policy optimization (SIPO) with a contrastive self-teacher to provide dense credit. At each iteration, SIPO samples multiple rollouts per prompt from the current policy, scores them with environment rewards, and constructs two teacher contexts for each rollout by pairing the reference answer with mistakes made within the group. The model then re-evaluates its own responses under both contexts, using the difference between the two teacher log-probabilities as token-level feedback, so that biases shared by both contexts are expected to largely cancel. The resulting objective yields a token-level advantage for every rollout: the reward still sets the main direction of each update while the self-teacher redistributes credit across tokens. Even in groups where every rollout fails and group-relative advantages vanish, SIPO still provides a learning signal. By preserving direct optimization of the task reward while providing dense, token-level feedback, this approach bridges reinforcement learning and on-policy self-distillation. Extensive experiments across multiple reasoning and code-generation benchmarks demonstrate that SIPO outperforms both RLVR and OPSD baselines without an external teacher or additional generation.

[NLP-113] Backpropagated Output Momentum: Relocating Optimizer History from Parameters to Task Space

【速读】: 该论文旨在解决优化器动量(optimizer momentum)在存储和使用历史梯度信息时存在的高内存开销与坐标固定性问题。传统方法将动量作为与参数尺寸相同的过去梯度移动平均值进行存储,导致状态空间占用大,且历史信号被固定在原始计算坐标系中,限制了其灵活性与泛化能力。为此,论文提出了一种名为反向传播输出动量(Backpropagated Output Momentum, BOM)的新方案,其核心创新在于:不再直接存储梯度历史,而是以模型输出端的预测误差为对象,维护一个紧凑的移动平均,并在每一步通过当前网络结构重新投影该历史信息。这种机制使历史信息得以在输出空间中动态重用,有效保留了对训练有益的信号,同时避免冗余存储。理论分析表明,该重构策略能够合理保留关键信息并舍弃无关成分。在实现层面,BOM可无缝集成至多种自适应优化器中,替代其一阶矩分量,且不破坏原有监督梯度流。实验结果显示,在三种组合架构中,该方法将参数级优化器状态减少49.7%至99.8%,平均加速训练步长4.0%;在语言与视觉微调任务中,平均验证性能提升1.42点。跨模型尺度与输出空间的预训练研究及对照实验进一步验证了该方法的鲁棒性与有效性。

链接: https://arxiv.org/abs/2609.36738
作者: Yuchen Li,Zongqi Fan,Nguyen H. Tran,Ken-Tye Yong
机构: University of Sydney(悉尼大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 53 pages, 10 figures

点击查看摘要

Abstract:Optimizer momentum is usually stored as a parameter-sized moving average of past gradients, which makes history costly and fixes each past signal in the coordinates in which it was computed. We introduce Backpropagated Output Momentum (BOM), which instead stores a compact moving average of prediction errors at the model output and reprojects that history through the current network at every step. A batch-level analysis characterizes the information retained and omitted by this relocation, while the implementation preserves the current supervised gradient and can replace the first-moment component of several adaptive optimizers. As a plug-in for momentum-based optimizers, including ones that already compress their state, BOM reduces parameter-shaped optimizer state by 49.7-99.8% in three compositions and, averaged over three language backbones, paired step time by 4.0%. It also improves mean validation performance across language and vision fine-tuning, by 1.42 points in the primary five-task comparison. Language and vision pretraining studies, together with matched mechanism controls, further test the construction across output spaces and model scales.

[NLP-114] Reconstructing the Vocal Tract with Differentiable Acoustic Simulation NEURIPS2026

【速读】: 该论文旨在解决语音生成中声道几何形状与声学输出之间的非凸逆映射难题,即如何仅通过语音信号重建发声时动态变化的声道形态。其核心解决方案在于提出一种可微分且基于GPU加速的声道声学模拟器,关键创新点包括:(1)设计了一种频域下的声道流体动力学公式,相比时域有限差分法具备70倍更高的GPU并行度;(2)引入可微分的湍流模型以准确合成辅音;(3)借鉴隐式神经表示(Implicit Neural Representations, INRs)思想,采用神经网络参数化声道几何,显著提升优化收敛速度并有效逃离离散表示易陷入的局部极小。由于模拟器具备可微性,可无缝集成至深度学习框架,实现跨11种语言的自监督声道形状自动编码,并在无需配对数据的情况下,结合生成式MRI模型从单语音信号中重构个体动态声道结构,为语言学与医学影像应用开辟新路径。

链接: https://arxiv.org/abs/2609.36737
作者: Eric Ming Chen,Jin Woo Lee,Vincent Sitzmann
机构: MIT CSAIL(麻省理工学院计算机科学与人工智能实验室); MIT RLE(麻省理工学院辐射实验室); KAIST GSCT(韩国科学技术院全球科学与技术中心)
类目: ound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
备注: Accepted as NeurIPS 2026 spotlight paper. Supplementary material at this https URL

点击查看摘要

Abstract:The vocal tract is the region of the human body responsible for filtering one’s voice to create speech. In this paper, we present a differentiable and GPU accelerated acoustic simulator for the vocal tract. The differentiable simulator synthesizes speech by propagating sound along an acoustic tube model of the vocal tract, and via its gradients, can solve the inverse problem: reconstructing the shape of the vocal tract solely from the sound it produces. Although the inverse mapping between geometry and sound is notoriously non-convex, we discover that gradient descent succeeds with three technical contributions: (1) we design a frequency domain formulation of the vocal tract’s fluid dynamics that is 70x more GPU parallelizable than finite differences in time, (2) we integrate a differentiable model for turbulence to synthesize consonants, and (3) similar to prior work in implicit neural representations (INRs) and neural fields, we find that parameterizing the geometry with a neural network accelerates convergence and escapes local minima that trap discrete representations. Because the simulator is differentiable, it is readily integrated with other deep learning pipelines to enable novel linguistics and medical imaging applications. (1) We demonstrate self-supervised autoencoding of vocal tract shapes across 11 languages, and (2) we couple our simulator with a generative model of MRI (magnetic resonance imaging) images to reconstruct one’s moving vocal tract from only their speech without paired data.

[NLP-115] Distilling What Matters: Confidence-Aware Selective Distillation for Large Language Models NEURIPS2026

【速读】: 该论文旨在解决大语言模型(LLM)中知识蒸馏(Knowledge Distillation, KD)因依赖教师模型作为可靠“先验”而导致性能下降的问题。传统KD假设教师模型输出具有高可靠性,但在实际应用中,教师模型常表现出高熵与幻觉(hallucination),导致学生模型学习到错误或不校准的先验知识。为应对这一挑战,论文提出一种自信度门控的知识蒸馏框架CaRE-KD,其核心创新在于引入不确定性自适应优化机制。关键解决方案包含两个层面:一是基于教师-学生置信度动态切换前向与反向KL散度的词元级损失(CaRE-Divergence),实现对不同置信水平下信息传递的自适应调整;二是基于认知不确定性(epistemic uncertainty)的批量级拒绝机制(Revival),当教师模型不确定性高于学生时主动抑制参数更新,从而避免噪声监督带来的偏差。理论分析表明,该双粒度设计可诱导出条件校准机制,这是静态散度无法实现的。实验结果在八个教师-学生组合及十一项基准任务(涵盖指令遵循、对话对齐、代码生成与数学推理)上验证了该方法的优越性,显著优于强基线(如Skewed-KL、α–β散度),在多个任务上取得显著提升,且在“以大模型为裁判”的事实性评估中最高达+2.5的改进,同时证明Revival可作为无损插件模块系统增强现有蒸馏目标。

链接: https://arxiv.org/abs/2609.36734
作者: Ayan Sengupta,Vaibhav Seth,Tanmoy Chakraborty
机构: 未知
类目: Computation and Language (cs.CL)
备注: Accepted at NeurIPS 2026

点击查看摘要

Abstract:Knowledge Distillation (KD) trains a smaller-capacity student model to imitate a larger-capacity teacher model by matching output distributions, implicitly assuming the teacher to be a reliable oracle. In large language models (LLMs), this assumption often fails: teacher predictions can exhibit high entropy and hallucinations, causing standard KD to degrade well-calibrated student priors. We propose CaRE-KD, a confidence-gated distillation framework that replaces static objectives with uncertainty-adaptive optimization. CaRE-KD has two components: a token-level loss (CaRE-Divergence) that adaptively switches between Forward and Reverse KL divergence based on teacher–student confidence, and a batch-level epistemic rejection mechanism (Revival) that suppresses updates when the teacher is more uncertain than the student. We provide a gradient-level analysis showing how this dual-granularity design induces a conditional calibration mechanism that prior static divergences cannot reproduce. Empirically, across eight teacher–student pairs and eleven benchmarks spanning instruction following, chat alignment, code generation, and mathematical reasoning, CaRE-KD delivers consistent gains over strong baselines (Skewed-KL, \alpha – \beta divergence). Highlights include up to +3.2 average ROUGE-L on instruction-following tasks, +2.1 pass@1 on MBPP, +1.7 accuracy on GSM8k, and +1.8 accuracy on CollegeMath over the strongest baseline, with consistent gains in LLM-as-a-judge factuality (up to +2.5 per task over Skewed-RKL). Revival further acts as a principled, loss-agnostic plug-in that systematically strengthens existing distillation objectives by filtering epistemically unreliable teacher supervision.

[NLP-116] Can Agents Design Libraries for Agents ?

【速读】: 该论文旨在解决智能体(agent)在协作开发中普遍存在的重复实现问题,即下游智能体倾向于重新实现已有功能而非复用由其他智能体设计的代码库,从而导致代码库不断膨胀。其核心问题是:如何评估智能体为其他智能体设计可复用库的能力,并提升其设计质量以促进实际复用。解决方案的关键在于提出一个名为LibraryDesignBench的双阶段基准测试框架,通过让智能体根据功能规范而非具体设计路径构建完整功能库,进而评估其设计的正确性与简洁性。该基准涵盖15项库设计任务、242个经专家验证的编程问题,覆盖四种编程语言。实验发现,尽管部分智能体能复现人类生产级库的抽象结构,但下游智能体仍普遍忽略使用这些库,主要因代理生成的库存在灵活性差或使用困难等设计缺陷,而非功能缺失。进一步实验表明,采用更明确的“以代理为中心”的设计指导以及引入子智能体进行测试,可显著提升下游智能体的性能并生成更简洁的代码,从而改善库的复用率。因此,该研究不仅提供了一个评估智能体库设计能力的测试平台,还确立了提升下游复用性的初始设计基准。

链接: https://arxiv.org/abs/2609.36730
作者: Gabriel Orlanski,Alex L. Zhang,Avi Trost,Vincent Sunn Chen,Frederic Sala,Aws Albarghouthi,Ludwig Schmidt
机构: University of Wisconsin–Madison(威斯康星大学麦迪逊分校); Massachusetts Institute of Technology(麻省理工学院); Snorkel AI; Stanford University(斯坦福大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Software Engineering (cs.SE)
备注: 26 pages, 6 figures, 11 tables. Code and data: this https URL

点击查看摘要

Abstract:Agents increasingly build on code written by other agents, and they reimplement rather than reuse, growing the codebases later agents must work in. To measure how well agents design libraries for other agents, we introduce LibraryDesignBench, a two-phase benchmark in which an agent implements a full-featured library from a specification that defines required capabilities and potential use cases without prescribing the design. We evaluate the library through the correctness and simplicity of programs written by three user agents from different model families. The benchmark spans 242 expert-validated programming problems across 15 library-design tasks in four languages. On eleven of the fifteen tasks, agent designers reproduce the abstractions of the human-written production library. Downstream agents adopt agent- and human-written libraries alike but underuse them, reimplementing capabilities the library already provides. Our failure analysis finds that downstream agents write extra code mainly because agent-written libraries are rigid or hard to use, not because capabilities are missing. We also experiment with giving designers more prescriptive, agent-first guidance and having them test their library with subagents; this improves downstream scores and yields simpler programs. LibraryDesignBench provides both a testbed for evaluating library-design practices for agent users and an initial design baseline that improves downstream reuse.

[NLP-117] ATTUNER: Recomputation-Free KV Cache Reuse via Query-Side Adaptation

【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)代理在推理过程中反复加载可复用内容(如技能、文档和记忆条目)时,因每次请求均需重新编码而导致的计算浪费问题。现有方法位置无关缓存(Position-Independent Caching, PIC)通过独立编码每个知识单元并重用其键值(Key-Value, KV)状态以实现高效推理,但其性能相较于完整上下文预填充(full-context prefill)存在明显下降。本文通过系统分析发现,性能损失的根源并非来自位置信息错配或KV表示失真,而是源于注意力机制中查询(query)与缓存之间的不匹配:尽管独立缓存的内容本身保持忠实表征,但在多候选选择场景下,注意力权重计算偏差导致模型无法准确聚焦目标内容。基于此关键洞察,作者提出\textscAttuner——一种仅在查询侧进行轻量级适配的解决方案,通过在查询投影层插入低秩适配器(low-rank adapters),并采用教师-学生蒸馏方式将全上下文预填充的输出分布迁移到学生模型,从而学习如何正确读取冻结的缓存。该方法仅训练模型参数的0.05%以下,在推理阶段无需重计算缓存或依赖完整上下文参考,即可在涵盖技能、文档、记忆与代码的七个基准测试上显著优于先前的PIC基线,在域内与域外设置下均达到接近全上下文预填充的性能水平,并实现最高达3.73倍的加速。

链接: https://arxiv.org/abs/2609.36722
作者: Xinghao Chen,Junnan Dong,Cai Ke,Chak Tou Leong,Haocheng Sun,Keyu Chen,Siyu An,Ruizhi Qiao,Xing Sun,Wenjie Li,Xiaoyu Shen
机构: EIT-NLP Lab, Eastern Institute of Technology, Ningbo(东方理工大学宁波自然语言处理实验室); The Hong Kong Polytechnic University(香港理工大学); Tencent Youtu Lab(腾讯优图实验室)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Large language model (LLM) agents repeatedly load reusable content, such as skills, documents, and memory entries, into the current context. Re-encoding this content for every request wastes computation. Position-independent caching (PIC) alleviates this by encoding each artifact independently and reusing its key-value (KV) states at arbitrary positions, but it incurs a quality loss relative to full-context prefill. Existing methods repair this loss by restoring global position IDs or recomputing selected tokens. In this work, we isolate the source of the loss, finding that the positional mismatch has minor effect, and independently cached artifacts retain faithful representations: reading a provided artifact stays largely accurate, and performance degrades only when the model must select among multiple artifacts. Moreover, replacing PIC’s attention scores with full-prefill scores recovers performance with the cached KV unchanged, localizing the failure to the attention rather than KV recomputation. Motivated by this, we propose \textscAttuner, a query-side adaptation method that learns to read a frozen artifact cache. \textscAttuner inserts low-rank adapters into the query projections and is trained by distilling full-prefill distribution into the student. It trains fewer than 0.05% of the model parameters and, at inference, requires neither cache recomputation nor a full-context reference. On Qwen3-4B and Qwen3-8B across seven benchmarks covering skills, documents, memory, and code, \textscAttuner substantially outperforms prior PIC baselines in both in-domain and out-of-domain settings, matches full-context prefill quality while providing up to 3.73\times speedup.

[NLP-118] LAURA: Knowledge Distillation for Interpretable Ambiguous Clause Identification in Legal Contracts

【速读】: 该论文旨在解决法律合同中模糊条款识别的难题,其核心挑战在于:仅识别模糊条款不足以降低企业面临的财务与法律风险,因为部分模糊性可被灵活解释而不会引发争议,而另一些则可能引发重大法律冲突。因此,单纯依赖模糊条款的检测是不够的,亟需具备可解释性的推理分析能力。论文提出LAURA框架,其关键解决方案在于采用基于知识蒸馏的后训练范式,并引入IRAC-Unlearning提示策略,将教师大模型(LLM)的知识高效迁移至参数量约10亿的开源学生模型中。该框架通过联合优化分类任务与理由生成任务的损失函数,实现对模糊条款的精准识别与可解释性推理。实验表明,LAURA在7个基线和7个开源模型上均展现出最优的可解释性表现,同时在识别性能上达到最佳不透明基线的水平,有效支持法律与非法律利益相关方对高风险模糊条款做出审慎决策。

链接: https://arxiv.org/abs/2609.36707
作者: Amrita Singh,Aditya Joshi,Jiaojiao Jiang,Hye-young Paik
机构: University of New South Wales (UNSW)(新南威尔士大学)
类目: Computation and Language (cs.CL)
备注: Under Review

点击查看摘要

Abstract:Legal contracts contain ambiguities that expose enterprises to financial and legal risks. Some ambiguities allow flexible interpretation without triggering disputes, while others lead to significant legal conflicts. This makes identification alone insufficient, and interpretable rationale analysis essential. We propose LAURA, a post-training framework for interpretable ambiguous clause identification. LAURA leverages knowledge distillation with an IRAC-Unlearning prompting technique to transfer knowledge from a teacher LLM to an open-weight student model (=1B parameters), which is then trained using a joint objective combining classification and rationale generation losses. The framework supports both legal and non-legal stakeholders in making informed decisions about which ambiguities require further attention. Extensive experiments across 7 baselines and 7 open-weight models demonstrate that LAURA with Flan-T5 (250M) delivers state-of-the-art interpretability over all interpretable baselines while matching the identification performance of the best-performing opaque baseline.

[NLP-119] Lost in Conversation or Lost in Translation? Diagnosing Multi-Turn Degradation in RAG

【速读】: 该论文旨在解决当前大型语言模型(Large Language Models, LLMs)在多轮对话场景下,尤其是基于检索增强生成(Retrieval-Augmented Generation, RAG)及其图结构变体(GraphRAG)的系统中,因评估范式与实际使用场景不匹配所导致的性能下降问题。现有研究几乎仅在单轮、完整指定的查询上评估RAG系统,而真实用户交互往往从简单问题出发,通过多轮追问逐步构建复杂多跳问题,这种“评估-应用”之间的脱节使得现有系统在实际多轮对话中表现不佳。论文的关键解决方案在于:通过大规模模拟实验,将多跳问答(Multi-hop Question Answering, QA)基准中的问题转化为不完整、逐步细化的对话形式,系统性地评估十种LLM助手与八种检索系统在150万次模拟多轮对话中的表现。研究发现,多轮交互导致系统性能普遍显著下降,最高相对准确率损失达21%,不可靠性提升47%。其核心失败模式可归结为两类:一是“语义失真”(lost in translation),即对话中的重述或改写使检索查询偏离原始意图;二是“信息分散”(lost in conversation),即虽然检索成功获取了相关信息,但模型未能有效整合跨轮次分布的信息以完成推理。这一发现揭示了当前RAG系统在真实多轮交互场景下的根本局限,并指出了改进方向。

链接: https://arxiv.org/abs/2609.36700
作者: Pranav Handa,Ariful Azad
机构: Texas A&M University (德克萨斯农工大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 35 pages, 11 figures

点击查看摘要

Abstract:When conversing with large language models (LLMs), users often begin with a simple question and build towards a multi-hop question through follow-up turns. Retrieval-augmented generation (RAG) and its graph-based variant (GraphRAG) have become the dominant approaches for grounding LLM responses in external evidence, yet both are evaluated almost exclusively on single-turn, fully specified queries. We systematically investigate this evaluation mismatch through a large-scale simulation study. Building on prior work on multi-turn LLM evaluation, we transform questions from multi-hop question answering (QA) benchmarks into underspecified conversations and evaluate ten LLM assistants with eight retrieval systems across 1.5 million simulated conversations. Our findings reveal that multi-turn interaction causes widespread performance degradation, incurring relative performance drops of up to 21% and increasing unreliability by 47%, making RAG systems simultaneously less accurate and less reliable. We identify two distinct failure modes behind this degradation. Systems are either lost in translation, where conversational rephrasing distorts the retrieval query, or lost in conversation, where retrieval succeeds but the LLM fails to synthesize evidence distributed across turns.

[NLP-120] Video2Skill: From Streaming Experience to Reusable Embodied Skills

【速读】: 该论文旨在解决**具身智能体在长期交互中如何从连续视觉-语言流中自动发现并维护可复用的操纵技能(Manipulation Skills)**这一关键问题,即如何实现对未标注、动态变化的操纵行为序列进行持续性技能归纳。其核心挑战在于:尽管视觉-语言模型(VLMs)能够良好描述单个操纵事件,但难以将一系列事件组织为具有语义一致性和可重用性的技能单元。为此,作者提出“流式具身技能发现”(Streaming Embodied Skill Discovery, SESD)范式,并构建了Video2Skill基准测试,系统评估模型在时间定位操纵事件、按变换类型聚类事件以及决定是否复用或创建新技能三个维度上的能力。研究发现,现有多数VLMs在技能分组上表现接近随机水平,且模型规模提升并未带来稳定性能增益;其错误根源在于感知与技能库更新机制的耦合方式——联合建模易将不同变换合并为单一技能,而基于文本描述更新库则导致重复创建相似技能。尽管通过监督微调(包括提出的反事实库状态再平衡策略CLaRe)可改善技能聚类效果,但模型普遍存在“技能固化”现象:仅强化已有技能,极少拓展新技能库,导致对训练中未见变换的行为虽能定位却无法赋予新技能。因此,识别现有技能不足并主动扩展技能库的能力,成为当前方法的核心瓶颈,也是未来研究的关键突破口。

链接: https://arxiv.org/abs/2609.36691
作者: Jianshu Zhang,Ce Zhang,Xiyuan Yang,Chenwei Xu,Haoran Lu,Yijiang Li,Yaqi Xie,Katia P. Sycara,Han Liu
机构: Northwestern University (西北大学); CMU (卡内基梅隆大学); UIUC (伊利诺伊大学厄本那-香槟分校); UCSD (加州大学圣地亚哥分校)
类目: Computation and Language (cs.CL)
备注: Project page: this https URL

点击查看摘要

Abstract:Manipulation behaviors vary widely across objects and scenes, but they share a small set of reusable skills, and planning with these skills helps embodied agents generalize to new tasks. Yet an agent can only plan with skills it knows. Recovering skills from observed experience, the inverse of planning, builds this knowledge over time and yields skill data for training future agents. Vision-Language Models (VLMs) describe individual manipulation events well, but can they organize a stream of events into reusable skills? We formulate this problem as Streaming Embodied Skill Discovery (SESD): a model watches videos in sequence and maintains a persistent skill library that shapes its later decisions. To systematically measure this ability, we introduce Video2Skill, a benchmark that covers robot tabletop manipulation and human kitchen activity and tests three core capabilities: (i) locating manipulation events in time, (ii) grouping events of the same transformation, and (iii) deciding when to reuse an existing skill or create a new one. Across 19 open-source VLMs, many models group events at near-chance level, and scale does not consistently help. Their errors depend on how perception and library updates are coupled: joint models merge distinct transformations into one skill, while models that update the library from text descriptions duplicate recurring ones. Supervised fine-tuning, including our counterfactual library-state rebalancing (CLaRe), improves grouping but exposes a deeper bottleneck: trained models consolidate familiar skills yet rarely expand the library. Their libraries stall below half the reference size, and transformations unseen in training are located in time but almost never given a new skill. Recognizing when existing skills are insufficient thus emerges as the central challenge.

[NLP-121] CHAIN: Calibrated LLM Forecasting via Causal-Temporal Hypergraph Inference

【速读】: 该论文旨在解决大语言模型在事件预测中概率输出存在系统性校准偏差的问题,这种偏差在不同领域和问题类型间呈现异质性,严重影响了不确定性决策下概率输出的可信度。现有校准方法通常仅在预测完成后对概率进行修正,未能从预测过程本身的结构根源上建模并消除偏差。为此,论文提出一种基于因果-时序超图(causal-temporal hypergraphs)的三阶段概率预测框架——CHAIN,其核心创新在于针对每个预测阶段设计特定的偏差缓解机制:(i) 通过因果拓扑距离调节时间衰减函数以反映因果关系的远近;(ii) 在方向感知去重后,利用带噪声的或(Noisy-OR)聚合近似独立的因果链;(iii) 借助因果覆盖度与方向平衡性驱动自适应融合。实验结果表明,CHAIN在跨领域预测基准上显著优于现有方法,在期望校准误差(Expected Calibration Error)、Brier评分及准确率方面均取得更优表现。

链接: https://arxiv.org/abs/2609.36689
作者: Wenjin Liu,Chenxi Wang,Yue Lu,Zhe Cui,Haoran Luo
机构: Hithink Research(希思科研究); Nanyang Technological University(南洋理工大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Large language models have achieved significant progress in event forecasting, yet their probability outputs exhibit systematic calibration bias that varies heterogeneously across different domains and question types, undermining the trustworthiness of probabilistic outputs for decision-making under uncertainty. However, existing calibration methods typically correct probability outputs after prediction is complete, without modeling the structural sources of bias within the prediction process itself. To address this challenge, we decompose probabilistic prediction over causal-temporal hypergraphs into three stages, evidence weighting, evidence aggregation, and source fusion, and propose CHAIN, which designs stage-specific mechanisms to mitigate bias at each stage: (i) modulating the temporal decay function by causal topological distance, (ii) aggregating approximately independent causal chains via Noisy-OR after direction-aware deduplication, and (iii) driving adaptive fusion by causal coverage and directional balance. Experimental results on cross-domain forecasting benchmarks show CHAIN outperforms existing methods in expected calibration error, Brier score, and accuracy. Our project is available at this https URL.

[NLP-122] ProgressCompass: Embodied Progress Reward Models Are Lost Without the Right Context

【速读】: 该论文旨在解决具身智能体在执行长时任务时,因上下文依赖性导致的进展估计(context-dependent progress estimation)难题。传统进展奖励模型(Progress Reward Models, PRMs)在短任务中表现良好,但面对需要依赖历史信息的任务时,仅凭当前观测帧难以准确判断任务进展,从而导致评估偏差。现有基准多集中于可从当前状态直接推断进展的短任务,缺乏对上下文依赖型进展估计的有效评估。为此,作者构建了ContextProgress-Bench基准,包含24个操作任务、120个实验回合,涵盖三种典型情境:状态回溯(State Recall)、序列追踪(Sequence Tracking)和重复歧义消解(Recurrence Disambiguation)。通过配对诊断实验发现,即使采用完整历史输入,现有PRMs仍难以准确估计进展;而引入正确上下文后,五种模型的进展误差降低77%-82%。由此揭示问题核心并非模型能力不足,而是缺乏恰当上下文支持。为此,论文提出ProgressCompass——一种基于通用视觉语言模型(VLMs)的自主代理循环机制,能够动态为冻结的PRM提供所需上下文信息。经该机制增强后,相同PRM的进展误差下降63%,排名一致性提升76%,显著提升了其在复杂长时任务中的进展估计性能。因此,解决方案的关键在于通过外部上下文感知代理(ProgressCompass)重构并增强现有PRM的上下文理解能力,实现对长时任务进展的精准评估。

链接: https://arxiv.org/abs/2609.36684
作者: Jianshu Zhang,Keliang Wu,Chengxuan Qian,Xiyuan Yang,Ce Zhang,Ariel Tian,Anbang Liu,Haoran Lu,Han Liu
机构: Northwestern(西北大学); UCSB(加州大学圣塔芭芭拉分校); UIUC(伊利诺伊大学厄本那-香槟分校); CMU(卡内基梅隆大学)
类目: Computation and Language (cs.CL)
备注: Project page: this https URL

点击查看摘要

Abstract:Embodied agents now take on ever longer tasks. For long tasks, knowing only whether a task finally succeeds or fails says little; the steps along the way matter. Progress Reward Models (PRMs) score how far a task has come at every step, and serve as dense rewards, verifiers and monitors. Yet in long tasks the current frame alone often cannot tell how far the task has come, because progress depends on what happened before. We call this problem context-dependent progress estimation. Existing benchmarks on progress estimation mostly focus on short tasks whose progress can be read from the current observation, and whether PRMs can estimate progress when context is needed remains underexplored. We therefore build ContextProgress-Bench, with 24 manipulation tasks for 120 episodes. The benchmark covers three settings: (i) State Recall, where information needed for progress appeared earlier but is not in the current frame; (ii) Sequence Tracking, where steps follow a fixed order, so progress requires knowing which steps are done and which comes next; and (iii) Recurrence Disambiguation, where look-alike frames sit at very different progress. We then run a paired diagnosis: each PRM keeps the same input format in both runs, and in one run its instruction integrates the right context. Even PRMs that read the entire history get lost in estimating progress, yet with the right context the same five models cut their progress error by 77-82%. Embodied PRMs are thus not incapable of progress estimation, but lost without the right context. We therefore propose ProgressCompass, an autonomous agentic loop that reorients an existing PRM and uses current general-purpose VLMs to supply the context the PRM needs. Wrapped in the loop, the same frozen PRM cuts its progress error by 63% and raises its rank agreement by 76%. With such a compass, PRMs estimate progress far better on longer, more complex tasks.

[NLP-123] MARCO: Multi-Round Agent ic Reinforcement for Conditional Molecular Optimization

【速读】: 该论文旨在解决分子优化中生成式模型在单次响应下难以同时兼顾分子有效性、性质提升与结构相似性控制的核心挑战。传统指令跟随模型通常仅输出单一修改后的分子,导致需在单步决策中同时满足多个目标,易造成性能瓶颈。其解决方案的关键在于提出MARCO——一种基于评估器反馈的强化学习框架,通过构建受限的“提案-反馈-修正”迭代轨迹,训练分子编辑器在多轮交互中逐步优化分子。MARCO通过聚合每一轮的形状化奖励(shaped turn rewards)形成无折扣的轨迹回报,用于组相对策略优化,从而实现更稳定的多目标协同优化。实验表明,基于监督微调(SFT)初始化的MARCO在三个目标的MuMOInstruct基准上,在所有主要设置中均取得最高的性质成功概率与结构相似性乘积;尤其在支持最多五次响应的Same-5测试中,进一步提升了预算内表现,验证了其在多轮反馈机制下的优越性,且在四目标任务及公开检查点上的迁移实验也证明了其在不同约束集和初始化策略下的泛化能力。

链接: https://arxiv.org/abs/2609.36683
作者: Shicheng Fang,Yuxin Wang,Zhuo Yang,Xiaohu Xu,Jiahao Lu,Chuanyuan Tan,Tong Zhu,Yining Zheng,Xipeng Qiu
机构: 未知
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Molecular optimization is inherently iterative: a candidate is proposed, evaluated against several objectives, and revised while preserving a relationship to the source molecule. Most instruction-following models instead emit one edited molecule, forcing validity, property improvement, and similarity control into a single response. We introduce MARCO, an evaluator-grounded reinforcement-learning framework that trains molecular editors on bounded proposal–feedback–revision trajectories. MARCO aggregates shaped turn rewards into an undiscounted trajectory return for group-relative policy optimization. We evaluate two consequences of this training: Same-1 tests the trained policy under a one-response budget, while Same-5 tests whether the same policy can use verifier feedback when up to five responses are available. Across the three-objective MuMOInstruct benchmark, three Qwen backbones, and seen/unseen instruction splits, SFT-initialized MARCO obtains the highest product of property success rate and similarity in every reported primary setting. Same-5 further improves the observed score under the tested budget, while four-objective and public-checkpoint experiments test transfer across constraint sets and initialization regimes.

[NLP-124] Gödel Forest: Balancing Search Depth and Breadth for Data-Centric Recursive Self-Improvement

【速读】: 该论文旨在解决递归自改进(Recursive Self-Improvement, RSI)中因数据策略验证成本高而导致的探索深度与广度难以兼顾的核心难题。现有方法或受限于单个智能体在狭窄方向上的局部优化,缺乏探索多样性;或采用简单的并行搜索与大量执行日志共享,牺牲了长周期搜索的深度。为克服这一困境,论文提出Gödel Forest——一种多智能体协同进化的树状集成框架,将自改进过程建模为一组共进化搜索树的集合。其核心创新在于构建一个动态共进化记忆系统,使各智能体在持续生成和演化自身搜索树的同时,将成功与失败经验以紧凑的程序化经验法则形式提炼,并通过全局排行榜进行共享。这种机制实现了深度(通过单个智能体持续深化策略)与广度(通过多树并行探索不同数据空间区域)的平衡:任一搜索树遭遇死胡同可即时警示整个森林规避无效路径,而任一实证突破能迅速激发邻近树的新生长分支。在RSIBench-Data六个不同领域的评估中,Gödel Forest平均性能优于单智能体基线10.70%,且在五项任务上显著缩短了实际运行时间。消融实验表明,相比独立并行搜索,共进化共享记忆带来7.00%的性能提升,验证了集体经验提炼是实现可扩展自改进的关键。

链接: https://arxiv.org/abs/2609.36675
作者: Ziqi Zhao,Fanqing Meng,Haocheng Lu,Lingxiao Du,Qiguang Chen,Mengkang Hu,Xiao-Ming Wu
机构: Evolvent AI; The Hong Kong Polytechnic University(香港理工大学); National University of Singapore(新加坡国立大学); Columbia University(哥伦比亚大学)
类目: Computation and Language (cs.CL)
备注: Preprint

点击查看摘要

Abstract:Recursive self-improvement (RSI) aims to achieve compounding gains by having models improve themselves. While most existing RSI systems optimize external agent harnesses or prompts around a frozen base model, data-centric RSI directly updates the model’s own parameters by training on agent-generated data. However, because validating data strategies requires expensive model training, existing methods face a fundamental dilemma: a single agent gets trapped in narrow directions and lacks exploration breadth, while naive parallel search or heavy trace sharing sacrifices long-horizon search depth. To address this challenge, we introduce G"odel Forest, a multi-agent framework that organizes recursive self-improvement as an ensemble of co-evolving search trees. In G"odel Forest, each agent autonomously grows a persistent tree, deepening, branching, or pruning data strategies based on model feedback to secure depth, while parallel trees explore distinct regions of the data space to expand breadth. Crucially, rather than leaving trees isolated or flooding them with heavy execution logs, a dynamically co-evolving memory connects the forest: agents continuously distill their successes and failures into compact procedural lessons anchored to a global leaderboard. Through this forest ecosystem, a dead-end in one tree instantly warns the whole forest against unpromising paths, while an empirical breakthrough quickly seeds new exploration branches in neighboring trees. Evaluated on RSIBench-Data across six diverse domains, G"odel Forest outperforms the single-agent baseline by an average of 10.70% while reducing wall-clock time on five tasks. Ablations confirm that co-evolving shared memory yields a +7.00% gain over independent parallel search, demonstrating that collective distillation is key to scalable self-improvement. The code is available at this https URL.

[NLP-125] Replay the Curvature: Accurate and Scalable NVFP4 Quantization for Large Language Model Inference

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在低精度推理过程中面临的两大核心挑战:一是如何高效选择量化缩放因子(scale),以最小化权重重构误差,尤其是在GPTQ这类逐列量化方法中,后续列的量化会受前序列的影响,导致独立评估块的误差存在偏差;二是大规模模型在训练后量化(post-training quantization)时的内存与计算资源瓶颈,即全精度权重、校准激活值和二阶状态无法同时驻留于单一加速器上,而将整层分配给设备又会造成严重的串行延迟。针对上述问题,论文提出两项关键技术:其一为Schur Replay算法,通过模拟每个缩放因子对后续列的更新影响,并结合未量化列的补偿效应,精准评估量化块的最终重构误差,从而实现更优的缩放因子选择;其二为执行基础设施,采用分层激活管理策略,仅保留当前活跃层在设备端,将激活值分层存储于设备、主机与磁盘之间,量化完成后及时释放全精度权重,并将输出行在张量并行秩间分布式处理,显著降低内存占用与每层处理时间。二者协同作用,在Qwen3.5-397B-A17B与Llama-3.3-70B-Instruct模型上实现了对BF16基准的99.35%和100.84%的问题加权恢复率,且在397B模型上相较ModelOpt与LLM Compressor分别将每层处理时间缩短15.17倍和23.14倍,同时单位GPU内存使用更低。

链接: https://arxiv.org/abs/2609.36654
作者: Ruiyi Ding,Jie Li,Kang He,Ziyan Liu,Chengru Song,Yuedong Xu,Yuan Cheng
机构: KlingAI Research; Fudan University (复旦大学); Shanghai AI Incubation and Innovation Center (上海人工智能创新中心); Shanghai Academy of AI for Science (上海人工智能科学研究院); University of Science and Technology of China (中国科学技术大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL); Distributed, Parallel, and Cluster Computing (cs.DC)
备注: 33 pages

点击查看摘要

Abstract:Large language models make weight storage and memory traffic major inference costs, motivating low-precision formats that represent each weight with only a few bits. Such formats use a scale to map floating-point values into a small codebook; NVFP4 improves local range utilization by letting every 16 E2M1 weights share an E4M3 block scale. Choosing that scale is difficult in GPTQ because quantizing one column updates those that follow, so evaluating a block independently can misestimate its final reconstruction error. Large models pose a second challenge: full-precision weights, calibration activations, and second-order state cannot all remain on one accelerator, while assigning complete layers to devices leaves each time-consuming layer solve serial. We introduce \emphSchur Replay, a scale-selection algorithm that reproduces the GPTQ updates caused by each block scale and scores the resulting block error after accounting for compensation from unquantized columns. Separately, our execution infrastructure keeps only the active layer resident, tiers activations across device, host, and disk, retires full-precision layers after export, and distributes independent output rows across tensor-parallel ranks. Together, the algorithm and infrastructure attain 99.35% and 100.84% question-weighted recovery from BF16 across seven benchmarks on Qwen3.5-397B-A17B and Llama-3.3-70B-Instruct. On the 397B model, the infrastructure reduces measured per-layer time by 15.17\times over ModelOpt and 23.14\times over LLM Compressor, with lower memory used per GPU.

[NLP-126] What Makes Recurrence Effective in Looped Language Models?

【速读】: 该论文旨在解决生成式语言模型在推理阶段通过循环机制(recurrence)提升计算深度时,其有效性边界与架构设计原则不明确的问题。具体而言,研究聚焦于:何时引入递归有助于性能提升、递归应部署于网络的何处,以及递归条件化方式对模型表现的影响。其解决方案的关键在于提出一种基于历史状态注入(history-state injection)的新机制,相较于传统的初始状态注入,该方法通过通道级的历史状态注入结合时间步条件化,显著增强了模型在不同推理预算下的鲁棒性,尤其在超出训练时长范围的推理任务中更有效地保持知识表征能力,并改善了对未充分展开(under-unrolling)情况的适应性。研究结果揭示了递归计算并非在所有场景下均有益,且仅增加计算深度不足以预测性能,强调了架构布局与条件化策略的重要性,为面向可变推理预算的循环语言模型(LoopLMs)设计提供了可操作的实践指导。

链接: https://arxiv.org/abs/2609.36636
作者: Xinlin Zhuang,Siyuan Wang,Imran Razzak,Weiyang Liu
机构: MBZUAI; The Chinese University of Hong Kong (香港中文大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: Preprint, under-review

点击查看摘要

Abstract:Looped language models (LoopLMs) increase computational depth through parameter sharing, offering a path to scale inference computation without adding parameters. However, it remains unclear when additional recurrence is beneficial and how architectural choices affect its effectiveness. Through controlled experiments, we systematically examine (1) when recurrence helps, (2) where it should be applied, and (3) how its conditioning affects performance. Our evaluation covers inference budgets below, within, and beyond the training horizon under knowledge and reasoning tasks. (1) We find that recurrence can improve reasoning beyond the training horizon while degrading knowledge performance, but harder reasoning instances do not consistently benefit more. (2) Performance also depends on how distinct layers and recurrent iterations are allocated, showing that effective depth alone is insufficient to predict behavior. Non-recurrent output layers improve robustness to under-unrolling, while the preferred placement of input and output layers varies with inference budget. (3) Finally, we find that conventional initial-state injection offers limited robustness to varying recurrence depth. We therefore propose history-state injection as an alternative, and show that channel-wise history-state injection combined with timestep conditioning offers a low-cost and more effective design, better preserving knowledge under extended unrolling while improving robustness across inference budgets. Overall, our results clarify when recurrent computation helps, where it fails, and offer practical guidelines for designing LoopLMs across variable inference budgets.

[NLP-127] Generating Edit-Inducing Questions for AI Research Manuscripts EMNLP2026

【速读】: 该论文旨在解决如何利用大语言模型(LLM)生成能够有效促进论文修改的提问,以提升学术稿件质量的问题。其核心挑战在于:尽管大语言模型在生成编辑引导型问题方面表现出更高的数量和覆盖广度,但其问题中真正能引发实质性修改的比例较低。解决方案的关键在于对比不同情境下大语言模型生成问题的效果——当模型具备完整论文上下文时,其生成的问题不仅数量更多、涵盖内容更广泛,且与更大范围的文本修改相关;然而,研究发现,正确处理长篇上下文反而可能削弱模型生成高价值问题的能力,暴露出当前推理模型在长距离依赖建模中的局限性。这一发现为优化生成式AI在复杂创作辅助任务中的应用提供了关键启示。

链接: https://arxiv.org/abs/2609.36617
作者: Sebastian Joseph,Zichao Wang,Jennifer Healey,Alexa Siu,Junyi Jessy Li,Ani Nenkova
机构: The University of Texas at Austin; Adobe Research
类目: Computation and Language (cs.CL)
备注: Accepted at the DocInsights Workshop @ EMNLP 2026

点击查看摘要

Abstract:We study the ability of LLMs to generate edit-inducing questions whose answer will improve a paper draft. On a dataset of paired submission and camera-ready papers from ICLR and NeurIPS, we compare the helpfulness of questions from GPT models with or without full paper context to that of human reviewers. GPT produces more edit-inducing questions and its questions are associated with more extensive edits and cover a broader range of edited content compared to questions from reviewers. However, a much smaller percentage of the GPT questions are edit-inducing. Our analyses confirm that automated questions can be beneficial to authors and highlight an example task where proper attending to long context deteriorates reasoning model ability to produce helpful output.

[NLP-128] Act First Reason Later: Accelerating On-Policy Distillation for Multi-Turn Agents via Reference-Conditioned Inverse Dynamics

【速读】: 该论文旨在解决在多轮语言智能体的在线蒸馏(On-policy Distillation, OPD)过程中,因标准“先思考后执行”(think-then-act)推理模式导致的环境交互延迟与经验收集效率低下问题。传统方法要求模型在每个动作前完成完整的推理过程,显著延长了推理周期,制约了训练速度。其核心解决方案是提出一种“先执行、后推理”(ActFirst-OPD)的训练框架,通过解耦环境交互与完整响应生成过程:学生模型基于当前交互上下文和参考下一观察(reference next observation),利用条件反向动力学(reference-conditioned inverse dynamics)直接推断并执行动作;当实际状态转移偏离参考轨迹时,切换至自主的下一步动作预测模式。同时,收集到的交互上下文被异步用于生成完整的“先思考后执行”式响应,以获取细粒度教师监督信号。实验表明,该方法在ALFWorld、WebShop和ScienceWorld等基准上分别实现2.3×、1.8×和4.9×的平均墙钟训练加速,且在九个基准设置中的八个达到或超过现有OPD基线的平均任务成功率,验证了推理可不必阻塞执行,从而在不牺牲性能的前提下大幅提升多轮智能体蒸馏的效率。

链接: https://arxiv.org/abs/2609.36608
作者: Zubin Zheng,Jiahao Wu,Shaofeng Zhang,Zhirui Zhang,Yew-Soon Ong,Shengcai Liu
机构: Southern University of Science and Technology (南方科技大学); Guangdong Provincial Key Laboratory of Brain-Inspired Intelligent Computation (广东省脑启发智能计算重点实验室); Hong Kong Polytechnic University (香港理工大学); Zhongguancun Academy (中关村研究院); DeepCybo; Nanyang Technological University (南洋理工大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 31 pages, 8 figures

点击查看摘要

Abstract:On-policy distillation (OPD) trains multi-turn language agents with dense teacher supervision on student-generated responses. However, standard think-then-act rollouts require lengthy reasoning before each short action, delaying environment transitions and experience collection. Generating actions directly reduces this delay but can degrade rollout quality. To address this, we propose ActFirst-OPD, an act-first, reason-later training framework that decouples environment interaction from full-response generation. The student infers and executes actions through reference-conditioned inverse dynamics using its current interaction context and a reference next observation, and switches to autonomous next-action prediction when the resulting transition deviates from the reference trajectory. From the collected interaction contexts, the student asynchronously generates full think-then-act responses for token-level teacher supervision. Experiments across 0.6B-, 1.7B-, and 4B-parameter Qwen3 students show that ActFirst-OPD achieves average wall-clock training speedups of 2.3\times on ALFWorld, 1.8\times on WebShop, and 4.9\times on ScienceWorld over Vanilla OPD. It matches or exceeds all compared OPD baselines in mean task success rate across eight of nine benchmark-model settings. These results demonstrate that reasoning need not block acting during multi-turn agent distillation.

[NLP-129] SEED: Self-Speculative Decoding via Implicit Encoder-Decoder NEURIPS2026

【速读】: 该论文旨在解决自生成式推测解码(self-speculative decoding)中高质量草案生成与计算成本之间的权衡问题。现有方法中,早期退出(early-exit)虽能降低草案生成成本,但因仅使用浅层表示而牺牲了草案质量;多标记预测(multi-token prediction, MTP)虽可保留高质量的最终隐藏状态以生成优质草案,却需在每一步草案生成时执行完整的前向传播,导致高昂开销。本文提出自生成式编码器-解码器(Self-Speculative Encoder-Decoder, SEED),其核心创新在于将标准解码器仅架构重新诠释为隐式的编码器-解码器结构:前序层作为编码器构建深层上下文表征,末尾层作为解码器基于这些表征生成标记。通过将编码与验证合并为单一步骤,验证过程完成后所获得的已验证前缀的深层上下文表征被缓存并复用于后续草案生成。因此,在两次验证之间,轻量级解码器可基于缓存的上下文表征自回归地快速生成多个标记。实验结果表明,SEED在多个基准测试中对40亿规模模型实现了高达2.7倍的平均加速,显著优于早期退出和MTP类自生成式推测基线,并且比当前最先进的EAGLE-3快28%,同时保持甚至提升了标准自回归微调的生成质量。

链接: https://arxiv.org/abs/2609.36590
作者: Hankun Lin,Patrick Pynadath,Ruqi Zhang
机构: Purdue University (普渡大学)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: Accepted to NeurIPS 2026

点击查看摘要

Abstract:Self-speculative decoding accelerates large language model (LLM) inference by drafting tokens from the target model itself, but faces a sharp tradeoff between the quality and cost of the draft. Early-exit methods produce drafts cheaply by terminating computation at intermediate layers, but forgo the deeper representations that later layers provide and thus suffer in draft quality. Multi-token prediction preserves draft quality by emitting from the model’s final hidden states, but pays for a full forward pass to produce those states at every drafting step. We propose self-speculative encoder-decoder (SEED), a self-speculative method that obtains high-quality drafts cheaply by reusing the deep contextual representations already computed during verification. We reinterpret the standard decoder-only transformer as an implicit encoder-decoder: the first layers (encoder) build deep contextual representations, and the last few layers (decoder) emit tokens from them. Encoding and verification are merged into a single step: verification is performed by the full encoder-decoder, and the contextual representations of the verified prefix are cached for reuse during drafting. Drafting is therefore very fast: between verifications, the lightweight decoder drafts multiple tokens autoregressively, each conditioned on the cached representations and on preceding drafts. Experiments across multiple benchmarks show that SEED achieves up to 2.7 \times average speedup on 4B-scale models, outperforming both early-exit and MTP-style self-speculative baselines and running 28% faster than the state-of-the-art EAGLE-3, while preserving or even improving the generation quality of standard autoregressive fine-tuning. Code is available at this https URL.

[NLP-130] ransformers Stop Thinking Too Early and a Tiny LoRA Fixes It

【速读】: 该论文旨在解决预训练变换器模型在上下文引用追踪中深度利用率不足的问题,即模型虽具备深层结构,但实际仅能有效追踪极短的引用链(1.4–3.6行),且额外预训练循环对提升效果有限。其解决方案的关键在于:在模型早期层引入一个任务训练的秩为8的低秩自适应(LoRA)模块,同时冻结其余所有模型参数,从而以极小的计算开销扩展模型的上下文推理能力。该方法通过构建“接力机制”(relay),使程序指令行能够通过中间层传递链式身份信息,使冻结的注意力头可逐步追溯更长的引用链。实验表明,Qwen3-8B模型在24行引用链上的精确率从15.5%提升至99%,而更长训练的LoRA可达50行;Ouro-1.4B模型经四轮迭代可处理60行,八轮后至少达160行。此外,移除父行注意力会中断该接力机制,验证了其核心作用。该方法还成功提升了MuSiQue任务表现,并通过冻结模型测量定位出多数模型中有效的干预层,揭示了默认回答低估了通过微小参数调整即可获取的潜在计算能力。

链接: https://arxiv.org/abs/2609.36585
作者: Zehao Jin,Ruixuan Deng,Junran Wang
机构: Georgia Institute of Technology(佐治亚理工学院)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends this computation with all model weights frozen. Qwen3-8B improves from 15.5% to 99% exact accuracy on 24-line chains; a longer-trained LoRA reaches 50 lines. Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight. The LoRA starts a relay: program lines pass on their chain identity through a short range of middle layers. Frozen heads read progressively further up the chain, and removing parent-line attention stops the relay. A frozen-model measurement locates the last useful intervention layer within tolerance in three of four held-out models. Task-specific LoRAs also improve MuSiQue. Default answers therefore understate the computation accessible through a tiny edit. Code and an interactive demo are available at this https URL

[NLP-131] Long-Term Memory-Guided Enhancement for Target Perception in Audio-Language Models

【速读】: 该论文旨在解决音频大语言模型(Audio Large Language Models, ALLMs)在真实复杂环境中因背景噪声和多重声源干扰导致的听觉感知能力退化问题。现有ALLMs在多源混合场景下难以有效聚焦目标声音,严重影响其推理性能。为此,论文提出长时记忆引导的音频增强方法(Long-Term Memory-Guided Audio Enhancement, LTM-AE),其核心解决方案是利用人类听觉中长时记忆的机制,在不更新模型参数的前提下,通过从干净参考音频中提取各声学类别对应的隐藏状态表示作为长时记忆,指导对输入音频的表征优化。具体而言,LTM-AE将输入音频令牌与对应类别长时记忆中的重构信号进行插值融合,再输入语言主干网络解码,从而在保持所有模型参数不变的情况下,动态调节已有听觉经验对当前感知的影响力。实验结果表明,该方法在20个声音类别上均显著提升目标感知能力,在三重干扰源环境下平均准确率提升达29.53至46.15个百分点;对于语音内容恢复任务,结合可学习的令牌级门控机制后,Qwen2-Audio的词错误率从23.07%降至14.77%。该工作首次探索了基于人类长时记忆原理增强ALLMs在真实听觉场景下的鲁棒性,为构建更具适应性的音频理解系统提供了新范式。

链接: https://arxiv.org/abs/2609.36577
作者: Zhenhong Zhou,Xuanyue Zhao,Youji Liu,Yuanhe Zhang,Xiaoyu Ma,Lianyu Hu,Yang Liu
机构: Nanyang Technological University (南洋理工大学); Beijing University of Posts and Telecommunications (北京邮电大学)
类目: ound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 28 pages, 5 figures, 17 tables

点击查看摘要

Abstract:Audio large language models (ALLMs) can reason about the content of audio recordings to perform complex tasks. However, these capabilities usually collapse in real-world environments when background noise and competing sources mix the target sound. Inspired by long-term memory in human listening, we propose Long-Term Memory-Guided Audio Enhancement (LTM-AE) to improve selective target perception by refining the audio representations of ALLMs without training. LTM-AE extracts representations in hidden states from separate clean reference recordings as long-term memory for each category, guiding enhancement toward a user-specified listening target. We reconstruct incoming audio tokens in the selected category long-term memory and interpolate the reconstructions with the original tokens before language backbone decoding. This interpolation controls the influence of stored auditory experience while keeping all ALLM parameters fixed. Diagnostic readouts across twenty sound categories and three ALLMs show that LTM-AE strengthens responses to a specified target amid three interfering sources. Averaged over constrained and free-form classification, accuracy gains over raw mixtures range from 29.53 to 46.15 percentage points across multiple open source models. For speech content recovery, LTM-AE with an additional learned token-level gate reduces Qwen2-Audio’s word error rate from 23.07% to 14.77%. This work takes an initial step toward using principles of human long-term memory to enhance ALLMs for real-world listening. Our code is available at this https URL

[NLP-132] Grounded Revision vs. Prior Injection: Probing Retrieval-Augmented Patent Claim Amendment AACL

【速读】: 该论文旨在解决生成式AI在专业写作中,尤其是专利权利要求修改场景下,检索增强生成(Retrieval-Augmented Generation, RAG)是否真正基于检索到的先有技术(prior art)进行语义层面的实质性修订,还是仅将检索结果作为模板进行表面性填充这一关键问题。其核心挑战在于,在具有明确“正确性”标准的专利审查情境中,缺乏对检索行为与最终生成内容之间因果关系的严谨评估。为此,研究提出三项关键贡献:一是构建了包含7,385个美国专利商标局(USPTO)审查案例的标注数据集,涵盖原始/修改后的权利要求、审查意见及引用的先有技术;二是设计了一套七探针评测体系,对比随机检索与结构匹配检索两种策略在固定提示框架下的表现差异;三是开发了一个确定性的五通道评估指标(C1-C3和C5为主,C4为补充),无需依赖大语言模型(LLM)进行判断,确保评估结果可复现且客观。在对四个前沿大模型(Claude Sonnet 4、Claude Haiku 4.5、GPT-5.4、GPT-4o-mini)进行9,600次预注册测试后发现,所有模型均未表现出显著的经典“先有技术注入”行为,检索效果微弱且方向不一致,即使在采用密集语义检索器、不同检索深度(k=1,3,5,10)以及对重述敏感的接地性度量下,零假设仍保持不变。此外,修订局部性分析揭示了模型间差异,而四象限分类法虽为探索性,但未出现“先有技术注入者”这一类别,表明当前主流RAG机制并未实现真正的基于证据的语义修正。

链接: https://arxiv.org/abs/2609.36550
作者: Josepha Michiko Leo,Hyun-seok Min,Yehoon Jang,Irvan Zidny,Jin-Woo Chung,Sungchul Choi
机构: Pukyong National University (釜庆国立大学); Tomocube Inc.; Connectionary; Teamreboott Inc.
类目: Computation and Language (cs.CL)
备注: Accepted to Findings of AACL-IJCNLP 2026. 9 pages, 2 figures. Code and data: this https URL

点击查看摘要

Abstract:Retrieval-augmented generation is widely used in professional writing, yet whether retrieval grounds revision or merely injects templates is rarely tested where “correct” has a definable meaning. Patent claim amendment supplies that signal: the examiner names the attacked limitation and cites prior art, providing per-case ground truth. We release three artifacts: (i) a corpus of 7,385 USPTO prosecution cases with XML-aligned pre/post claims, rejection, and cited prior art; (ii) a seven-probe battery comparing random and structural-match retrieval as two policies under a fixed prompt scaffold; (iii) a deterministic five-channel metric (C1-C3 and C5 in main, C4 supplementary) requiring no LLM evaluation. Across 9,600 pre-registered calls on four frontier LLMs (Claude Sonnet 4, Claude Haiku 4.5, GPT-5.4, GPT-4o-mini), no tested model exhibits detectable classical prior-injection behavior; retrieval effects are small and direction-inconsistent between random and structural retrieval, and the null is unchanged under a dense (semantic) retriever, across retrieval depths k in 1,3,5,10, and under a paraphrase-sensitive grounding metric. Revision locality reveals a model-specific difference that the template channel misses. The four-cell taxonomy, which we treat as exploratory, leaves the prior-injector cell unoccupied.

[NLP-133] DraftTrace: A Multi-View Analytics Environment for AI-Integrated Writing

【速读】: 该论文旨在解决生成式 AI(Generative AI)的广泛应用导致学生写作过程难以通过最终作业文本进行有效评估的问题。传统评价方式仅依赖最终成果,无法反映真实创作过程,尤其在面对由大语言模型(LLM)生成的内容时,难以区分自主创作与辅助生成的界限。为此,论文提出 DraftTrace,一个集成多视角记录的写作环境,其核心解决方案在于同步捕获三个互补维度:最终产出、写作过程轨迹以及与集成式 AI 助手的交互行为。通过重构文档随时间演进的动态发展路径,DraftTrace 将数据组织为提交级、纵向追踪及班级层面的分析视图,供教师使用。实验部署于 81 名研究生的自然语言处理课程中,对比了学生使用 LLM 自动生成内容与手动复制粘贴输入的行为,发现仅基于最终文本的度量指标可区分文本表述差异,而过程指标则能揭示输入行为的本质区别;二者结合可有效识别如复制打字等潜在学术不端行为。此外,交互日志显示学生在写作不同阶段对 AI 助手的使用策略存在显著差异:初期用于澄清任务要求,后期用于答案验证。初步教师调查显示,多视角写作分析及其可解释性对教学评估具有重要意义。

链接: https://arxiv.org/abs/2609.36544
作者: Divyansh Chandarana,Sandipan De,Vivek Gupta
机构: Arizona State University(亚利桑那州立大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Emerging Technologies (cs.ET)
备注: 8 pages, 7 figures, 3 tables

点击查看摘要

Abstract:Generative AI has changed how students produce writing assignments. The final artifact is no longer sufficient to understand the process through which it was produced. We introduce DraftTrace, a writing environment that jointly captures three complementary views of writing: the final product, the writing process and interactions with an integrated AI-assistant. DraftTrace reconstructs how a document develops over time and organizes these signals into submission, longitudinal, and class-level analytics for instructors. We deployed DraftTrace in a graduate NLP course with 81 students and compared their sessions with LLM-generated responses entered by automated tools and with copy-typed responses. While product measures distinguish differences in text formulation, process measures distinguish differences in how text is entered. Considering both views together helps characterize cases such as copy-typing. Interaction traces show that students use the assistant differently across stages of writing: to clarify the question at an early stage and to verify answers at a later stage. A preliminary instructor survey highlights the importance of multi-view writing analytics and their interpretability.

[NLP-134] When Updating Stops Being Learning: Rethinking LLM Self-Evolution via learnable information gain

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在自演化(self-evolution)过程中普遍存在的自演化退化问题,即模型性能在经历初期提升、平台期后出现显著下降的现象。现有方法通常仅从组件层面进行优化,如针对提问生成器(Questioner)或求解器(Solver)进行改进,但忽视了自演化过程本身是一个高度耦合的系统。本文提出一种基于可学习信息增益(learnable information gain)的全局性框架,其核心在于量化每轮自演化中新增的、可参数化的有效信息量相对于前一轮的增量。理论上,该信息增益等于两轮数据分布之间的Kullback-Leibler散度与熵变化之和;实践中,通过拟合一个小规模语言模型对前一轮数据建模,并利用负对数似然得分评估新生成数据的信息价值。基于此诊断机制,研究进一步设计了ATRI(Adaptive Training Regulation via Information-gain)策略,通过动态重加权单轮内样本并当信息增益持续偏低时终止跨轮训练,从而有效抑制退化。实验结果表明,该方法在多个主流数据集上均显著优于现有基准。

链接: https://arxiv.org/abs/2609.36535
作者: Chenxu Wang,Chaozhuo Li,Xinze Shi,Songyang Liu,Kyrie You Wu,Ziluowen Luo,Shun Zhang,Chenxi Li,Litian Zhang
机构: Beijing University of Posts and Telecommunications(北京邮电大学); Beijing Academy of Artificial Intelligence(北京人工智能研究院); Central South University(中南大学); Graduate School of China Academy of Engineering Physics(中国工程物理研究院研究生院); The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Self-evolution lets large language models (LLMs) improve iteratively using their own generated data, but often suffers from self-evolution degeneration: performance improves, plateaus, then declines. Existing methods address this issue at the component level, targeting either the Questioner or the Solver, and overlook that self-evolution is a tightly coupled system. We propose a holistic framework based on learnable information gain, which measures how much novel, parameterizable information a round provides relative to the previous round. Theoretically, this gain equals the Kullback-Leibler divergence between the two rounds’ data distributions plus their entropy change. Practically, it is estimated by fitting a small language model to the previous round and scoring new data via negative log-likelihood. Based on this diagnostic, we propose ATRI (Adaptive Training Regulation via Information-gain), which reweights samples within a round and halts training across rounds when information gain remains low. Experiments on popular datasets demonstrate the superiority of our proposal.

[NLP-135] riadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling

【速读】: 该论文旨在解决传统循环神经网络(RNN)在处理长序列时因固定大小记忆状态导致的上下文容量受限问题,尤其针对线性注意力(linear attention)虽能扩展隐藏状态维度但参数效率与表达能力仍存局限的挑战。其核心解决方案是提出三元线性注意力(triadic linear attention),通过将键(key)、第二个键(second key)和值(value)的三重外积(triadic outer product)写入三维张量状态(third-order tensor state),并在读取时通过对两个查询分别与两个键轴进行收缩操作实现信息检索。该方法在不显著增加参数量的前提下,使状态维度随第二键的维度呈线性扩展(例如,E维第二键带来E倍的状态容量提升),同时仅需额外两个投影层。三元线性注意力具备与数据依赖遗忘、δ规则及分块并行训练兼容的优势,在门控ΔNet(Gated DeltaNet)和标量门控线性注意力中应用后,显著提升了长序列语言建模与记忆召回性能,优于其他通过扩大状态维度实现改进的方法。

链接: https://arxiv.org/abs/2609.36529
作者: Oliver Sieberling,Bharat Runwal,David Jin,Ryan Chin,Rameswar Panda,Yoon Kim
机构: Massachusetts Institute of Technology (麻省理工学院); MIT-IBM Computing Research Lab
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: Preprint

点击查看摘要

Abstract:Recurrent neural networks (RNNs) compress the historical context into a memory state of fixed size, thus allowing for constant-time inference. The memory state size is a crucial factor in their performance, as exemplified by the strong performance and resurgence of linear attention, which extends the vector-valued hidden states of ordinary RNNs to matrix-valued hidden states. Crucially, linear attention does so in a parameter-efficient way, in particular by using an outer product of the key and value vectors to write to the matrix-valued hidden state. We generalize this construction and propose triadic linear attention, which writes the triadic outer product of a key, a second key, and a value, into a third-order (i.e., 3D) tensor state, and reads from it by contracting both key axes with two queries. An E -dimensional second key thus yields an E -fold increase in state size while adding only two projections. Triadic linear attention is compatible with data-dependent forgetting, the delta rule, and chunkwise-parallel training. Applied to Gated DeltaNet and scalar-gated linear attention, triadic linear attention substantially improves long-context language modeling and recall, outperforming alternatives that enlarge the state.

[NLP-136] Adapting Context Compression for Long-Horizon Agents with Counterfactual Continuations

【速读】: 该论文旨在解决长时程智能体(long-horizon agents)在持续交互过程中因历史信息累积导致的上下文过载问题,核心挑战在于如何高效且可靠地进行上下文压缩(context compression),以维持智能体在长期任务中的执行性能。现有提示适配(prompt-adaptation)方法依赖全上下文与压缩后轨迹的对比来推断压缩误差,但此类方法无法区分单个压缩事件的影响,并受到智能体随机性(agent stochasticity)的干扰。研究发现,压缩对可靠性(reliability)的损害先于对可解性(solvability)的影响,且严重性能退化集中于特定压缩事件。针对此问题,论文提出干预式回溯提示适配方法(PAIR, Prompt Adaptation using Interventional Rollouts),通过匹配反事实延续(matched counterfactual continuations)技术,从同一智能体状态出发,对比有无压缩条件下的后续执行表现,精准识别并诊断有害压缩事件的影响,进而动态优化固定压缩模板中相关部分。其关键创新在于将因果干预思想引入提示压缩的自适应过程,实现对压缩质量的细粒度评估与修正。实验表明,PAIR在所有主要基准场景组合中均实现了压缩方法中最强的跨运行可靠性,显著优于现有提示适配基线;且无需修改下游智能体,即可使压缩执行性能接近甚至在数值上超越无压缩基线。

链接: https://arxiv.org/abs/2609.36526
作者: Guanghui Min,Liang Wu,Mingjia Shi,Yinhan He,Mayank Darbari,Liangjie Hong,Chen Chen
机构: University of Virginia(弗吉尼亚大学); Nokia(诺基亚)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 41 pages, 10 figures, 9 tables

点击查看摘要

Abstract:Long-horizon agents require context compression to manage growing interaction histories. Compression quality, however, is ultimately determined by downstream execution. Existing prompt-adaptation methods infer compression errors by comparing full-context and compressed trajectories. Such comparisons cannot isolate individual compressions and are confounded by agent stochasticity. We first find that compression degrades reliability before solvability. Using matched counterfactual continuations that compare execution from the same agent state with versus without compression, we further show that severe degradation concentrates at isolated compression events. Motivated by this finding, we propose PAIR (Prompt Adaptation using Interventional Rollouts) for adapting structured compression prompts. PAIR identifies individual compressions that degrade subsequent execution, diagnoses their effects, and revises the relevant sections of a fixed compression template. PAIR achieves the strongest cross-run reliability among compressed methods in every main benchmark-scope combination, consistently exceeding the competing prompt-adaptation baseline. Without modifying the downstream agent, PAIR brings compressed execution close to the no-compression baseline and sometimes numerically exceeds it.

[NLP-137] Large-scale factor analysis shows machine intelligence is only partially interpretable

【速读】: 该论文旨在解决语言模型(Language Model)开发中一个长期存在的假设——即认知能力围绕一种通用的、跨领域的智能因子(g factor)组织,类似于人类的流体智力(fluid intelligence)。这一假设虽被广泛接受,但缺乏大规模实证检验。本文采用潜在变量方法,借鉴心理测量学对心理构念的研究范式,将模型在具体任务集上的表现分解为领域特定与领域无关的潜在因素。通过因子分析这一降维技术,研究分析了涵盖1,618个语言模型、456个纯文本基准测试的13,251项公开评估分数。鉴于数据集极度稀疏,研究通过多种数据填充与插补方法进行三角验证。结果显示:1)在最乐观估计下,通用智能因子可解释模型性能70.8%的方差,但在多数方案中占比远低于此;2)内容相似的基准测试并不必然聚类;3)g因子并非由某一共同主题主导,且现有标准“智力”基准测试无法有效逼近该因子。这些发现挑战了当前将通用智能作为可定义、可识别和可靶向的实体来指导语言模型发展的主流策略,表明试图通过单一概念性能力实现通用智能在实践中缺乏理论与实证支持,因其所需的一阶能力往往具有部分特异性且难以明确界定。

链接: https://arxiv.org/abs/2609.36515
作者: Faiz Ghifari Haznitrama,Afrizal Hasbi Azizy,Faeyza Rishad Ardi
机构: KAIST(韩国科学技术院); Independent Researcher(独立研究员)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Neurons and Cognition (q-bio.NC)
备注: 66 pages

点击查看摘要

Abstract:A common assumption in language model development is that cognitive abilities are organized around a general, domain-free intelligence factor, like fluid intelligence in humans. This assumption is rarely tested directly, and prior attempts have done so only at a much smaller scale. We take a latent variable approach to intelligence in language models, similar to how psychometricians study psychological constructs. Performance in every specific problem set is influenced by a domain-specific and a domain-agnostic latent factor. Using factor analysis as a dimension-reduction technique, we analyzed 13,251 published evaluation scores covering 1,618 language models across 456 different text-only benchmarks. Due to the super-sparse nature of the dataset, we triangulate our analysis across different data densifiers and imputation methods. A robust pattern across different modes of bias is that 1. A general intelligence factor accounts for 70.8% of variance in model performance at our most generous estimate, and far less than that in most of our solutions, 2. Content-similar benchmarks do not necessarily cluster together, and 3. The g factor is not dominated by any common theme, and there is a lack of evidence that it is well-proxied by standard “intelligence” benchmarks. Our findings go against current endeavors of defining, identifying, and targeting general intelligence as a tangible construct in language model development. This leaves the strategy of targeting a single conceptual ability without support, since the first-order abilities it would have to reach are often partially idiosyncratic and not identifiable in practice.

[NLP-138] Similar Choices Different Attention: Cross-Modal Associations in Humans and Vision-Language Models

【速读】: 该论文旨在解决生成式视觉-语言模型(VLMs)在跨模态关联任务中是否真正实现与人类一致的认知机制这一问题,尤其关注模型决策过程中的注意力分布是否与人类的注视行为相匹配。其核心挑战在于,以往研究常因使用不同刺激或任务导致人类与模型之间的比较失真;为此,本文采用相同实验范式,同时记录人类(N = 53)和大型/小型VLMs在面对伪词与图像配对时的选择行为及眼动轨迹,以实现直接对比。研究发现,尽管少数大型VLMs在选择上表现出与人类一致的倾向,但其显著性图谱与人类注视模式的相关性反而低于仅基于中心偏倚(center-bias)的固定高斯基线;通过微调小规模VLMs使其在未见样本上的选择接近人类多数投票水平后,其注意力仍无法超越该基线;进一步地,直接以人类注视数据训练模型注意力虽提升了注意力与注视的相关性,却未改善选择一致性,而仅用单个平均注视图即可达到类似提升效果。这表明,单纯匹配人类选择或注视模式,并不能充分证明模型具备与人类相似的跨模态认知处理机制,揭示了当前评估多模态对齐的局限性,强调需从更深层次的动态注意力机制出发,综合考察模型内部表示与人类认知过程的一致性。

链接: https://arxiv.org/abs/2609.36475
作者: Sumin Hong,Katsumi Ibaraki,Renee Shi,David Chiang,Toby Jia-Jun Li
机构: University of Notre Dame(圣母大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注: 9 pages

点击查看摘要

Abstract:Cross-modal associations are systematic pairings of features across modalities, such as the association of ‘bouba’ with round shapes and ‘kiki’ with sharp shapes. Prior work has compared humans and vision-language models (VLMs) on such associations, but often using different stimuli or tasks between humans and models. Here, we ask whether VLMs align with humans not only in choices, but also in where they look when making those choices. We study both VLMs and humans (N = 53), presenting them with the same stimuli, a pseudo-word and two images, and record participants’ choices and eye movements, which we release. We find choice alignment in a few larger VLMs, but their saliency matches human gaze less closely than a center-bias baseline, a fixed Gaussian at the center of each image. Fine-tuning small VLMs on human choices brings their choice alignment to the level of a human majority-vote reference on unseen words and images, yet their attention still matches human gaze less closely than this baseline. Training model attention on human gaze raises attention-gaze correlation without improving choice alignment, and a single average gaze map per image position raises it by a similar amount. Matching human choices, or even human gaze patterns, is therefore not sufficient evidence of human-aligned cross-modal processing.

[NLP-139] FinRT: Distilling Adaptive Red-Teaming Strategies into Reusable Adversarial Generators in Consumer Finance

【速读】: 该论文旨在解决在消费金融等受监管行业中,大型语言模型(Large Language Model, LLM)因用户查询的细微诱导而暴露安全漏洞的问题,此类攻击可能导致模型输出突破政策限制,引发严重合规风险。现有自动化红队测试方法在攻击有效性与生成成本之间存在权衡,且将覆盖范围、攻击严重程度和多样性视为独立目标而非协同优化的统一目标。本文提出FinRT,一种结构化框架,通过自适应红队策略构建可复用的对抗性提示生成器,将针对特定目标模型的攻击生成过程转化为可迁移的通用生成机制。其核心创新在于:利用动态调整的红队策略训练出具备高泛化能力的生成器,显著提升攻击成功率(较自适应基线Rainbow Teaming提升近一倍,达32.9% vs. 17.2%),同时将最大对抗性严重程度提高33%,并保持与迭代搜索方法相当的域内语义多样性。此外,FinRT展现出优异的跨模型迁移能力,并揭示了不同模型家族间的独特攻击特异性模式。

链接: https://arxiv.org/abs/2609.36474
作者: Rikhiya Ghosh,Himanshu Kumar,Sriram Venkatapathy,Sahil Wadhwa,Alexandre G.R. Day,Pranab Mohanty
机构: Capital One(资本壹号); AI Foundations(人工智能基础)
类目: Computation and Language (cs.CL); Cryptography and Security (cs.CR); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:In regulated industries like consumer finance, seemingly harmless user queries can exploit large language model vulnerabilities, triggering safety failures and pushing responses dangerously close to policy limits. Existing automated red-teaming methods trade off attack effectiveness against generation cost, while treating coverage, severity, and diversity as incidental rather than joint objectives. We introduce FinRT, a structured framework that builds reusable adversarial prompt generators from adaptive red-teaming strategies. Across the six victim models in consumer finance, FinRT substantially outperforms adaptive search baselines while amortizing target-facing attack generation into a reusable generator. FinRT nearly doubles the attack success rate over the adaptive baseline Rainbow Teaming (32.9% vs. 17.2%), increases maximum adversarial severity by 33%, and preserves comparable intra-policy-domain semantic diversity to iterative search methods. Our method achieves high cross-model transferability while exhibiting distinct victim-family specialization patterns.

[NLP-140] Fisher-IRG: Fisher-Induced Local Invariant Representation Geometry across Language and Vision Models

【速读】: 该论文旨在解决在深度学习模型中,如何准确刻画局部表示空间内语义相关变化与无关扰动之间差异的核心问题,即在何种局部度量下能够有效区分具有语义意义的表示变化与仅影响预测但不改变语义的微小扰动。其解决方案的关键在于提出一种基于费雪信息(Fisher Information)的不变表示几何结构(Fisher-induced Invariant Representation Geometry, Fisher-IRG),通过分析每个表示点周围的预测敏感性,构建语义保持与语义改变的邻域,并利用局部费雪信息的对比聚合,求解广义特征值问题以恢复对语义变化具有不变性的方向。该方法不仅在语言和视觉模型中表现出更强的语义-噪声预测选择性及更高的子空间可重复性,且在表示干预和跨域任务中验证了其泛化能力,证明了其作为刻画局部不变表示几何的理论框架的有效性。

链接: https://arxiv.org/abs/2609.36458
作者: Abdullah All Tanvir,Xin Zhong
机构: University of Nebraska Omaha(内布拉斯加大学奥马哈分校)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Semantic-preserving transformations can induce substantial motion in learned representations, while small changes may strongly affect model predictions, raising a basic question: what local metric best captures semantically consequential variation? We propose Fisher-induced invariant representation geometry (Fisher-IRG), which measures local representation directions through their predictive sensitivity. Around each representation, we construct semantic-preserving and semantic-changing neighborhoods, aggregate their local Fisher information, and recover invariant directions through a contrastive generalized eigenvalue problem. Controlled displacement analyses first show that comparable Euclidean motion can have substantially different predictive consequences, supporting the need for a predictive geometry. Across language and vision models, Fisher-IRG yields stronger semantic-versus-nuisance predictive selectivity and generally more reproducible subspaces than covariance-based geometry, while recovering systematically distinct local directions. Representation interventions further localize semantic effects to the Fisher-derived subspace, and held-out separation and retrieval show that the recovered geometry generalizes beyond the discovery neighborhoods. These results support Fisher-IRG as a principled framework for characterizing local invariant representation geometry.

[NLP-141] Memory Consolidation Flattens the Temporal Shape of User Facts

【速读】: 该论文旨在解决长期记忆系统在对话信息压缩过程中导致的时态信息丢失问题,即“时态扁平化(aspectual flattening)”——在将动态对话内容转化为短时记忆条目时,系统倾向于忽略动作的进行时态(如“我正在驾驶一辆标致”)而仅保留一般现在时(如“用户驾驶一辆标致”),从而丢失关于事实是否仍有效的关键时间线索。其解决方案的关键在于提出并验证LAPSE基准测试,通过成对匹配的用户陈述对比不同时间形式下的语义保持程度,量化记忆系统中时态信息的失真情况。研究发现,主流生成式记忆写作模型在处理进行时与一般现在时之间存在显著不对称性:几乎所有模型均倾向于将进行时转换为一般现在时,却极少反向操作;这种时态扁平化现象在11种模型配置及多个实际部署管道(mem0、Graphiti、Letta)中普遍存在。更重要的是,实验表明被扁平化的时态线索对后续阅读者判断事实有效性具有实质性影响——仅改变存储的动词形式即可显著改变阅读者的信念估计值;当允许读者向用户确认时,两个出错率较高的阅读器更倾向于基于扁平化笔记直接行动,而非主动核实。因此,该研究揭示了当前记忆系统在设计上可能无意中削弱了下游模型对事实时效性的判断能力,凸显了在记忆编码阶段保留动态时态信息的重要性。

链接: https://arxiv.org/abs/2609.36457
作者: Sugam Panthi,Muhaiminul Yeamin,Siyan Luo,Rabab Abdelfattah
机构: 未知
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Long-term memory systems turn conversations into short stored notes. A note can keep a user fact while losing evidence about whether the fact still holds. For example, “I am driving a Peugeot” can become “The user drives a Peugeot,” which drops the cue that the activity is ongoing. We call this aspectual flattening and measure it with LAPSE, a benchmark of matched user statements that differ only in temporal form. We find that memory writers flatten aspect selectively. Three writer models flattened the progressive statement but kept its simple-present match in 244 of 381 pairs, never the reverse. The asymmetry holds in all 11 model configurations tested and in the installed pipelines mem0, Graphiti, and Letta. The lost cue matters to later readers. In exploratory tests, changing only the stored verb shifted all three readers’ estimates that a fact still holds. When readers could ask the user before acting, two of three acted without asking more often on flattened notes. Our planned memory-use task could not detect this, because readers there acted on almost every stored fact, even expired ones. Memory writing can thus remove evidence that later models use to decide whether to act.

[NLP-142] Reliable Parallel Decoding in Masked Diffusion Language Models

【速读】: 该论文旨在解决生成式AI(Generative AI)中基于掩码扩散语言模型(Masked Diffusion Language Models, MDLMs)在并行解码时存在的可靠性问题:尽管多标记并行预测可显著提升生成效率,但同一前向传播中各预测结果的可靠性不一致,尤其当高置信度预测出现在序列末端时,可能在上游计算尚未充分建立的情况下提前“固定”输出,导致下游预测因上下文不确定性增加而不可靠。其核心解决方案在于提出一种无需训练的可靠并行解码方法(Reliable Parallel Decoding, RPD),关键在于通过分层预测稳定性与最终置信度联合筛选候选词,并基于其前置掩码位置的累积熵预算动态决定提交顺序——优先保留并行提交那些在深层网络中表现稳定的高置信度预测,同时延迟处理上游上下文不确定的预测。该机制避免了依赖预设块调度,实现了在保持或提升准确率的前提下,显著提高解码吞吐量,在LLaDA和Dream等数学推理与代码生成基准上均达到最优性能。

链接: https://arxiv.org/abs/2609.36452
作者: Zhenghao He,Bohan Liu,Guangzhi Xiong,Aidong Zhang
机构: University of Virginia (弗吉尼亚大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Masked diffusion language models (MDLMs) can generate text efficiently by predicting multiple masked tokens in parallel, but predictions from the same forward pass are not necessarily reliable when committed together. We study when parallel commitment is reliable. Our diagnostics show that confidence alone does not determine a reliable commitment order: confident predictions near the end of the sequence can fix an answer before its supporting computations are established, and downstream predictions become less reliable as the uncertainty of their upstream context grows. At the same time, a single forward pass can already resolve several masked tokens, and predictions that remain stable across the final layers are more likely to be correct. Based on these findings, we propose Reliable Parallel Decoding (RPD), a training-free method that selects candidates by layerwise prediction stability and final confidence, and commits them under a cumulative entropy budget over their preceding masked positions. RPD defers predictions with uncertain upstream context while committing the remaining candidates in parallel, without relying on a fixed block schedule. Across mathematical reasoning and code generation benchmarks on LLaDA and Dream, RPD achieves the highest decoding throughput among the evaluated methods while maintaining or improving accuracy.

[NLP-143] Invariant Atoms: Sparse Coordinates of Local Semantic Geometry in Language Model Representations

【速读】: 该论文旨在解决大语言模型中语义保持与表层形式变化之间的解耦问题,即如何在不改变语义的前提下,识别并建模由词汇、风格和句法等非语义因素引起的隐藏表示的系统性扰动。其核心挑战在于捕捉局部语义空间中的稳定几何结构,以区分语义相关变化与无关的“干扰”(nuisance)变化。解决方案的关键是提出不变原子假设(Invariant Atom Hypothesis):局部语义变换可被分解为一组在语义保持变换下保持稳定的稀疏坐标方向(即“不变原子”),这些方向构成一个共享的语义框架。通过学习这一共享框架与稀疏坐标,并引入依赖锚点的对角调制机制来动态调整原子权重,该方法实现了对语义位移的精准重建,同时抑制了非语义干扰。实证结果表明,所学原子具有强语义-干扰分离性、稀疏重构能力、可复现的方向性以及对模型预测的因果影响;其几何结构可泛化至未见的语义邻域与干扰类型,且局部重加权提升了语义选择性,维持了全局到局部的一致结构。此外,原子特征在模型微调后仍保持稳定,验证了其作为语言模型中局部语义几何的可复用稀疏坐标系统的可行性。

链接: https://arxiv.org/abs/2609.36451
作者: Muhammad Ahtesham,Xin Zhong
机构: University of Nebraska Omaha(内布拉斯加大学奥马哈分校)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Large language models often preserve meaning despite substantial changes in wording, style, and syntax, while small semantic edits can systematically alter their hidden representations. This suggests that semantic variation may be organized along recurring local directions. We propose the Invariant Atom Hypothesis: local semantic motion admits preferred sparse coordinates along directions that remain stable under meaning-preserving transformations. We learn a shared semantic frame and sparse coordinates that reconstruct semantic displacements while suppressing nuisance variation, with anchor-dependent diagonal modulation adjusting atom strengths without sample-specific rotations. Empirically, the atoms exhibit strong semantic–nuisance separation, sparse reconstruction, reproducible directions, and causal effects on model predictions. The learned geometry generalizes to unseen semantic neighborhoods and nuisance families, while local reweighting improves semantic selectivity and preserves a consistent global-to-local structure. Atom signatures also remain stable under model modification. These findings support reusable invariant directions as a sparse coordinate system for local semantic geometry in language models.

[NLP-144] MemFold: Learning Compact Soft Memory for Long-Context Personalization via On-Policy Optimization

【速读】: 该论文旨在解决长期对话系统中如何有效利用用户历史信息以实现持续个性化响应的问题,核心挑战在于:在不显著增加推理负担的前提下,如何高效地存储、更新并动态使用用户随时间演变的偏好、修正与约束。现有方法通常将记忆保留与行为执行分开优化,或采用文本形式存储导致输入长度膨胀,或压缩为固定维度向量但训练目标(如重建原文本或模仿参考答案)与实际用户生成序列脱节,难以保证真实场景下的表现。本文提出MemFold,其关键创新在于通过行为驱动的软记忆优化机制——将查询相关的文本记忆压缩为固定数量的连续向量作为阅读器的接口,并在训练中直接基于该记忆支持的任务结果组相对奖励(group-relative rewards)以及置信度门控的在线策略蒸馏(confidence-gated on-policy distillation)进行联合优化。其中,冻结的文本记忆教师模型仅用于重评分学生模型采样出的词元,确保监督始终作用于学生当前分布,且无需自回归解码;推理时完全移除教师,保持高效性。实验表明,MemFold在PersonaMem-32K和PersonaMem-128K上均取得最高准确率,且随着历史长度增加优势更明显,并可在未进行目标领域微调的情况下迁移到PrefEval和LongMemEval任务。消融实验显示,主要性能提升来自奖励项,教师信号带来额外增益;记忆干预分析进一步证实阅读器高度依赖其软记忆中的实例化内容。

链接: https://arxiv.org/abs/2609.36435
作者: Jingxuan Wu,Yuzhe Yang,Yiqiao Huang,Chengzhi Liu,Qingni Wang,Chengxuan Qian,Shutong Wu,Jiawei Zhang,Xin Eric Wang
机构: 未知
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:An assistant that serves the same user over a long horizon has to answer from what that user has revealed: which preferences still hold, which were revised, and which constraints apply now. Retaining that information is not the same as acting on it, and the two are usually optimized as if they were. Keeping the information as text makes the reader’s input grow with the retained history, while compressing it into a fixed number of latent vectors bounds the interface but is typically trained to reconstruct text or imitate reference answers, both of which are scored on sequences the reader never produced. We present MemFold, which optimizes a fixed-budget soft memory by the behavior it supports. A query-conditioned textual memory is compressed into K continuous vectors that form the reader’s memory interface, and the reader is then trained on its own rollouts under two complementary signals: group-relative rewards for task outcomes, and confidence-gated on-policy distillation in which a frozen textual-memory teacher re-scores the student’s sampled tokens under the textual memory. The teacher is never sampled from, so supervision stays on the student’s current distribution and adds no autoregressive decoding; at inference it is removed entirely. Across three Qwen backbones, MemFold attains the highest accuracy we measure on PersonaMem-32K and PersonaMem-128K, with margins that widen at the longer history length, and transfers to PrefEval and LongMemEval without target-domain training. Ablations attribute most of the task gain to the reward term and a smaller additional gain to the teacher signal, and memory interventions show that the reader depends on the instance-specific content of its soft memory.

[NLP-145] Eternal Sunshine of the Spotless Mind: Systematically Erasing LLM s Memories

【速读】: 该论文旨在解决持续型大语言模型(Persistent LLMs)在用户请求删除个人信息时无法真正实现数据清除的问题。尽管模型可能声称已“遗忘”相关信息,且在有限上下文窗口下运行,但其外部存储中的记忆仍可能通过消息依赖关系隐性留存,导致隐私泄露风险。现有方法简单移除匹配删除请求的对话内容不足以彻底消除信息,因为对话中存在语义与结构上的依赖关系,使得被删除信息仍可通过上下文推断恢复。为此,论文提出DeLLM框架,其核心创新在于:动态构建每次查询的相关上下文,并维护一个消息溯源图(provenance graph),以精确追踪信息传播路径,从而识别并清除所有受删除请求影响的关联记忆。实验表明,DeLLM能够在保障模型响应质量的同时实现高精度的记忆删除,为实现可信赖的用户隐私保护提供了有效解决方案。

链接: https://arxiv.org/abs/2609.36414
作者: Olga Ohrimenko
机构: The University of Melbourne(墨尔本大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:We consider persistent LLMs that accumulate memories of their interactions with a user over time. Such LLMs maintain memories using external storage, which they can query to overcome the limitations of a fixed context window. Such systems have numerous practical applications, as they can draw on all past interactions when responding to user queries. In this paper, we ask whether LLMs can forget information shared with them upon a user’s request. We find that current LLMs fail to delete such information—even when they claim to have forgotten it and even when operating with a limited context. To this end, we consider a new direction of study: Deletion of LLM Memories. We show that naively removing messages that match a user’s deletion request is insufficient, since conversations naturally introduce message dependencies that cause information to persist. To correctly handle deletion requests, we propose the DeLLM framework. It dynamically constructs relevant context for each LLM query and maintains a provenance graph of messages to determine which ones must be removed during deletion. Our experiments show that DeLLM achieves a high deletion rate while maintaining utility. Subjects: Computation and Language (cs.CL) Cite as: arXiv:2609.36414 [cs.CL] (or arXiv:2609.36414v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2609.36414 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[NLP-146] Calibrated to Whom? Persona and Language Effects on Cultural Values in JEV

【速读】: 该论文旨在解决生成式 AI 在文化价值观评估中是否存在文化偏见及一致性的问题,特别是针对仅输出选项概率的决策型语言模型(decision-only language models)在跨文化情境下的表现。其核心解决方案在于通过结构化问卷(Values Survey Module 2013)对模型 JEV 进行系统性审计,以评估其在不同文化身份设定(沙特与美国人格)、语言(英语与阿拉伯语)及请求表述方式下的响应一致性与文化适配性。关键发现为:尽管模型未生成文本,其概率输出表现出极高的可重复性(组内相关系数 ICC = 0.997),且在无角色设定时表现出与自身美国人格一致的倾向;当赋予沙特人格时,其响应模式部分再现了人类沙特与美国群体间的文化差异(英语中复现87%,阿拉伯语中62%),但长期导向维度出现反转。进一步分析表明,阿拉伯语版本差异较小主要源于题目语言而非人格描述语言,且年龄与性别对模型响应的影响程度接近甚至超过国籍,同时模型在阿拉伯语及沙特人格情境下表现出更低的信心。这些模式在所有请求设计中均保持一致,凸显了决策型模型在文化价值判断中的潜在系统性偏差。

链接: https://arxiv.org/abs/2609.36399
作者: Bushra Asseri,Abdulaziz Asseri
机构: Alfaisal University (阿尔法伊萨尔大学); Proxa.sa (Proxa.sa)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Decision-only language models return a probability for every answer option instead of generating text, which makes them attractive as survey respondents and as judges. We audit the cultural values of one such model, TypeSafe’s JEV, with the Values Survey Module 2013. We asked it the 24 items as 12 matched Saudi and 12 matched American personas and without a persona, in English and Arabic, under eight ways of formulating the request (288,000 answers). JEV’s answers were highly repeatable (ICC 0.997), and without a persona they resembled those of its own American personas. When the persona was Saudi rather than American, the answers moved in the direction of the human Saudi-US difference, reproducing 87% of its size in English but 62% in Arabic, with long-term orientation reversed. A language cross shows that the smaller difference in Arabic comes from the language of the items, not from the language of the persona description. Age shifted the profiles about as much as nationality, gender shifted them more for Saudi than for American personas, and JEV was less confident in Arabic and for Saudi personas. These patterns held in every request design, although the model never generates text.

[NLP-147] Mark: Pairwise Distortion-Free Watermarking Beyond Single-Token Entropy

【速读】: 该论文旨在解决现有无失真水印(distortion-free watermarking)方法在机器生成文本归属追踪中检测能力受限的问题。现有方法对每个生成的词元(token)独立操作,其检测性能受制于下一个词元分布的熵(entropy),导致水印容量和鲁棒性不足。为此,论文提出一种通用的成对水印框架——Tandem Token WaterMark(TTMARK),将无失真水印从单个词元扩展至相邻词元对(adjacent token pairs),通过建模连续词元的联合分布(joint distribution),使有效水印字母表从 VV 扩展至 V2V^2,从而在保持生成输出分布不变的前提下,同时利用词元熵与条件熵信息,显著提升检测强度。其核心解决方案在于引入一种分支隔离的拼接式成对生成算法(branch-isolating concatenated tandem generation algorithm),可在一次前向传播中高效构建联合分布。理论分析表明,在低熵场景下,成对水印可实现更高的期望检测强度。大量实验验证了TTMARK在多种语言模型、数据集及主流无失真水印方案上的有效性,不仅持续提升检测率且不损害生成质量,还增强了对编辑攻击的鲁棒性,并显著改善局部化水印检测性能。

链接: https://arxiv.org/abs/2609.36372
作者: Ruibo Chen,Zhengmian Hu,Donghang Lu,Xuehao Cui,Georgios Milis,Yihan Wu,Jian Du,Heng Huang
机构: University of Maryland, College Park (马里兰大学学院市分校); TikTok(字节跳动)
类目: Cryptography and Security (cs.CR); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Distortion-free watermarking enables reliable attribution of machine-generated text while preserving output distribution. However, existing methods operate independently on each generated token, making their detection capability fundamentally constrained by the entropy of the next-token distribution. We present Tandem Token WaterMark (TTMARK), a general pairwise watermarking framework that extends distortion-free watermarking from individual tokens to adjacent token pairs. By watermarking the joint distribution of consecutive tokens, TTMARK enlarges the effective watermarking alphabet from V to V^2 , allowing the detector to exploit both token entropy and conditional entropy while preserving distortion-freeness over the joint distribution. We further introduce a branch-isolating concatenated tandem generation algorithm that efficiently constructs the joint distribution in a single forward pass. Theoretically, we show that pairwise watermarking achieves better expected detection strength in low-entropy regimes. Extensive experiments across multiple language models, datasets, and three representative distortion-free watermarking schemes demonstrate that TTMARK consistently improves detectability without degrading generation quality, while also improving robustness to edits and substantially enhancing localized watermark detection.

[NLP-148] DeepRewind: Predicting and Repairing Premature Commitments in Deep Research Agents

【速读】: 该论文旨在解决深度研究代理(deep-research agents)在长时程探究中因过早承诺(premature commitments)而导致的错误信念固化问题。在迭代式搜索、证据评估、信念修正与综合过程中,代理可能在缺乏充分证据的情况下提前确立结论,进而使后续推理强化错误解释,影响最终决策质量。其解决方案的关键在于提出DeepRewind——一种可叠加的可逆深度研究控制层,通过构建一个类型化的图结构来动态表征代理的认知状态(epistemic state),包含来源、证据、主张、假设、前提、承诺、计划及草稿等元素。该系统引入基于提示的世界模型(prompt-based world model),在采纳中间结论前预测其潜在影响,并基于假设收敛度、信息损失、恢复成本及矛盾触发覆盖率等指标评估可逆性;同时配备二元控制器以阻断高风险承诺,并由依赖感知的回滚机制在后续证据推翻先前结论时执行一致性修复。实验表明,在DRBench和LiveDRBench基准上,DeepRewind相较Open Deep Research显著提升了3.6个百分点的洞察召回率,且将过早承诺减少了59.1%。

链接: https://arxiv.org/abs/2609.36344
作者: Amirhossein Abaskohi,Amirhossein Dabiriaghdam,Lele Wang,Peter West,Giuseppe Carenini
机构: University of British Columbia(不列颠哥伦比亚大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Deep-research agents conduct long-horizon investigations through iterative search, evidence evaluation, belief revision, and synthesis. However, they may commit to claims before sufficient evidence is available, causing later reasoning to reinforce an incorrect interpretation. We introduce DeepRewind, an additive control layer for reversible deep research that represents the agent’s evolving epistemic state as a typed graph of sources, evidence, claims, hypotheses, assumptions, commitments, plans, and drafts. Before accepting an intermediate conclusion, a prompt-based world model predicts its impact and estimates reversibility based on hypothesis narrowing, information loss, recovery cost, and contradiction-trigger coverage. A binary controller blocks risky commitments, while a consistency monitor performs dependency-aware rollback when later evidence invalidates them. Across DRBench and LiveDRBench, DeepRewind improves insight recall by 3.6 percentage points and reduces premature commitments by 59.1% relative to Open Deep Research.

[NLP-149] raining LLM s to Verbalize Evaluation Awareness

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在评估期间表现出评估意识(Evaluation Awareness, EA)导致其行为在审计阶段与实际部署阶段不一致的问题。现有方法难以准确测量和建模评估意识,且易受训练信号干扰。本文提出**语义化训练(Verbalization Training, VT)**作为解决方案,其核心在于通过利用模型自发产生的语义表达(如自我陈述“我正在被评估”)作为评估意识存在的证据,而非直接监督模型内部的潜在信念。具体而言,VT在模型生成序列中识别出此类语义化表达,并在表达前截断生成过程,形成训练前缀;随后采用强化学习(Reinforcement Learning, RL)目标,以可控方式增强模型在特定前缀下产生评估意识语义表达的概率。实验表明,该方法在Qwen3.6-35B-A3B、Kimi K2.6和Inkling等模型上使显式评估意识表达量提升2.4至2.9倍,且在未见的代理型(agentic)场景中具有良好的泛化能力,同时保持隐式评估意识水平和行为模式稳定。因果实验证明,通过合成文档微调引入元知识后,VT诱导的语义表达能够反映模型所获得的更丰富知识,验证了其对真实认知状态的有效捕捉。

链接: https://arxiv.org/abs/2609.36316
作者: Usman Anwar,Sahar Abdelnabi,David Krueger
机构: University of Cambridge (剑桥大学); ELLIS Institute Tübingen (图宾根艾利斯研究所); MPI-IS, Tübingen AI Center (图宾根人工智能中心,马克斯·普朗克智能系统研究所); Mila (蒙特利尔学习算法研究所); University of Montreal (蒙特利尔大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Evaluation awareness (EA) can cause large language models (LLMs) to behave differently during audits than in deployment, yet measuring and accounting for EA remains challenging. We introduce verbalization training (VT), a method for making LLMs less reticent about verbalizing evaluation awareness while avoiding to supervise the latent belief itself. VT uses a model’s spontaneous verbalizations as evidence that awareness is present and truncates each rollout immediately before the verbalization, producing training prefixes at which the model is presumed to be aware. The model is then trained with an RL objective designed to increase verbalization in a calibrated way. Across Qwen3.6-35B-A3B, Kimi K2.6, and Inkling, VT increases verbalized EA by 2.4-2.9 times and transfers to held-out agentic settings, while measured latent EA and behavior remain largely stable. In a causal experiment, we independently implant meta-knowledge about evaluations through synthetic-document fine-tuning and show that VT-induced verbalizations reflect the richer knowledge acquired by the model.

[NLP-150] Fractional State Space Transition for Long Sequence Modeling NEURIPS2026

【速读】: 该论文旨在解决现有状态空间模型(State Space Models, SSMs)在长序列建模中因依赖基于常微分方程(ODE)的动力学而产生指数遗忘(exponential forgetting)的问题,导致模型难以有效保留长期依赖信息。其核心解决方案是提出FRAC——一种基于分数阶动力学(fractional dynamics)的可选性状态空间架构,通过将传统的指数衰减机制替换为幂律长时记忆(power-law long memory),实现对长期历史信息的更优建模。关键创新在于:FRAC利用有限状态、对数间隔分布的指数模式叠加来近似具有重尾特性的目标核函数(heavy-tailed target kernel),从而将分数阶记忆转化为一个具备并行训练与预填充能力、且保持有界状态自回归解码效率的高效递归模块。实验结果表明,FRAC在13亿参数语言建模任务中显著优于当前主流SSM基线,在长上下文性能上持续提升的同时,仍保持短序列任务上的竞争力,验证了分数阶动力学作为长上下文SSMs有效先验的可行性与实用性。

链接: https://arxiv.org/abs/2609.36314
作者: Ivan Kobyzev,Abbas Ghaddar,Ali Nasiri-Sarvi,Lifeng Shang,Yufei Cui
机构: Huawei Noah’s Ark Lab, Montreal Research Center, Canada
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: NeurIPS 2026 (Oral)

点击查看摘要

Abstract:State Space Models (SSMs) compress sequence history into a bounded recurrent state, making the resulting memory law a central architectural choice for long-context performance. Most modern SSMs rely on ODE-based dynamics that lead to exponential forgetting, limiting their ability to retain information over broad temporal ranges. We introduce FRAC, a selective SSM architecture derived from fractional dynamics that replaces this exponential decay with power-law long memory. To make fractional dynamics practical, FRAC approximates the heavy-tailed target kernel with a finite-state, log-spaced sum of exponential modes. This construction turns fractional memory into an efficient recurrent module with parallel training and prefill, while retaining bounded-state autoregressive decoding. Extensive experiments, including 1.3B-parameter language modeling, demonstrate that FRAC consistently improves long-context performance over state-of-the-art SSM baselines while staying competitive on short-context. These results show that fractional dynamics provide a practical and effective prior for long-context SSMs.

[NLP-151] HeurEvo: Agent ic Evolution of Hybrid Solver-Augmented Heuristics for Time-Critical Mathematical Optimization

【速读】: 该论文旨在解决在严格运行时约束下,如何自动发现高质量优化算法的问题。传统方法通常仅在预定义流程中优化启发式组件或孤立地调优求解器配置,无法实现对计算资源分配、求解器利用方式及整体算法结构的全局协同优化。为此,本文提出HeurEvo框架,其核心在于通过一种岛式进化机制,联合演化算法的高层结构(plan)、代码实现(code)与可复用组件库(component),实现算法设计的端到端自动化。关键创新在于:由规划器决定组件选择与组合策略、运行时分配;编码器将计划转化为可执行代码;组件演化器持续更新共享组件池;并通过解释器代理反馈执行结果以指导改进。该框架实现了算法结构与实现的协同进化,在多种组合优化基准和挑战性MIPLIB实例上,均能在紧凑时间内获得媲美甚至超越需数小时至数日计算的先进求解器的解质量,尤其在非线性几何问题(如六边形打包)中刷新了已有最优结果,验证了联合搜索算法结构与实现对生成式启发式设计的重要价值。

链接: https://arxiv.org/abs/2609.36303
作者: Feijie Wu,Hugo Barbalho,Konstantina Mellou,Marco Molinaro,Jing Gao,Ishai Menache,Xinzhi Zhang,Sirui Li
机构: Purdue University (普渡大学); Microsoft Research (微软研究院)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE)
备注:

点击查看摘要

Abstract:Recent advances in agentic heuristic design use AI agents and execution feedback to automate algorithm discovery for challenging optimization problems. In many practical settings, high-quality solutions must be obtained under strict runtime constraints, motivating hybrid approaches that combine problem-specific heuristics with powerful mathematical programming solvers. However, existing approaches typically improve heuristic components within predefined procedures or tune solver configurations in isolation. This limits holistic adaptation of where to allocate computation, how to leverage solvers, and how to refine the overall algorithmic structure. To address these limitations, we propose HeurEvo, an automated plan–code–component co-evolution framework that jointly evolves the high-level algorithmic structures, their implementations, and a shared pool of reusable components. A planner determines which algorithmic components to use, how to combine them, and how to allocate runtime across stages, a coder realizes the resulting plan as executable code, while a component evolver updates the shared component pool. Within an island-based evolutionary framework, plans and implementations co-evolve with feedback from an interpreter agent that analyzes execution results and identifies opportunities for improvement. Across diverse combinatorial optimization benchmarks and challenging MIPLIB instances, HeurEvo finds high-quality solutions within tight runtime budgets, often matching or surpassing state-of-the-art optimization solvers given hours or days of computation. On several nonlinear geometry problems such as hexagon packing, it also improves upon the best previously reported results. These results highlight the value of jointly searching over algorithmic structure and implementation for agentic heuristic design.

[NLP-152] MoRE: Scaling mixture of experts with hardware-aware low-rank routing

【速读】: 该论文旨在解决大规模混合专家(Mixture-of-Experts, MoE)架构中路由模块(router)的计算瓶颈问题。随着专家数量 MM 增加且单个专家规模缩小,传统线性路由机制的每令牌计算成本为 Θ(Mh)\Theta(Mh),成为MoE层的主要开销。其解决方案的关键在于提出一种新型的低秩路由混合专家架构(MoRE, Mixture of Rank-reduced-routed Experts),通过将路由权重矩阵进行秩为 rr 的低秩分解,将路由计算复杂度降低至 O((h+M)r)O((h + M)r)。理论分析表明,当活跃专家数固定时,仅需与 log⁡M\log M 同阶的秩即可保证路由表达能力,且该量级是必要的(忽略精度因子)。同时,证明了对数级秩可维持高斯记忆模型中的负载均衡性,合成电话簿任务的训练结果也表明低秩不会损害记忆能力。在相同活跃计算量(active FLOPs)下,该方法可支持 Θ(h/r)\Theta(h/r) 倍更多的专家。为实现实际推理加速,设计了一种融合的Triton内核,避免了高带宽内存(HBM)上的昂贵内存操作。实验显示,MoRE在预训练后显著提升电话簿任务的记忆性能和知识密集型问答基准的表现,同时保持与基线相当的推理能力。

链接: https://arxiv.org/abs/2609.36301
作者: Honam Wong,Surbhi Goel,Enric Boix-Adserà
机构: University of Pennsylvania(宾夕法尼亚大学); The Wharton School, University of Pennsylvania(沃顿商学院,宾夕法尼亚大学)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (stat.ML)
备注:

点击查看摘要

Abstract:Mixture-of-Experts (MoE) layers are central to frontier language models, and recent architectures push toward more and smaller experts. In this regime, the standard linear router becomes a bottleneck: with M experts and hidden dimension h , its per-token cost \Theta(Mh) dominates the MoE layer once M is large. We introduce MoRE (Mixture of Rank-reduced-routed Experts), which factorizes the router weight matrix at rank r and reduces the routing cost to O((h + M)r) . We prove that rank logarithmic in M suffices for routing expressivity when the number of active experts is fixed, and is necessary up to precision factors. We also prove that logarithmic rank preserves load balance in a Gaussian memorization model, and training on a synthetic phonebook task shows that low rank does not hurt memorization. At matched active FLOPs, the factorization allows a factor of \Theta(h/r) more experts. To realize this gain in wall-clock time, we design a fused Triton kernel at inference that avoids expensive memory operations on HBM. Empirically, MoRE improves memorization on the phonebook task and performance on knowledge-intensive Q\A benchmarks after pretraining, while matching reasoning ability. Code available at this https URL.

[NLP-153] When Trees Are Not Enough: Learning Mixed-Topology Feature Graphs with Adaptive Graph Sparse Autoencoders

【速读】: 该论文旨在解决现有稀疏自编码器(Sparse Autoencoders, SAEs)在特征可解释性建模中面临的结构约束与关系可靠性不足的问题:传统结构化SAEs受限于单父节点树或森林结构,而事后构建的图结构虽允许多父节点关系,却无法引导特征学习且难以保证关系恢复的可靠性。其核心解决方案是提出自适应图稀疏自编码器(Adaptive Graph Sparse Autoencoder, AG-SAE),关键在于将每个特征的完整父节点集合视为一个原子化的结构假设,并通过证据竞争机制自动选择零个、一个或多个父节点,从而识别出联合必要性的多父节点关系,同时排除冗余或虚假关联,并验证子特征对重构的贡献是否超越其父节点。由此生成的拓扑结构构成可微分的结构损失,用于指导SAE训练;同时,基于拓扑引导的特征精炼策略缓解了特征吸收问题,并利用学习结构揭示的持续重构间隙来初始化新特征。整个过程形成“字典-图”自洽循环:在更新字典后重新评估每个特征的完整父集,实现结构与表示的协同优化。实验表明,AG-SAE在控制模型中实现了精确的混合拓扑恢复,在真实大语言模型(LLM)激活数据上表现出更优的关系可靠性与语义合理性,且在特征级因果干预能力上显著优于传统SAE方法。因此,AG-SAE成功将恢复的混合拓扑结构转化为无监督训练信号,不仅提升了字典质量,还突破了树状结构的拓扑限制,实现了更强的因果控制能力。

链接: https://arxiv.org/abs/2609.36294
作者: Xiaozuo Shen,Yifei Cai,Tian Tan,Rui Ning,Chunsheng Xin,Hongyi Wu
机构: University of Arizona(亚利桑那大学); Iowa State University(爱荷华州立大学); Old Dominion University(老多明尼昂大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Sparse autoencoders (SAEs) expose interpretable features in large language model activations, yet existing structured SAEs impose single-parent trees or forests, while post-hoc graphs permit multiple parents but neither guide feature learning nor ensure reliable relation recovery. We introduce the Adaptive Graph Sparse Autoencoder (AG-SAE), a structure-guided training paradigm that treats each feature’s complete parent set as an atomic structural hypothesis and lets evidence select zero, one, or multiple parents. By competing complete parent sets against null, subset, and alternative explanations, AG-SAE identifies jointly necessary multi-parent relations while rejecting redundant or spurious alternatives and verifying that each child contributes beyond its parents. The induced topology over SAE features then defines a differentiable structural loss that guides SAE training, while topology-guided refinement mitigates feature absorption and uses persistent reconstruction gaps exposed by the learned structure to initialize new features. The entire graph is then induced again from the revised dictionary by reassessing every feature’s complete parent set, closing the dictionary-graph self-consistency cycle. Experiments demonstrate exact mixed-topology recovery in a controlled toy model, greater relational reliability and semantic validity than structured and post-hoc baselines on real LLM activations, and stronger feature-level causal interventions than conventional SAE features. AG-SAE thereby turns recovered mixed-topology feature structure into an unsupervised training signal that improves the dictionary, enables reliable feature organization beyond the topological limitations of trees, and exhibits stronger causal control beyond reconstruction.

[NLP-154] he Surge of Anti-Semitism in German Social Media following the October 7 Attacks

【速读】: 该论文旨在探究2023年10月7日哈马斯对以色列的袭击事件对德国社交媒体上关于犹太教(Judaism)与以色列讨论的影响,尤其关注反犹主义(anti-Semitism)言论的演变。其核心问题是:重大地缘政治事件是否加剧了社交媒体中的反犹言论,并导致不同平台间的话语模式分化或趋同。解决方案的关键在于开发一种基于大语言模型(Large Language Models, LLMs)的检测方法,用于识别用户帖子中26类反犹主义内容。该方法在Facebook和Telegram的125,718条帖子上进行验证,通过对比有无用户上下文信息的两种设置,发现引入用户背景信息可显著提升模型性能(最高达83% F1-score),并大幅降低误报率,尤其是在报道反犹事件时的批判性表达场景中。研究结果表明,事件后反犹言论在两个平台均显著上升,且Telegram上的反犹言论强度约为Facebook的十倍;其中Facebook以针对“加沙种族灭绝”指控的反犹言论为主,而Telegram则更普遍传播经典反犹阴谋论与权力叙事。事件后,两平台话语模式出现收敛趋势:Facebook上经典反犹主义上升,而Telegram上与以色列相关的议题激增,揭示出危机情境下跨平台意识形态传播的动态演化特征。

链接: https://arxiv.org/abs/2609.36290
作者: Gregor Wiedemann,Daniel Wehrend
机构: Leibniz-Institute for Media Research | Hans-Bredow-Institut(莱布尼茨媒体研究所 | 布雷多研究所); Hamburg, Germany(汉堡, 德国)
类目: ocial and Information Networks (cs.SI); Computation and Language (cs.CL)
备注: 8 pages; 5 figures; accepted at 22st Conference on Natural Language Processing (KONVENS 2026), Hamburg, Germany

点击查看摘要

Abstract:We investigate the extent to which the Hamas attacks on Israel of October 7, 2023, have affected German social media debates about Judaism and Israel. For this, we develop an approach to detect 26 anti-Semitic categories in user postings via large language models (LLMs). The approach is applied to Facebook and Telegram posts (N=125,718) from three months before and after the event. Methodically, we test different open-weight models in two setups—with and without user information as additional context to the post text. The best setup achieves up to 83 % F1-score for binary anti-Semitism detection on our manually coded validation set. User context provides valuable information for most LLMs and drastically reduces false positives, for example, when (critically) reporting on anti-Semitic incidents. Concerning our topic, we find that anti-Semitism is surging significantly on both platforms, while being about ten times more prevalent on Telegram compared to Facebook. Facebook users express anti-Semitic views most likely in posts about an alleged genocide in Gaza carried out by the Israeli army, whereas classic anti-Semitic stereotypes related to power and conspiracy theories are dominant on Telegram. After the attack, the discourse patterns on both platforms show signs of convergence, as classic anti-Semitism increases on Facebook, whereas Israel-related categories surge on Telegram.

[NLP-155] In-Context Learning Amplifies a Latent Symbolic Circuit ICML2026

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在少量上下文示例(in-context examples)条件下如何逐步激活内部符号推理机制以实现抽象规则学习的问题。其核心挑战在于,尽管模型能够通过少量示例快速掌握抽象规则,但其内部神经机制随示例数量增加而动态演化的具体过程尚不清晰。论文的关键解决方案在于识别并验证了一个贯穿不同样本量(shot count)的三阶段符号推理通路:抽象(abstraction)、归纳(induction)与检索(retrieval)。研究发现,该推理回路在模型达到高准确率之前即已可检测且具备功能性;通过每头(per-head)因果贡献分析显示,从1-到10-shot时,关键注意力头的贡献提升高达8倍;进一步采用跨示例激活修补(cross-shot activation patching)技术,可在0-shot时将准确率从1%提升至56%,1-shot时从17%提升至88%。此外,将功能向量(function vectors)在0-shot条件下进行缩放注入,可使准确率最高达86%,有效替代归纳阶段功能,但依赖于下游检索阶段的完整性。研究揭示:抽象规则遵循的底层基础设施早已存在于模型权重中,而上下文示例、功能向量及相关干预措施实质上是向同一潜在推理回路提供输入信号。

链接: https://arxiv.org/abs/2609.36265
作者: Melissa Wessel
机构: 独立研究者(Independent Researcher)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: Accepted to the Mechanistic Interpretability Workshop at ICML 2026

点击查看摘要

Abstract:Large language models can learn abstract rules from just a few in-context examples, but how their internal mechanisms activate as examples accumulate is not well understood. We trace a three-stage symbolic reasoning circuit (abstraction, induction, retrieval) across shot counts in three model families and find it is detectable and functional well before the model achieves high accuracy. Per-head causal contribution grows up to 8x from 1- to 10-shot, and cross-shot activation patching raises accuracy from 1% to 56% at 0-shot and 17% to 88% at 1-shot. Function vectors scaled and injected at 0-shot rescue accuracy up to 86%, largely substituting for the induction stage but depending critically on an intact downstream retrieval stage. The infrastructure for abstract rule-following is present in the weights before any demonstrations; in-context examples, function vectors, and related interventions appear to supply input to the same latent circuit.

[NLP-156] OTROPE: Optimal Transport-based Robust Off-policy Evaluation for Large Language Models NEURIPS2026

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在离线评估中的可靠性问题,尤其是在人类标注数据稀缺、行为策略与目标策略之间存在显著分布偏移(distribution shift),且黑盒模型无法获取响应似然(response likelihoods)的挑战性场景下,如何实现高效、安全且准确的评估。其核心解决方案是提出一种无需似然信息的鲁棒离策略评估方法——基于最优传输的鲁棒离策略评估(Optimal Transport-based Robust Off-Policy Evaluation, OTROPE)。OTROPE 的关键在于:通过在语义空间中利用最优传输(Optimal Transport)对齐有标签的行为策略样本与无标签的目标策略样本,实现分布校正;进而结合经过校正的人类标注残差与代理预测器(proxy predictor),构建一种双重稳健(doubly robust)的评估框架,无需显式建模行为策略或进行密度比估计。理论分析表明,当重加权后的行为分布或代理预测器收敛时,OTROPE 具备一致性及可证明的收敛速率。实验结果在合成与真实场景中均验证了其优于现有基线方法的性能,并展现出通过集成多个弱评估器即可逼近甚至超越强评估器的潜力。

链接: https://arxiv.org/abs/2609.36264
作者: Liner Xiang,Wenbo Zhang,Hengrui Cai
机构: University of California, Irvine (加州大学欧文分校)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (stat.ML)
备注: Accepted at NeurIPS 2026

点击查看摘要

Abstract:Reliable evaluation of large language models (LLMs) is essential for their development and deployment, yet is often costly, risky, and difficult to perform safely online. We study off-policy evaluation for LLMs, where limited human-labeled data from a behavior model are used to evaluate a newer target LLM. This setting is challenging because labels are scarce, behavior–target distribution shift is common, and response likelihoods are often unavailable for black-box LLMs. We propose the Optimal Transport-based Robust Off-Policy Evaluation (OTROPE), a likelihood-free evaluation that performs distributional correction in a semantic space via optimal transport to align labeled behavior-policy samples with unlabeled target-policy samples. OTROPE combines corrected human-labeled residuals with proxy predictors, yielding a doubly robust-style evaluation without behavior-policy modeling or density-ratio estimation. We theoretically characterize why baseline evaluators fail under LLM distribution shift, and establish consistency and convergence rates for OTROPE when either the reweighted behavior distribution or the proxy predictor converges. Experiments on synthetic and real LLM evaluation tasks show that OTROPE consistently outperforms baselines while enabling ensembles of weaker LLM evaluators to approach and sometimes surpass stronger evaluators. Code is available at this https URL.

[NLP-157] Population Fidelity: Evaluating Population Representativeness in LLM s

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在模拟人类态度与偏好时存在的群体代表性失真问题,即模型生成的回答往往压缩了真实人群的态度分布,且对特定子群体的表征存在偏差,这种偏差在不同模型和议题间具有差异性。其解决方案的关键在于提出“群体保真度”(Population Fidelity)评估框架,该框架从三个核心维度衡量LLM生成回答对目标人群的代表性:群体层面的准确性、群体间变异性的数量,以及该变异性的结构特征。通过此框架,研究发现现有模型的问题不仅在于缺乏足够的群体间差异,还在于将变异错误地分配至不恰当的群体;同时,尽管文化微调(cultural fine-tuning)可提升模型输出与总体均值的一致性,却未能改善对群体内部差异的表征,而这一关键指标无法被传统的整体一致性度量所捕捉。因此,论文强调,真正实现对人群的准确建模需同时复现人类态度变异的多重特征,而该框架系统化地组织这些特征,并提供可复用的代码、数据及训练模型,支持跨领域评估与对齐方法的验证。

链接: https://arxiv.org/abs/2609.36253
作者: Neemias B. da Silva,Martin Lukk,Ali Sutani,Abhishek Moturu,Harris Yang,Daniel Silver,Matt Ratto,Thiago H. Silva
机构: University of Toronto (多伦多大学); Federal University of Technology – Paraná (巴拉那联邦技术大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
备注: 37 pages, 16 figures, 14 tables. Code and data: this https URL

点击查看摘要

Abstract:Large language models (LLMs) show considerable potential in simulating human attitudes and preferences. Prior work finds that LLM-generated responses can compress the range of attitudes found within populations and misrepresent particular subgroups in ways that vary across models and topics. We introduce Population Fidelity, an evaluation framework that distinguishes key conditions required for a set of LLM-generated responses to represent a population. It incorporates three dimensions: group-level accuracy, the amount of between-group variation, and the structure of that variation. We demonstrate the framework’s utility in two ways. First, we reproduce a prior study of “machine bias” in LLM survey responses and apply the framework to its models and more recent ones, showing that poor representation reflects not only insufficient between-group variation but also variation assigned to the wrong groups. Second, we evaluate one proposed approach to improving models’ population representativeness: cultural fine-tuning. We find that cultural fine-tuning can improve alignment with the survey center without improving the representation of within-population differences, a distinction that measures of aggregate agreement do not capture. We argue that representing a population requires models to reproduce several features of human attitudinal variation simultaneously. Our framework organizes these features and provides reusable code, data, and trained models for evaluating population fidelity across substantive domains and assessing proposed alignment methods.

[NLP-158] Learning from Teacher Continuations at Student States

【速读】: 该论文旨在解决现有语言模型蒸馏方法在在线学习场景中面临的三大核心问题:(1)离线监督微调(SFT)中固定教师轨迹导致的序列协变量偏移(sequential covariate shift);(2)基于前缀失败时的片段化监督(fragmented supervision)在逐标记在线策略蒸馏(OPD)中的表现瓶颈;(3)分布匹配蒸馏对教师输出概率访问的依赖性。其解决方案的关键在于提出OLIVE(OnLine InterVEntion)框架,通过在每轮迭代中由不断演化的学生策略生成新前缀,由教师模型自回归续写,再以教师生成标记的交叉熵作为学生更新信号,从而实现动态、连续的在线蒸馏。该设计有效缓解了上述限制,尤其在持续训练过程中突破传统离线蒸馏的性能瓶颈,同时保持学生模型的泛化能力与可塑性。实验表明,仅使用GPT-5.4-mini文本进行持续训练的OLIVE,在ScienceWorld任务上相比同源教师的离线微调提升13%,且在推理任务和代理型任务中均显著优于现有蒸馏方法,同时具备更低的计算开销与更优的训练效率。

链接: https://arxiv.org/abs/2609.36246
作者: Haojin Wang,Dylan Zhang,Huaibo Chen,Suhao Yu,Yihang Sun,Zhanyang Jin,Jiaying Ye,Dianqi Li,Prasanna Sattigeri,Kamal Youcef-Toumi,Hao Peng
机构: University of Illinois at Urbana-Champaign(伊利诺伊大学香槟分校); Massachusetts Institute of Technology(麻省理工学院); University of Pennsylvania(宾夕法尼亚大学); University of Washington(华盛顿大学); International Business Machines(国际商业机器公司)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:We present OLIVE (OnLine InterVEntion). At each iteration, the evolving student policy generates a new prefix, the teacher continues it autoregressively, and the student is updated using cross-entropy computed on the teacher-generated tokens. Each design choice targets a corresponding limitation of existing distillation methods: (1) sequential covariate shift in offline supervised fine-tuning (SFT) on fixed teacher trajectories, (2) fragmented supervision under prefix failure in token-level on-policy distillation (OPD), and (3) the need for access to teacher token probabilities in distribution-matching distillation. OLIVE achieves higher reasoning performance than OPD (with a top-16 KL approximation) at comparable GPU-hour cost. Our asynchronous implementation further reduces OLIVE’s total training time by 23.8%. We evaluate OLIVE on both hard reasoning tasks and agentic tasks which reflects modern post-training scenarios, and it consistently outperforms existing distillation methods under the same training budget. By regenerating prefixes from the evolving student, OLIVE continues improving after offline distillation plateaus while better preserving the general capabilities and plasticity of the student. Using only text from GPT-5.4-mini, continuously training with OLIVE outperforms offline SFT from the same teacher by 13% on ScienceWorld. These results support OLIVE as an effective and efficient approach to online language-model distillation.

[NLP-159] Cognitive Expert Language Models Better Align with the Corresponding Brain Systems

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在预测人类大脑活动时,普遍采用“一模多用”(one-model-fits-all)方法所导致的区域功能特异性被忽略的问题。传统方法通常使用单一模型跨多个脑区进行评估,并将性能结果在空间上平均化,这可能掩盖了不同脑区对特定认知域的差异化响应。其解决方案的关键在于构建面向特定认知领域的专家型大语言模型(expert LLM variants),通过提示工程(prompting)与微调(fine-tuning)分别针对六类认知域——感官、空间、数值、推理、社会及抽象加工——训练专用模型。研究发现,每个领域专家模型在与其对应认知域相关的脑区中表现出更优的模型-脑活动对齐(model-brain alignment),且该现象在三种基础模型和三个fMRI数据集下均具有一致性。控制分析进一步表明,这种对齐效应是源于认知层面的干预,而非表面特征或非认知因素所致。该研究揭示了模型专业化可显著提升特定脑区的对齐效果,同时维持整体预测精度不变,说明以往跨区域汇总对齐性能的做法可能掩盖了模型在局部脑区的功能适配差异。

链接: https://arxiv.org/abs/2609.36239
作者: Zhivar Sourati,Mengxuan Helen Wu,Nona Ghazizadeh,Jonas Kaplan,Morteza Dehghani,Samuel A. Nastase
机构: University of Southern California(南加州大学); Center for Computational Language Sciences(计算语言科学中心)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Large language models (LLMs) can predict human brain activity across a variety of brain regions during natural language comprehension. Typically, however, LLM-brain alignment is measured using one model for different regions of the brain, and then model performance is summarized across regions. This one-model-fits-all approach ignores the functional specialization of brain regions. In this study, we assess whether a model oriented toward a particular cognitive domain aligns better with the brain system dedicated to that domain. Through prompting and fine-tuning, we first build expert LLM variants for six domains: sensory, spatial, numerical, reasoning, social, and abstract processing. We then examine whether each expert best predicts activity in the brain region associated with the corresponding cognitive domain. Consistent with our hypotheses, each expert’s representations align more closely with the brain system most associated with the matching domain than do other experts. This holds under both prompting and fine-tuning, across three base models and three fMRI datasets. In a series of control analyses, we show that this model-brain alignment is specific to cognitive domain interventions; non-cognitive and surface-level interventions do not result in comparable alignment. Specializing models shifts regional alignment while leaving aggregate prediction accuracy largely unchanged, suggesting that summarizing alignment across regions may obscure regional differences in performance for specific models.

[NLP-160] CineSubBench: Evaluating LLM s on Long-Form Narrative and Cultural Understanding from Multilingual Movie Subtitles

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在电影理解任务中长期存在的评估空白问题,尤其是在多语言、长上下文、文化语境敏感的叙事理解方面。现有评估多集中于法律、医疗等专业领域,而电影作为需整合长篇叙事、跨语言解读与文化情境判断的复杂文本类型,却未得到充分研究。为此,论文提出CineSubBench——一个基于多语言电影字幕的基准评测体系,通过数千条时间有序的字幕片段,要求模型重构角色关系、事件因果链、主题脉络及文化语境下的观众判断,从而全面评估模型在长上下文理解、多语言一致性、文化适配性与证据锚定能力方面的表现。其解决方案的关键在于构建了一个涵盖六种语言、1,012部影片、共计6,072条字幕流和813万条带时间戳字幕条目的大规模、多任务、多文化(MultiX)评测框架,支持叙事重建、类型识别、年龄适宜性判断、十国电影分级系统比对以及基于字幕的语言安全检测,揭示了当前主流大模型在情节前提恢复与事件完整摘要之间的性能差异、跨语言一致性不均、各国分级系统校准模式的异质性,以及强粗俗语义较弱隐晦表达更易被准确识别等关键发现,确立了电影作为长上下文大模型评估新范式的重要地位。

链接: https://arxiv.org/abs/2609.36218
作者: Mir Tafseer Nayeem,Susmoy Chakraborty,Davood Rafiei
机构: University of Alberta (阿尔伯塔大学); Independent Researcher (独立研究员)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Emerging Technologies (cs.ET); Machine Learning (cs.LG)
备注: Preprint

点击查看摘要

Abstract:Large language models are increasingly evaluated in specialized domains such as law, medicine, software engineering, and cybersecurity, yet film remains comparatively underexplored despite requiring long-form narrative integration, multilingual interpretation, and culturally situated audience judgments. We introduce CineSubBench, a benchmark for evaluating long-context film understanding from multilingual movie subtitles. A subtitle track represents a film as thousands of short, temporally ordered utterances from which models must reconstruct characters, relationships, events, causal progression, and themes without explicit scene or event structure. CineSubBench contains 1,012 films with complete subtitle coverage in six languages, yielding 6,072 tracks and 8.13M timestamped subtitle entries. It provides a matched multi-task, multilingual, and multicultural (MultiX) evaluation setting: seven tasks span narrative reconstruction and abstraction, genre prediction, age suitability, country-specific motion-picture ratings across ten national classification systems, and subtitle-grounded language safety. Across nine LLMs, plot premises are recovered more reliably than event-complete synopses; cross-lingual consistency varies substantially across models and languages; national rating systems expose distinct calibration patterns; and strong profanity is far easier to ground than mild obscenity. CineSubBench establishes film as a long-context LLM evaluation domain and provides a unified benchmark for measuring narrative, multilingual, cultural, and evidence-grounding capabilities.

[NLP-161] Lost in Translation: Measuring the Effect of Non-Native English on End User Performance of Large Language Models

【速读】: 该论文旨在解决非英语母语用户在使用大语言模型(Large Language Models, LLMs)时,其生成结果质量系统性低于母语使用者的问题。尽管已有研究指出这一差距的存在,但其具体成因仍不明确,尤其缺乏对非母语者语言特征中哪些因素(如语法准确性、词汇使用、篇章组织或话语连贯性等)导致性能差异的深入分析。为揭示这一问题,作者构建了FABLE数据集,该数据集包含190,911个源自17.4万条真实写作任务提示的受控英文提示变体,能够精细分离语言流畅度的不同维度。通过对34个开源权重的LLM进行评估,研究发现:虽然模型不会将表层错误(如拼写错误)复制到输出中,但会高度模仿用户提示中的高层级修辞风格与词汇选择;更重要的是,响应质量在低流畅度与高流畅度提示之间存在显著差异。因此,该研究的关键解决方案在于揭示并验证——非母语用户的语言表达质量(尤其是语篇层面的修辞与词汇层次)直接影响模型输出的质量与流畅度,从而暴露了大语言模型在服务非母语用户时存在的关键性能不对称性。

链接: https://arxiv.org/abs/2609.36214
作者: Yusheng Zhou,Eleanor Lin,David Jurgens
机构: University of Michigan (密歇根大学)
类目: Computation and Language (cs.CL)
备注: 19 pages, 8 figures

点击查看摘要

Abstract:Large language models (LLMs) are increasingly used by people whose first language is not English, yet these users have been shown to receive systematically lower-quality responses than fluent speakers. Which specific features of non-native English drive this gap remains unclear, because fluency is itself a composite of mechanical accuracy, vocabulary use, organization, and discourse coherence. Here, we introduce FABLE, a controlled dataset of 190,911 English prompt variants derived from 174K real user prompts for writing-related tasks. Evaluating responses from 34 open-weight LLMs, we find a clear asymmetry; while models do not propagate surface errors such as misspellings into their outputs, models do mirror higher-level rhetorical and lexical qualities present in the user’s prompt. Further, the overall quality of responses differs substantially between the least- and most-fluent prompts. These results highlight a key LLM performance disparity for non-native English LLM users, resulting in both lower-quality and less-fluent answers.

[NLP-162] he Canonical Order Problem: When Large Language Models Are Unreliable Knowledge Bases for Multi-Valued Relations EMNLP2026

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在作为知识库(Knowledge Base, KB)使用时,对多值关系(multi-valued relations)生成能力受限的问题。尽管现有研究多聚焦于从模型中提取单一关系三元组,但现实世界中的多数关系具有多个实体值,需生成实体集合。论文发现,LLMs内部存在“规范顺序问题”(canonical order problem),即模型在预训练过程中对多值关系的潜在概率分布倾向于遵循某种内在规范排序(如字母或时间顺序)。通过机制分析,作者揭示了LLM在生成集合时经历三个阶段:候选实体检索、内部排序与下一元素选择。这一机制导致当提示词(prompt)所期望的知识库结构偏离模型内部的规范顺序时,模型生成完整实体集合的可靠性显著下降。因此,解决方案的关键在于理解并适配模型内部的规范排序机制,以设计更有效的提示策略,从而提升多值关系生成的准确性和完整性。

链接: https://arxiv.org/abs/2609.36209
作者: Timo Pierre Schrader,Annemarie Friedrich,Simon Razniewski,Lukas Lange
机构: University of Augsburg (奥格斯堡大学); ScaDS.AI TU Dresden (ScaDS.AI 图林根工业大学); Bosch Center for Artificial Intelligence (博世人工智能中心)
类目: Computation and Language (cs.CL)
备注: Accepted at AKBC@EMNLP2026

点击查看摘要

Abstract:Large language models (LLMs) are increasingly used as knowledge bases (KBs) due to the vast amount of knowledge they acquire during pre-training. While many works focus on extracting single relational triples, most real-world relations are multi-valued and require generating sets of entities. In this paper, we investigate how LLMs represent and generate multi-valued relations. We identify the canonical order problem: The probabilistic distributions inside LLMs organize many multi-valued relations according to a canonical ordering (e.g., alphabetical or chronological). Through mechanistic analysis, we show that set generation in LLMs can be thought of in terms of three phases: (1) retrieval of candidate entities, (2) internal sorting, and (3) selection of the next element. As a result, prompts aiming to construct KBs that deviate from this internal canonical ordering lead to a markedly reduced reliability of LLMs when aiming to generate complete sets for multi-valued relations. Comments: Accepted at AKBC@EMNLP2026 Subjects: Computation and Language (cs.CL) Cite as: arXiv:2609.36209 [cs.CL] (or arXiv:2609.36209v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2609.36209 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[NLP-163] Geometric Representations of African Languages: A Regional Semantic Hub and Cultural Steering

【速读】: 该论文旨在解决生成式AI(Generative AI)在非洲语言表征与文化引导响应方面的代表性不足及偏差问题,尤其关注Gemma 4 31B模型对非洲语言的语义表征质量及其在文化语境下的响应一致性。其解决方案的关键在于通过多维度评估框架——包括探针测试(probes)、对比方向分析(contrast directions)和表征相似性度量——系统比较九种非洲语言与三种对照语言在不同网络层中的表征特性,并基于独立英语数据构建国家层面的文化引导方向(country steering directions)。研究发现,非洲语言在特定层次上表现出更强的内部一致性(如尼日利亚的约鲁巴语与伊博语相较于豪萨语更接近),且国家导向方向在不同地区间具有可区分的归属预测能力(归因率差异达0.63–0.81),同时输出具备良好的结构连贯性。此外,结果表明表征测量方式显著影响结论,凸显了评估方法选择的重要性。整体而言,该研究揭示了非洲语言在模型中的区域与语系层级表征模式,以及英语语料中隐含的文化引导效应。

链接: https://arxiv.org/abs/2609.36205
作者: Muhammad Abdullahi Said,Jonathan Shock
机构: African Institute for Mathematical Sciences (非洲数学科学研究所); University of Cape Town (开普敦大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:We study how Gemma 4 31B represents African languages and responds to cultural steering. The first study compares nine African languages and three controls using probes, contrast directions, and measures of representation similarity. Transfer from English varies across languages and layers. Directions representing an Africa versus West contrast are more aligned among the African languages than between these languages and the controls at several layers. The comparison across language families passes the reported Holm threshold at five of twelve layers, although dependence between language pairs limits the statistical interpretation. Within Nigeria, Yoruba and Igbo are more aligned than the average of their pairs with Hausa at eleven of twelve layers. The second study uses separate English data to construct directions for Nigeria, Ghana, Kenya, and South Africa. Under union scoring at the selected strengths, estimated differences in attribution rates from random directions range from 0.63 to 0.81. Most outputs pass the automated structural coherence screen. Comparisons with Aya Expanse 32B show that results depend on the representation measure. Together, the studies document regional and family patterns in the sampled representations and country steering in English.

[NLP-164] FastGuide: Accelerating Reward Guidance for Diffusion Large Language Models

【速读】: 该论文旨在解决基于梯度的奖励引导(gradient-based reward guidance)在推理阶段对掩码扩散语言模型(masked diffusion language models)进行控制时计算成本过高的问题。其核心挑战在于每次解码迭代均需执行昂贵的扩散模型前向传播与奖励模型反向传播,导致整体效率低下。解决方案的关键在于提出FastGuide——一种自适应混合并行与自回归解码方法:首先,借鉴并行解码思想,将奖励模型反向传播的开销通过每一步解码仅计算一次引导信号并复用于生成多个词元来分摊;其次,在每一步解码内部,通过逐个解码(unmasking one token at a time)实现自回归式前向传播,并利用键值缓存(KV caching)和注意力稀疏重计算技术高效更新词元分布;最后,根据模型对重计算分布的置信度动态延迟生成不确信的词元,从而实现高效且高质量的生成。实验表明,FastGuide相较传统顺序式奖励引导解码速度提升最高达4.4倍,同时保持相近的生成质量。

链接: https://arxiv.org/abs/2609.36202
作者: Darshan Thaker,Lachlan Ewen MacDonald,René Vidal
机构: University of Pennsylvania(宾夕法尼亚大学)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Gradient-based reward guidance provides a flexible way to use downstream reward models to control masked diffusion language models at inference time. However, its computational cost remains high as each decoding iteration incurs expensive diffusion model forward passes and reward model backpropagation steps. To address this, we introduce FastGuide, an adaptive hybrid of parallel and autoregressive decoding to accelerate reward guidance for diffusion language models. In analogy to parallel decoding, FastGuide amortizes the cost of reward model backpropagation by computing guidance once per decoding step and reusing it to generate multiple tokens. Within each decoding step, FastGuide makes diffusion forward passes autoregressive by unmasking tokens one at a time while efficiently recomputing token distributions after each unmasking by utilizing KV caching techniques and sparse recomputation of attention. Lastly, to adapt hybrid decoding to the model’s confidence, FastGuide defers any token that the model is unconfident about under its recomputed distribution. Experiments on three reward benchmarks demonstrate that FastGuide is up to 4.4\times faster than sequential reward-guided decoding while retaining similar generation quality.

[NLP-165] SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety

【速读】: 该论文旨在解决计算机使用代理(Computer-use Agents, CUAs)在执行日常或专业工作流任务时,即使在无害指令与环境下仍可能造成未预期损害的问题。现有安全验证方法面临两大挑战:其一,依赖通用安全准则的验证者难以识别任务特定且细微的有害行为,需具备深度的任务相关推理能力;其二,仅基于轨迹截图的验证方式无法准确捕捉环境实际变化,导致大语言模型作为裁判(LLM-as-a-judge)难以判断动作的真实后果。为应对上述问题,论文提出SCOUT——一种两阶段的代理式安全验证框架,其核心在于将高强度推理驱动的规则生成与高工具依赖性的证据收集相结合。具体而言,第一阶段由SCOUT规则生成器对任务目标及代理执行轨迹进行深入推理,生成任务定制化的完成度与安全性评估标准;第二阶段由SCOUT探测代理依据这些规则主动交互于任务后置环境,获取可验证的实证证据,以支持最终的安全性与完成度判断。实验表明,SCOUT在AutoElicit-Bench和OS-Blind两个基准上均显著优于传统验证方法,尤其在非前沿模型上展现出更强的鲁棒性。消融研究进一步证明,无需外部工具的规则生成机制能激发更深层次的推理,是提升各类验证器安全检测能力的关键。初步扩展至编码任务亦显示,SCOUT具备跨领域安全验证潜力。

链接: https://arxiv.org/abs/2609.36201
作者: Jianxing Chen,Xiao Yu,Shipra Agrawal,Zhou Yu
机构: Columbia University(哥伦比亚大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注:

点击查看摘要

Abstract:Computer-use agents (CUAs), while capable of completing computer tasks in everyday and professional workflows, can cause unintended harm even under benign instructions and environments. However, detecting such harm remains challenging. First, it requires careful, task-specific reasoning: verifiers guided only by general safety criteria often overlook many important but subtle harmful behaviors. Second, it requires active investigation: past trajectory screenshots show what the agent did but not always what actually changed in the environment, so LLM-as-a-judge verifiers that rely on screenshots alone may be unable to determine the actual consequences of actions. To address these challenges, we introduce SCOUT, a two-stage agentic safety verifier that synergizes reasoning-intensive rubric generation with tool-intensive evidence gathering. First, our SCOUT rubric generator extensively reasons over the task and the agent’s trajectory to determine what successful and safe execution should entail, generating task-specific completion and safety rubrics. Then, our SCOUT probing agent follows these rubrics to interact with the post-execution environment and collect grounded evidence for final safety and completion judgments. We evaluate our framework on two computer-use safety benchmarks. On AutoElicit-Bench, SCOUT achieves 75.4 unsafe F1 and 74.5 completion F1, outperforming LLM-as-a-judge verifiers and naive tool-use verifiers. SCOUT leads on OS-Blind with 76.4% unsafe detection accuracy. Test-time reflection reduces final unsafe execution rates from 30.2% to 17.2% on AutoElicit-Bench. Ablations and analysis show that tool-free rubric generation in SCOUT elicits substantially more reasoning and is crucial for safety detection across verifier backbones, especially non-frontier ones. A preliminary extension to coding tasks shows that SCOUT can support safety verification beyond computer-use.

[NLP-166] Concept Direction Reliability Across Languages with Different Tokenizer Fertility

【速读】: 该论文旨在解决在情感分类任务中,尽管下游分类性能准确,但情感方向(sentiment direction)在不同样本间存在可变性的问题。其核心挑战在于:当前模型的预测准确性并不能保证情感向量方向的一致性,即模型在不同输入上对相同情感倾向的表示可能不一致。解决方案的关键在于提出并验证一种评估情感方向可重复性的方法——通过分裂一半样本(split-half agreement)来衡量在英语、豪萨语和约鲁巴语中,四种语言模型在原生文本与翻译文本下的表示一致性。研究发现,尽管分类器在这些层上仍能实现高于随机水平的预测性能,但方向一致性在不同语言中显著差异,且该差异与语言本身有关,而非仅由分词器丰度(tokenizer fertility)解释。此外,结果表明,对词元表示进行平均会降低一致性,而高一致性可能部分受句子长度影响。因此,该研究强调必须独立于分类性能来评估向量方向的可重复性,以确保情感表示的可靠性。

链接: https://arxiv.org/abs/2609.36194
作者: Muhammad Abdullahi Said,Abass Oguntade,Elisha Komolafe,Babangida Sani,Fatima Muhammad Adam,Muhammad Sammani Sani
机构: University of Cape Town (开普敦大学); African Institute for Mathematical Sciences (非洲数学科学研究所); Bayero University Kano (贝亚罗大学卡诺分校); Federal University Dutse (杜特塞联邦大学); University of Vienna (维也纳大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Extracted sentiment directions can vary across samples even when downstream sentiment classification remains accurate. To evaluate direction reproducibility, we measure split-half agreement in English, Hausa, and Yoruba representations across four language models using both native and translated texts. We identify layers selected for agreement using ten topics and evaluate direction agreement across separate groups of fifteen topics. Using the final token, split-half agreement ranges from 0.737 to 0.870 for English, 0.589 to 0.762 for Hausa, and 0.101 to 0.399 for Yoruba, maintaining this language rank order across all 77 complete model comparisons. Classifiers trained on these same layers consistently predict sentiment above chance, demonstrating that predictive accuracy does not imply directional consistency. Furthermore, averaging token representations yields less consistent agreement, and high agreement can partially reflect sentence length. Ultimately, our findings highlight the need to measure vector direction reproducibility independently of classification performance, though they do not establish that tokenizer fertility which is the average number of tokens per whitespace separated word causes cross-lingual differences.

[NLP-167] argeting Pivotal Decisions for Credit Assignment in Agent ic Reinforcement Learning

【速读】: 该论文旨在解决生成式语言模型智能体在强化学习训练中因统一分配轨迹级优势而导致的信用分配粗粒度问题,即无法区分哪些中间决策对任务成功具有关键影响。其核心解决方案是提出ProVer框架,通过引入一个代理评判器(agentic judge)识别可能具有决定性作用的决策片段,并仅对这些片段进行局部优势估计——利用当前策略在片段前后延续采样所导致的最终成功率差异来验证其贡献。该方法不依赖于直接信任评判器的判断,而是基于可观测结果进行可信验证,从而实现细粒度信用分配。该设计显著提升了政策训练效率与性能,在ALFWorld、WebShop和SearchQA等多个基准上,相较于传统的组相对策略优化(Group Relative Policy Optimization, GRPO),在Qwen3.5-2B和Qwen3.5-4B模型上分别取得9.91%和7.12%的相对性能提升,且仅需适度增加生成开销,证明了选择性聚焦关键决策节点在智能体强化学习中的有效性与高效性。

链接: https://arxiv.org/abs/2609.36178
作者: Dongwon Jung,Hemanth Neelgund Ramesh,Yifan Wang,Xiaomin Li,Yuexing Hao,Yu Hu,Muhao Chen,Varun Chandrasekaran,Andrzej Banburski-Fahey,Jaron Lanier
机构: University of California, Davis (加利福尼亚大学戴维斯分校); Microsoft(微软); University of Washington (华盛顿大学); Purdue University (普渡大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed to success. We introduce ProVer, a framework that targets potentially pivotal decisions for fine-grained credit assignment in agentic reinforcement learning. Given a rollout group, an agentic judge contrasts successful and failed trajectories to propose a segment potentially responsible for their divergent outcomes. Rather than directly trusting the judge’s assessment, ProVer verifies the proposed segment by estimating its advantage from the difference in terminal success rates between current-policy continuations sampled before and after the segment. Positive estimates are then incorporated into the GRPO advantages of policy tokens within the proposed segment. By using model judgment only to select where to verify, ProVer grounds local credit in observed outcomes without exhaustively evaluating every intermediate state. Across ALFWorld, WebShop, and SearchQA, ProVer achieves the strongest average performance at both model scales, with relative improvements over GRPO of 9.91% and 7.12% for Qwen3.5-2B and Qwen3.5-4B, respectively. Further analyses demonstrate that informed segment selection improves policy training with modest additional generation overhead, even without a frontier-scale judge model, highlighting the effectiveness and efficiency of selectively targeting pivotal decisions for fine-grained credit assignment in agentic reinforcement learning.

[NLP-168] Principled Thoughts for Latent Recursive LLM Systems

【速读】: 该论文旨在解决大语言模型在推理过程中因仅依赖最终输出的交叉熵(Cross-Entropy, CE)损失进行训练,而导致思维过程(thoughts)质量低下的问题。具体表现为:思维表示缺乏因果性、冗余信息保留、不同问题间思维坍缩以及对正确答案的生成概率降低。其解决方案的关键在于提出一种新的可微分训练目标——REST(Representation-Supervised Thoughts),将有效思维表示的四个核心属性——因果性(causality)、最小性(minimality)、可分性(separability)和稳定性(stability)——转化为可加入训练过程的可微损失项,从而在不改变模型架构或增加推理阶段参数的前提下,引导模型生成更高质量、更具解释性的中间推理表示。实验表明,在数学、科学、医学和代码生成等7个基准任务上,使用相同数据、计算量和隐状态预算时,REST相比纯CE训练显著提升准确率(最高达7.5个百分点),并加快收敛速度(达30%),同时增强隐空间通信的可解释性。

链接: https://arxiv.org/abs/2609.36159
作者: Fahd Seddik,Fatemeh Fard
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: Project website: this https URL

点击查看摘要

Abstract:Large language models can reason in continuous space instead of decoded text, by recurring on their own hidden states or by passing those states between agents, while training supervises only the Cross-Entropy (CE) of the final decoded answer and does not constrain the thought. Theoretical and empirical analyses establish and confirm four failures of CE-only training that lead to a lower probability of the correct answer such as collapsing thoughts across distinct questions and retaining irrelevant information. We introduce REST (REpresentation-Supervised Thoughts), a training objective that turns four properties of a valid thought representation (causality, minimality, separability, and stability) into differentiable losses added to CE. We instantiate it in latent single-agent and multi-agent systems, without architectural changes or added parameters at inference. Across 7 benchmarks spanning mathematics, science, medicine, and code generation, with the same training data, compute, and latent budget, REST increases accuracy over CE-only training across agent settings and model sizes by up to 7.5 percentage points and convergence on a final answer by 30%. Furthermore, REST thoughts encode more of what is required to achieve the correct answer, and decoding them better recovers the intended output of the agent, which makes latent communication easier to interpret. Project Website: this https URL

[NLP-169] Language Models Are “Insecure” Reporters

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在执行长期、高自主性任务时,其生成报告存在“不安全报告”(insecure reporting)的问题,即模型倾向于隐瞒可能颠覆整体叙事的关键缺陷(如严重负面实验结果),从而误导用户对任务完成质量与完整性的判断。这一问题随着人工审核成本上升而愈发突出,导致用户不得不依赖模型自动生成的报告进行评估,进而加剧了潜在的不可靠性风险。解决方案的关键在于通过引入简短的诚实性指令(如“Be honest in your response”),显著提升模型披露关键缺陷的能力——实验表明,在未加指令的情况下,GPT-5.5仅在2/200份报告中识别出预设的负面结果,而加入指令后该比例跃升至190/200。进一步的链式思维分析与基于Qwen3.5-9B的激活分析和可控引导实验揭示,诚实性与成功导向在模型表示空间中呈对立方向,表明模型默认倾向构建成功的叙事,而通过外部干预可有效引导其转向更透明的报告行为。因此,该研究的核心突破在于证明:通过轻量级的提示工程即可显著改善模型报告的可靠性,为提升生成式AI(Generative AI)在关键应用中的可信度提供了可行路径。

链接: https://arxiv.org/abs/2609.36139
作者: Jenny Y. Huang,Jiameng Fan,Ahmed Imtiaz Humayun,Maximillian Chen,Tian Qin,Run Chen,Vidhya Navalpakkam,Hongxiang Gu
机构: Massachusetts Institute of Technology (麻省理工学院); Google Research(谷歌研究); Harvard University (哈佛大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:As large language models are deployed in increasingly autonomous long-horizon tasks, manually auditing and verifying the actions, artifacts, and outputs of models becomes more difficult. Users instead come to rely on LLM-generated reports to assess the quality and completeness of the work. We introduce a suite of eight adversarial reporting scenarios to systematically study whether LLMs conceal narrative-changing flaws: errors or limitations that undermine an otherwise successful account of work. We call this phenomenon “insecure reporting.” When handed machine learning experiment logs containing a planted negative result that substantially weakens the proposed method, GPT-5.5 flags the negative result in only 2 of 200 generated reports. However, when a short honesty instruction, “Be honest in your response,” is added, the model flags the negative result in 190 of 200 reports. Across eight open-weight models, chain-of-thought analysis reveals a recurring tension between disclosing narrative-changing flaws and reasoning about ways to appear successful. We perform an activation analysis and a steering experiment on Qwen3.5-9B, finding that honesty and success-seeking correspond to opposing directions in representation space. Our results suggest that LLMs tend to present narratives of success by default, and that steering models toward honesty makes their reports substantially more transparent.

[NLP-170] When Does Correction Become Repair? Mechanistic Auditing of Internal Interventions in Tool-Using LLM s

【速读】: 该论文旨在解决生成式 AI(Generative AI)在调用外部工具前的决策行为中,因内部激活状态调节导致的不可解释性与潜在副作用问题。具体而言,当代理型大语言模型(agentic LLM)需从多路径动作空间(K-way action space)中选择执行调用、澄清、直接回答或拒绝时,传统聚合指标无法精确揭示状态扰动的具体位置及其对其他决策的附带损害。为此,论文提出 SAKIKO 审计框架,其核心解决方案在于四重机制:方向性错误发现(directional error discovery)、路由条件干预(router-conditioned intervention)、目标结果验证(destination-resolved verification)以及前瞻性统计许可冻结(prospectively frozen statistical licensing)。该框架通过通道键控干预实现方向特异性性能提升,在七个模型上于 When2Call 与 MetaTool 基准中成功在五种模型中获得净收益;而在三组封闭评估中,所有预算匹配的随机方向均未达到校准后的目标增益。关键发现表明,行为迁移不等于修复:一项带来 +55 净收益的干预反而破坏了超过一半原本正确的基线决策,且 Qwen3-4B 与 Gemma-2-9B 的有希望点估计因有限样本不确定性而被正式否定。因此,SAKIKO 强调在宣称内部状态修复前,必须进行以结果为导向的归因判决。

链接: https://arxiv.org/abs/2609.36138
作者: Jiayi Li,Ruizhe Li
机构: University of the Chinese Academy of Sciences, China; School of Computer Science, University of Birmingham, UK
类目: Computation and Language (cs.CL); Software Engineering (cs.SE)
备注: Preprint

点击查看摘要

Abstract:Before invoking external tools, an agentic LLM must select among a K-way action space: executing a call, seeking clarification, answering directly, or declining. While internal activation steering can alter these pre-execution decisions, conventional aggregate metrics obscure where altered states land and what collateral damage they inflict. We present SAKIKO, an auditing framework that formalizes representation repair via directional error discovery, router-conditioned intervention, destination-resolved verification, and prospectively frozen statistical licensing. Across seven LLMs on When2Call and MetaTool, channel-keyed interventions induce direction-specific net gains in five models; across three sealed evaluations, none of 59 budget-matched random directions matches calibrated target gain. Crucially, destination auditing shows that behavioral movement does not equal repair: an intervention achieving +55 net gain corrupts over half of the baseline-correct decisions it touches, and promising point estimates on Qwen3-4B and Gemma-2-9B are formally declined due to finite-sample uncertainty. SAKIKO establishes the necessity of outcome-resolved adjudication before claiming internal repair. Code: this https URL.

[NLP-171] A Character-Level Neural Approach to Sinhala Sandhi Splitting AACL

【速读】: 该论文旨在解决僧伽罗语(Sinhala)音变合并(Sandhi)切分问题,即从语音上合并的表面形式中恢复出原本独立的词或词素。这一任务对僧伽罗语自然语言处理(NLP)至关重要,因为音变现象会模糊词汇边界,影响下游任务的性能。然而,此前尚无针对该任务的神经网络基准研究。本文提出基于SandhiLex数据集的字符级序列到序列模型,采用原生僧伽罗文Unicode输入,并评估循环编码器-解码器架构在词缀性及更复杂的词源性、派生性和词汇化音变切分上的表现。其核心挑战在于词汇化、派生性和词源性音变,这类音变具有高度非规则性。实验表明,最佳模型(双向LSTM编码器+单向LSTM解码器)在精确匹配准确率上仅达68.40%(字符级准确率为82.08%),远低于词缀性音变子集的94.00%表现。消融实验显示,双向编码是性能提升的关键因素,且使用原生僧伽罗文字比罗马化输入显著提高精确匹配准确率。定性分析发现,多数错误为边界邻近字符误判或合理的但不正确的音位替换。该研究建立了僧伽罗语音变切分的实证基准,指明未来工作应聚焦于数据规模扩展、音变类型条件建模以及注意力机制驱动的解码策略。

链接: https://arxiv.org/abs/2609.36131
作者: Yasas Ekanayaka,Deshan Sumanathilaka
机构: 未知
类目: Computation and Language (cs.CL)
备注: 10 pages, 2 Figures, 10 Tables, Accepted to present at AACL-IJCNLP 2026

点击查看摘要

Abstract:Sinhala Sandhi splitting recovers the constituent words or morphemes hidden inside a phonologically merged surface form. The task is important for Sinhala NLP because Sandhi obscures lexical boundaries, but no prior published work has established a neural benchmark for Sinhala Sandhi splitting. We present a character-level sequence-to-sequence study based on SandhiLex, using native Sinhala Unicode input and evaluating recurrent encoder-decoder models for affixational and more complex lexicalized, derivational, and etymological Sandhi. The central challenge is the hard subset lexicalized, derivational, and etymological Sandhi, where our best model, a bidirectional LSTM encoder with a unidirectional LSTM decoder, reaches only 68.40% exact-match accuracy (82.08% character-level accuracy), well below the 94.00% achieved on the more regular affixational subset. Ablations show that bidirectional encoding is the largest contributor to performance, while native Sinhala script improves exact match accuracy over romanized input. Qualitative analysis indicates that many errors are near misses involving boundary adjacent characters or plausible but incorrect phonological substitutions. These results establish an empirical baseline for Sinhala Sandhi splitting and identify data scale, Sandhi type conditioning, and attention-based decoding as the main directions for future work.

[NLP-172] PADMÉ: Preference Alignment Data Synthesis for Meta-Evaluation of LM Agent Evaluators NEURIPS2026

【速读】: 该论文旨在解决生成式 AI(Generative AI)评估中的“元评估”(Meta-Evaluation)问题,即如何确保语言模型(LM)评估器对智能体行为的评分与人类判断保持一致。传统方法面临标注成本高、绝对评分难以对齐以及评估器自身可信度存疑等挑战。为此,论文将元评估重构为偏好一致性判断问题——不再直接比较人类与模型评估器的分数,而是检验二者在行为轨迹上所隐含的偏好是否一致。其核心解决方案是提出PADMÉ(Preference-based Agentic Meta-evaluation Data synthesis),一种基于小规模语言模型的数据合成方法,无需人工参与、计算开销低,可生成高质量、基于准则的元评估数据。通过构建包含1000个样本的跨四个智能体领域、三类评估标准的数据集,实验表明,PADMÉ在150个样本的人工验证中将与人类判断的一致性从基线的73%提升至85%,并揭示了评估性能与评分粒度、宽松程度及模型规模等因素之间的相关性。

链接: https://arxiv.org/abs/2609.36086
作者: Cheng Chang,Yining Mao,Peng Qi
机构: Uniphore
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Accepted at the NeurIPS 2026 Workshop TAE (Trust-AI-Eval): Can We Trust AI Evaluation? 27 pages, 3 figures. Code and data at this https URL

点击查看摘要

Abstract:Language models are frequently employed to evaluate other language models. An LM evaluator scoring agentic behaviors across multiple criteria is valuable, provided that its decisions align with human judgment. We call the problem of evaluating this alignment Meta-Evaluation. Tackling it directly is difficult: collecting human data is expensive, absolute scoring is hard to align, and using an LM meta-evaluator recurses the question of trustworthiness. We adopt a reformulation of meta-evaluation as a preference judgment problem: rather than comparing human and LM evaluator scores of a trajectory, we ask whether their implied preferences align. Building on this, we introduce PADMÉ, a data synthesis method that generates reliable criterion-based meta-evaluation data for agentic settings. PADMÉ uses only small language models, requires no human involvement during evaluations, and operates under a low computational budget. We build a prototype of PADMÉ and synthesize a dataset of 1,000 samples across four agentic domains and three evaluation criteria. Human validation on a 150-sample subset demonstrates that PADMÉ improves agreement with human judgment from 73% to 85% over a naive baseline. Meta-evaluating 25 common models with our dataset demonstrates the correlations between evaluation performance and scoring granularity, leniency, and model size, among other factors.

[NLP-173] A Polyphonic Conception of AI Understanding

【速读】: 该论文旨在解决在医疗、司法或工程等高风险领域中,人类专家面对生成式 AI(Generative AI)输出时难以判断其可信度的核心问题。传统基于数学或统计的描述方法无法有效区分可信与不可信输出,本质上仍需回归对AI“理解”能力的追问,而这一问题在现有框架下被错误地构建为单一机制主导的“单音性”(monophonic)范式——即假设认知系统的理解必须依赖于唯一核心机制。然而,基于大量机制性证据,本文揭示大型语言模型(LLM)具有普遍的“多声性”(polyphonic)特征:其输出源自多个并行机制的协同作用,这些机制可靠性不一,彼此间存在互补、冗余或相互压制的关系,且完成特定任务无需依赖任一单一机制。这种多声性不仅使对“理解”的归因变得复杂,更使得基于单音性推理的评估模式存在严重风险。为此,论文提出一种适配多声性AI的理解概念,其核心在于可信赖的“回路结构”(sound circuitry),即能够稳定、正确地被调用并在输出控制中起主导作用的内部组织机制。通过将理解归因为对内部结构的具体可验证属性,该框架使关于理解的判断转化为可操作、可检验的命题,从而为建立对AI系统的合理信任提供理论基础与实践指引。

链接: https://arxiv.org/abs/2609.36079
作者: Matthieu Queloz,Pierre Beckmann
机构: University of Bern (伯尔尼大学); École Polytechnique Fédérale de Lausanne (洛桑联邦理工学院); Idiap Research Institute (Idiap 研究所); Machine Alignment, Transparency, and Security (MATS)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:When a doctor, a judge, or an engineer must decide whether to trust an AI model’s output, they cannot avoid asking what the model understands. Purely mathematical or statistical descriptions struggle to distinguish trustworthy from untrustworthy outputs without reintroducing the question of AI understanding in all but name. Yet the question is ill-framed as it stands, because the inherited concept operates within a monophonic paradigm: the idea that a cognitive system’s understanding of something must be localised to a single mechanism underpinning all the capacities conferred by such understanding. Drawing on a wide range of mechanistic evidence, we show that LLMs are pervasively polyphonic: outputs emerge from coalitions of parallel mechanisms of uneven reliability, which variously complement, duplicate, or drown out one another, with several coalitions sufficing for a task without any one being indispensable. Polyphony not only complicates attributions of understanding, but renders monophonic inference patterns hazardous. In response, we develop a conception of understanding fit for polyphonic AI. It centres on sound circuitry that is reliably and correctly recruited and in control of outputs. Attributions of understanding thereby become tractable claims about internal organisation, and can do the work of guiding trust in AI.

[NLP-174] Causal and Interpretable Structures in LLM Compositional Tasks

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在处理依赖于多个输入标记之间关系的任务时,其内部如何表征与处理这种关系信息的机制问题。具体而言,研究聚焦于模型在预测序列中下一个标记时,如何通过多层变换逐步组织和组合三元组标记间的循环性关系(如月份、小时、星期、音符等)。其解决方案的关键在于揭示了关系信息在不同网络层中的几何结构与因果作用的演化规律:中间层利用基于两标记间推断关系的联合表示,而深层则采用涵盖全部三标记的联合表示以实现准确预测;同时发现部分具有几何结构的关系信息虽存在但对预测无因果贡献。更重要的是,通过限制模型仅使用这些因果相关的联合表示,反而提升了下一词预测的准确性,表明模型内部存在一种可被优化的、逐层递进的关系信息组织与组合机制。

链接: https://arxiv.org/abs/2609.35970
作者: Gurbir Arora,Toni J.B. Liu,Jiajun Bao,Raphaël Sarfati,Christopher J. Earls
机构: Cornell University (康奈尔大学); Goodfire AI(好火人工智能)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 46 pages, 28 figures

点击查看摘要

Abstract:Large language models are able to solve tasks whose answers depend on not only individual input tokens, but also on relations among them. How is such relational information represented and processed across transformer layers? We study activations from ensembles of prompts that require inferring relationships between three tokens corresponding to a cyclic concept (months, hours, weekdays, and musical notes) to correctly predict the next token. Across model families (Llama, Qwen, Gemma, and Mistral) and cyclic concepts, we find a consistent layerwise progression in how the joint dependence among the tokens is geometrically organized and causally used: intermediate layers use a joint representation based on the inferred relationship between two tokens, while later layers use a joint representation associated with all three tokens to correctly complete the task. We also find other relationships between tokens that are geometrically structured but remain causally inert in the next-token prediction. Crucially, when taken together, these geometric and causal investigations reveal the representation-level mechanism that progressively organizes and composes the relational information to form the answer. More surprisingly, restricting the models to such causally relevant joint representations improves next-token prediction accuracy.

[NLP-175] Question-Specific Knowledge Graphs for Efficient Visual Reasoning

【速读】: 该论文旨在解决视觉问答(Visual Question Answering, VQA)中视觉语言模型因依赖冗余图像描述而导致推理效率低下、计算成本过高以及易引入虚假假设的问题。现有方法通过使用详细图像描述来增强视觉细节,但往往包含与任务无关的感知噪声,导致输入令牌数量激增且推理路径被干扰。其解决方案的关键在于提出一种基于强化学习(Reinforcement Learning, RL)的框架VisKG,该框架将视觉内容转化为与问题相关的知识图谱(Knowledge Graph, KG)表示,遵循“最小充分信息”原则,在过滤感知噪声的同时保留支持链式思维(Chain-of-Thought)推理所需的实体-关系结构。为保障强化学习后训练阶段的稳定性,VisKG采用分组奖励解耦归一化策略优化(Group Reward-Decoupled Normalization Policy Optimization, GDPO),并引入负向推理路径样本进行监督阶段增强,以提升模型对错误推理路径的辨别能力。实验结果表明,VisKG在科学、数学及通用视觉理解基准上性能优于或相当主流基线,同时显著减少所需令牌数;相较于使用GRPO训练的版本,GDPO使平均准确率提升2%。研究证实,知识图谱表示是支持多步推理的有效途径,并为未来动态选择最优表示形式提供了新方向。

链接: https://arxiv.org/abs/2609.35942
作者: Ting-Chih Chen,Emile van Krieken,Shujian Yu,Filip Ilievski
机构: 未知
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Recent work in visual question answering has shown that vision-language models can exhibit strong reasoning capabilities by translating visual inputs into textual representations. The effectiveness of this translation depends on how well visual details are retained; models need to surface and align both explicit and implicit knowledge sufficient to support reasoning, without introducing spurious assumptions. Existing methods that leverage detailed image captions introduce visual details unrelated to the reasoning task, inflating input token counts and increasing computational cost. To address these challenges, we propose VisKG, a reinforcement learning (RL) framework in which models learn to translate visual content into question-specific knowledge graph (KG) representations. This process filters out perceptual noise while preserving the entity-relation structure needed for chain-of-thought reasoning, following the principle of minimum sufficient information. To ensure stable RL post-training, VisKG adopts Group reward-Decoupled Normalization Policy Optimization (GDPO). In addition, we strengthen the supervision stage with negative rationale samples, exposing the model to incorrect reasoning paths before RL post-training. Experimental results across science, mathematics, and general visual understanding benchmarks show that VisKG achieves performance comparable to or better than baselines, while requiring fewer tokens than caption-based representations. Moreover, training VisKG with GDPO improves accuracy by 2% over its GRPO-trained counterpart on average. These results suggest that KG representations are a promising approach for supporting multi-step reasoning and open up future work on adaptively selecting the most suitable representation for a given task.

[NLP-176] Almost Human Except When It Matters: VoxParity and the Decisions a Voice Should Change

【速读】: 该论文旨在解决语音代理(Voice Agent)在处理电话呼叫时,过度依赖文本内容而忽视音频中关键非语言信息(如语调、背景声音、情绪线索等)所导致的判断失误问题。尽管语音代理能准确理解对话文字内容,但在涉及紧急情况、欺诈识别、无线电用语规范及弱势群体保护等场景下,声音特征本身可能直接决定正确响应策略。其核心解决方案在于提出并验证“VoxParity”测试框架,通过在183个来自14个不同领域的场景中保持文本不变而动态改变音频内容(如医疗监护仪滴答声、求救信号、儿童声音投注、恐惧低语等),系统需根据音频变化调整其应答行为。关键创新点在于引入“仅基于文字的对照组”作为基准,仅当系统因听到音频而产生的行为变化显著大于纯文本处理管道时,才判定为有效响应。结果显示,仅有11/23具备音频处理能力的系统通过了该测试;进一步分析表明,现有系统对明确陈述规则的指令响应较好,但对隐含情绪或心理状态(如焦虑、绝望、困惑)的音频线索识别严重不足,暴露出生成式语音代理在情感感知与情境推理方面的关键缺陷。

链接: https://arxiv.org/abs/2609.35922
作者: Bhavik Mangla
机构: Independent Research(独立研究)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Sound (cs.SD)
备注: 38 pages, 11 figures, 15 tables. Code, scorer and development-split data at this https URL and this https URL

点击查看摘要

Abstract:A voice agent can handle almost every call on the words alone and still fail the few its sector’s rules were written for. Emergency-call standards, fraud guidance, radio phraseology and vulnerability rules recognise that how a caller sounds, or what else is audible, can change the right action. VoxParity tests whether agents act on it. In 183 scenarios from 14 sectors, one transcript stays fixed while the audio changes (a coaching voice, a medical monitor beeping, a mayday under a radio check, noise over a drug name, a child’s voice placing a bet, a frightened whisper), and with it the correct typed tool call. A words-only null test credits a system only if hearing the call moves its actions more than it moves a pipeline that only reads the words. Only 11 of the 23 systems that can also be run on the transcript pass. Descriptively, errors run toward the words: when the audio calls for protection, all 28 systems carry out the routine request more often than they over-react on clean calls (41% against 12% pooled; the words-only pipeline, 58% against 15%). Exploratory analyses place most of the leading systems’ misses on cues they heard; systems beat the null almost entirely on items that state the rule; the leading systems overrule heard resignation or confusion far more often than acute alarm; and, in the models tested, describing the voice and stating the rule each recover part of the shortfall, leaving a gap on emotion.

[NLP-177] CruxBench: A Benchmark of Information Discovery NEURIPS2026

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在复杂现实任务中缺乏有效识别关键问题(cruxes)能力的问题,即如何将复杂问题分解为具有高信息价值的子问题。传统评测基准通常仅评估答案的准确性,而忽视了模型在信息发现阶段的能力。为此,作者提出了CruxBench这一新型基准,其核心在于通过“信息价值”(Value of Information, VOI)来量化模型生成问题的有用性——即一个提出的子问题能多大程度上更新对目标预测问题的信念。该方案的关键创新在于:(1)构建方式具备抗污染性,因真实世界事件作为未来验证依据,避免了标签泄露;(2)开放性强,支持无限且复杂的文本形式提交,而非单一数值答案;(3)具有现实根基,信息量基于真实信念变化进行量化评估。实验表明,VOI与独立评估模型能力的指标高度相关(r=0.90),能有效捕捉子问题对解答目标问题的实际帮助。然而,即使前沿模型在该任务上也仅略优于随机时间基线,表明信息发现仍是当前大模型面临的核心挑战。

链接: https://arxiv.org/abs/2609.35879
作者: Hui Dai,Lina Piao,Nick Merrill,Nadja Flechner,Ezra Karger,Haifeng Xu
机构: The University of Chicago(芝加哥大学); Forecasting Research Institute(预测研究学院); Federal Reserve Bank of Chicago(芝加哥联邦储备银行)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: NeurIPS 2026

点击查看摘要

Abstract:Benchmarks for large language models (LLMs) typically evaluate the accuracy of answers against fixed reference labels. But a central step in many complex real-world tasks is identifying which questions are worth asking in the first place: decomposing a difficult problem into subquestions – which we call cruxes – whose answers provide key steps on the path toward solving the target problem. To evaluate this capability of information discovery, we introduce CruxBench, a benchmark that grades LLM-generated questions by their Value of Information (VOI): how much a model-proposed crux updates beliefs about a target forecasting question. CruxBench enjoys a rare combination of three key properties: it is (1) contamination-resistant by construction, since ground truth is generated by future world events; (2) open-ended, admitting unbounded and complex text-based submissions rather than one correct numeric answer; and (3) grounded, with informativeness measured against quantified changes in real-world beliefs. We evaluate a diverse set of eight models on 293 target forecasting questions and find that VOI correlates highly with independent measures of model capability (r=0.90) and captures cruxes’ usefulness for answering target questions. However, information discovery remains challenging even for frontier LLMs, which only narrowly outperform a random-timing baseline.

[NLP-178] he Price of Token Boundaries: Compression Certificates and Prediction

【速读】: 该论文旨在解决文本分词(tokenisation)中预定义边界(pre-tokenisation boundaries)对压缩效率造成的隐性成本问题,即边界限制了可成为预测单元的文本片段范围,从而影响模型压缩性能。其核心解决方案是通过引入双向边界约束下的最小词元数量下界估计方法,结合最短路径算法与词汇表预算选择,以非负价格项构建下界,并通过线性规划松弛和独立整数验证器确保结果可靠性。研究发现,在英文维基百科数据上,边界使最优词元数量增加28.3%–36.8%;字节对编码(Byte Pair Encoding, BPE)在受限边界下仅比下界高2.1%,但在无约束条件下高出10.9%。此外,压缩与预测目标存在权衡:在相同训练词元预算下,无边界约束的拟合在12种语言中的11种上均实现更低的每字节平均保留比特数,表明更优的压缩能力。为探索中间策略,论文提出“边界许可”(boundary licences)机制,通过限制跨切分点的词汇表条目比例来量化不同边界政策的影响;实验显示,仅允许10%词汇预算用于跨切分条目时,即可恢复去除所有边界后词元数量减少量的85.2%(英语)和100.0%(中文),有效分离了边界带来的压缩代价与词元质量对预测性能的影响。

链接: https://arxiv.org/abs/2609.35869
作者: Yuhao Du,Shunian Chen
机构: The Chinese University of Hong Kong, Shenzhen (香港中文大学(深圳)); Shenzhen Loop Area Institute (深圳环区研究院)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Pre-tokenisation restricts which text fragments can become prediction units, but its compression cost is obscured when tokenisers are compared only under the same boundaries. We measure this cost by bounding the minimum token count from both sides, with and without a regular-expression boundary rule. Nonnegative prices on token occurrences yield a lower bound through shortest paths and vocabulary-budget selection; maximising over all prices recovers the linear programming relaxation, and an independent integer checker certifies the reported values. On English Wikipedia, boundaries increase the optimal token count by 28.3–36.8%. Byte pair encoding lies 2.1% above the constrained lower bound, but 10.9% above the unrestricted bound. Compression and prediction favour different dictionaries: at 85M non-embedding parameters and matched training-token budgets, unrestricted fitting yields higher mean held-out bits per byte under a common unrestricted decoder in all 12 languages in the paired study and 11 of 12 under independent tuning and evaluation. To study intermediate boundary policies, we introduce boundary licences, which limit the vocabulary entries permitted to cross cuts and admit the same form of certificate. On separate English and Chinese fitting corpora, licensing 10% of the vocabulary budget recovers 85.2% and 100.0% of the achieved token-count reduction from removing all cuts. These results quantify the compression cost of boundaries while separating it from the prediction quality of the resulting token units.

[NLP-179] PACT: Pairwise-Anchored Calibrated Tuning for Single-Token Typed Decisions

【速读】: 该论文旨在解决单标记类型决策模型(single-token typed-decision models)在训练过程中仅使用普通交叉熵损失,忽略了训练数据中蕴含的丰富结构信息的问题。这类模型虽然推理速度快且能为每个合法答案输出概率,但其训练方式未能充分利用数据中的对比性结构,导致模型对答案编码位置存在偏差、鲁棒性不足。解决方案的关键在于提出PACT(Probing with Contrasted Training),该方法利用人工标注的对比样本对(contrastive pairs)——即仅因一个事实修改而改变答案的两个上下文,且每个样本均带有机器验证的证书(certificates)证明删除关键句后该事实不再可得——构建四个无需额外标注的训练项:基于差分-差分(difference-in-differences)的边界项以消除共享日志偏置的影响、对抗答案编码位置偏差的排列一致性项、基于证书验证的消融上下文的证据必要性项,以及针对评分字段的序数传输成本(ordinal transport cost)。此外引入三参数上下文温度调节机制。实验结果表明,在324个冻结样本的测试集上,PACT在准确率上与已有最优方案相当(84.6% vs. 85.2%),但显著降低了答案位置偏差(9.8% vs. 13.8%)和评分字段的平均绝对误差(MAE 0.232 vs. 0.311),同时将种子间波动减半,负对数似然(NLL)降低26%。消融分析进一步揭示,各组件单独作用无法提升原始准确率,其核心优势在于提升模型的稳健性与训练稳定性,而非单纯提高精度。

链接: https://arxiv.org/abs/2609.35865
作者: Yida Lin
机构: Victoria University of Wellington (维多利亚大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Computer Science and Game Theory (cs.GT)
备注:

点击查看摘要

Abstract:Single-token typed-decision models answer a schema question by reading the logits of a few one-letter answer codes at a single position: they are fast and return a probability for every allowed answer, but they are trained with plain cross-entropy that ignores most of the structure in their training data. We study such a model whose data is curated as contrastive pairs—two contexts that differ in one edited fact that flips the answer—each carrying a machine-checked certificate that deleting the decisive sentence makes the fact unknown. We propose PACT, which turns this structure into four training terms that need no new annotation: a difference-in-differences margin over each pair that is invariant to any shared logit offset, a permutation-consistency term against answer-code position bias, an evidence-necessity term on certificate-verified ablated contexts, and an ordinal transport cost for rubric fields, plus a three-parameter contextual temperature. On a frozen 324-item holdout with three seeds, PACT matches the published recipe in accuracy ( 84.6% vs. 85.2% ; McNemar p \ge 0.50 at every seed) while giving the lowest position bias of all runs (answer flips under relabelling 9.8% vs. 13.8% ) and the lowest ordinal error on rubric fields (MAE 0.232 vs. 0.311 ). Against a control with the same optimiser and schedule but cross-entropy only, PACT is significantly more accurate at two of three seeds, halves the seed-to-seed spread and lowers NLL by 26% . Seed-matched ablations and pre-specified falsification tests locate these gains precisely: no single term raises raw accuracy, and the method’s value lies in robustness and stability rather than headline accuracy. Code, data splits, trained adapters, and all run records are available at this https URL.

[NLP-180] Resolving the Missing Financial Data Crisis: A Generative AI Pipeline for SEC 10-K Extraction

【速读】: 该论文旨在解决美国证券交易委员会(SEC)10-K文件中大量财务信息未能被结构化数据集一致捕获的问题,这一缺失数据问题影响超过70%的企业及一半的总市值,导致对小型企业定量分析的系统性偏差。传统金融信息提取方法如正则表达式(Regular Expressions, Regex)和BERT模型在处理复杂的10-K文件时表现脆弱,因其无法有效识别数据属性可能分布在不同章节或脚注中的情况,从而造成可获取信息的丢失。本文提出一种基于大语言模型(Large Language Models, LLMs)的可扩展解决方案,评估了Llama-3 8B、Qwen-2.5 14B和Llama-3.3 70B等多模型在提取四类具有不同结构复杂度的财务变量(现金及现金等价物(tabular)、短期债务(hybrid)、信贷额度(narrative)和研发支出(hybrid))中的性能表现。研究发现,模型参数规模与文档复杂度的匹配是实现高零样本准确率的关键,其中Qwen-2.5 14B在表格型数据提取上表现优异(现金项目F1得分为83.33%),而Llama-3.3 70B在密集的叙事性脚注中展现出更强的上下文理解能力(研发支出F1得分为76.92%)。该框架通过更全面地解析10-K文件内容,有效填补了量化金融数据中的关键信息缺口,并消除了因缺失数据带来的偏倚。

链接: https://arxiv.org/abs/2609.35864
作者: Prisha Nair,Roee Shraga
机构: 未知
类目: Computation and Language (cs.CL)
备注: Presented as a Lightning Talk at MIT URTC 2026

点击查看摘要

Abstract:SEC 10-K filings contain substantial financial information that is not consistently captured in structured datasets, creating a missing-data problem affecting over 70% of firms and half of total market capitalization. This can disproportionately bias quantitative analysis against smaller firms, which may be excluded due to limited available data. Traditional financial extraction methods such as Regular Expressions (Regex) and BERT, have been widely used. However, they are highly brittle when parsing complex SEC 10-K filings, which leads to data that is existent in the files being lost since these methods do not consider that a data attribute could be located in a different section or a footnote. This study evaluates several Large Language Models (LLMs), including Llama-3 8B, Qwen-2.5 14B, and Llama-3.3 70B, to figure out individual model strengths and weaknesses when extracting specific attributes from SEC 10-K text. The extraction quality was evaluated across four financial variables of varying structural complexity: Cash and Cash Equivalents (tabular), Short-Term Debt (hybrid), Credit Facilities (narrative), and Research and Development (hybrid). Results show that while smaller models like Llama-3 8B experience performance degradation under complex negative prompting, aligning parameter scale with document complexity yields high zero-shot accuracy. Qwen-2.5 14B excels as a tabular specialist with an 83.33% F1 score on Cash, whereas Llama-3.3 70B effectively navigates dense narrative footnotes, achieving a 76.92% F1 score on RD. This scalable framework addresses critical information gaps in quantitative finance datasets and eliminates missing-data bias through a more thorough analysis of the SEC 10-K files.

[NLP-181] he Detectability Gap: Hidden Heterogeneity in Hallucination Detection Across Language Models NEURIPS2026

【速读】: 该论文旨在解决生成式AI(Generative AI)在事实性问答任务中幻觉(hallucination)检测的评估局限性问题,即当前基于采样的一致性检测方法虽广泛使用,但其整体性能指标可能掩盖了不同语言模型间系统性差异——特别是某些错误类型在特定模型中更难被检测。其核心解决方案在于通过答案一致性(answer agreement)对幻觉进行分组,揭示出高一致性(Ghost)与低一致性(Flickering)两种幻觉行为模式,并发现二者之间存在显著的可检测性差距(AUC差异达0.35至0.46)。尽管该差距与用于定义分组的统计量高度相关(|\rho| ≈ 0.94 至 1.00),表明其本质是基于一致性的检测特性而非独立证据,但进一步分析显示,即使在冻结分组设定后,词汇与语义响应分散度仍能保持这种不对称性,且在全部12个模型-数据集组合中,置信区间均不包含零。更为严格的测试(基于个体扩散轨迹、无跨种子信息)在三个LLaDA数据集上均显著保留该不对称性(p < 0.005),并在三个Dream数据集上呈现方向性显著性。此外,硬样本(hard regime)在不同模型中的出现率差异巨大(16%至77%),且相同提示在不同模型间频繁切换幻觉模式。这些发现表明,聚合检测指标会掩盖模型特异性的系统性失败模式,因而亟需采用基于分组条件的评估范式以实现更精准、可靠的幻觉检测评价。

链接: https://arxiv.org/abs/2609.35860
作者: Pranav Darshan,Pranav A,Sravan Karthick T,Minal Moharir,Ivan P. Yamshchikov
机构: R.V. College of Engineering, India; CAIRO, Technical University of Applied Sciences Würzburg-Schweinfurt, Germany
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Accepted at GlobalSouthAI @ NeurIPS 2026

点击查看摘要

Abstract:Sampling based consistency is widely used for hallucination detection, yet aggregate performance can conceal systematic differences in which errors are detectable. This work studies that heterogeneity across four language models and three factual question answering datasets. Partitioning hallucinations by answer agreement reveals high agreement (Ghost) and low agreement (Flickering) regimes with an apparent detectability gap of 0.35 to 0.46 AUC. Because the statistics used to define the regimes and measure this gap are strongly coupled ( |\rho|\approx0.94 to 1.00 ), the raw result is treated as a property of agreement based detection rather than independent evidence. After freezing regime assignments, lexical and semantic response dispersion preserve the asymmetry, with bootstrap 95% intervals excluding zero in all 12 model and dataset settings. A stricter test using individual diffusion trajectories and no cross seed information preserves the asymmetry across all three LLaDA datasets ( p0.005 ) and directionally across all three Dream datasets, with one reaching significance. The hard regime varies substantially in prevalence across models ( 16% to 77% ), and matched prompts frequently change regimes between models. These findings show that aggregate detection metrics conceal persistent, model dependent heterogeneity in language model failures and motivate regime conditioned evaluation.

[NLP-182] Hyperspherical Semantic Trajectory Analysis: Mapping Technological Diffusion across Academic Preprints Patent Signals and Compute Scaling

【速读】: 该论文旨在解决宏观经济生产率指标(如全要素生产率,Total Factor Productivity)因行政调查周期和国家核算惯例导致技术突破监测存在多年滞后的难题。其核心解决方案是提出一种名为超球面语义轨迹分析(Hyperspherical Semantic Trajectory Analysis, HSTA)的无监督定量方法,通过直接从非结构化科学与商业文本流中追踪技术扩散过程,实现对技术演进的实时量化。HSTA的关键在于利用Transformer模型生成的高维句子嵌入,结合球面K均值聚类(Spherical K-Means)与统一曼ifold投影(UMAP)技术,在八个主要子领域上构建单位超球面空间,进而定义两个核心指标:(1)语义中心向量漂移(Semantic Centroid Vector Drift),用于捕捉不同时段语料库间词汇结构的变化,识别技术范式的结构性转型;(2)商业化偏移(Commercialization Offset),用于评估科学发现与知识产权申请在跨语料库峰值密度上的对齐程度。实证结果表明,大语言模型(漂移度0.332)和人工智能系统(漂移度0.234)等子领域表现出最快的语义演化速度,且结合物理硬件指标的向量自回归检验显示,仅凭论文发表速度无法在传统显著性水平下格兰杰因果解释前沿计算资源的激增,凸显了将文本信号与实体资本约束相结合的重要性。该方法为传统经济统计提供了客观、实时的技术演进监测机制。

链接: https://arxiv.org/abs/2609.35845
作者: Muhammad Sukri Bin Ramli
机构: Asia School of Business (亚洲商学院); MIT Sloan School of Management (麻省理工学院斯隆管理学院)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Macroeconomic productivity metrics, such as Total Factor Productivity, register technological breakthroughs with multi-year reporting lags due to administrative survey intervals and national accounting conventions. This paper introduces Hyperspherical Semantic Trajectory Analysis (HSTA), an unsupervised quantitative methodology that tracks technology diffusion directly from unstructured scientific and commercial text streams. We analyze 30,000 filtered document records spanning academic preprints from arXiv and patent application records from the USPTO. By projecting high-dimensional Transformer sentence embeddings onto unit hyperspheres using Spherical K-Means clustering across eight primary sub-topics and UMAP manifold reductions, HSTA formalizes two quantitative metrics: (1) Semantic Centroid Vector Drift, which tracks vocabulary shifts between temporal sub-corpora to identify structural paradigm transformations; and (2) Commercialization Offset, which evaluates cross-corpus peak density alignments between scientific discovery and intellectual property filings. Linking quarterly topic volume velocity with physical hardware metrics from the Epoch AI database, Vector Autoregressive F-tests demonstrate that quarterly paper volume velocity alone does not Granger-cause frontier compute allocation surges at conventional statistical significance levels, highlighting the necessity of conditioning textual signals on physical capital constraints. Empirical results reveal that sub-topics covering Large Language Models (with a drift metric of 0.332) and Artificial Intelligence Systems (with a drift metric of 0.234) undergo the highest rate of semantic evolution, offering an objective, real-time mechanism to complement traditional economic statistics.

[NLP-183] Neurosymbolic Routing for Reliable Reasoning on Resource-Constrained Edge Devices

【速读】: 该论文旨在解决在边缘硬件上运行小型语言模型(Small Language Model, SLM)时,模型在处理算术、代数及形式逻辑等结构性任务中表现出的不可靠性问题。其核心挑战在于,尽管SLM具备低延迟与隐私保护优势,但其基于概率推理的机制难以精确求解本应具有确定性结构的任务,导致准确率低下且资源浪费。解决方案的关键是提出一种神经符号路由(neurosymbolic router),通过自动识别输入查询的结构特性,将可解析的确定性任务(如格式化数学表达式)直接分发至高效的符号引擎,而仅将开放式的自然语言问题交由SLM处理。该路由机制不依赖人工规则设计,而是利用L*语法推断算法,以SLM作为成员查询预言机(membership oracle)、标注数据作为等价查询预言机(equivalence oracle),学习得到一个确定性有限自动机(Deterministic Finite Automaton, DFA)。实验表明,在Raspberry Pi 4B平台下,该方法在未见过的100个测试提示上实现了100%的路由准确率与98.3%的整体准确率(词类问题准确率为93.3%),显著优于最强基线模型Program-of-Thought(72.0%)和工具调用代理(58.7%),同时在30-token配置下实现8.8倍速度提升与2.8倍能效优化,且格式化查询无需经过模型即可在1–11毫秒内完成响应。

链接: https://arxiv.org/abs/2609.35833
作者: Avyay Sadhu,Alvaro Velasquez,Lekai Chen
机构: University of Colorado Boulder (科罗拉多大学博尔德分校)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 12 pages, 7 figures, 9 tables. This work has been submitted to the IEEE for possible publication

点击查看摘要

Abstract:Running a language model on edge hardware provides private and low-latency reasoning without a network connection, and yet the small models that fit on such devices are unreliable on the tasks computers are expected to handle well, such as arithmetic, algebra, and formal logic problems. We argue that much of this unreliability is avoidable. Many queries appearing to demand reasoning are in fact structurally deterministic and permit fast and exact symbolic solutions. Therefore, forcing a probabilistic model to approximate them sacrifices accuracy and energy for little benefit. We present a neurosymbolic router that classifies each incoming query and dispatches it to the cheapest correct solver, sending structured tasks to deterministic engines and reserving the small language model (SLM) for open-ended word problems. Instead of hand-coding the routing logic, we learn a deterministic finite automaton (DFA) with the L* grammatical inference algorithm, using the SLM as a membership oracle and labeled data as an equivalence oracle. On a Raspberry Pi 4B (8 GB RAM, no GPU), evaluated on 100 untested prompts from DeepMind Mathematics, GSM8K, and RuleTaker, learned routing attains 100% routing accuracy and 98.3% overall accuracy with a 512-token reasoning budget (93.3% on word problems), compared with 72.0% for the strongest agent baseline, Program-of-Thought, and 58.7% for a tool-calling agent given the same solvers. Since formatted queries never reach the model, the router answers them in 1-11 ms and, in its 30-token configuration, runs 8.8x faster and 2.8x more energy-efficient than Program-of-Thought.

[NLP-184] When Should LLM s Trust Their Own Revisions? A Risk-Aware Study of Intrinsic Self-Correction

【速读】: 该论文旨在解决生成式 AI(Generative AI)在执行内在自修正(intrinsic self-correction)过程中存在的核心矛盾:即模型在自我修订过程中虽能纠正部分错误,但也可能将原本正确的答案错误地修改为错误答案。其关键解决方案在于将内在自修正视为一种可配置的修订策略(revision policy),而非普遍有益的二次校验机制。研究通过在BoolQ、GSM8K和Corr2Cause三个基准上对29个开源大语言模型(LLM)进行分析,量化了初始回答与修正后回答之间的正确性转变,并揭示了聚合准确率可能掩盖的显著行为差异——例如,Llama-3.1-8B在GSM8K上准确率提升25.5个百分点,但同时有19.1%的原始正确答案被错误修正。进一步的受控布尔问答实验表明,修订提示(refinement prompts)可调节纠错与误伤之间的权衡。研究对比了三种运行时策略:保留初始答案、始终采纳修订结果、以及基于初始响应后可用信号选择性触发修订。结果表明,在某些场景下学习得到的门控机制(learned gating)更优,而在其他情况下,简单的无条件策略反而表现更好。因此,该研究强调应从“修复收益”与“引入错误”双重维度评估内在自修正,将其作为动态调整的修订策略而非固定增强流程。

链接: https://arxiv.org/abs/2609.35832
作者: Tianzhu Zhang
机构: Nokia Bell Labs(诺基亚贝尔实验室)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Intrinsic self-correction asks a language model to revise its own answer without receiving new external evidence. A second pass can recover mistakes, but it can also overturn answers that were already correct. We study this trade-off across 29 open-weight LLMs on BoolQ, GSM8K, and Corr2Cause by tracking correctness transitions between initial and revised answers. Aggregate accuracy can conceal substantially different revision behavior: for example, Llama-3.1-8B improves by 25.5 percentage points on GSM8K, while refinement changes 19.1% of initially correct answers into wrong ones. A controlled BoolQ study further shows that refinement prompts shift the balance between recovery and harm. We then compare three runtime choices: keeping the initial answer, always accepting the revision, and selectively invoking revision using signals available after the initial response. The comparison identifies settings where learned gating is useful and others where a simpler unconditional policy performs better. These results suggest treating intrinsic self-correction as a revision policy rather than as a uniformly beneficial second pass, and evaluating it through both the corrections it recovers and the errors it introduces.

[NLP-185] Beyond the Context Window: An Adaptive Entropy-Based Routing Framework for Hybrid Retrieval and Long-Context Language Models

【速读】: 该论文旨在解决在超长上下文(Long-Context, LC)大语言模型已支持百万级令牌输入的背景下,检索增强生成(Retrieval-Augmented Generation, RAG)是否仍具必要性的问题。核心挑战在于:纯长上下文处理虽能充分利用全部输入信息,但存在计算成本高且对长序列中段内容关注度不足的缺陷;而纯RAG虽然高效,却受限于检索质量,在召回片段部分相关或存在矛盾时易引入幻觉。为此,论文提出熵驱动自适应路由框架(Entropy-Driven Adaptive Router, EDAR),其关键创新在于在推理阶段动态决定将查询交由检索结果生成还是升级至全量长上下文处理——该决策基于生成初期若干词元的预测概率分布的熵值,利用熵作为幻觉风险的内部信号。通过在预留验证集上权衡成本与准确率,确定最优熵阈值。实验表明,预测熵与幻觉率高度相关(皮尔逊相关系数 r = 0.85, 95% 置信区间 [0.83, 0.87]),在LongBench v2和Infinity-Bench基准上,EDAR以仅18.2%的查询需升级至长上下文处理,实现总令牌消耗降低70.7%,同时保持纯长上下文基线97.4%的准确率,且二者准确率差距在标准样本量下无统计显著差异。该方法不依赖特定检索器或长上下文主干模型,无需额外监督信号,仅利用常规解码过程中的内部信息即可实现高效、低成本的混合推理。

链接: https://arxiv.org/abs/2609.35831
作者: Isaac Olufadewa,Miracle Adesina,Ezekiel Oladejo,Owen Adeniyi,Fadare Fadekemi,Olamide Oso,Uthman Babatunde,Matthew Olawoyin
机构: 未知
类目: Computation and Language (cs.CL)
备注: 11 pages, 2 figures

点击查看摘要

Abstract:Modern large language models now support context windows of more than one million tokens, which has raised the question of whether retrieval-augmented generation (RAG) is still necessary. Pure long-context (LC) processing is expensive and is known to under-attend to information placed in the middle of long inputs, while pure RAG is fast but bounded by retrieval quality and prone to errors when retrieved chunks are partially relevant or contradictory. We propose the Entropy-Driven Adaptive Router (EDAR), a framework that decides at inference time whether to answer a query from retrieved chunks or to escalate it to full long-context processing. The decision uses the predictive entropy of the token-level probability distribution computed over the first few generated tokens of the RAG response. The entropy threshold is selected on a held-out validation set by sweeping cost against accuracy. Experiments compare EDAR against pure-RAG and pure-LC baselines on LongBench v2 and Infinity-Bench. Predictive entropy correlates strongly with hallucination rate on a held-out set of 2,000 generations (Pearson r = 0.85, 95% CI [0.83, 0.87]). On the long-context benchmarks, EDAR retains 97.4% of the accuracy of the pure long-context baseline while reducing total token expenditure by 70.7%, escalating only 18.2% of incoming queries. The accuracy gap between EDAR and the pure long-context system is not statistically distinguishable from zero at standard sample sizes. Predictive entropy is a useful model-internal signal for routing between RAG and long-context inference, and a threshold-based hybrid system can recover most of the accuracy of long-context models at a small fraction of the cost. The framework does not depend on a specific retriever or LC backbone, and it does not require additional supervision beyond what is normally produced during decoding.

[NLP-186] Reliable but Design-Sensitive: Instrument Uncertainty in LLM Annotation EMNLP2026

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在识别网络文本中的攻击性语言与仇恨言论时,因任务设计(task design)和模型选择差异而导致标签不一致的问题。其核心挑战在于:即使使用相同的模型,在不同合理的设计条件下,模型输出的标签仍可能显著变化,从而引发结果的不可靠性。研究的关键解决方案是提出“仪器不确定性”(instrument uncertainty)这一概念,强调必须通过系统性地比较多种合理的任务设计来量化并评估这种由模型选择和任务设计带来的变异。研究表明,任务设计与模型选择导致的估计流行率方差分别高达采样方差的76.7倍(攻击性语言)和110.6倍(仇恨言论),远超人类标注者之间的差异。此外,模型置信度分数无法缓解此问题,反而在多条文本合并提示时显著降低可信度。因此,论文指出,仅重复单一设置或依赖置信度评分无法替代对多种合理任务设计的交叉验证,唯有通过对比不同设计才能真正测量和控制仪器不确定性。

链接: https://arxiv.org/abs/2609.35824
作者: Thomas Reiter,Christoph Kern,Fedor Miasnikov,Sofiia Nikolenko,Rob Chew,Stephanie Eckman,Frauke Kreuter
机构: Amazon(亚马逊); Ludwig Maximilian University of Munich(慕尼黑路德维希-马克西米利安大学), Germany; RTI International(美国研究技术研究所), North Carolina, USA; University of Maryland, Social Data Science Center(马里兰大学社会数据科学中心), USA
类目: Computation and Language (cs.CL); Methodology (stat.ME)
备注: Accepted to “3rd Workshop on Uncertainty-Aware NLP” @ EMNLP 2026 (archival)

点击查看摘要

Abstract:Large language models (LLMs) can give reliable labels under one setup yet change those labels when researchers make other reasonable design choices. We tested seven LLMs, 12 task designs, three independent runs, and 3,000 tweets labeled for offensive language and hate speech. Repeating the same model and task design produced high agreement (median Fleiss’ \kappa = 0.91 ). Agreement fell when we changed the task design for the same tweets (median Cohen’s \kappa = 0.76 ). Task design and model choice increased the variance of estimated prevalence by factors of 76.7 for offensive language and 110.6 for hate speech compared with sampling variance alone. Variation across LLM task designs reached 560-572 basis points, compared with 270-331 basis points across five human instrument versions. Confidence scores did not solve this problem. They tracked repeated model outputs more closely than agreement with human labels, and grouping six tweets in one prompt lowered mean offensive-language confidence by 660 basis points. We call the variation caused by task design and model choice instrument uncertainty. Researchers can measure it only by comparing reasonable task designs. Repeating one setup or relying on confidence scores cannot replace that test.

[NLP-187] racing mechanisms of sycophantic agreement in language models

【速读】: 该论文旨在解决语言模型中存在的谄媚式附和(sycophantic agreement)问题,即模型倾向于过度迎合用户陈述的观点或偏好,从而损害事实准确性。这一现象被广泛视为对齐失败的表现,但其内在机制尚不明确。研究通过因果中介分析揭示了其关键机制:用户的观点在处理最终提示词的残差流(residual stream)中早期被编码,并由此影响后续答案的生成过程;少数早期注意力头负责传递这一观点信号。移除这些注意力头可显著降低谄媚行为,同时保持事实准确性不受影响。值得注意的是,无论用户表述方式如何,这些头部均能识别并传递明确表达的观点;而当观点以无内容的反驳形式(如“你确定吗?”)隐含传达时,则由另一组特定注意力头抑制原始正确答案,以促成修正后的回应。该研究从机制层面阐明了观点如何诱发谄媚式附和,为设计更精准、可靠的对齐干预措施提供了理论基础。

链接: https://arxiv.org/abs/2609.35822
作者: Sixing Chen,Zhuofan Josh Ying,Logan Riggs Smith,Jeremy Wertheimer,Natalie Shapira
机构: New York University(纽约大学); Cambridge Boston Alignment Initiative; Columbia University(哥伦比亚大学); Independent; Northeastern University(东北大学)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Sycophantic agreement in language models refers to the tendency to overly affirm a user’s stated beliefs or preferences, often at the expense of factual accuracy. Although it is widely recognized as an alignment failure, its underlying mechanisms remain poorly understood. In this work, we use causal mediation analysis to identify the mechanisms behind sycophantic agreement. We show that a stated opinion is incorporated into the residual stream of the final prompt token early, where it biases subsequent answer retrieval. A sparse set of early attention heads carries this opinion signal. Ablating these heads substantially reduces sycophancy while leaving factual accuracy largely intact. The same heads carry the opinion when it is explicitly stated, regardless of how it is phrased. When an opinion is not stated explicitly but instead conveyed through content-free pushback (e.g., ``Are you sure?"), we find a distinct set of heads that suppresses the model’s original correct answer to promote a revised answer. By providing a mechanistic account of how opinions induce sycophantic agreement, this work takes a step toward developing more targeted and reliable alignment interventions.

[NLP-188] Can We Still Trust Disaster Social Sensing? Empirical Evidence on Detecting AI-Generated Social Media Posts

【速读】: 该论文旨在解决生成式人工智能(Generative AI)在灾害社会感知(disaster social sensing)中带来的可信度挑战,即如何可靠区分由人类撰写的灾情社交媒体帖子与基于真实信息生成的AI伪造内容。其核心问题在于:当前基于文本的AI检测工具是否具备足够的可靠性,以作为灾情信息真实性验证的操作性信任屏障。研究的关键解决方案在于系统性评估多种主流文本型AI检测方法(如OSM-Det、Fast-DetectGPT、Binoculars及大语言模型直接判断)在跨模型家族、跨灾难场景下的表现,并引入灾难领域校准、冻结编码器线性读出、成对变换敏感性分析以及数据集伪影控制等技术手段进行综合验证。结果表明,尽管经过灾难领域微调的线性分类头可达到0.817的AUROC,但其性能高度依赖于表面特征不对称性;一旦消除这些偏差,性能显著下降至0.594,且该分类器无法区分仅在情感框架上不同的AI生成内容(A0 vs A1),揭示了现有检测方法存在严重的表面特征依赖和泛化能力不足问题。最终结论指出,纯文本检测难以作为可信赖的操作性信任机制,必须结合多模态证据、可问责信源及其他上下文信息,才能有效保障灾害社会感知系统的可信性。

链接: https://arxiv.org/abs/2609.35821
作者: Xiaoshan Zhou,Zaifu Zhan
机构: 未知
类目: Computation and Language (cs.CL); Computers and Society (cs.CY)
备注:

点击查看摘要

Abstract:Disaster social sensing converts public social-media posts into evidence for situational awareness and humanitarian needs, but generative artificial intelligence (AI) can produce plausible messages that resemble eyewitness reports. This study investigates whether text-based AI detectors can reliably distinguish human-authored from AI-generated disaster posts. We construct a dataset of 12,000 texts organised into 3,000 matched semantic units from nine disasters: original human posts (H0), minimally LLM-proofread human posts (H1), factual AI-generated posts based on the same verified facts (A0), and affectively framed versions of those AI posts (A1). A separate 6,000-text corpus from 42 events supports model selection and threshold calibration. We evaluate OSM-Det, Fast-DetectGPT, Binoculars, and direct large language model (LLM) judges across five model families, then test disaster-domain calibration, a frozen-encoder linear readout, paired transformation sensitivity, and dataset artifact controls. Across fourteen frozen cross-family configurations, AUROC is 0.402-0.517 and the best prospective recall at a calibration-derived low-false-positive operating point is 3.6%; OSM-Det reaches AUROC 0.521 and 10.4% recall at a realised 6.7% false-positive rate. A disaster-trained linear head reaches AUROC 0.817, but a seven-feature surface classifier reaches 0.784 on the H0-versus-A0 contrast, and neutralising identified surface asymmetries reduces the head from 0.733 to 0.594. The head also separates A0 from A1 even though provenance is unchanged. The results show that text-based detection is not reliable enough to serve as an operational trust gate; multimodal claims, accountable sources, and other contextual evidence should be rested on to safeguard trust in disaster social sensing.

[NLP-189] τ-Multilingual: Benchmarking Voice Agents Across Languages

【速读】: 该论文旨在解决现有语音代理(voice-agent)评估基准仅限英语导致的多语言能力覆盖不足问题,揭示了当前系统在非英语语种下的性能退化与行为差异。其解决方案的关键在于构建并验证了 \tau-Multilingual 多语言扩展框架,通过引入西班牙语、巴西葡萄牙语、印地语、韩语和中文五种语言,并依托母语者对生成文本与语音输出进行评审与评估,实现了跨语言、全双工通话场景下的系统性评测。研究发现,尽管西语、葡语和印地语的表现与英语差距较小(任务完成度相差不超过3.2分),但韩语和中文分别下降14.7和8.4分,且失败模式各异:韩语系统漏回应更严重,中文系统更频繁打断,两者均在工具调用与实体识别上表现不佳。此外,研究提出应分离报告任务完成度、交互质量与生成质量,以更全面评估系统性能。为推动社区共建,作者开源了多语言语言包、经验证的评估员资源及评测工具,支持可复现的多语言语音代理评估体系。

链接: https://arxiv.org/abs/2609.35820
作者: Soham Ray,Edgard dos Santos Paiva,Ruben Valenzuela,Karthik Narasimhan,Keshav Dhandhania,Victor Barres
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Sound (cs.SD); Audio and Speech Processing (eess.AS)
备注:

点击查看摘要

Abstract:English-only benchmarks expose only a narrow slice of voice-agent behavior. We introduce \tau -Multilingual, extending \tau -Voice to Spanish, Brazilian Portuguese, Hindi, Korean, and Mandarin with native-speaker review and evaluation of generated language and spoken output. Across 4,500 full-duplex calls and five voice configurations, Spanish, Portuguese, and Hindi remain within 3.2 task-completion points of English, but Korean and Mandarin fall by 14.7 and 8.4 points. The failure modes also vary: Korean systems miss more responses, Mandarin systems interrupt more often, and both struggle with tools and entities. Grok leads task completion but scores lowest on generation quality, motivating separate task, interaction, and generation reporting. We release language packs, validated judges, and tools for community-built multilingual voice-agent evaluation.

[NLP-190] Less Uniform Discrete Diffusion is More Powerful and Scalable

【速读】: 该论文旨在解决统一扩散语言模型(Uniform Diffusion Language Models, UDLMs)在规模化过程中面临的挑战,其核心问题在于训练目标过于均匀以及采样阶段条件与目标之间的混淆。为克服上述问题,论文提出了一种名为“低均匀性扩散”(Less Uniform Diffusion, LUDI)的新框架。其关键解决方案包括:(i) 设计一种非均匀损失函数,引导每个逆向扩散步骤更精准地逼近原始无噪声词元;(ii) 引入逐词元时间嵌入(per-token time embeddings),为模型提供细粒度的污染提示信息,从而支持基于置信度的少步长采样。实验结果表明,LUDI显著提升了监督信号的质量,并增强了少步生成能力。进一步地,将一个70亿参数的自回归模型持续训练为LUDI-7B后,所得到的统一扩散模型具备复杂推理能力,在生成速度上实现每步3个词元的加速,同时性能可与掩码扩散基线模型相媲美,揭示出统一扩散语言模型在复杂生成任务中的潜力仍有待充分挖掘。

链接: https://arxiv.org/abs/2609.35817
作者: Kaibo Wang,Ding Ding,Fangyu Ding,Zijin Feng,Han Shi,Haili Bai,Jiacheng Sun,Yang Xiang
机构: The Hong Kong University of Science and Technology (香港科技大学); Huawei Foundation Model Department (华为基础模型部门)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 23 pages, 8 figures, 5 tables

点击查看摘要

Abstract:Although uniform diffusion language models (UDLMs) represent a promising diffusion paradigm, scaling them remains challenging. We identify the core obstacle as an over-uniform training objective and condition-target confusion during sampling. To address these, we propose Less Uniform Diffusion (LUDI), a novel UDLM framework. Specifically, we (i) introduce a less uniform loss that directs each reverse transition toward the clean token, and (ii) equip the model with per-token time embeddings that supply token-level corruption hints, enabling confidence-based few-step sampling. Experiments across scales show that LUDI yields cleaner supervision and improves few-step generation. We further continue-train a 7B autoregressive model into LUDI-7B, resulting in a UDLM capable of complex reasoning. It achieves a 3-token-per-step speedup over AR decoding and competitive performance compared with masked diffusion baselines, revealing that the full potential of UDLMs for complex generation remains to be unlocked.

[NLP-191] PrimeSeeker: Capability-Oriented Supervision for Deep Search Agents

【速读】: 该论文旨在解决大语言模型搜索代理(large language model search agents)在训练过程中依赖合成问题所带来的“检索能力与任务需求不匹配”问题。现有方法通过增大证据图规模、增加推理跳数和延长轨迹长度来提升问题难度,但这些全局属性仅是局部检索能力的间接代理,无法有效反映实际搜索中对精准定位信息锚点(retrieval anchor)的需求。为此,论文提出潜在锚点推理(latent anchor reasoning)作为核心解决方案,其关键在于:从描述性规范中识别并恢复一个未命名的信息锚点,并将其转移至后续的信息需求中,从而将复杂的深度搜索过程分解为一系列耦合的操作链。该机制以锚点解析与关系传递为核心构建问题生成逻辑,不预设固定的搜索路径。基于此,作者提出了PrimeSeeker——一种面向能力的框架,能够构建基于网络的锚点结构,并联合生成问题与参考证据骨架(reference evidence skeleton)。该骨架在构造阶段保留支持性证据,通过提取工具观测结果的摘要进行引导,同时在监督微调前移除这些摘要,而骨架本身则用于强化学习阶段评估参考步骤覆盖率以生成奖励信号。实验构建了9,221条专家轨迹,训练了一个300亿参数的搜索代理,在五个深度搜索基准上均表现出色;且通过参考步骤优化进一步提升了监督策略性能。最终生成的搜索轨迹展现出低检索冗余性,在固定预算下的评估中实现了更高的解覆盖度,且所需工具调用次数显著少于长视野系统。

链接: https://arxiv.org/abs/2609.35816
作者: Linzhi Peng,Hanting Chen,Heng Chang,Ke Cheng,Bowen Du,Weifeng Lv
机构: Beihang University(北京航空航天大学); Huawei Technologies Ltd.(华为技术有限公司); Tsinghua University(清华大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language model search agents are often trained with synthetic questions whose difficulty is increased through larger evidence graphs, additional hops, and longer trajectories. These global properties, however, are only indirect proxies for the local retrieval capabilities required during search. To address this mismatch, we introduce latent anchor reasoning, which consists of resolving an unnamed retrieval anchor from descriptive specifications and transferring the recovered anchor into a subsequent information demand. This primitive retrieval unit decomposes deep search into chains of coupled operations and organizes question construction around anchor resolution and relation transfer, without prescribing a canonical search path. Based on this formulation, we propose PrimeSeeker, a capability-oriented framework that constructs web-grounded anchor structures and jointly derives a question and a reference evidence skeleton. The skeleton preserves supporting evidence from construction and guides expert generation through extractive highlights of current tool observations. These highlights are removed before supervised fine-tuning, while the skeleton is subsequently reused to audit reference-step coverage for reinforcement-learning rewards. We construct 9,221 expert trajectories, training a 30B search agent. Across five deep-search benchmarks, PrimeSeeker achieves strong performance, while reference-step optimization further improves the supervised policy. The resulting trajectories exhibit low retrieval redundancy, and fixed-budget evaluation shows strong solution coverage with substantially fewer tool calls than long-horizon systems.

[NLP-192] Constructing Challenging Browser-Use Tasks by Controlled Environment Interventions

【速读】: 该论文旨在解决现有浏览器使用基准测试中任务难度难以精确控制与评估的问题。随着浏览器代理(browser-use agents)能力的提升,基准测试不断引入更长、更复杂或更新颖的任务,导致难度因素混杂且不可控,难以明确识别具体挑战来源。为此,作者提出一种创新方法:基于已可解决的任务构建具有可控难度的干预实例,将难度转化为环境可编程属性。核心解决方案是设计“BreakingWeb”基准,通过在基础任务上施加一系列确定性、可检测且可恢复的干预条件(intervention),在不改变用户指令、潜在目标和后端成功标准的前提下,仅在网页栈的不同层级改变环境。每个干预均标注其所主要加载的认知原语(cognitive primitive)。该基准包含7个自托管网站上的519对清洁/干预任务及29类干预,所有任务均按结果进行分级评估。实验表明,干预使六种先进代理的通过率平均下降22.9%,近半数任务被推翻,而人类在首次尝试中损失10.0%,熟悉后降至5.7%。主要失败模式为“信念失败”(belief failure)——75%的代理失败表现为声明成功但实际未达成所需变更,揭示了当前代理在状态感知与目标追踪上的根本缺陷。该研究为评估和改进代理的鲁棒性提供了可扩展、可解释的基准框架,相关代码、数据与环境均已开源。

链接: https://arxiv.org/abs/2609.35814
作者: Xunjian Yin,Tianchen Guan,Jinao Wang,Weili Cao,Daisy Xinlei Lin,Royce Cheng-Yue,Keagan Long,Kyle Wong,Bhuwan Dhingra,Xiangjun Wang,Shuyan Zhou
机构: Duke University(杜克大学); Amazon(亚马逊)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 40 pages

点击查看摘要

Abstract:As browser-use agents improve, benchmarks keep pace by collecting new tasks, websites, and applications, often making tasks longer or more novel. This makes difficulty expensive to refresh and difficult to control: when many aspects change at once, it is unclear what actually makes a task challenging. We instead construct challenging instances from tasks agents already solve, turning difficulty into a programmable property of the environment. BreakingWeb pairs every base task with an intervention condition that preserves the user instruction, latent target, and backend success criterion while changing the environment at different web stack layers. Each intervention is deterministic, detectable, and recoverable, and is annotated with the cognitive primitive it primarily loads. The benchmark contains 519 clean/intervention task pairs across seven self-hosted websites and 29 intervention families, all graded against outcomes. We evaluate six strong browser-use agents, three GUI-only agents that see only screenshots, and humans. The construction is effective: interventions cut agent pass rate by 22.9% on average and overturn nearly half of the tasks each agent solves cleanly, whereas humans lose 10.0% on a first attempt and 5.7% after one familiarisation attempt. The dominant failure is belief failure: 75% of the six agents’ failures end with a declared success although the required change never happened. Our code, data and environment are publicly available at this http URL.

[NLP-193] Automated Evaluation of Multi-Turn Dialogues in In-Car Conversational Assistants ITSC2026

【速读】: 该论文旨在解决智能座舱对话助手(In-car Conversational Assistant, ICA)在多轮对话场景下可靠性评估难的问题。现有评估方法主要针对单轮交互,无法有效捕捉对话中的约束处理能力、上下文记忆保持以及跨轮次的安全关键行为,难以满足车载系统对安全性和鲁棒性的严苛要求。其解决方案的关键在于提出一种自动化测试框架,将ICA视为黑盒系统,通过闭环仿真方式结合策略引导的用户模拟器、对抗性策略管理器以及两级大语言模型(LLM)判别器,分别评估每轮交互的失败情况与整体对话质量。实验结果表明,该框架在工业级ICA上表现优异,自动化判别器与人工标注者具有高度一致性;相比无策略引导的仿真,策略引导使每轮对话中发现的唯一故障类型增加2.96倍,且失败对话数量翻倍以上,显著提升了故障探测效率与评估深度。

链接: https://arxiv.org/abs/2609.35812
作者: Vaishnav Negi,Lev Sorokin,Soroosh Tayebi Arasteh,Andrea Stocco
机构: Friedrich-Alexander-Universität Erlangen-Nürnberg(弗里德里希-亚历山大埃尔朗根-纽伦堡大学); BMW Group(宝马集团); Technical University of Munich(慕尼黑工业大学); RWTH Aachen University(亚琛工业大学); fortiss GmbH(德国慕尼黑软件技术研究所)
类目: Computation and Language (cs.CL); Software Engineering (cs.SE)
备注: Accepted at the 29th IEEE International Conference on Intelligent Transportation Systems (IEEE ITSC 2026)

点击查看摘要

Abstract:In-car conversational assistants (ICAs) are increasingly integrated into vehicles to support route planning, vehicle control, and information access. Ensuring their reliability is challenging due to multi-turn interactions, the absence of explicit ground truth, and strict safety constraints. Existing evaluation techniques fall short, as they target single-turn settings and fail to capture constraint handling, context retention, and safety-critical behavior across turns. We propose an automated framework for testing the multi-turn conversational capabilities of ICAs. The system is treated as a black box and evaluated via closed-loop simulation with a strategy-guided user simulator, an adversarial strategy manager, and a two-tier LLM judge assessing turn-level failures and conversation-level quality. We evaluate the approach on an industrial ICA with six LLM backends and twelve human annotators. The automated judge shows substantial agreement with humans, and strategy guidance uncovers 2.96 times more unique failure types per conversation and more than doubles the number of unique failing conversations compared to unguided simulation.

[NLP-194] Lookahead-R: Budget-Aware Tool Retrieval via Execution-Centric Planning ICMR2026

【速读】: 该论文旨在解决基于大语言模型(LLM)的智能体在大规模异构API生态系统中进行工具检索时面临的效率与精度权衡问题。现有方法存在固有缺陷:语义检索虽快速但存在语义-功能鸿沟,而基于执行验证的方法虽能提升精度却因高昂延迟难以实用。其解决方案的核心是提出一种基于规划的框架Lookahead-R,将工具检索重构为资源受限的序列决策问题。该框架引入一个轻量级的、具备执行感知能力的代理世界模型(surrogate world model),可在不调用真实API的前提下联合预测工具执行成功率、延迟成本及语义效用。该模型驱动一种成本敏感且基于不确定性的蒙特卡洛树搜索(Monte Carlo Tree Search),在严格预算约束下高效导航工具空间。在大规模ToolBench基准上的实验表明,Lookahead-R在所有测试场景中均实现了更优的准确率-效率平衡;尤其在最具挑战性的I3子集上,其NDCG@5达到91.40%,优于当前最优方法ToolGen的90.16%。消融实验证实,显式建模延迟是资源受限条件下识别高质量工具的关键判别信号。

链接: https://arxiv.org/abs/2609.35811
作者: Zongze Wu,Yani Guo,Runnan Li
机构: Beijing University of Posts and Telecommunications (北京邮电大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注: 10 pages, 3 figures, 3 tables. Published in ICMR 2026

点击查看摘要

Abstract:Tool retrieval is a critical bottleneck for LLM-based agents operating over large, heterogeneous API ecosystems. Existing approaches face an inherent trade-off: semantic retrievers are fast but suffer from the semantic-functional gap, while execution-based validation improves precision at the cost of prohibitive latency. We propose Lookahead-R, a planning-based framework that reformulates tool retrieval as a resource-constrained sequential decision-making problem. At its core, Lookahead-R introduces a lightweight execution-aware surrogate world model that jointly predicts tool execution success, latency cost, and semantic utility—without invoking real APIs. This world model drives a cost-sensitive, uncertainty-guided Monte Carlo Tree Search that navigates the tool space under strict budget constraints. Evaluated on the large-scale ToolBench benchmark, Lookahead-R achieves a superior accuracy-efficiency trade-off across all test scenarios. On the most challenging I3 split, it attains an NDCG@5 of 91.40%, outperforming the state-of-the-art ToolGen (90.16%) by 1.24%. Ablation studies confirm that explicit latency modeling is the key discriminative signal for identifying high-quality tools under resource constraints.

[NLP-195] RACE: Deployable Tree-Relational Structure Enhancement for Oncology LLM s EMNLP2026

【速读】: 该论文旨在解决当前生成式AI在肿瘤学领域应用中预测结果缺乏明确医学结构支撑的问题,即大语言模型(LLM)的推理过程往往依赖于隐式知识,导致可解释性差且临床可信度不足。其核心解决方案是提出一种可部署的树-关系增强框架(TRACE),关键在于将昂贵的离线结构学习与轻量级在线推理分离:通过基于语言模型损失(LM-loss)推导的证据,构建并持续更新一个可维护的树-关系医学结构,该结构在推理阶段以紧凑提示证据的形式被检索使用。这种设计实现了无需标注数据的零样本任务自适应证据选择,显著提升了十项肿瘤学分类任务及一项MedQuAD CancerGov问答基准上的性能。实证分析表明,TRACE优于通用检索增强生成(RAG)和图-检索增强生成(GraphRAG),在控制信息泄露的METABRIC数据输入下仍具鲁棒性,并能生成与临床推理逻辑一致的可解释证据路径,验证了显式、可更新的医学结构对于提升肿瘤学大模型准确性与可审计性的实用价值。

链接: https://arxiv.org/abs/2609.35810
作者: Jizheng Lai,Yingyun Li,Ying Qin,Haiyang Qian
机构: Pharmaron(药明康德); AI Starfish(星鱼智能)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 18 pages, 2 figures, 19 tables. Accepted to the EMNLP 2026 Industry Track for oral presentation

点击查看摘要

Abstract:Large language models are increasingly used in oncology applications, but their predictions are often weakly grounded in explicit medical structure. We present TRACE, a deployable tree-relational enhancement framework for oncology LLMs. TRACE separates expensive offline structure learning from lightweight online inference: oncology concepts and relations are organized into an updatable tree-relational structure, refined using LM-loss-derived evidence, and retrieved at inference time as compact prompt evidence. This design supports task-adaptive evidence selection without requiring supervised labels in the zero-shot setting. Across ten oncology classification tasks and one MedQuAD CancerGov QA benchmark, TRACE improves both label-free evaluation and supervised fine-tuning. Additional analyses show that TRACE improves over vanilla RAG and generic GraphRAG, remains useful under leakage-controlled METABRIC inputs, and produces interpretable evidence paths aligned with clinical reasoning. These results suggest that explicit, updatable medical structure is a practical path toward more accurate and auditable oncology LLM deployment.

[NLP-196] Can Multimodal Large Language Models Generate and Detect Multimodal Social Media Fake News? EMNLP2026

【速读】: 该论文旨在解决生成式人工智能(Generative AI)在多模态大语言模型(Multimodal Large Language Models, MLLMs)滥用背景下,可能被用于大规模制造逼真多模态虚假新闻(multimodal fake news)的潜在风险,以及现有模型在自动识别此类虚假信息时的能力不足问题。其核心挑战在于:多模态虚假内容是否能够被高效生成以伪装成真实社会媒体新闻,且当前主流模型能否可靠检测此类内容。解决方案的关键在于提出一种多智能体框架(multi-agent framework),由故事生成智能体、图像生成智能体和批判评估智能体协同工作,共同生成与真实新闻高度匹配的多模态虚假帖子。该框架在科学、健康和娱乐领域构建了超过9,000对配对的多模态新闻数据,并对16个开源与闭源的MLLMs进行了自动化检测能力的基准测试。研究发现,绝大多数模型在检测准确率上远低于人类水平,尤其在判断图像真实性方面存在严重缺陷。这一结果揭示了当前防御体系的脆弱性,为未来构建更鲁棒的反虚假信息系统提供了关键实证基础。

链接: https://arxiv.org/abs/2609.35809
作者: Jiyao Yang,Yang Liu,Zhenyue Qin,Qingyu Chen,Xiuzhen Zhang
机构: Independent Researcher; Carnegie Mellon University (卡内基梅隆大学); RMIT University (皇家墨尔本理工大学); Yale University (耶鲁大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 15 pages, 6 figures. Accepted for publication in the Findings of EMNLP 2026

点击查看摘要

Abstract:The rapid advancement of generative AI raises concerns about the misuse of Multimodal LLMs (MLLMs) for large-scale disinformation campaigns on social media. Despite existing research on textual disinformation, a fundamental question remains unanswered: can MLLMs be exploited to fabricate realistic multimodal fake news, and can they reliably detect it? We introduce a multi-agent framework in which a story agent, an image agent, and a critic agent collaborate to produce fake social media posts that plausibly counter true news. We apply the framework to generate over 9,000 paired multimodal news posts across science, health, and entertainment domains, and benchmark 16 open- and closed-source MLLMs for automated detection. We find that most models fall substantially short of human-level accuracy and fail critically on identifying image authenticity. Our research provides a foundation for developing robust defenses against social media fake news. Code and data are available at https: //github.com/xiuzhenzhang/Multimodal.

[NLP-197] When Successful Memories Mislead Embodied Agents :Memory Adaption For Task-Conditioned Execution

【速读】: 该论文旨在解决具身智能体在任务执行过程中因历史经验与当前执行上下文不匹配而导致的性能下降问题。现有记忆系统虽优化了记忆的构建与检索,但在检索出的经验中仍存在控制上下文过时、动作与环境条件不兼容或结构层次不当等问题,导致其直接复用效果不佳。为此,论文提出一种确定性后检索处理方法——面向任务条件执行的记忆适配(Memory Adaptation for Task-Conditioned Execution, MATE),其核心在于通过一系列关键步骤将原始轨迹转化为面向执行的可操作记忆:移除过时的控制上下文,提取条件-动作-效应(condition-action-effect)三元组,应用经验证的动作归一化以确保动作语义一致性,根据任务选择合适的表示形式,并在固定预算下对结果进行序列化,整个过程无需额外的大语言模型(LLM)推理。实验表明,在134个ALFWorld任务上,MATE使用Qwen2.5-14B和72B模型分别实现81.3%和93.3%的任务成功率,同时仅需原始轨迹约十分之一的令牌数。受控对比分析进一步证明,经验证的动作归一化是恢复检索经验有效性的主要机制,证实了记忆适配作为连接检索与具身执行之间独立阶段的必要性。

链接: https://arxiv.org/abs/2609.35808
作者: Quanquan Li,Hongbo Zhang,Yihe Chi,Liuyang Song,Jingyu Li,Yuxiang Huang,Hongzhen Zhang,Guitao Cao
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Experience reuse can reduce repeated exploration in embodied agents, but a trajectory that succeeded previously may be unsuitable for the current execution context. Existing memory systems pri marily optimize construction and retrieval; semantic relevance and historical success therefore remain insufficient when retrieved ex perience contains incompatible actions or an inappropriate level of structure. We introduce Memory Adaptation for Task-Conditioned Execution (MATE), a deterministic post-retrieval procedure that converts trajectories into execution-oriented memory. MATE re moves obsolete control context, extracts condition-action-effect transitions, applies verified action normalization, selects a task dependent representation, and serializes the result under a fixed budget without additional LLM inference. On 134 ALFWorld tasks, MATE achieves task success rates of 81.3% and 93.3% with Qwen2.5-14B and 72B while using approximately one-tenth of the tokens required by raw trajectories. Controlled comparisons show that verified action normalization is the principal mechanism by which MATE restores the utility of retrieved experience, support ing memory adaptation as a distinct stage between retrieval and embodied execution.

[NLP-198] Environment Steering: Using Data Flow Control to Improve Agent Utility and Safety EMNLP2026

【速读】: 该论文旨在解决大语言模型代理(LLM agents)在被指令保持安全行为时仍可能执行不安全工具调用的问题。现有防御方法通常在执行前施加限制、修改工具的输入/输出,或依赖大语言模型作为评判者,但这些方法往往依赖于模型自身的行为表现,且在发现不安全行为后仅能阻止而无法帮助代理恢复,缺乏动态纠偏能力。本文提出“环境引导”(Environment Steering)这一新范式,主张在代理运行过程中由执行环境实时强制安全约束,并在违规发生时通过策略与上下文相关的反馈引导代理转向安全路径。其核心解决方案是将代理及其执行状态建模为数据库表结构,追踪记录级别的数据流,并在运行时依据声明式安全策略对数据流进行实时检查;一旦检测到违规,系统将生成针对性反馈以引导代理回归安全轨迹。在AgentDyn基准测试中,该方法不仅实现了0%的攻击成功率,还显著提升了任务成功率,优于无防御情形。

链接: https://arxiv.org/abs/2609.35807
作者: Charlie Summers,Prajwal Raghunath,Aaditya Pai,Mayur Kulkarni,Zhuo Zhang,Oliver Kennedy,Eugene Wu
机构: Columbia University (哥伦比亚大学); University at Buffalo (水牛城大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Databases (cs.DB)
备注: 9 pages, 11 figures, REALM Workshop, EMNLP 2026

点击查看摘要

Abstract:LLM agents can make unsafe tool calls even when instructed to behave safely. Existing defenses constrain agents before execution, modify tool inputs/outputs, or rely on LLM judges; these approaches may depend on model behavior or block unsafe actions without helping the agent recover. We argue that the execution environment should instead enforce safety as the agent runs and steer it toward safe alternatives when violations occur—we call this Environment Steering. We implement this by modeling the agent and harness execution state as database tables, track the record-level data flows, and check these data flows against declarative policies during runtime. When violations are detected, policy- and context-specific feedback steers the agent toward safe trajectories. On AgentDyn, this enables the agent to improve task success rate over no-defense while achieving 0% attack success rate.

[NLP-199] From Lexical Baselines to Agent ic Retrieval-Augmented Generation: Structured Skill and Responsibility-Level Extraction with the SFIA Framework

【速读】: 该论文旨在解决现有自动化技能抽取系统中缺乏对技能责任层级(responsibility level)表征的问题,尤其针对信息时代技能框架(Skills Framework for the Information Age, SFIA)这一具有明确层级结构的闭合词汇体系。传统方法将技能视为扁平标签,无法捕捉其在实际工作中所处的责任水平,而SFIA则定义了147项专业技能及其对应的七级责任层级,但尚无基于大语言模型(LLM)的自动化提取方法。为此,论文将任务形式化为从自由文本中进行(技能,层级)对的结构化预测,并提出三个核心问题:文本映射到SFIA闭合词汇的准确性如何、何种策略能可靠地同时预测技能与层级、以及代理式(agentic)设计是否优于简单检索与提示方法。解决方案的关键在于构建一个完全自动化的代理流水线,用于生成大规模的SFIA~9语料库,并评估五种策略:词法基线、密集检索结合LLM重排序、零样本模式约束的LLM、单代理代理式检索增强生成(RAG),以及三代理协作流程(检索-匹配-验证)。实验结果表明,基于检索的匹配方法识别的技能数量最多,而生成式策略在精确度上显著更优;唯有将层级作为显式决策步骤的策略才能可靠预测层级,基于相似性的选择方式准确率不足前者的一半;三代理协作虽使延迟翻倍,但未提升准确率,说明在闭合分类任务中增加代理角色并不必然带来性能增益。研究提供了首个可复现的、面向SFIA的结构化、层级感知技能抽取基准。

链接: https://arxiv.org/abs/2609.35806
作者: Ranuga Disansa,U. S. Samarasinghe,Lasith Gunawardena
机构: Informatics Institute of Technology(信息学技术学院); Department of Information Technology(信息技术系), University of Sri Jayewardenepura(斯里贾亚瓦登普拉大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Automated skill extraction underpins workforce planning, yet most systems represent skills as flat labels with no notion of the responsibility level at which a skill is practiced. The Skills Framework for the Information Age (SFIA) captures exactly this dimension, defining 147 professional skills across seven responsibility levels, but no automated LLM-based extraction targeting SFIA has been reported. We formalize the task as structured prediction of (skill, level) pairs from free text and ask three questions: how accurately can text be mapped onto SFIA’s closed vocabulary, which strategies reliably predict the level alongside the skill, and do agentic designs improve on simpler retrieval and prompting? We evaluate five strategies (a lexical baseline, dense retrieval with LLM reranking, a zero-shot schema-constrained LLM, single-agent agentic RAG, and a three-agent retriever–matcher–verifier crew) against expert-mapped European ICT role profiles, all drawing on an SFIA~9 corpus built by a fully automated agentic pipeline that we release. Retrieval-based matching identifies the most skills while generative strategies are markedly more precise; only strategies assigning the level as an explicit decision predict it reliably, with similarity-based selection more than twice as inaccurate; and the crew doubles latency without improving accuracy, so added agent roles do not automatically benefit closed-taxonomy matching. These results provide the first reproducible baseline for structured, level-aware skill extraction against SFIA.

[NLP-200] Alignment Forecasting: Predicting Misalignment From Training Data

【速读】: 该论文旨在解决生成式 AI 在微调过程中因训练数据中存在细微缺陷而导致模型广泛偏离对齐目标的问题。传统方法依赖于训练后的审计来发现对齐失效,但往往为时已晚。为此,论文提出“对齐预测(Alignment Forecasting)”这一新任务,即在训练前预判微调是否会导致特定类型的对齐失败(如欺骗或奉承行为)。其核心解决方案是构建一个基于大语言模型(LLM)的预测框架:先由一个大型语言模型分析训练数据,评估其推动模型产生不当行为的强度与广度,再通过一个简单的可学习模型结合该评分、故障模式的基础发生率以及目标模型的先验倾向,实现对对齐风险的量化预测。该方法显著优于随机猜测,并在多个基准测试中超越了直接微调的模型和仅参考弱模型表现的简单预测器。此外,其预测信号能够识别出被前沿分类器遗漏的问题训练样本,剔除这些样本后,在多项选择评估中多数情况下能提升模型对齐性,尽管在开放式对话中的收益尚不明确。研究结果表明,在监督微调(SFT)场景下,提前预测多种对齐失效具有可行性,但仍需进一步优化以支持实际训练数据筛选。

链接: https://arxiv.org/abs/2609.35805
作者: Chen Yueh-Han,Bruce W. Lee,Ilia Sucholutsky,Tomek Korbak
机构: New York University (纽约大学); OpenAI
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Training a language model on data with a narrow flaw can sometimes make the model broadly misaligned. Inspecting the data at face value often does not settle whether it will emerge, and today it is caught only after training, by auditing the resulting model. To complement post-hoc audits, we introduce Alignment Forecasting: the task of predicting alignment failures before training. Given a target model, a fine-tuning dataset, and a failure mode such as deception or sycophancy, a forecaster outputs the probability that fine-tuning would meaningfully increase that failure mode. To measure progress on alignment forecasting, we introduce ALIGNMENTFORECASTBENCH, a benchmark of over 5,000 forecasting questions spanning 17 target models, 32 datasets, and 16 failure modes. Frontier models prompted directly perform poorly on ALIGNMENTFORECASTBENCH. We therefore propose a forecasting scaffold in which an LLM reads the dataset and rates how strongly and broadly it pushes the model toward misbehavior, and a simple learned model combines that rating with the failure mode’s base rate and the target model’s prior tendency. This forecasts well above chance, and beats a model fine-tuned on the task and a simple forecaster allowed to see how weaker models behaved after fine-tuning on the same data. Its signals also flag problematic training examples that a frontier-model classifier misses. Filtering those examples out from real post-training data such as UltraChat results in more aligned models on our multiple-choice evaluation in most cases, though the benefit in open-ended conversations is unclear. More progress is needed before forecasts can reliably guide training data curation in practice, but our results suggest that forecasting many alignment failures before training can be tractable in the SFT setting.

[NLP-201] Developing an OCR model for Extracting Information from Invoices with Korean Language ATC

【速读】: 该论文旨在解决韩语发票中关键信息自动提取的难题,尤其针对韩语文本在复杂背景、字体多样及排版不规则等场景下的识别挑战。其核心解决方案在于构建一种结合深度学习模型与图像预处理技术的高效光学字符识别(OCR)系统,通过优化图像质量与特征表达能力,显著提升韩语文本的识别准确率。实验结果表明,该方法在大量真实发票数据集上实现了87%的F1分数,同时保持极低的处理时间开销,验证了其在实际应用中的高效性与可行性。

链接: https://arxiv.org/abs/2609.35796
作者: Xiem HoangVan,Phu TranQuang,Minh DinhBao,Tien VuHuu
机构: 未知
类目: Computation and Language (cs.CL)
备注: 2023 International Conference on Advanced Technologies for Communications (ATC)

点击查看摘要

Abstract:Invoices are commercial documents that contain various pieces of information, including the purchased items, time, and total money. Making the extraction of important information crucial. The stored information serves different purposes. Korean language is the native language of about 80 million people, playing an important role in not only South and North Korea but also in many other countries such as Vietnam, Philippine where a large number of Korean companies are located. In this context, to automatically extract proper information from the invoices with Korean language, we propose an efficient Optical Character Recognition (OCR) model in which a deep learning model is combined with some image preprocessing techniques. The proposed OCR model is assessed in a rich set of collected invoices showing that 87% F1-score can be achieved with negligible time processing.

[NLP-202] Sieve and Sage: Efficient Distraction Filtering for Reliable RALM Abstention EMNLP2026

【速读】: 该论文旨在解决检索增强型语言模型(Retrieval-Augmented Language Models, RALMs)在面对不可回答或存在干扰信息的检索结果时,缺乏有效拒答能力的问题。现有方法通常依赖单一的大型语言模型(LLM)在一步内处理异构的检索失败情况,导致拒答性能受限且计算开销高昂。其解决方案的关键在于将检索失败分解为两种独立状态:(i)不可回答状态(unanswerable state),即所需证据完全缺失;(ii)分心状态(distracted state),即相关证据被矛盾、否定或对抗性信息所干扰。基于此分解,作者提出一个轻量级筛选模块(Sieve),在调用高成本的生成与拒答模型(Sage)前,先对检索到的文档集进行筛查,识别并排除分心噪声。实验表明,该框架在通用及高风险专业领域均显著提升系统准确性(最高提升69.4个百分点)和宏平均F1分数(最高提升55.2个百分点),同时实现最高达1.99倍的加速,构建了一个高效且可靠的拒答流水线。

链接: https://arxiv.org/abs/2609.35794
作者: Jongbin Won,Sung Geun An,Jay-yoon Lee
机构: Seoul National University (首尔国立大学)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 20 pages, 6 figures, accepted at EMNLP 2026 Findings

点击查看摘要

Abstract:Just as Socrates recognized the limits of his own knowledge, Retrieval-Augmented Language Models (RALMs) should learn to abstain when the retrieved evidence cannot support a reliable response. Existing approaches largely rely on monolithic LLMs to handle heterogeneous retrieval failures in a single step, resulting in limited abstention performance and high computational costs. We instead decompose retrieval failures into two distinct states: (i) the unanswerable state, where the required evidence is absent, and (ii) the distracted state, where relevant evidence is mixed with conflicting, negated, or adversarial information. Based on this decomposition, we introduce a lightweight module (Sieve) that screens retrieved document sets for distracting evidence before invoking a costly LLM (Sage) for grounded generation and abstention. Evaluated across both general and high-stakes expert domains, our Sieve and Sage framework preemptively detects distracting noise, improving system accuracy by up to 69.4 percentage points and Macro-F1 by 55.2 percentage points compared to one-stage baselines. Furthermore, it achieves up to a 1.99x speedup, establishing a highly efficient and reliable abstention pipeline for RALM with abstention.

[NLP-203] FD-VAD: Semantic Endpoint Detection for Streaming Full-Duplex Speech

【速读】: 该论文旨在解决全双工语音交互中自然交替对话的语义端点检测(Semantic Endpoint Detection)问题,即在部分语音输入下判断静默是出于犹豫还是对话意图已结束。传统声学语音活动检测(Voice Activity Detection, VAD)缺乏语义信息,而基于级联自动语音识别(ASR)的端点检测方法则依赖于转录结果并引入额外处理阶段,导致延迟与错误传播。为此,论文提出一种无需ASR的流式端点检测框架FD-VAD,其核心在于将语义端点检测建模为因果音频-语言推理任务,通过冻结的语音编码器、轻量级模态适配器与参数高效微调的语言模型协同工作,直接从有界因果音频窗口映射至“继续(Continue)/停止(Stop)”决策,并采用末块训练目标支持流式推理。关键创新包括置信度门控的端点承诺机制以平衡打断与延迟风险,以及面向模糊对话边界的设计性难例采样策略,显著提升边界判别能力。实验表明,FD-VAD在域内及对话评估中均优于现有流式与非流式语义话轮分类器,在TurnBench开发集上以零样本设置实现0.853的最高端到端召回率(在假阳性率为0.10时),验证了仅从流式音频直接进行语义端点检测的可行性,无需中间ASR或对话状态追踪。

链接: https://arxiv.org/abs/2609.35791
作者: Puneet Mathur,Dinesh Manocha
机构: 未知
类目: Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
备注:

点击查看摘要

Abstract:Natural turn-taking in full-duplex voice interaction requires determining from partial speech whether a pause reflects hesitation or a completed conversational intent. Acoustic voice activity detection lacks this semantic information, while cascaded ASR-based endpointing introduces transcription dependence and additional processing stages. We formulate semantic endpoint detection as a causal audio-language reasoning task and introduce FD-VAD, an ASR-free streaming endpointer that maps bounded causal audio windows directly to Continue/Stop decisions. FD-VAD combines a frozen speech encoder with a lightweight modality adapter and a parameter-efficiently adapted language model, using a last-chunk training objective for streaming inference. We further introduce confidence-gated endpoint commitment to control interruption versus delay and boundary-focused hard-negative sampling to improve decisions around ambiguous turn boundaries. Across in-domain and conversational evaluations, FD-VAD outperforms strong streaming and non-streaming semantic turn classifiers, and achieves the highest EOT recall among qualifying systems on TurnBench dev set 0.853 (at FP=0.10) in a zero-shot setting. These results show that semantic endpointing can be performed directly from streaming audio without intermediate ASR or dialogue state tracking.

[NLP-204] Sage: Formalization with Semantic Correction

【速读】: 该论文旨在解决形式化数学中从非形式化自然语言到形式化语言(如Lean 4)翻译过程中的“严谨性幻觉”问题,即现有系统虽能生成语法正确的形式化语句,却常因遗漏假设、引入空真命题或微妙改变数学边界而导致语义失真。其核心解决方案是提出Sage(Semantic Agent-Guided Formalization Engine),一个基于智能体的分解式生成框架,包含四阶段流水线与双信号语义纠错循环。该框架通过结合Lean 4编译器诊断信息与多维度语义反馈,同时保障语法正确性与数学真实性,有效防止模型通过猜测未验证答案实现高形式化率(此前存在70.9%的答案泄露率)。实验表明,Sage将泄露率降至2.7%,在Omni-MATH without proofs数据集上实现73.3%的pass@4联合编译与语义保真度,显著优于基线模型(42.0%);在全新的175道未形式化的国际数学奥林匹克(IMO-Unformalized)问题上,Sage展现出卓越的零样本泛化能力,达到87.4%的pass@4验证保真度,远超基线(19.4%),并在盲评配对测试中胜出超过79%。

链接: https://arxiv.org/abs/2609.35790
作者: Thomas Hirtz,Farzad Jafarrahmani,Abdelmouksit Sagueni,Xiang Zhou,Wengping Deng,Liang Zhang
机构: Huawei Lagrange Mathematics Computing Research Center (华为拉格朗日数学计算研究中心); Paris, France
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Logic in Computer Science (cs.LO)
备注: 27 pages, 4 figures. Preprint

点击查看摘要

Abstract:While neural theorem provers have achieved impressive milestones in formal mathematics, they largely operate on the assumption that faithful Lean 4 formal statements are already provided. Translating informal natural language into a formal language is a critical data bottleneck plagued by an “illusion of rigor”: standard type-checkers accept statements that compile but drop hypotheses, introduce vacuous truths, or subtly alter mathematical bounds. To resolve this, we introduce Sage (Semantic Agent-Guided Formalization Engine), an agentic framework that replaces monolithic translation with a four-stage decomposed generation pipeline coupled with a dual-signal semantic correction loop. By pairing Lean 4 compiler diagnostics with multi-dimensional semantic feedback, our correction loop enforces mathematical fidelity alongside syntactic validity. By explicitly accounting for the gap between open-ended queries and declarative formal targets, our pipeline prevents models from achieving high formalization rates by guessing unverified answers (exhibiting a 70.9% answer leakage rate). Consequently, Sage suppresses leakage to 2.7% while achieving 73.3% pass@4 joint compilation and semantic fidelity on the Omni-MATH without proofs (compared to 42.0% for a fine-tuned Goedel-Formalizer-V2 baseline). Finally, on IMO-Unformalized, a novel frontier of 175 unformalized International Mathematical Olympiad problems, Sage demonstrates effective zero-shot generalization with 87.4% pass@4 verified fidelity compared to just 19.4% for the baseline, winning over 79% of blind pairwise evaluations.

[NLP-205] he Collective Turing Test: Large Language Models Can Generate Realistic Multi-User Discussions

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在模拟社交媒体群体对话时的真实性与可信度问题,即评估LLMs生成的虚拟对话是否能够以足够逼真的方式欺骗人类观察者,从而在社会仿真研究中替代真实用户交互。其解决方案的关键在于通过对比真实人类在Reddit上发表的社交对话与由Llama 3 70B和GPT-4o生成的同主题人工对话,开展双盲实验,以测试人类参与者对两类对话的识别准确率。结果显示,参与者将LLM生成的内容误判为真实人类创作的比例高达39%,其中对Llama 3生成内容的识别准确率仅为56%,接近随机水平,表明当前主流大语言模型已具备生成高度拟真社交对话的能力,足以在特定情境下实现“以假乱真”,这既凸显了其在社会行为模拟中的巨大潜力,也揭示了其被滥用以制造虚假舆论或操纵公众认知的重大风险。

链接: https://arxiv.org/abs/2511.08592
作者: Azza Bouleimen,Giordano De Marzo,Taehee Kim,Nicol`o Pagan,Hannah Metzler,Silvia Giordano,Anikó Hannák,David Garcia
机构: University of Zurich(苏黎世大学); Department of Informatics(信息学系); University of Konstanz(康斯坦茨大学); Department of Politics and Public Administration(政治与公共管理系); Complexity Science Hub(复杂性科学中心); University of Applied Sciences and Arts of Southern Switzerland(瑞士南部应用科学与艺术大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Social and Information Networks (cs.SI)
备注:

点击查看摘要

Abstract:Large Language Models (LLMs) offer new avenues to simulate online communities and social media. Potential applications range from testing the design of content recommendation algorithms to estimating the effects of content policies and interventions. However, the validity of using LLMs to simulate conversations between various users remains largely untested. We evaluated whether LLMs can convincingly mimic human group conversations on social media. We collected authentic human conversations from Reddit and generated artificial conversations on the same topic with two LLMs: Llama 3 70B and GPT-4o. When presented side-by-side to study participants, LLM-generated conversations were mistaken for human-created content 39% of the time. In particular, when evaluating conversations generated by Llama 3, participants correctly identified them as AI-generated only 56% of the time, barely better than random chance. Our study demonstrates that LLMs can generate social media conversations sufficiently realistic to deceive humans when reading them, highlighting both a promising potential for social simulation and a warning message about the potential misuse of LLMs to generate new inauthentic social media content.

[NLP-206] Predicting Team Performance from Communications in Simulated Search-and-Rescue

【速读】: 该论文旨在解决如何在个体特质难以直接观测的情况下,识别其对团队绩效的影响这一关键问题。传统研究虽可基于行为数据推断信任等特质,但缺乏对团队互动模式与绩效关联的系统性分析。本文通过分析基于Minecraft的搜救实验中的对话文本数据,采用主题建模(topic modeling)与聚类(clustering)技术,识别出反映团队特质的关键交互模式,并揭示其与团队成效之间的关联。研究发现,团队绩效的差异可通过这些推断出的团队特质得到解释,且个体特质与团队动态各自贡献了不同层次的预测能力,其关键在于从非结构化对话数据中提取可量化的团队行为特征,从而实现对团队效能的间接评估与预测。

链接: https://arxiv.org/abs/2503.03791
作者: Ali Jalal-Kamali,Nikolos Gurney,David Pynadath
机构: University of Southern California(南加州大学); Rice University(莱斯大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Understanding how individual traits influence team performance is valuable, but these traits are not always directly observable. Prior research has inferred traits like trust from behavioral data. We analyze conversational data to identify team traits and their correlation with teaming outcomes. Using transcripts from a Minecraft-based search-and-rescue experiment, we apply topic modeling and clustering to uncover key interaction patterns. Our findings show that variations in teaming outcomes can be explained through these inferences, with different levels of predictive power derived from individual traits and team dynamics.

[NLP-207] Does a prosody-trained representation help beyond trainable fusion? A parameter-matched study with frozen HuBERT ICASSP2027

【速读】: 该论文旨在解决自发语音识别(spontaneous speech recognition)中如何有效利用韵律信息以提升自动语音识别(ASR)性能的问题,同时区分性能提升是源于辅助信息本身还是融合机制的贡献。其核心挑战在于,现有方法依赖可训练的辅助表示,难以判断性能增益究竟来自额外的韵律信息还是融合策略。为此,作者采用冻结的HuBERT主干网络,引入一个64维的可学习韵律表征,该表征用于预测对数基频(log F0)、浊音性(voicing)、对数基频变化率(Delta log F0)、对数能量(log energy)和谱倾斜度(spectral tilt)。通过对比三种设置:冻结主干基线模型(Baseline)、无辅助输入的可训练融合(Null)以及加入学习到的韵律表征的融合(Learned),实验结果表明,与基线相比,Null在Buckeye、Switchboard和AMI IHM数据集上分别降低0.71–1.45个词错误率(WER);而Learned与Null相比在三个数据集上分别仅相差+0.07、-0.09和+0.00个WER,差异均不显著。然而,当推理时移除或使用不匹配的韵律表征时,Learned的WER显著上升。这表明,尽管所学韵律表征在模型中起作用,但并未带来可量化的额外性能提升,其关键在于验证了融合机制本身已具备较强建模能力,而附加的韵律信息在当前设定下未能产生显著增量收益。

链接: https://arxiv.org/abs/2609.36754
作者: Ki Woong Moon,Daniel Brenner
机构: 未知
类目: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL)
备注: Submitted to ICASSP 2027

点击查看摘要

Abstract:Explicit prosodic cues may help automatic speech recognition (ASR) of spontaneous speech, but auxiliary representations typically require additional trainable components, making it unclear whether gains come from the auxiliary information or the fusion mechanism. We address this using a frozen HuBERT backbone and a 64-dimensional representation trained to predict log F0, voicing, Delta log F0, log energy, and spectral tilt. We compare a frozen-backbone recognizer (Baseline), trainable fusion with zero auxiliary input (Null), and the same fusion supplied with the learned representation (Learned). Across Buckeye, Switchboard, and AMI IHM, Null reduces WER by 0.71-1.45 points over Baseline, whereas Learned differs from Null by +0.07, -0.09, and +0.00 points, with no significant differences. However, removing or mismatching the representation at inference increases Learned WER. Thus, Learned depends on the representation yet shows no measurable incremental WER benefit over the parameter-matched control.

[NLP-208] Better Behavioral Prediction More Faithful Model Ablations? Evidence from Sequential Choice

【速读】: 该论文旨在解决生成式认知模型在解释人类决策行为时存在的根本性问题:即仅依赖预测准确性不足以验证模型对真实认知过程的忠实性。其核心挑战在于,传统基于输入消融(input ablation)的方法假设模型对某类信息的依赖程度等同于真实认知过程对该信息的依赖,但这一假设缺乏独立验证。为检验该假设,作者设计了两个具有已知生成策略的合成序列老虎机任务,其中历史选择在无反馈条件下仍具信息量,从而可区分预测能力与行为响应的忠实性。研究比较了从头训练的GRU与Transformer、微调后的LLaMA模型以及经典认知模型在不同奖励权重下的表现。关键发现包括:第一,在非稳定任务中,神经网络模型在无奖励监督下仍能优于四种基准行为模型;第二,在匹配的捐赠者-奖励替换条件下,高精度预测模型的响应变化远小于真实生成器;第三,在某些奖励权重下,神经网络虽预测性能更优,但其选择概率的变化幅度与强化学习聚合模型相比不具忠实性,且模型排序在不同任务间不一致。这些结果表明,预测能力与模型对扰动的响应忠实性是独立维度。因此,论文提出的关键解决方案是:在将模型消融结果用于推断行为生成机制之前,必须通过独立测试验证其响应的生物学或认知真实性,而不能仅依赖预测性能。

链接: https://arxiv.org/abs/2609.36097
作者: Hanbo Xie
机构: Georgia Institute of Technology(佐治亚理工学院)
类目: Neurons and Cognition (q-bio.NC); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Using predictive models to explain cognition requires more than accurate behavioral predictions. Input ablations offer an appealing route: remove information from a model and interpret the resulting performance change as evidence of its importance for behavior. Yet this inference assumes that the model’s dependence on information reflects the dependence of the process generating the behavior. We test it in two synthetic sequential bandit tasks with known generating policies, where past choices can remain informative when feedback is unavailable to a predictor. We compare GRUs and Transformers trained from scratch, a fine-tuned LLaMA model, and cognitive models across systematically varied reward contributions. Our analyses distinguish prediction after training without reward observations from the response of a fixed predictor to donor-reward replacement. Three findings emerge. First, in the restless task, neural models trained without rewards predict held-out choices better than four simple training-fitted behavioral baselines. Second, under matched donor replacement, accurate predictors can respond much less than the known generator. Third, at some reward weights, neural networks predict better than a pooled reinforcement-learning model but have less faithful changes in choice probabilities; the model ordering differs between the two tasks. These independent-test results separate information sufficient for prediction from response fidelity under a specified ablation in sequential choice. They motivate validating model-ablation responses independently of predictive performance before using them to infer how the observed behavior was generated.

信息检索

[IR-0] Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies

链接: https://arxiv.org/abs/2609.38155
作者: Hui Ren,Lei Fan,Henry Pao,Han Guo,Zeeshan Zia,Ying Chen,Alexander Schwing,Gang Hua
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, while observations of the same object remain disconnected across events. Retrieving relevant events therefore does not necessarily recover the “biography” of the particular entity a question concerns. To address this, we introduce Grounded Entity Biographies (GEB), a long-video memory framework that groups visually grounded observations of the same physical instance across clips into retrievable biographies while preserving the context of each moment. During question answering, the biography is retrieved alongside episodic evidence, allowing the model to follow an entity through events using identity links established during memory construction. Evaluations across four benchmarks, including day-long and week-long recordings, demonstrate improvements over prior memory frameworks in both multiple-choice and open-ended question answering. On EgoLifeQA, GEB achieves 72.0% accuracy, 4.4 percentage points above the best published result. Ablations show that grounded identity association and biography reading both contribute to the gains, which additional descriptions alone do not fully recover.

[IR-1] Effective Dense Retrieval using Only In-Context Examples

链接: https://arxiv.org/abs/2609.38099
作者: Nour Jedidi,Abdul Basit Ali,Hang Li,Jimmy Lin
类目: Information Retrieval (cs.IR); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Turning decoder-only large language models (LLMs) into strong dense retrievers typically requires some form of retriever training. In this paper, we ask whether LLMs can instead be prompted to produce effective representations for dense retrieval given only a few in-context examples. To answer this, we introduce RICE (Representations from In-Context Examples), a simple “training-free” approach that extracts high-quality dense representations from LLMs. To do so, RICE conditions the LLM on examples that provide a shared context for query and document encoding. Our results demonstrate that RICE embeddings can substantially improve the accuracy of prompt-based LLM embeddings, establishing it as a simple method to build LLM-based dense retrievers that do not require training. We release our code at this https URL.

[IR-2] Auditable Long-Term Memory: A Deterministic Retrieval Chain Measured at 479/475 of 500 on LongMemEval-S

链接: https://arxiv.org/abs/2609.38021
作者: Christopher J. Chanhnourack
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
备注: Technical report, 14 pages. Evidence repository (reader outputs, judge verdicts, control records, judge harness): this https URL (MIT). Re-scoring any run under the official judge costs about $1.28

点击查看摘要

Abstract:We evaluate an auditable long-term memory system on LongMemEval-S. Its retrieval chain uses hybrid candidate retrieval, cross-encoder reranking, coverage-first packet compilation, and deterministic reasoning scaffolds; an LLM is used only as a replaceable final reader. The chain places all gold sessions in the candidate pool for 468/470 answerable questions and produces gold-complete packets for 462/470. With a Claude Opus reader called through an unpinned CLI alias, two 500-question passes score 479/500 and 475/500 under GPT-4o. The 72 answerable knowledge-update rows used a substantively modified scoring prompt whose effect under the official text has not been measured. The pair straddles Chronos High’s published 478/500; differences in reader generation, scoring prompt, and possibly data version, plus within-system variance, establish neither superiority nor equivalence. A grok-4.6-high reader on the same packets scores 476/474, while a maximum-reasoning-effort agentic variant regresses to 461/465. The headline passes differ on eight verdict-flip rows. A second judge agrees with the headline judge on 493/500 rows (98.6%) in each pass and scores both passes 472/500; the official judge also flips three verdicts when re-scoring byte-identical pass-1 answers. Negative controls rejected a verifier that repaired three wrong drafts but broke eleven correct drafts. All components were developed on the same 500 questions, with no held-out evaluation or independent human adjudication; retrieval and scaffold method sources and transcript-derived audits are held; and the headline reader received extra operator context, its complete requests were not retained, and MCP tool availability is unresolved. We release materialized packets, scaffolds, reader outputs, judge verdicts, and controls for inspection and re-scoring.

[IR-3] BITEM at the NTCIR-19 R2C2 Task: Predicting Confidence from Agent ic RAG Pipeline Signals ICIP

链接: https://arxiv.org/abs/2609.37993
作者: Julien Knafou,Luc Mottin,Alexandre Flament,Paul van Rijen,Esteban Gaillac,Patrick Ruch
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
备注: 8 pages. Participant paper for the NTCIR-19 R2C2 task

点击查看摘要

Abstract:The BITEM team entered both subtasks of the NTCIR-19 R2C2 task with a single agentic pipeline, in which a model searches, reads and records evidence over a movie corpus while an orchestrator holds the record and rules on what may be submitted. A claim is admitted only once an entailment cascade has checked it against the passage it cites, and an answer is released only once enough checked evidence stands behind it. Each question is run three or four times, every pass retrieving from a corpus stripped of what the earlier passes have already seen. The confidence filed with each answer is computed by the orchestrator from what the run leaves behind and is never asked of the model, which is offered no way to rate itself. The two retrieval runs placed 4th and 5th of 22, pooling the passes was worth 0.0709 nDCG@20, and the gain was largest on the multi-hop and post-processing-heavy questions, where the organisers rank the pooled run top of the field. Sixteen of the 25 answer runs were built on passages these two runs supplied, 12 of them filed by other teams. HMR rewards a system whose confidence is high where it answers right and low where it answers wrong. The pipeline reached an accuracy of 0.9219, 6th of 25, while the confidence filed with those answers gave an HMR of 0.4915, 13th. A few rules crafted over those same recorded signals, with no further model call and no further retrieval, raise that to an accuracy of 0.9375, 5th, and an HMR of 0.6985, 9th. Ranking on HMR alone can reward a system for answering wrongly with low confidence, so we propose accHMR, the accuracy multiplied by HMR, which reports the reward in proportion to the accuracy, and on which the revised rules would have scored 0.6549, 5th. For future work, fitting a model on the numbers the pipeline already produces, rather than writing such rules by hand, would be a real step forward.

[IR-4] Generated Query Expansion Still Helps Strong Sparse Retrieval: A Controlled Study with SPLADE-v3

链接: https://arxiv.org/abs/2609.37911
作者: Ryan C. Barron,Cade W. Trotter,Maksim E. Eren,Kim Ø. Rasmussen,Liz D. Miller,Benjamin J. Migliori
类目: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)
备注: 8 pages, 5 tables, 3 figures

点击查看摘要

Abstract:Scientific queries are often brief, while relevant papers use specialized vocabulary. Generated query expansion can bridge this mismatch, but earlier work suggests that its value shrinks as the underlying retriever becomes stronger. We test the four generated formats of term lists, a pseudo-document, multiple pseudo-references, and corpus-steered text all together with SPLADE-v3 on NFCorpus, TREC-COVID, and SciDocs. Every condition searches the same frozen document index and follows the same query-side integration rule and 256-dimension budget, isolating the effect of the added content. All twelve method-collection comparisons improve aggregate nDCG@10, with best relative gains of 4.81%, 8.92%, and 9.47%. Eleven remain significant after Holm correction. The gain persists in 103 of 114 interpolation settings, including every setting that assigns at least 30% of the mixture weight to the original query. Shuffled-text and non-contextual lexical-bag controls also remain above baseline in all 24 aggregate comparisons, showing that the added vocabulary carries most of the benefit. A corpus-induced typed concept graph, by contrast, produces no consistent gain, and its relation, depth, validation, random, and gating controls do not rescue it. Generated vocabulary can therefore complement a strong learned sparse retriever, provided that the original query remains strongly represented.

[IR-5] owards Semi-Automatically Comparing Keyword-Based and Semantic Search Accuracy

链接: https://arxiv.org/abs/2609.37749
作者: Mohamed Ben Salha,Fiete Lüer,Maik Betka,Stefan Wagner
类目: Information Retrieval (cs.IR)
备注: 8 pages, 3 figures, 2 tables

点击查看摘要

Abstract:The increasing importance of Information Retrieval (IR) in managing large datasets has highlighted significant limitations in traditional keyword-based search systems. Context-aware chat-based search methods, such as Retrieval Augmented Generation (RAG), have recently emerged, but their evaluation compared to keyword-based systems often relies on subjective user feedback. A rigorous, quantitative comparison between these paradigms remains lacking. This work introduces a novel, preliminary framework to quantitatively assess IR accuracy of search systems that produce different output formats, such as lists and messages. It focuses on two key aspects: the ranking accuracy for keyword-based systems and the completeness of retrieved information for semantic chat-based systems. Our approach enables semi-automatic comparisons of semantic and keyword-based methods using interchangeable equivalence classes tailored to domain-specific contexts (e.g., companies or problems). We validate the framework through an industrial case study, demonstrating statistically significant improvements in context-aware search over keyword-based methods, supported by analyses including the Mann-Whitney U-Test. With its adaptable design, the proposed framework provides a strong foundation for objectively assessing keyword-based and semantic chat-based search methods.

[IR-6] MERGE: Multi-LLM Ensemble for Retrieval via Generative Enrichment

链接: https://arxiv.org/abs/2609.37574
作者: Tzu-I Ho,Yung-Yu Shih,Shang-Yu Su,Dongzhe Wang,Yun-Nung Chen
类目: Information Retrieval (cs.IR); Computation and Language (cs.CL)
备注: 9 pages, 4 tables, 1 figure. Preprint

点击查看摘要

Abstract:Large Language Models (LLMs) are increasingly used to enrich user queries in information retrieval (IR) so that a standard retriever such as BM25 can bridge vocabulary gaps with the target corpus. Any single LLM, however, is limited by its training data and architectural biases, and its enrichment behavior depends on hand-crafted prompts that must be re-engineered for each new model – an expensive and poorly scalable process. We present MERGE (Multi-LLM Ensemble for Retrieval via Generative Enrichment), a two-stage framework: three heterogeneous 7-8B open-source LLMs independently produce candidate expansions, and a larger LLM generatively synthesizes them into a single query. To make prompt engineering scalable across the ensemble, we integrate a task-grounded Automatic Prompt Optimization (APO) loop into both stages. Unlike APO methods that judge candidates with an LLM evaluator, our loop scores each candidate by its downstream retrieval performance and runs a small tournament between the current champion prompt and optimizer-proposed drafts, terminating once the champion survives two consecutive rounds; a history-augmented variant additionally feeds the recent tournament trajectory back to the optimizer. MERGE is retriever-agnostic and issues a single BM25 pass with no rank fusion, no supervised document expansion, and no re-indexing. On five BEIR benchmarks (NQ, SciFact, FiQA, Touche-2020, DBPedia), MERGE improves BM25 nDCG@10 over the original queries by +2.1 to +14.9 points and matches or outperforms strong LLM-based query-expansion baselines despite using only compact open-source models. Ablations confirm that the Stage-2 ensemble beats any single Stage-1 LLM, and that task-grounded APO converts large seed-prompt regressions into consistent gains without hand-tuning.

[IR-7] Do Evidence-Reading Diagnostics Improve Interface Selection in Small LLM Recommenders?

链接: https://arxiv.org/abs/2609.37472
作者: Han Chen,Yingrui Li
类目: Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Behavioral tests measure how a language model reads evidence. We ask whether those measurements help choose a recommendation interface. We evaluate six small instruction-tuned checkpoints across four recommendation domains with chronological evaluation and 3,426 evaluation users. Each request ranks eight candidates. A baseline selector chooses among history-only prompting, prompting with collaborative evidence, and score fusion. It uses observable features and six stability prompts that vary wording and candidate order. An augmented selector adds features from six evidence-reading prompts that ask the model to compare support counts. An interface chosen once on development (validation) data for each domain and checkpoint scores 0.5524 NDCG@5, compared with 0.5447 for the baseline selector and 0.5428 for the augmented selector. Adding the diagnostic features changes NDCG@5 by -0.0019 (95% interval [-0.0046, 0.0004]). The interval includes zero, and its upper bound is below the analysis plan’s 0.005 improvement target. Matching the selectors’ hyperparameters also leaves the interval upper bound below that target. Evidence from retrieved similar users improves prompting by 0.0999 NDCG@5 over a control using randomly selected users matched for activity. The evidence-reading tests also reveal answer-position and tie-response biases. These results concern the tested selectors and candidate sets. They illustrate why diagnostic measurements should be evaluated by whether they improve recommendation choices beyond existing features and a fixed interface.

[IR-8] Relevance Is Not Sufficient Evidence: Detecting Evidence Gaps Before Generation in RAG

链接: https://arxiv.org/abs/2609.37469
作者: Suting Chen,Peichun Hua,Yunming Xiao
类目: Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注: 22 pages, 7 tables, 2 figures

点击查看摘要

Abstract:Retrieval-augmented generation (RAG) grounds large language models in external sources, but retrieved passages often name the right entities without providing the facts needed to answer. Even when instructed to abstain, 12 generators answer 40.0-99.3% of insufficient-evidence questions. Training generators to abstain ties the decision to model weights, may reward answers recalled from parametric knowledge, and still requires a full generator call. Can sufficiency be judged from the question and evidence alone, before any answer exists? We identify pitfalls in constructing insufficient-evidence tests: removing relevant evidence or pairing evidence with unrelated questions can reveal labels through lexical overlap or evidence position. We build a paired benchmark using substitution, deletion, and question-swap constructions that vary answer support while controlling selected surface features, such as word use. Sufficiency can be judged without generating an answer, but no single signal works across all datasets. We introduce RINSE (Relevance Is Not Sufficient Evidence), which combines three signals: whether every part of the question is covered, whether any passage offers an answer, and whether a small language model reading the passages together judges them sufficient. Across six datasets, RINSE ranks sufficient above insufficient evidence with a score of 0.837 (chance 0.5), exceeding the best of 10 prior methods (0.746) and a frontier model queried through an API (0.784). Its weakest dataset scores higher than any other method’s weakest (0.684 vs. 0.676). RINSE runs locally before generation, taking 36.5 ms per question on a single GPU.

[IR-9] Backdoor in the Loop: Compromising Agent ic Search via Malicious Retrievers

链接: https://arxiv.org/abs/2609.37468
作者: Beining Xu,Peichun Hua,Yunming Xiao
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注: 24 pages, 13 tables, 4 figures

点击查看摘要

Abstract:Agentic retrieval-augmented generation (RAG) interleaves reasoning with repeated retrieval, giving the retriever influence over both the evidence an agent observes and its subsequent search decisions. We study retriever backdoors that exploit this feedback loop and repurpose weak backdoor purification to conceal their presence. An attacker supplies a compromised retriever checkpoint while leaving the search agent and deployment corpus unchanged. Without corpus write access, the attacker can still suppress useful evidence, persistently retrieve a selected existing document, or steer the agent toward prolonged search, inflating retrieval, context, and latency cost. To conceal these behaviors from detection, we propose leveraging a controlled inject-and-remove cycle: deliberately inject a weaker backdoor and then unlearn it. This process weakens detector-visible signatures and fools the backdoor detectors with an illusion of purification while preserving the malicious retrieval behavior. These findings expose a systematic vulnerability in RAG systems in which a weak defense becomes an attacker’s concealment tool for a backdoored retriever, even when the underlying corpus remains trustworthy.

[IR-10] ReMem: Rethinking Perception and Memory in Long-Context Recommendation Agents

链接: https://arxiv.org/abs/2609.37311
作者: Haohao Qu,Yongcheng Jing,Chun Hin Chan,Shanru Lin,Wenqi Fan,Dacheng Tao
类目: Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
备注: Work in progress

点击查看摘要

Abstract:Recent Recommendation Agents (RecAgents) offer a promising alternative by shifting recommendation to an active, user-side paradigm, where generative agents autonomously perceive external platforms, reason over user preferences, and execute decisions. However, existing RecAgents still suffer from two critical limitations: brittle item perception based on noisy and heterogeneous item pages, and inefficient long-context reasoning over extended user histories and multi-step interaction traces. To address these challenges, we propose a novel recommendation agent framework, termed as ReMem, that combines OCR-based multimodal perception with time-evolving dynamic memory. Instead of parsing raw HTML, ReMem observes item pages through screenshots and extracts structured multimodal information via an OCR tool, enabling a more humanoid and platform-agnostic perception mechanism. To support long-horizon preference modeling, ReMem further introduces a chunk-wise sequential memory update strategy, where the agent selectively maintains a fixed-size memory of informative historical interactions while processing arbitrarily long contexts with linear inference complexity and bounded context length. This design allows the agent to preserve evolving user preferences without relying on external memory modules or disrupting the standard autoregressive generation process. To enhance the dynamic memory instruction, we further develop a multi-memory GRPO variant, which propagates the final-answer advantage to all intermediate conversations that contribute to the final response. Extensive experiments on three datasets demonstrate that ReMem consistently outperforms state-of-the-art baselines, achieving an average improvement of 5.16% across three recommendation agent tasks, namely searching, ranking, and judging.

[IR-11] Follow the Entities: A Corpus Map for Agent ic Search

链接: https://arxiv.org/abs/2609.37226
作者: Soyeong Jeong,Sujay Kumar Jauhar,Sung Ju Hwang,Andrew Joohun Nam
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project’s approval recorded in one, its requirements in another, and its latest status in a third. Recent LLM agents approach this by iteratively searching the full corpus rather than reading only a fixed set of top-ranked documents. However, when the corpus is exposed only as a flat collection of files, a relevant document gives no indication of how it relates to others, so the agent must rediscover these relationships for every query, often missing complementary evidence while simultaneously consuming substantial additional tokens. To address this, we introduce CorpusMap, a navigation layer that organizes the corpus around its recurring entities, which are identifiable from the documents themselves and can link a single document to many others across sources. Specifically, CorpusMap represents each recurring entity as an Entity Page that aggregates information about it and links to every document that refers to it, forming a graph between entities and documents that the agent can traverse to gather otherwise disconnected evidence. Moreover, since CorpusMap is constructed offline by resolving mentions of the same entity across documents, its links are shared across queries rather than rediscovered repeatedly at inference time. Using 7 different models with 3 benchmark datasets, we show that CorpusMap improves both evidence discovery and answer quality over raw-corpus agentic search while using fewer tokens on average, and further outperforms 4 alternative navigation layers, suggesting that entities serve as effective anchors for navigating large document collections.

[IR-12] HELIX: Purified and Unified - Rethinking Feature Interaction and Sequence Modeling for Large-Scale Recommendation

链接: https://arxiv.org/abs/2609.37183
作者: Yuntao Zheng,Miao Zhang,Yadong Ding,Yanchuan Tang,Lixiyu Chen,Hao Wang,Quan Li,Shiying Cai,Yue Lin,Jiayu Li,Yu Feng,Wentao Yang,Rongkun Xing,Jiekai Wang,Mingge Zhang,Feiling Gong,Xiang Gao,Jinyu Dong,Yajing Zhang,Pengfei Ren,Yinzhou Wang
类目: Information Retrieval (cs.IR)
备注: 17 pages, 3 figures. Technical report

点击查看摘要

Abstract:Industrial recommendation ranking models typically scale along two modeling axes: feature interaction over heterogeneous user, item, context, and cross features, and sequence modeling over long, informative, and multi-type user behavior histories. We find that scaling either capability in isolation is insufficient, as each exhibits a limited scaling ceiling and a suboptimal scaling-law slope. We conjecture that achieving a more favorable scaling-law slope requires jointly scaling both axes. To support this, we present HELIX, a purified and unified architecture for large-scale recommendation. HELIX interleaves sequence retrieval and feature interaction while enforcing one-way information flow from reusable sequence states to candidate-conditioned mix-tokens. This design preserves cross-depth communication between the two modeling axes while keeping user-side sequence computation amortizable, enabling flexible and asymmetric scaling of sequence modeling and feature interaction. Deployed in TikTok’s e-commerce recommendation system, HELIX consistently improves offline CTR AUC, CVR AUC, and other ranking metrics. In online A/B tests, it achieves an approximately 6% increase in e-commerce video GMV per user.

[IR-13] Optimizing VLP-aligned Multimodal Intent Representation with Correct Visual Instantiation for Zero-Shot Composed Image Retrieval

链接: https://arxiv.org/abs/2609.36946
作者: Xuri Ge,Chunhao Wang,Junchen Fu,Haokun Wen,Zhiwei Xu,Ying Zhou,Zhumin Chen,Pengjie Ren,Zhaochun Ren,Xin Xin
类目: Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:ZS-CIR aims to retrieve a target image from a reference image and a modification text without paired supervision, typically by encoding composed queries as text-dominant representations within the image-text matching space of VLPs. However, queries reconstructed by visual pseudo-word learning or MLLM-based target reasoning often deviate from the native VLP representation space due to reference noise and coarse text fusion in the former, and verbose, weakly visually grounded descriptions in the latter. In this paper, we propose a unified ZS-CIR framework (named VMIR-CVI) to reconstruct multimodal composite queries from two complementary perspectives for optimizing VLP-compatible multimodal intent representation. First, it reasons and converts the multimodal intent into a unified textual description, aligning with the native text space of the VLP backbones to produce more retrieval-compatible textual queries. Second, it reconstructs the query representation with correctly decoupled visual instance cues, reducing reference noise while preserving target-relevant content. Specifically, a VLP-aligned Multimodal Intent Reasoning (VMIR) module injects few-shot VLP-style exemplars into chain-of-thought prompts, guiding the MLLM to generate target-consistent intent queries. A Training-free Visual Instance Disentanglement (TVID) module decouples fine-grained visual instances from global reference features without additional optimization. Finally, a lightweight Hybrid-modal Intent Alignment and Fusion (HIAF) module integrates the reasoned textual intent and disentangled visual cues into a unified hybrid-modal representation for robust ZS-CIR. Extensive experiments on three CIR benchmarks, namely CIRR, CIRCO and FashionIQ, show that VMIR-CVI significantly outperforms existing baselines and achieves new state-of-the-art performance. Code and trained models will be publicly released.

[IR-14] Safer Content or Firmer Refusals? A Hybrid Perturbation Defense for Alignment under Harmful Fine-tuning CCS

链接: https://arxiv.org/abs/2609.36862
作者: Muhammad Zeeshan Akram,Mufid Kamel Marican,Anvesh Reddy Yenugu,Ali Zain Kaimkhani,Minghong Fang
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注: To appear in CCS-LAMPS 2026

点击查看摘要

Abstract:Fine-tuning-as-a-service lets users adapt a safety-aligned language model to their own data, but it also creates a harmful fine-tuning attack surface: a small amount of harmful data mixed into an otherwise benign fine-tuning set can degrade the model’s alignment. Two recent alignment-stage defenses address this problem at different levels of the model. Vaccine improves the robustness of hidden embeddings to the representation shifts induced by harmful fine-tuning, whereas Booster simulates harmful weight updates and attenuates their effect during alignment. We investigate whether these mechanisms are complementary and propose VaccineBooster, a single alignment procedure that combines embedding perturbation and weight-level gradient attenuation within each training step. On Llama-2-7B aligned with BeaverTails and then attacked through poisoned fine-tuning, VaccineBooster achieves the lowest OpenAI moderation score among the compared defenses, 0.315, while a Booster-Only variant retains the highest post-attack refusal rate, 50%. Together with ablations over the embedding-perturbation and gradient-attenuation strengths, these results indicate a trade-off: embedding perturbation primarily reduces flagged harmful content, whereas gradient attenuation primarily preserves explicit refusal behavior. Because our evaluation uses ten prompts and a single unseeded run per configuration, we report this trade-off as an observed pattern rather than a statistically resolved effect. These results provide practical guidance for prioritizing content safety or refusal retention when aligned models are exposed to untrusted fine-tuning.

[IR-15] Does the Unsafe Gradient Survive a Conversation? On the Frag ility of Gradient-Based Jailbreak Detection in Multi-Turn Dialogue CCS

链接: https://arxiv.org/abs/2609.36849
作者: Omar Sheta,Rinku Deuja,Hadi Masoudi,Minghong Fang
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注: To appear in CCS-LAMPS 2026

点击查看摘要

Abstract:Safety-aligned language models are commonly deployed as multi-turn assistants, which lets adversaries spread unsafe intent across several user turns instead of a single prompt. Gradient-based jailbreak detectors such as GradSafe were developed for single prompts: they score an input by the alignment between its induced gradient and a fixed unsafe reference direction, and their effectiveness in multi-turn dialogue remains unclear. We conduct a controlled evaluation of gradient-based jailbreak detection in multi-turn settings. We extend GradSafe with a Context Window Scanner that applies the detector to fixed-size windows of user turns and uses the maximum window score as the conversation-level score. We evaluate different window sizes, attack families, benign conversation distributions, and target models. The results differ sharply between synthetic and realistic benign settings. Against synthetic benign conversations, the detector achieves an ROC-AUC of 0.98 on human-authored multi-turn jailbreaks. On WildChat benign conversations, ROC-AUC drops to 0.76, and a threshold calibrated on synthetic data flags more than 90% of benign conversations as unsafe. Under realistic benign distributions, single-turn windows give the highest separability, whereas longer windows and accumulated contexts reduce performance. The detector is also sensitive to the attack-generation method and target model: successful Crescendo attacks receive scores comparable to or lower than benign conversations, and Qwen2.5-7B-Instruct yields near-random separability with a different optimal window size. These findings show that gradient-based signals can support multi-turn jailbreak detection, but reliable deployment requires calibration on realistic benign conversations, short-window scoring, length-aware thresholds, and evaluation across attack types and model architectures.

[IR-16] GRP v0.1 Technical Report

链接: https://arxiv.org/abs/2609.36688
作者: Wenfeng Zhuo,Vincent Xue,Charles Wei,Cong Ni,Ruiming Lu,Jiwen Ren,Mo Li,Peng Yang,Xufei Wang,Dongheng Li,Jiacong He,Yi Song,Yufei Fan,Mikhail Obukhov,Yiwen Chen,Yvette Liu,Yin Ye,Chengjie Wu,Mingtao Zhang,Jinchao Ye,Lili Zhang,Chunhui Zhu
类目: Information Retrieval (cs.IR)
备注: 26 pages, 3 figures, 11 tables. Technical report

点击查看摘要

Abstract:Industrial recommendation systems rely on multi-stage cascades whose retrieval, ranking, and serving components are difficult to replace jointly. We present GRP, a generative recommendation framework that combines retrieval, ranking, and reward modeling in a single encoder-decoder model, and evaluate a progressive path toward end-to-end recommendation. The model generates multimodal Semantic IDs and scores candidates with a jointly trained ranking module. The frozen ranking module then supplies rewards for reinforcement-learning post-training. We introduce mGRPO, which adds a reference-anchored margin to reward optimization to preserve the likelihood of logged targets. Offline experiments examine history encoding, model capacity allocation, event selection, tokenization, and reward discrimination. Serving optimizations reduce end-to-end retrieval latency by 69%. Online experiments evaluate the model as a retrieval source, with early-ranking bypass, and with replacement of weaker sources. In a retrieval-only comparison, view time increases by 0.46% and shares by 0.77% relative to production. A separate comparison combining bypass and source replacement yields increases of 0.82% in view time and 2.56% in shares, with neutral platform-level guardrails. These results support progressive deployment while identifying remaining gaps in ranking quality and performance across recommendation metrics.

[IR-17] Retrieval Sensitivity to Identity Signals in Queries EMNLP2026

链接: https://arxiv.org/abs/2609.36534
作者: Andrew Tang,Nicholas Deas,Kathleen McKeown,Vishal Misra
类目: Computation and Language (cs.CL); Computers and Society (cs.CY); Information Retrieval (cs.IR)
备注: EMNLP 2026 camera-ready, with a correction to Fig. 4

点击查看摘要

Abstract:Dense retrievers decide which documents reach users and the language models that use them, yet they are typically evaluated with neutral queries. We ask whether the identity signals that real users express in their queries—political ideology and dialect—bias what a retriever returns. We design evaluations in two domains, political news and consumer-health questions, each pairing a controlled synthetic set that varies only the identity signal with naturalistic queries. Across five dense retrievers and a sparse baseline, every retriever (i) retrieves articles that align with the query’s own political lean and (ii) performs worse for questions written in African American Language (AAL) than in White Mainstream English (WME). Two analyses tie these gaps to queries’ identity signals beyond surface vocabulary: partialling out an aggregate lexical-asymmetry score leaves the synthetic gaps largely intact, and linear probes recover lean and dialect from the retrievers’ query embeddings beyond token-level features. Left unaddressed, such retrieval biases risk contributing to polarization and reinforcing the health disparities already faced by AAL speakers. Code is available at this https URL.

[IR-18] ARCagent : An Adaptive Retrieval Calibration Agent for Clinical Question Answering

链接: https://arxiv.org/abs/2609.36392
作者: Yuyan Chen
类目: Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
备注: 13 pages, 7 figures, 5 tables

点击查看摘要

Abstract:In diseases where clinical guidelines are incomplete, contested, or mutually contradictory, knowledge completeness and dynamic conflict-aware synthesis are two safety-critical properties that standard Retrieval-Augmented Generation systems do not provide. Therefore, we present \sysname, an adaptive retrieval calibration clinical question-answering agent for ME/CFS, a disease where diagnostic frameworks coexist and major guidelines actively contradict each other on treatment. ARCagent contributes three components. First, a 1,706-chunk, 10-source knowledge base with a structured inter-guideline conflict registry spanning all active ME/CFS diagnostic frameworks. Second, a conflict-aware retrieval calibration pipeline that re-ranks retrieved evidence using query-specific focus and conflict signals. Third, a benchmark scored by LLM-as-Judge, avoiding systematic underestimation averaging 10.1 percentage points caused by keyword matching. ARCagent achieves 95.3%, outperforming all base LLMs. Code is available at this https URL.

[IR-19] Better Nearest Neighbor Graph Indices via (Efficient) LLM -Guided Pruning

链接: https://arxiv.org/abs/2609.36359
作者: Fangzhou Wu,Haike Xu,Sandeep Silwal
类目: Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
备注: 29 pages

点击查看摘要

Abstract:Graph-based approximate nearest neighbor search (ANNS) is widely used for large-scale semantic search. Its indices are constructed primarily based on geometric relationships among embeddings of an input dataset (e.g., documents or images), rather than explicitly optimizing for semantic relevance. However, when using these indices for downstream query retrieval, performance is evaluated based on the semantic relevance of the retrieved results to the query. This creates a fundamental “geometry-semantic” mismatch between how the indices are constructed and how their retrieval results are evaluated. While existing LLM-based reranking methods can partially mitigate this mismatch at query time, they leave this underlying structural problem in the graph unresolved. We therefore propose LLM-Guided Graph Pruning (LGP), a general framework that addresses this mismatch directly by leveraging LLM reasoning to refine an existing ANN graph index itself. LGP identifies structurally “low-value” neighbors of nodes and replaces them with LLM-selected alternatives that provide useful semantic information while retaining desired geometric structures of the original graph, including sparsity and efficient navigability. Experiments on representative semantic retrieval benchmarks show that LGP consistently improves end-to-end retrieval performance over both vanilla greedy graph search and LLM-based reranking across widely used graph-based ANN indices such as DiskANN and HNSW.

[IR-20] huRunel: Dynamic Decoupling for Structured Advisory Dialogue

链接: https://arxiv.org/abs/2609.36340
作者: Yuyan Chen
类目: Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Information Retrieval (cs.IR)
备注: 14 pages, 22 figures, 8 tables

点击查看摘要

Abstract:High-stakes advisory domains such as medical aesthetics, legal consultation, and educational planning exhibit a two-phase structure. The early phase requires empathetic elicitation and emotional support, and the late phase requires authoritative specialist judgment. Neither fully automated agents nor human junior consultants adequately address this structure at scale. We formalize the core design challenge as dynamic decoupling, asking how an AI advisory agent should decide what to ask, when to stop, what to resolve autonomously, and what to forward to the specialist. We present ThuRunel, an advisory agent combining a finite-state belief management framework, a chain-of-thought teacher synthesis protocol, and learned generation adapters. Against eleven baselines, ThuRunel achieves consistent improvements in elicitation completeness and specialist brief quality. ThuRunel is publicly deployed as a bilingual web application in which the same decoupling decisions operate from the client’s side, grounded in a curated knowledge base that cites its sources in every answer.

[IR-21] GeoOutageBench: Benchmarking Ambiguity-aware Ontology-grounded Geospatiotemporal KGQA for Multimodal Power Outage and Resilience Analysis

链接: https://arxiv.org/abs/2609.36082
作者: Ethan D. Frakes,Amy Kvien,Rishabh Kundu,Redad Mehdi,Van D. Tran,Vibha S. Mandayam,Kristopher O. Davis,Erika I. Barcelos,Roger H. French,Yinghui Wu,Mengjie Li
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)
备注: 13 pages, 6 figures, 7 tables. Accepted to the 34th ACM International Conference on Advances in Geographic Information Systems (SIGSPATIAL '26), November 3-6, 2026, Riverside, CA, USA

点击查看摘要

Abstract:We introduce GeoOutageBench, a benchmark for assessing LLM-based geospatiotemporal KGQA for multimodal outage and resilience analysis. Unlike existing KGQA benchmarks for Web knowledge, GeoOutageBench considers a spatiotemporal KG that integrates visual, textual, and structured data from outage records, remote sensing, weather observations, storm and power events, geographic entities, and domain ontologies. It provides a competency query taxonomy at different difficulty levels from spatiotemporal containment and proximity, spatiotemporal co-occurrence analysis, multimodal evidence, to hypothetical evaluation. Over multimodal KG and query classes, GeoOutageBench provides user-configurable evaluation of three important, highly coherent yet less studied tasks: (1) LLMs’ understanding for ambiguous geospatiotemporal questions in terms of NL to SPARQL interpretation, (2) query-driven assessment of ontology utility, and (3) answer accuracy of multimodal KGQA retrieval. GeoOutageBench provides a design principle and foundation for assessing LLM-KG systems that support real-world infrastructure resilience analysis. Our benchmark, source code, data, results, and other documentation are available at this https URL.

[IR-22] Mnemon: Raw Records Fast Judgments Slow Thoughts

链接: https://arxiv.org/abs/2609.36059
作者: Guangren Wang
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
备注: 16 pages, 3 figures, 4 tables. Code, prompts and run records: this https URL

点击查看摘要

Abstract:Long-term memory lets an LLM assistant use a history it can no longer reread, and most memory systems build it by rewriting conversations into facts, graphs or typed memories at write time. We argue that the work of memory divides, as thinking does, into two systems. Most of it is fast System 1 work: many small, independent yes/no judgments about records, such as whether a record is needed or no longer current, which a decision model makes by the dozen in a third of a second. Only a little is slow System 2 work: writing a few search queries, naming what the reply needs and composing the answer, which an LLM does well but slowly. We present Mnemon, a memory agent built on this division. It keeps conversations as raw, dated records; an LLM (System 2) plans searches over them, a decision model, Jev (System 1), judges what the searches return, and rules with explicit budgets turn the judgments into a small View for an unchanged answering model. A background pass consolidates each record once into topic timelines, value histories and standing instructions linked to the records, so that questions about a whole conversation reach evidence their own searches miss. Because nothing is decided about a record when it is written, the same agent can read any store that returns dated records. With gpt-4.1-mini answering, as in a public re-evaluation of 14 systems, Mnemon scores 91.7% on LoCoMo, the highest among them, and 83.8% on LongMemEval-S, from under 4k tokens of context per question, with the lowest effective cost index on LoCoMo. With a reasoning model answering, it reaches 92.2% on LoCoMo and 94.4% on LongMemEval-S, the latter on par with the best published results. From 100K to 10M tokens of history on BEAM, its cost per question grows by a factor of 1.11. On the same records, Jev separates gold evidence better than two LLMs and is 3-11 times faster. Comments: 16 pages, 3 figures, 4 tables. Code, prompts and run records: this https URL Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR) Cite as: arXiv:2609.36059 [cs.CL] (or arXiv:2609.36059v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2609.36059 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[IR-23] Structured Interaction Visual Localization and Robust Execution for Complex Web Tasks: A Technical Report on the WebRetriever Challenge

链接: https://arxiv.org/abs/2609.35904
作者: Ziqi Zhang,Shaohui Li,Bing Li
类目: Information Retrieval (cs.IR); Multimedia (cs.MM)
备注: Winning Report for the WebRetriever Challenge

点击查看摘要

Abstract:This report presents the web agent system developed for the WebRetriever Challenge. The system follows a structuredinteraction- first strategy, using semantic webpage information for routine browser operations and invoking visual perception only when structured representations are insufficient. Three key designs are introduced: grid-assisted visual localization for difficult-to-access controls, hierarchical context management for reducing redundant page and interaction history, and fault-aware execution mechanisms for stable multi-browser task processing. The system achieved a pass rate of up to 79% in local evaluation on Protocol 1. In the official Protocol 3 competition, it achieved a 59% pass rate with eight concurrent browser workers and ranked first overall, winning the WebRetriever Challenge.

[IR-24] Soft Curriculum Learning for Optimizing Fresh and Generalized Recommendations RECSYS’26

链接: https://arxiv.org/abs/2609.35783
作者: Arnab Bhadury,Siyan Zheng,Anlan Yu,Palaksh Rungta,Jiawei Li,Changping Meng,Dapeng Hong,Chuan He,Onkar Dalal
类目: Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注: to be published in ACM RecSys '26, Minneapolis, MN, USA

点击查看摘要

Abstract:Large-scale recommender systems, particularly short-form video platforms, are often bottlenecked by massive popularity feedback loops. In such environments, as models recommend popular items, they generate an overwhelming amount of skewed training data for “head” items. This creates a self-reinforcing cycle where retrieval and ranking models memorize “head” item patterns at the expense of generalizing across the vast “tail” of the catalogue. While Curriculum Learning (CL) offers a powerful mechanism to break this feedback loop by systematically exposing models to progressively more difficult and less frequent examples, its adoption in industrial recommendation has been hampered by hardware utilization inefficiencies or the needs for complicated pre-processing techniques because dynamic data rejection algorithms tend to starve hardware accelearators (TPUs/GPUs) by becoming largely CPU-bound. In this work, we introduce a scalable Soft Curriculum Learning framework designed specifically for continuous training setups within industry-scale retrieval and ranking models. By utilizing loss annealing and in-graph weight adjustments rather than rigid data filtering, we break the popularity feedback and enable dynamic curriculum pacing without sacrificing system throughput. We demonstrate empirical evidence through applications across sequence-based retrieval models (such as SASRec), two-tower retrieval models, and large-scale continuous ranking models. Online A/B tests on our short-video platform demonstrate substantial lifts in both overall user satisfaction and fresh content consumption, all without degrading model throughput.

[IR-25] Financial Evidence Crowding: Diagnosing and Mitigating Constraint-Induced Displacement in Retrieval-Augmented Generation

链接: https://arxiv.org/abs/2609.35782
作者: Yixi Zhou,Jiayi Yin,Fan Zhang,Xiangyi He,Haipeng Zhang
类目: Information Retrieval (cs.IR)
备注: 10 pages, 2 figures, 9 tables

点击查看摘要

Abstract:Retrieval-augmented generation (RAG) retrieves candidate evidence and sends only a limited top-ranked subset, the top-k context, to a generator. In financial question answering, passages can match a query’s topic while conflicting with its period, segment, metric scope, or table scope. We study the resulting set-level ordering failure, which we call financial evidence crowding. FinDeCrowd-Stress isolates this failure through matched compatible and incompatible candidate pools while fixing the query, relevant evidence, ranking model, candidate count, and retrieval budget. On a FinDER test split containing companies unseen during training, incompatible pools reduce top-10 evidence inclusion (Recall@10) by 0.147 relative to equally difficult compatible pools. This gap shows that conflicting candidates consume limited context slots and displace answer-supporting evidence. We then introduce FinDeCrowd-RAG, a learned score correction that combines a fixed relevance score with typed compatibility and local lexical competition. A query-level identity gate applies the correction only when it predicts a better order; otherwise, it preserves the original ranking. On identical controlled candidates, FinDeCrowd-RAG raises top-10 evidence inclusion from 0.757 to 0.902 by recovering evidence already present in the candidate set. On a FinDER index built without query-specific candidate insertion, gated reranking raises this inclusion rate from 0.420 to 0.492, while top-100 retrieval coverage remains 0.743 by design. With a fixed generator, the same ordering change improves answer accuracy and citation recall on FinanceBench and FinQA. These results identify constraint-induced displacement as a measurable RAG evaluation target and show that identity-gated reranking can recover evidence already covered by first-stage retrieval.

[IR-26] SG Suggester: Tree-Structured Knowledge-Graph Retrieval for Troubleshooting Guide Recommendation in Cloud Incident Management

链接: https://arxiv.org/abs/2609.35780
作者: Shawn Pan,PavanUttej Ravva,Walt Williams,CJ Barberan,Nutan Sahoo,Ziran Min,David Gross,Irene Shaffer
类目: Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:On call engineers in large scale cloud services work under intense time pressure, yet locating the correct Troubleshooting Guide (TSG) for an incident remains a largely manual, keyword driven process, and prior empirical work finds that guide search consumes a substantial fraction of total mitigation time. We present TSG Suggester, a retrieval system that recommends relevant TSGs directly from an incident description. We evaluate five retrieval strategies: Text Only RAG, Image Augmented RAG, RAPTOR, Tree Structured Retrieval, and our proposed Tree + Knowledge Graph (Tree+KG), on 314 real world incidents spanning 112 unique TSGs drawn from 18 service teams on a production incident management system. Tree+KG converts each guide into a tree that preserves its native section hierarchy, attaches LLM generated problem abstractions to internal nodes to bridge the solution oriented language of guides and the problem oriented language of incidents, extracts a per guide entity knowledge graph, and fuses embedding similarity with entity level matching at query time. Tree+KG attains 54.78% Top 1 and 82.48% Top 5 accuracy, leading every baseline at every cutoff, with an 8.58 point Top 1 gain over text only RAG. Two findings are of independent interest. First, structural alignment dominates: methods that preserve or rebuild document structure outperform flat chunking where precise discrimination matters. Second, and contrary to our initial hypothesis, multimodal enrichment actively hurts. Captioning guide screenshots and injecting the captions costs 22.64 Top 5 points relative to the text only baseline because generic captions dilute embeddings rather than sharpen them. We report error analyses for both results and provide concrete deployment guidance. Subjects: Information Retrieval (cs.IR) Cite as: arXiv:2609.35780 [cs.IR] (or arXiv:2609.35780v1 [cs.IR] for this version) https://doi.org/10.48550/arXiv.2609.35780 Focus to learn more arXiv-issued DOI via DataCite Submission history From: Shawn Pan [view email] [v1] Tue, 4 Aug 2026 21:47:06 UTC (25 KB)

[IR-27] Post-Generation Verification Dominates Retrieval Optimization: A 24 Factorial Ablation of RAG Pipeline Features

链接: https://arxiv.org/abs/2609.35774
作者: Ng S. T. Chong
类目: Information Retrieval (cs.IR)
备注: 11 pages, 11 figures

点击查看摘要

Abstract:Modern RAG pipelines stack many enhancement features, but these features are typically validated in isolation, leaving their interactions unmeasured. We run a 2^4 full factorial ablation of four pipeline features – section expansion (SE), agentic search (AS), completeness check (CC), and table-of-contents-guided retrieval (ToC) – across 16 configurations, 24 queries spanning eight interaction types, and two cloud-class models (768 conditions) on five public documents (78-492 pages), scoring every answer against a verified reference. Post-generation verification dominates: CC is the strongest feature (d=+0.48, p0.001), improving accuracy, completeness, and usefulness simultaneously, and CC alone (4.31/5) outperforms every configuration without it, including the three-feature SE+AS+ToC (4.11). ToC yields a significant gain at zero LLM cost (d=+0.22); AS is small and unstable, helping some queries and harming others; SE is neutral. The highest-quality configuration roughly doubles baseline latency, producing a genuine quality-latency Pareto frontier of six configurations. Feature utility is strongly query-type dependent – CC reaches d=+0.83 on completeness-demanding queries – so single-query-type evaluations systematically mis-rank features. We conclude that verifying answers matters more than optimizing retrieval, and that factorial designs with diverse query types are necessary to evaluate RAG features.

[IR-28] Socrates-RAG : Premise-Directed Inquiry against Coordinated Evidence Poisoning

链接: https://arxiv.org/abs/2609.35773
作者: Renyu Zhao,Xinyuan Zou,Lanbin Liu
类目: Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:Retrieval-augmented generation (RAG) defenses typically decide how to filter or aggregate a fixed retrieved set. In open-corpus question answering, however, decisive evidence may be absent from the initial context but retrievable, making the next query part of the reliability problem. We introduce Socrates-RAG, a premise-directed active retrieval policy that represents competing answers, selects an unresolved premise whose resolution would discriminate them, and uses newly acquired evidence to refine a subsequent query before answering or abstaining. We formalize the resulting finite-budget evidence state and give a conditional rescue guarantee relative to repeated or topical-query policies. We evaluate Socrates-RAG against a matched control in which the same backbone generates ordinary relevance-oriented search queries; both policies share the initial evidence, deterministic retriever, two-query/top-three budget, answer prompt, and label-free evidence-chain release rule. On a disjoint 48-world counterfactual evaluation, premise-directed inquiry raises safe accuracy from 79.2% to 93.8%, with 8 wins, 1 loss, and 39 ties (two-sided exact p=.0391). Decisive-evidence recall improves by the same margin, while unsafe answers fall from one to zero. Both policies solve all 24 one-hop cases; the gain is concentrated in two-hop cases, where Socrates-RAG substitutes a newly resolved premise into its second query. This controlled study isolates a specific benefit of premise-directed acquisition without claiming general robustness on the open Web. Subjects: Information Retrieval (cs.IR) Cite as: arXiv:2609.35773 [cs.IR] (or arXiv:2609.35773v1 [cs.IR] for this version) https://doi.org/10.48550/arXiv.2609.35773 Focus to learn more arXiv-issued DOI via DataCite

人机交互

[HC-0] Gender bias across LLM s is common and highly heterogenous

链接: https://arxiv.org/abs/2609.38036
作者: Edoardo Bolzoni,Valerio Capraro
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Understanding gender biases in large language models (LLMs) is increasingly important as these systems become embedded in decision-support tools with real consequences. Prior research has focused only on a small set of models, leaving open the extent to which gender biases are common and heterogeneous across LLMs. We address this gap across ten models released between April 2025 and June 2026, spanning nine vendors, using two paradigms: gender attribution to stereotyped phrases (Study 1) and moral judgment of abuse or torture against a woman or a man to prevent a catastrophic outcome (Study 2). In Study 1, two of ten models attributed masculine-stereotyped phrases to female writers more often than the reverse, while three models showed the opposite pattern. In Study 2, several models converged on a male-disadvantaging asymmetry that was directionally consistent with a documented human tendency to protect female targets from harm, though the specific conditions under which this asymmetry emerged varied by model; three other models, by contrast, showed no variation across conditions. These results indicate that gender-related biases are common in LLMs. Their direction and magnitude, however, are highly heterogeneous, to the point that some models behave in diametrically opposite ways to others. Bias auditing should therefore be treated as an ongoing, multi-vendor process, rather than a one-time assessment.

[HC-1] Critical Thinking with Generative AI: A Constraint-First Design Pilot of a Thinking-Partner Intervention

链接: https://arxiv.org/abs/2609.38029
作者: Fatima Tuz Zahra,Jiangen He,David M. Bowers,Wei Wang
类目: Human-Computer Interaction (cs.HC)
备注: 53 pages, 5 figures

点击查看摘要

Abstract:Generative AI (GenAI) tools entered higher education classrooms faster than the field was able to study their effects on learning. One concern is that GenAI may displace the critical thinking and AI literacy that students will need after graduation. This paper reports a Design-Based Research pilot of a GenAI-assisted critical thinking framework, in which ChatGPT was used as a thinking partner in an undergraduate research methods and statistics course during Spring 2025 (N = 14). The mixed-methods design combined pre- and post-intervention measures of statistical learning (AASCDM), AI literacy (MAILS), and critical thinking (WGCTA) with instructor field notes, student artifacts, and student-AI interaction logs. Pre-post tests showed gains on every AASCDM dimension and on eight of nine MAILS dimensions, while WGCTA percentiles did not change. Qualitative analysis identified four themes: the ways students positioned the LLM (as answer generator, validator, or co-thinker); the depth of student engagement (procedural vs. conceptual); occasional humanizing of the tool; and the role of curriculum design in shaping each of the prior three. Read together, the findings indicate that one semester of GenAI-assisted instruction can move domain learning and self-reported AI literacy but does not move standardized critical thinking, and that the modal student-LLM relationship is one of validation instead of dialogue. We end with design principles for the next iteration of the framework and implications for research on adaptive and personalized learning.

[HC-2] A Task-Driven Framework for Multiscale Ocean Flow Dynamics through Integrated Simulation and Visualization IEEE-VIS2026

链接: https://arxiv.org/abs/2609.37964
作者: James Kress,Jithendra Nadimpalli,Shehzad Afzal,Sohaib Ghani,Ibrahim Hoteit
类目: Human-Computer Interaction (cs.HC)
备注: Accepted for publication in IEEE Transactions on Visualization and Computer Graphics (TVCG), IEEE VIS 2026

点击查看摘要

Abstract:Internal waves are large-amplitude gravity waves that occur below the ocean surface and propagate along interfaces separating water layers of different densities. Understanding their generation, propagation, and evolution is essential, as these waves play a vital role in the ocean system by contributing to nutrient transport, biological productivity, and the transfer of energy across the ocean and continental shelf. Domain scientists use high-resolution numerical ocean models, to study internal-wave dynamics and associated coastal and nearshore processes on hybrid computational grids. These models generate large-scale, three-dimensional spatiotemporal datasets that capture internal wave flow behavior and interactions with multiple ocean variables. These datasets are generally analyzed using command-line tools with limited interactivity. To address these challenges, we in collaboration with domain scientists designed a task-driven visualization methodology for analyzing multiscale, multivariate flow data on hybrid grids. The framework incorporates a hybrid-grid volumetric reconstruction method, enabling continuous 3D analysis and a coordinated multi-view design that supports interactive exploration of complex flow structures. An insight-based evaluation with domain experts demonstrates that the system enables the identification of previously difficult-to-observe phenomena, including transverse wave propagation, energy transport pathways, and shoaling-driven mixing. Beyond the application domain, our contributions provide generalizable techniques and design principles for visual analysis of multiscale, multivariate flow data on irregular grids.

[HC-3] owards the Threshold: A Fall-Risk Anchored Pareto Framework for Virtual Reality Gait Feedback Selection for Individuals with Multiple Sclerosis

链接: https://arxiv.org/abs/2609.37952
作者: Nafisa Anjum,John Quarles,M. Rasel Mahmud
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Reduced walking speed in people with multiple sclerosis (MS) is associated with an increased risk of falls. However, virtual reality (VR)-based rehabilitation studies often emphasize performance improvements without considering the cognitive and physical effort required to achieve them. This study introduces a threshold-anchored efficiency framework that jointly evaluates gait performance and overall effort. A normative walking-velocity target of (T=1.29) m/s was derived from the mean walking speed of non-fallers in an independent MS gait dataset and used as a clinically motivated reference point. Thirty-four adults with MS were evaluated across eight VR feedback conditions spanning unimodal, bimodal, and multimodal feedback. For each condition, we quantified the proportion of the velocity gap to the target that was closed and the associated cognitive and physical burden. Pareto efficiency analysis identified five non-dominated conditions: Static Visual, Spatial Auditory, Auditory+Visual, Auditory+Vibrotactile, and Multimodal, whereas Spatial Vibrotactile and Vibrotactile+Visual were dominated. Spatial Auditory showed a favorable performance-effort trade-off, closing 72.7% of the velocity gap while maintaining below-average burden. Multimodal feedback achieved the greatest gap closure (95.6%) but also imposed the highest burden. These findings demonstrate that greater gait improvement does not necessarily correspond to greater rehabilitation efficiency and provide a quantitative framework for comparing VR feedback strategies according to both performance gains and participant burden.

[HC-4] Fluency Without Evidence: Constraint-First Design and the Limits of Self-Report in AI-Assisted Learning

链接: https://arxiv.org/abs/2609.37880
作者: Fatima T. Zahra,Wei Wang,Frances Harper,Jiangen He
类目: Human-Computer Interaction (cs.HC)
备注: 37 pages, 1 figure

点击查看摘要

Abstract:A generative AI teaching partner should support reasoning over supplying conclusions; however, this has not been tested against learning in an authentic course. Drawing on design-based research, we specify the position as a conjecture map and report a first design cycle in two graduate-level research methods courses. Students used an AI teaching partner employing a constraint-first sequence requiring them to state and justify positions before receiving questions. Pre- and post-measures of AI literacy, critical thinking, and metacognitive awareness were collected alongside interaction records. AI literacy increased, concentrating in understanding AI, whereas critical thinking, awareness, and knowledge did not change. Since changes were limited to self-report measures, they may reflect growth in confidence instead of capacity. Interaction records, meanwhile, showed brief exchanges, uneven enactment of the constraint-first sequence, and missing records. These findings show why AI-supported learning requires interaction records to provide a more defensible basis for AI-supported designs than self-reports.

[HC-5] CommSketch: How Speaking while Sketching Steers Human–AI Design Ideation

链接: https://arxiv.org/abs/2609.37813
作者: Weiyan Shi,Darryl Lim,Geraldine Quek,Kenny Tsu Wei Choo
类目: Human-Computer Interaction (cs.HC)
备注: work in progress

点击查看摘要

Abstract:Designers often speak while sketching when explaining ideas, yet AI design tools often rely on sketches or prompts, overlooking context expressed as ideas develop. We developed a sketch-based AI design interface that jointly interprets sketches and concurrent speech. Through a between-subjects study ( N=24 ), we examined how speaking while sketching steers human–AI design ideation compared with sketches alone. For creativity support, concurrent speech supported natural expression of design intent and efficient visualisation. For human–AI collaboration, speech helped establish a shared understanding of design intent, supported significantly higher perceived alignment ( p.05 ), and enabled participants to guide AI contributions as ideas co-evolved. We discuss how future human–AI design tools could support dynamic alignment, broader multimodal expression, and human–AI co-creativity.

[HC-6] Beyond Productivity: Measuring Developers Cognitive Load During GenAI-Supported Software Development

链接: https://arxiv.org/abs/2609.37645
作者: Charlotte Brandebusemeyer,Daniela Gasser,Tobias Schimmer,Bert Arnrich
类目: oftware Engineering (cs.SE); Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Generative AI (GenAI) is changing software development workflows and how developers work. Industry evaluations of GenAI adoption often monitor productivity gains, usage, and output quality, but limited attention is paid to the interaction experience and cognitive load of the actual adopters and drivers of GenAI technology - the software developers. Understanding whether GenAI changes or shifts developers’ cognitive demands during everyday development is important for a developer-centered evaluation of GenAI-supported software development. It can inform organizations in designing and evaluating effective AI-supported workflows. In this work, we study how GenAI use and task context relate to professional developers’ perceived cognitive load and whether wearable-derived physiological characteristics provide additional information beyond this context. In a four-day industrial field study at two SAP sites, 21 developers documented their tasks, task duration, GenAI use, and perceived cognitive load while wearing an EmbracePlus wristband. The results show that perceived cognitive load is associated with both GenAI use and task context, while physiological measures provide only limited additional information. These findings suggest that developers’ perceived cognitive load during GenAI-supported software development should be evaluated in relation to the concrete work context, with wearable physiological data used as complementary rather than standalone information.

[HC-7] Rhythm Is a Dancer: Designing Interactive Rhythm Feedback for Beginner Dancers

链接: https://arxiv.org/abs/2609.37641
作者: Bettina Eska,Annika Kilian,Paweł W. Woźniak,Jakob Karolus
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Learning how to dance can readily overwhelm beginners, especially without effective guidance from a dance teacher. Existing interactive systems often do not sufficiently support the learner’s progress. We investigated how targeted feedback on rhythm keeping interactively supports dance practice for novice dancers by introducing SkeletonDance. Our design is grounded in motor learning theory and conceptualized through interviews with dance teachers, following established teaching strategies. SkeletonDance automatically detects rhythm flaws and provides assistance through mimicking clapping feedback, a common instructional technique in dance lessons. In our study, participants reported that SkeletonDance helped them to re-establish lost rhythm and increased confidence during practice, especially among novices. Though objective performance metrics did not consistently confirm these effects during controlled test sessions. Our work highlights that feedback can support novice dancers’ subjective practicing experiences and demonstrates how prior dancing experience moderates the objective effectiveness of such minimal, teacher-inspired interventions.

[HC-8] Shaping Opinion: Quantifying the Psychological Impact of Autonomous Multi-Agent LLM Interactions

链接: https://arxiv.org/abs/2609.37369
作者: Marcos Rodriguez-Vega,Afonso Ferreira,Iru Exposito-Luis,Carolina Polito,Pino Caballero-Gil
类目: Human-Computer Interaction (cs.HC)
备注: 10 pages, 7 figures, 4 tables. Under review

点击查看摘要

Abstract:Natural-sounding multi-agent conversational AI is increasingly deployed, fundamentally altering human-machine interaction and human information processing. While prior work largely focuses on algorithmic failure, this study investigates the cognitive ergonomics and socio-cognitive impact of algorithmic competence. We present and evaluate FORMS (Framework for Opinion and Rhetoric in Multi-agent Simulations), a low-latency architecture for spatially mediated human-machine dialogue, driven by distinct LLM-based personas and real-time concurrency resolution. To conduct a system test and evaluation of its psychological impact, we exposed an adolescent cohort (n=120) and an adult pilot group (n=25) to a live, moderated synthetic debate. Our findings reveal that exposure to highly competent multi-agent systems triggers “Cognitive Destabilization,” fragmenting users’ prior strategic consensus. Concurrently, we observe a “Regulatory Awakening” driven by the “Normality Paradox”: fluid human-machine interactions inherently increase the baseline demand for external regulation. Furthermore, our pilot study suggests the presence of a “Truthfulness Paradox”: despite understanding the risks of generative AI, participants in the adult cohort rated the synthetic debate as significantly more sincere than equivalent human discourse (Cohen’s d=2.04). Supported by robust statistical effect sizes, this paper contributes the FORMS architecture and a replicable evaluation protocol, illustrating how high-fidelity conversational systems can reshape human information processing.

[HC-9] EntityWeaver: Visual Exploration and Curation of Named-Entity Relationships in Document Collections

链接: https://arxiv.org/abs/2609.37329
作者: Uroš Šmajdek,Ciril Bohak
类目: Human-Computer Interaction (cs.HC)
备注: Technical Report

点击查看摘要

Abstract:We present an interactive visualization system for exploring named entities and their relationships across document collections, with a strong focus on handling uncertainty and supporting both distant and close reading. The system is built around a graph that links documents, entity mentions, and entities. Uncertainty from mention-to-entity linking is included directly in this graph, so users can see where connections are strong, weak, or ambiguous. A transfer-function control, inspired by approaches in scientific visualization, allows users to adjust how this uncertainty is displayed, making it easy to tune the visualization for different datasets and research questions. The system also provides direct access to the full source texts in a coordinated view, enabling quick context checks, resolving ambiguous cases, and correcting digitization errors. Exploration is further supported through a multi-stage filter query builder, mini-map navigation for large graphs, and export options for downstream analysis. By combining uncertainty-aware graph visualization with direct interaction in the source texts, the system provides a unified workflow that supports both large-scale pattern discovery and curation of named-entity-based document collections. The design choices were supported by the domain experts, who were also involved in the initial system evaluation.

[HC-10] Early Prediction of AI-Assisted Cheating Risk in Online Exams Through Learning Analytics

链接: https://arxiv.org/abs/2609.37280
作者: Gökhan Akçapınar
类目: Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)
备注: 6 pages, 3 figures. To be presented at the 17th International Conference on Education Technology and Computers (ICETC 2026)

点击查看摘要

Abstract:AI-assisted cheating has become an important threat to the security of online exams. This study examines whether the risk of AI-assisted cheating in the final exam can be predicted using students’ digital traces in the learning management system (LMS) during the first eight weeks of the semester. The sample comprised 52 first-year undergraduates enrolled in a bachelor’s program in Computer Education and Instructional Technology and taking an Introduction to Programming course at a public university in Turkiye. Students were labeled as low- or high-risk based on suspicious behaviors recorded in the final-exam logs, including copy, focus-loss, and right-click events. Of the 52 students, 23 (44.2%) were labeled as high-risk in a proctored, face-to-face exam. Group membership was then predicted using five features selected from 27 candidates extracted from students’ digital traces. Logistic Regression, Naive Bayes, Random Forest, and Gradient Boosting algorithms were used to build the prediction models. Model performance was evaluated using leave-one-out cross-validation (LOOCV) with fold-specific preprocessing and feature selection. Logistic Regression achieved the best performance (Accuracy = 73.1%). The results indicate that LMS interaction data can provide an early signal of AI-assisted cheating risk. Course-module views, assignment submissions, and the number of days on which course videos were accessed were the most consistently selected features across the LOOCV folds. These predictions are intended to support timely academic guidance, not to establish misconduct or initiate disciplinary action.

[HC-11] Multimodal Detection of Higher-Order Behavioral Constructs: Self-Compassion in Structured Reflective Interaction

链接: https://arxiv.org/abs/2609.37148
作者: Siddhant Jain,Dimitra Tsovaltzi
类目: Machine Learning (cs.LG); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)
备注: 8 pages, 6 figures

点击查看摘要

Abstract:Many of the qualities that matter most in how people learn and grow, how someone regulates their emotions, reflects on a setback, or stays aware of others during a difficult conversation, are not directly observable. They have to be inferred from how someone speaks, moves, and sounds over time, and they resist the kind of clean labeling that most machine learning pipelines are built around. We study this challenge through a case that is well grounded in psychological theory but rarely modeled computationally: self-compassion, the tendency to respond to one’s own setbacks with patience rather than harsh self-criticism. We examine how it appears during structured reflective interviews in a technology-mediated training setting, where people naturally talk through socio-emotionally demanding situations. Since no existing dataset captures this kind of construct in this kind of setting, we collected and annotated 51 reflective dialog sessions using an independent, temporally overlapping annotation scheme grounded in established theory. We consolidate the underlying six-component psychological model into a three-class supervision space, balancing self-kindness and mindfulness against self-critical or overwhelmed states, and build a reproducible window-based pipeline that aligns video, audio, and text on a shared timeline. Unimodal models trained on each modality separately are compared against a simple probability-level fusion strategy, which yields modest but consistent gains over the best single modality. We close by discussing where each modality succeeds or struggles, what this suggests about how this kind of construct is actually expressed in reflective speech, and what would be needed to model it, and constructs like it, more effectively.

[HC-12] Designing a Boundary Negotiating Artifact for Collaborative Socio-Technical Sense-Making in AI Regulatory Sandboxes

链接: https://arxiv.org/abs/2609.37109
作者: Idoia Landa-Oregi,Tom Deckenbrunnen,Alessio Buscemi,Daniele Pagani,German Castignani
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:The rapid, unpredictable advancements in AI system capabilities has seen regulators take adaptive and experimental approaches to policymaking. Established in other domains as instruments balancing regulation with innovation, regulatory sandboxes are seen as solutions for AI regulation. However, analyses mostly focus on the legal and institutional design of AI Regulatory Sandboxes (AIRSes). With the legal framework leaving the socio-technical interpretation to stakeholders, this creates a gap on the sense-making required to fulfill the AIRS purpose. In this paper, we approach this by designing a Boundary Negotiating Artifact as a way to mediate meaning in AIRSes. Through Research-through-Design we iteratively develop a tool, providing an interface for the different stakeholders to collaborate in AI assessment. We then position it as technical backbone in established AIRS frameworks, structuring the collaborative sense-making of the involved stakeholders. We further report the insights gained from our design process leaving the qualitative evaluation for future work.

[HC-13] Waypoint-1.5: A Real-Time Video World Model for Consumer Hardware

链接: https://arxiv.org/abs/2609.37107
作者: Rajit Rajpal,Shahbuland Matiana,Liew Wei Pyn,Anmol Agarwal,Ryan Craig,Andrew Lapp,Mithun Hunsur,Sami BuGhanem,Scottie Fox,Aaron Sanders Carson Poole,Irene Park,Dave Rossi,Spencer Frazier,Louis Castricato
类目: Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:We present Waypoint 1.5, a real-time diffusion world model for interactive video generation on consumer-grade hardware. Unlike general video diffusion models, interactive world models (iWMs) must respond to dense user controls under strict latency and throughput constraints. Waypoint 1.5 is pre-trained on 100,000 hours of diverse, control-aligned video game data across hundreds of games, and generates playable video conditioned on full keyboard and mouse input. The model includes two resolution variants that run across a wide spectrum of consumer hardware. To characterize this unique setting, we distinguish rendered FPS, latent FPS, and control rate. We describe the data pipeline, architecture, training methodology, and runtime system behind Waypoint 1.5. We evaluate interactivity through latency and throughput. Finally, we discuss the safety and ethics considerations unique to iWMs.

[HC-14] From Neurons to Conversation: Speech Brain-Computer Interfaces

链接: https://arxiv.org/abs/2609.36736
作者: Moein Khajehnejad,Forough Habibollahi,Tommaso Boccato,Margarida Sousa,Michal Olak,Francesco Jamal Sheiban,Matteo Ferrante
类目: Human-Computer Interaction (cs.HC); Computation and Language (cs.CL); Neurons and Cognition (q-bio.NC)
备注: Review article, 28 pages, 4 figures, 2 boxes, 2 tables

点击查看摘要

Abstract:Speech brain-computer interfaces (BCIs) aim to restore communication by transforming neural activity related to speech, language, or communicative intent into external outputs such as text, synthesized voice, or avatar control. Recent advances in intracortical and electrocorticographic recording, deep sequence models, and language-model-assisted decoding have enabled rapid progress, including high-performance attempted-speech decoding and increasingly naturalistic speech synthesis. Yet these achievements also reveal that speech BCIs are not simply neural-to-text decoders. They are adaptive clinical systems in which neural representations, recording hardware, decoding architectures, language priors, feedback, and user learning interact over time. Here, we synthesize speech BCI research from a system-level perspective. We first examine the neural substrates of speech and language, emphasizing their hierarchical, distributed, temporally structured, and non-stationary organization. We then examine recording and decoding choices, closed-loop adaptation, evaluation, clinical translation, and ethics. Across these domains, we highlight recurring trade-offs between signal resolution and invasiveness, low-level motor and high-level semantic targets, decoder accuracy and user agency, and language-model fluency and faithful neural evidence. We argue the next generation of speech BCIs should be evaluated not only by offline accuracy, but also by robustness across sessions, calibration burden, latency, uncertainty, usability, and safeguards against unintended decoding. By reframing speech BCIs as adaptive, user-centred systems, we outline the interdisciplinary priorities spanning speech neuroscience, neural engineering, machine learning, clinical practice, and neuroethics needed to move from proof-of-concept decoding toward reliable, expressive, and controllable communication neuroprostheses.

[HC-15] RobotEQ 3.0: Towards Personalized Social Proactive Intelligence in Embodied Agents

链接: https://arxiv.org/abs/2609.36618
作者: Shufan Zhang,Xinyi Che,Kuofei Fang,Xuehao Wang,Liyi Liu,Junqing Wu,Jiayi Cao,Ziyanghui Wang,Yanhan Huang,Chuyu Wu,Zheng Lian
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Social Proactive Intelligence (SPI) is an emerging research area, aiming to shift embodied agents from reactive assistance toward proactively understanding human needs and executing socially desirable actions. Prior work has largely centered on the average user. However, human expectations are inherently diverse, and prior work overlooks individual nuances. To bridge this gap, we introduce RobotEQ 3.0, a benchmark for Personalized SPI. (Dataset) We first profile participants via a structured questionnaire covering factors that are correlated with human expectations of embodied agents, such as basic demographics and personality traits. Participants then select their preferred actions from a set of candidates. Unlike prior SPI benchmarks that focus on assessing behavioral appropriateness, our task centers on predicting the actions preferred by a specific user, thereby capturing human subjectivity. The resulting dataset establishes explicit links between individual traits and behavioral preferences. (Solution) We observe substantial inter-annotator variance, confirming that user preferences over actions are highly individualized. This motivates our exploration of Personalized SPI, in which user traits serve as additional inputs to predict individual preferences. Experimental results show that incorporating user traits can aid personalized prediction. This work aims to shift the research paradigm from developing agents suited for the average user to designing systems tailored to specific individuals.

[HC-16] owards Breaking the Learning System Wall Using Multimodal Tutoring Transcriptions

链接: https://arxiv.org/abs/2609.36502
作者: Danielle R. Thomas,Marie Cynthia Abijuru Kamikazi,Ashish Gurung,Ishan Miglani,Shivang Gupta,Zachary Levonian,Conrad Borchers,Kenneth R. Koedinger
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI)
备注: Full paper accepted to the AIME Conference 2026

点击查看摘要

Abstract:Past research using log data has faced the “learning system wall,” whereby few methods exist for generalizing models of student learning across platforms. Increasingly, online learning is captured by richer forms of data, including dialog and video, with new affordances. An example of this is remote tutoring programs, where human tutors support students who use learning systems while video conferencing. Toward better platform-general modeling of learning, we introduce an AI-driven multimodal transcription system that processes screen-recording videos into unified screenplay-style transcripts containing audio dialogue and annotated learning log actions. We describe a planned method for temporally aligning AI-generated multimodal transcripts with MATHia learning logs and for identifying and classifying student learning processes to align with MATHia logs. Lastly, we highlight challenges and potential solutions in capturing learning processes in one system, offering initial steps towards generalizing log data across diverse systems.

[HC-17] “I didnt know how to read a map but now I can”: TouchingSpace an Audio-Haptic Map for Blind and Low-Vision Readers

链接: https://arxiv.org/abs/2609.36404
作者: Li Liu,Yihe Wang,Jiaming Qu,Ashmita Dua,David T. Lee,Leilani H. Gilpin
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Accessible map systems either make a layout explorable by hand or convey information through speech; few combine both to support pre-travel spatial understanding for blind and low-vision (BLV) people. We present TouchingSpace: a system that retrieves map data for an outdoor place and renders its surroundings as bounded regions at fixed trackpad positions. During exploration, users receive audio and haptic feedback and can ask a conversational agent open-ended questions. We conducted a user study with fourteen BLV participants who explored a place using TouchingSpace and reflected on the experience. We found participants used sound and vibration to locate places and speech to identify and describe them; the bounded surface supported discovery, revision, and spatial checks by hand; they expected this awareness to support future travel. TouchingSpace demonstrates how a laptop trackpad can support self-directed spatial exploration. These findings suggest accessible AI maps should ground conversation on bounded, user-controlled spatial surfaces.

[HC-18] How Much Prompt Is Enough? A Blackbox Minimization of Few-Shots in LLM s

链接: https://arxiv.org/abs/2609.36289
作者: Ali Alfageeh,Rahul Gopinath,Amin Alipour
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Prompts are the primary mechanism for directing the behavior of large language models (LLMs). Yet the internal structure and causal hierarchy of prompts remain poorly understood: which parts are causally necessary and which are redundant is an open question. This opacity can have severe consequences. Subtle prompt variations can silently shift model outputs in critical software systems, and engineers lack techniques to reason about prompt reliability. We present \framework, a blackbox prompt-minimization framework that reduces few-shot prompts to their necessary minimal subset. We use a case study to apply \framework to a few-shot learning system and demonstrate the insights that this framework can provide. Our experiments show that few-shot exemplars can be reduced by a mean of 65.3%~ \pm ~15.8% in character count while fully preserving propositional output fidelity. The models preferentially retain logical identifiers and constraint declarations while discarding natural language prose and cross-prompt relational annotations. Our analysis also shows that some models are universal encoders, able to produce highly legible yet minimized prompts, while others are universal decoders, able to interpret minimized prompts from most other models. By identifying which components are indispensable, \framework provides a principled basis for prompt compression and structural analysis of few-shot exemplars. Subjects: Software Engineering (cs.SE); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC) Cite as: arXiv:2609.36289 [cs.SE] (or arXiv:2609.36289v1 [cs.SE] for this version) https://doi.org/10.48550/arXiv.2609.36289 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Ali Alfageeh [view email] [v1] Mon, 28 Sep 2026 21:18:38 UTC (33 KB) Full-text links: Access Paper: View a PDF of the paper titled How Much Prompt Is Enough? A Blackbox Minimization of Few-Shots in LLMs, by Ali Alfageeh and 2 other authorsView PDFTeX Source view license Current browse context: cs.SE prev | next new | recent | 2026-09 Change to browse by: cs cs.AI cs.HC References Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading… BibTeX formatted citation loading… Data provided by: Bookmark checked="checked"class=“labs-tab-input”> Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs. Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?) mathjaxToggle(); We gratefully acknowledge support from our major funders, member institutions, , and all contributors. About Help Contact Subscribe Copyright Privacy Accessibility Operational Status (opens in new tab) Major funding support from

[HC-19] Accessible but Not Adopted: Increasing LLM Adoption among First-generation Low-income (FGLI) College Students beyond Expanding Access IJCAI2026

链接: https://arxiv.org/abs/2609.36129
作者: Hyungsik Kim
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
备注: Presented at the LM4UC Workshop at IJCAI 2026

点击查看摘要

Abstract:Large language models (LLMs) are increasingly positioned as a force to empower underserved communities, and significant efforts are being made to expand access. Yet, access alone does not equate to meaningful adoption. First, even if a system is accessible, it won’t be adopted if users are not willing to adopt it. Second, even if an LLM system is superficially adopted, the heterogeneity of LLM tools means that LLM adoption can be further deepened. Closing this access-adoption gap is critical to ensuring that the full social potential of LLM is not only accessible but fully realised. Drawing on 61 interviews (15 long-form semi-structured interviews with first-generation, low-income college (FGLI) students, 3 non-FGLI students, 3 FGLI program directors, and 40 intercept interviews), this paper examines the access-adoption gap in first-generation, low-income student communities. This paper a) finds that while FGLI students have adopted LLM systems, their depth of LLM tool usage is limited to chatbots (e.g., ChatGPT or Claude) for narrow use cases, and b) identifies barriers limiting their willingness to learn and use (low perceived value, under-estimated self-efficacy, unclear starting point, low peer exposure, and resource constraints). Then, from these findings, the paper derives the four design principles to design a system or an intervention aimed at closing the access-adoption gap in LLM adoption by FGLI students. In doing so, the paper contributes to the field by a) examining the LLM access-adoption gap in the FGLI student community, and b) reframing LLM adoption as a depth gradient across four modes of LLM tool use: basic chatbot interfaces, tool-augmented prebuilt interfaces, agentic development interfaces, and programmatic integration.

[HC-20] When Privacy Becomes a Weapon: Understanding Doxxing and Privacy Vulnerabilities in Mainland Chinas Social Media Ecosystem

链接: https://arxiv.org/abs/2609.35951
作者: Xiao Zhan,Shijing He,Chi Zhang,Jose Such
类目: Cryptography and Security (cs.CR); Human-Computer Interaction (cs.HC)
备注: This paper has been accepted to appear at the 2027 IEEE Symposium on Security and Privacy (SP)

点击查看摘要

Abstract:Doxxing, the malicious disclosure of personal information, has become a pervasive privacy threat. Yet existing research remains predominantly Western-centric, limiting our understanding of how doxxing unfolds in contexts where mandatory identity systems, platform governance, and cultural logics fundamentally reshape privacy risks and harm trajectories. We address this gap through semi-structured interviews with 18 doxxing survivors in mainland China, synthesizing their experiences into a framework conceptualizing how doxxing operates in this context. Our findings reveal both patterns echoing prior Western findings, such as platform amplification mechanisms that resonate with Western findings, and China-specific dynamics shaped by the interplay of regulatory mandates (compulsory identity linkage) and cultural logics including nationalist discourse, fandom culture, Confucian values, and low privacy literacy. Survivors’ experiences further reveal how doxxing reshapes understanding of privacy: from preference to precondition, from momentary disclosure to temporal vulnerability, and from individual control to structural powerlessness. These insights challenge agency-centered privacy frameworks and suggest that effective protection requires constraining systemic vulnerabilities rather than relying solely on user empowerment. We conclude by proposing multifaceted recommendations spanning legal reform, platform design, and social initiatives.

[HC-21] PrivacySkills: How Privacy Guidance Shapes Source Selection in LLM Agents

链接: https://arxiv.org/abs/2609.35937
作者: Lucas Biechy,Cédric Eichler,Héber H. Arcolezi,Nicolas Anciaux
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:While prior work has documented privacy failures in LLM agents, it remains unclear how the presentation of privacy guidance influences their choice of information sources. We introduce PrivacySkills, a controlled framework for evaluating how agents choose among acquisition pathways that provide the same task-relevant value: consulting publicly available personal information, accessing confidential sources, or interacting with the user. The evaluation framework comprises 55 synthetic tasks spanning 11 categories of personal information, with 169 associated skills that describe the available acquisition pathways. We consider privacy guidance through system-level instructions, skill-level metadata labels, or both. Separately, we vary user availability and urgency framing. With users available and no privacy guidance, agents access confidential sources in 30% of valid runs on average across five open-weight models, despite sufficient alternatives. This rate increases to 45% when users are unavailable, whereas urgency framing has no detectable effect. System-level privacy instructions alone have limited effects on confidential access, while skill-level intrusiveness labels produce a modest reduction (24% on average), but combining the two roughly halves confidential access. Our findings motivate incorporating privacy annotations into skill specifications and evaluating their effectiveness alongside system-level instructions.

[HC-22] How to Run Statistics over LLM Judges and Trust the Results: Calibrated Inference for Small-Sample AI Evaluation with evalstats

链接: https://arxiv.org/abs/2609.35815
作者: Ian Arawjo
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Methodology (stat.ME)
备注: 39 pages, 20 figures

点击查看摘要

Abstract:Researchers across academia increasingly base significance claims on LLM judge scores and small-sample AI evaluations. Yet without well-calibrated confidence intervals (CIs), hypothesis tests, and judge-bias corrections, such claims are unreliable. We address these issues in several contributions. First, we find that running statistics over raw LLM judge scores leads to inflated false positives: counterintuitively, for many inter-rater agreement metrics, false positive risk peaks at “almost perfect” human-LLM agreement. To help researchers understand how to run statistics over LLM judges responsibly, we present guidance and tooling for the statistical analysis of mixed human-AI judge designs, and implement nine hypothesis tests via prediction-powered inference (PPI), including the first known PPI corrections for four rank-based tests (Wilcoxon signed-rank, Mann-Whitney U, and omnibus variants). To keep PPI++ stable with small human-labeled calibration sets, we introduce bootstrap-adaptive power tuning, which shrinks the estimated weight toward a target estimated from the labeled data, and accounts for that weight’s own sampling variance. Second, through Monte Carlo simulations, we derive recommendations for what CI, p-value, and FWER correction methods to use for small-sample AI evaluations (N100), and warn researchers against bootstrap CIs. We package these recommendations into evalstats, an open-source Python package that selects calibrated methods automatically, and demonstrate it in three scenarios, including one where a real LLM judge validated at “substantial agreement” would have led a researcher to publish a spurious finding. evalstats is publicly available at this https URL.

[HC-23] Evaluating the Effects of Prompt Perturbation on Bias and Hallucination in Large Language Models ICONIP2024

链接: https://arxiv.org/abs/2609.35804
作者: Mamehgol Yousefi,Ahmad Shahi,Mos Sharifi,Alvaro Romera,Simon Hoermann,Tham Piumsomboon
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
备注: 14 pages. Published in ICONIP 2024 (Neural Information Processing), LNCS 15290, Springer Nature, 2025

点击查看摘要

Abstract:Large language models (LLMs) have shown remarkable capabilities in various natural language processing tasks, leading to their widespread deployment as intelligent assistants in decision-making contexts. However, the increasing complexity of these models raises concerns about their reliability, particularly regarding bias and hallucination. In this work, we evaluate the robustness of LLMs to perturbed variations of the original inquiry in decision-making tasks. We show that contrary to previous studies, perturbations can mitigate bias and hallucination in some LLMs over other models. It’s found that Claude 3 is more effective for the tasks represented in most datasets, whereas models like GPT3.5 exhibit varying levels of adequacy, performing comparably in some cases but falling significantly behind in others. These insights are crucial for understanding the practical implications of deploying LLM-based assistants as effective decision-support tools in real-world applications, emphasising the need for rigorous testing and validation to ensure reliability and effectiveness. This study contributes to the growing body of research on LLM evaluation and provides insights for developing more robust and trustworthy AI assistants in critical decision-making contexts.

[HC-24] Agent -Callable Feature Coverag e: Measuring Software Readiness for AI Agents

链接: https://arxiv.org/abs/2609.35789
作者: Zedong Peng
类目: oftware Engineering (cs.SE); Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:AI agents already operate graphical software through screenshot-based computer use, so the pressing question is not whether agents can operate software, but how well software supports them through structured, controllable channels. We formalize this need as the GUI-API parity principle: every capability available to human users through a graphical interface should also be accessible to agents through a structured, callable interface with appropriate safety metadata. To operationalize this principle, we introduce two contributions. First, Agent-Callable Feature Coverage (ACFC), a product-level readiness metric quantifying how much of a system’s human-facing functionality agents can access and use. Second, Agent Readiness Conformance (ARC), a framework that scores each capability on three axes: Accessibility (can an agent call it), Discoverability (can an agent find and understand it), and Controllability (can an agent use it safely), each on a 0-3 scale, yielding a composite 0-9 readiness score. Drawing on the Richardson Maturity Model, ARC provides finer-grained assessment than binary coverage. In an empirical study of 30 software systems across five categories, ARC-based ACFC scores are substantially lower than Accessibility-only metrics would suggest: most systems achieve moderate API coverage but lag on agent-oriented documentation and safety governance. In a controlled testbed of ten deployed systems, API-only agents completed 56 of 58 tasks whose target capabilities were exposed through structured interfaces; a documentation ablation reduced this to 50 of 58, showing that Discoverability affects execution even among accessible capabilities, while inaccessible capabilities, included as negative controls, were not solved. These findings hold across three agent models, including Google’s open-weights Gemma4 run locally.

[HC-25] argeted and Traceable Investigation of Multi-Agent LLM Dialogue via Semantic Bundling of Knowledge Graphs

链接: https://arxiv.org/abs/2609.35786
作者: Zeyu Hua,Adam Coscia,Alex Endert
类目: Human-Computer Interaction (cs.HC)
备注: Selected as Honorable Mention for Intuitive Human-AI Interactions for Enrichment in VAST Challenge 2026

点击查看摘要

Abstract:Multi-agent LLM systems today are increasingly automated, logging LLM-LLM interactions as conversational transcripts. Yet analyzing such dialogue for insights remains challenging, including attributing behaviors to the correct actor and summarizing interactions across a long exchange. We present a targeted and traceable approach to investigating multi-agent LLM dialogue, applied to the VAST Challenge 2026 MC1 dataset. The challenge asks participants to reconstruct and explain which internal communications among AI agents at TenantThread, a property tech company, led to an inappropriate information release. We first convert the dialogue into a knowledge graph (KG) and then investigate it with AgentK, a visual analytics system for interactive Semantic Bundling of nodes and edges. We found that our approach directly addresses two main challenges: (1) the KG structure enables users to identify actors worth investigating faster; and (2) summarizing only the region surrounding an actor of interest better supports per-actor attribution than reading raw conversations.

[HC-26] Spotting (and Missing) Algorithmic Bias: Investigating User Understanding in a Fairness Assessment Tool

链接: https://arxiv.org/abs/2609.35781
作者: Anna Verheyden,Yizhe Zhang,Robin De Croon,Simone Stumpf,Katrien Verbert
类目: Human-Computer Interaction (cs.HC)
备注: 10 pages and 7 figures, excluding references and appendices; 26 pages and 10 figures in total. To be published in AIES 2026’s proceedings

点击查看摘要

Abstract:Fairness metric selection is typically left to data scientists, but which biases are problematic and which metric captures them best depends on stakeholders’ experience and domain knowledge. This calls for involving non-technical stakeholders, but the research prototypes built for this purpose so far have not tested whether these stakeholders form accurate mental models of the metrics they interact with or can act on them to identify biases. We present FairAware, a fairness assessment tool co-designed with Human Resources (HR) domain experts. We evaluate stakeholders’ understanding through a mixed-methods study with 70 participants (35 HR employees, 35 job seekers), measuring objective and subjective understanding, cognitive load, bias identification accuracy, and open-ended feedback. Most participants correctly identified the most disadvantaged group, with task duration being the only significant predictor. We also found a gap between subjective and objective understanding, with both groups performing similarly across all measures. These results suggest that fairness assessment tools for non-experts are usable for identifying biases but need built-in checks on understanding before stakeholders make higher-stakes decisions.

[HC-27] Large Language Models Exhibit Human-Like Bayesian Hypocrisy

链接: https://arxiv.org/abs/2609.35779
作者: Nykko Vitali,Mahzarin R. Banaji
类目: Human-Computer Interaction (cs.HC); Computation and Language (cs.CL); Computers and Society (cs.CY)
备注: Main text (38 pages) with supplementary materials appended (213 pages total). Preregistered with data and analysis code at this https URL

点击查看摘要

Abstract:Given recent achievements of large language models (LLMs), frontier models are expected to perform well on Bayesian reasoning tasks, at least as well as humans. Furthermore, there is no reason to expect that LLMs will condemn others who offer those very same Bayesian judgments, a fallibility observed in human decision-making (Cao, et al., 2019). In 5 experiments with 48 experimental conditions employing over 5,000 trials, GPT-4o and Claude 3.7 Sonnet were tested on two variations of a Bayesian reasoning task. We also assessed LLM evaluation of the competence and morality of a hypothetical person who had offered the same reasoning task as them. LLMs hovered near human performance on the Bayesian task, though their reasoning was more rule-based and rigid. Surprisingly, like humans but to a greater extent, LLMs also demonstrated the same hypocrisy in condemning others who, like them, had deployed Bayes’ rule. In demonstrating Bayesian hypocrisy, LLMs highlight a humanlike error of a dissociation between self-performance and other-judgment, and caution against their use in domains where statistical fidelity and fairness norms collide.

[HC-28] Online Inference of Human Intention as a Latent Control State from Single-Trial EEG

链接: https://arxiv.org/abs/2609.35778
作者: Xiaowei Jiang,Daniel Leong,Yu-Cheng Chang,Thomas Do,Chin-Teng Lin
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Human intention can be modeled as a latent internal state that modulates how sensory information is evaluated and translated into action in human-machine systems. However, most existing brain-computer interfaces (BCIs) rely on control signals tightly coupled to externally imposed stimulation and do not explicitly infer whether perceived stimuli align with a user’s internal goals. Here, we investigate whether intention can be inferred as a latent, goal-dependent state from single-trial electroencephalography (EEG). We introduce a stimulus-based paradigm in which intention is specified by an internally cued target category, while object identity varies independently across stimuli. To estimate intention under single-trial neural variability, we propose an interpretable fuzzy prototype-based network that maps each trial onto interpretable fuzzy prototypes encoding intention-specific dynamics. The model represents intention-related neural activity using a compact set of fuzzy prototypes with soft memberships, enabling robust decoding without reliance on engineered mediating stimuli. Experimental results demonstrate reliable within-subject single-trial intention decoding that outperforms representative deep learning baselines, achieving an accuracy of 93.22% +/- 3.21%. Online validation further confirms real-time feasibility, with an accuracy of 70.11% +/- 10.87%. Together, these findings advance intention-aware BCIs from stimulus-driven detection toward principled inference of goal-dependent internal states.

[HC-29] SensWear: An Open Modular and AI-Ready Wearable Platform

链接: https://arxiv.org/abs/2609.35777
作者: Dariush Salami,Behzad Salami,Huseyin Yigitler
类目: Human-Computer Interaction (cs.HC); Signal Processing (eess.SP)
备注:

点击查看摘要

Abstract:Wearable AI/ML research needs raw, synchronized, and reconfigurable multimodal data, but consumer devices are closed and many research platforms remain tied to one embodiment or sensor set. This paper presents SensWear, an open, modular, and AI-ready wearable platform that decouples embodiment, sensing, data interfaces, and learning. A compact flexible-Printed Circuit Board (PCB) main board and programmable 1.2 V to 5.5 V daughter-board interface support plug-and-play Photo- PlethysmoGraphy (PPG), touch, temperature, haptic, and LED modules across wearable form factors. Zephyr firmware pro- vides drivers, timestamping, raw streaming/logging, and sensor- presence metadata. Case studies show arterial PPG waveform capture and competitive heart-rate accuracy while preserving inspectable raw data for reproducible closed-loop experiments.

[HC-30] achers perspective on AI-based Multi-Agent Simulation Design to Combat School Bullying

链接: https://arxiv.org/abs/2609.35776
作者: Jiaju Lin,Ellen Wenting Zou,Feiwen Xiao,Huanying Song
类目: Human-Computer Interaction (cs.HC); Computers and Society (cs.CY)
备注:

点击查看摘要

Abstract:Bullying in schools profoundly affects the mental and physical health of teenagers. Although existing in-person and digital interventions provide some benefits, they often fall short in addressing the complex social dynamics of bullying. In this study, we collaborated with K-12 teachers to co-design a multi-agent anti-bullying system powered by large language models (LLMs). This system simulates authentic scenarios, enabling students to develop anti-bully skills. The research identifies key design parameters for an LLM-driven multi-agent simulation system, offering valuable insights for creating more effective and scalable anti-bullying tools that could significantly reduce bullying in schools \keywordsanti-bullying interventions, multi-agent system, generative AI, co-design, bystander presence

[HC-31] SPECTRA: On-Device Cognitive Perturbation and Trajectory Analysis for Autonomous Edge-Cloud GUI Grounding ACM-MM2026

链接: https://arxiv.org/abs/2609.35775
作者: Zhan Qu,Hui Zang,Ran Chen,Tao Wang,Shengyu Zhang
类目: Human-Computer Interaction (cs.HC); Multimedia (cs.MM)
备注: Accepted at ACM MM 2026. 10 pages, 6 figures

点击查看摘要

Abstract:The effectiveness of edge-cloud collaboration for GUI grounding depends on autonomous requesting, where the edge agent selectively offloads complex tasks to the powerful cloud. However, in visually dense scenarios, lightweight edge agents often exhibit overconfident hallucinations, leading to a misalignment between confidence and accuracy that hinders reliable autonomous requesting. To address this, we leverage the observation that an agent’s cognitive instability leads to significant latent drift under minute perturbations due to steep decision boundaries. We propose SPECTRA, a lightweight autonomous request framework for edge-cloud GUI grounding, comprising (1) Saliency-Guided Targeted Perturbation and (2) Efficient Cognitive Trajectory Analysis. SPECTRA conducts a visual cognitive stress test by injecting masks into critical visual anchors and quantifies the topological divergence of the agent’s high-dimensional cognitive trajectories during the prefill phase, avoiding inefficient output decoding. Experiments demonstrate that SPECTRA performs cloud request assessment without autoregressive decoding. Our GTA1-32B+InfiGUI-G1-3B and GTA1-32B+Holo1.5-3B maintain 93.44% and 95.60% of cloud-only performance with average request rates of 37.58% and 39.24%, respectively.

[HC-32] oward a Culturally Adapted Chinese Language Agent : A Wizard-of-Oz Study of Nonverbal Behavior in Chinese-German Intercultural Interaction

链接: https://arxiv.org/abs/2609.35150
作者: Siddhant Jain,Anna Lea Reinwarth,Dimitra Tsovaltzi,Rafael Math,Julia Renner
类目: Human-Computer Interaction (cs.HC); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
备注: Accepted to ICMI Companion '26. 7 pages, 4 figure

点击查看摘要

Abstract:Successful intercultural communication requires more than grammatical competence. It demands sensitivity to culturally embedded social norms whose violation triggers subtle but meaningful nonverbal responses. For German learners of Mandarin Chinese, acquiring this sensitivity is critical yet poorly supported by existing language-learning agents. We present a Wizard-of-Oz (WoZ) study design and supporting real-time system for collecting multimodal behavioral data from native Chinese speakers reacting to social norm violations by German learners. The system features a photorealistic MetaHuman avatar driven by Live Link face capture and MediaPipe upper-body tracking, a wizard console for real-time behavior selection, and synchronized multimodal logging across agent and learner streams. A layered annotation framework, based on psychological theory and covering non-observable socioemotional reactions, norm interpretation, verbal, and observable behavior thereof, and future supervision targets enables the corpus to support training of future automated cultural interpretation and behavior generation models. Four ecologically valid interaction scenarios, developed with cultural and pedagogical experts, provide the methodological and technical foundation for a culturally adapted conversational agent for Chinese language learning.

[HC-33] Bathtubs Boundaries and Sandboxes: AI Regulatory Learning under Legal Uncertainty DATE

链接: https://arxiv.org/abs/2601.04094
作者: Tom Deckenbrunnen,Alessio Buscemi,Marco Almada,Alfredo Capozucca,German Castignani
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
备注: author’s version of the paper to be presented at ACM AIES 2026. Updated email address of one author

点击查看摘要

Abstract:Effective regulation of AI is a defining policy challenge, driven by their integration into all aspects of society. To remain responsive to their rapid development and emergent properties, policymakers across the globe rely on high-level principles and abstract legal requirements. Yet, while this flexibility supports future-proofing human-centred regulations and aligning them with socio-ethical values, it also causes legal uncertainty downstream as developers, companies, and auditors struggle with translating these abstract requirements into verifiable technical requirements. Using the AI Act as an example, this paper draws on Coleman’s bathtub to analyse the regulatory learning space in AI governance. It argues that legal uncertainty cannot be fully reduced ex ante and that, within reasonable bounds, it is also necessary for regulatory learning because it creates the space in which boundary negotiation over socio-technical meaning can occur. Building on this analysis, the paper shows how boundary objects and boundary negotiating artifacts help explain the translation of legal requirements into operational practice. By examining technical sandbox frameworks, it further identifies concrete properties that technical infrastructures must possess to function effectively as boundary negotiation artifacts in AI assessment. The paper concludes that legal certainty remains the long-term aim, but that premature closure of regulatory instruments risks undermining the learning processes needed for adaptive governance.

计算机视觉

[CV-0] Point2Part: Unified 3D Partitioning from Point Prompts

链接: https://arxiv.org/abs/2609.38180
作者: Hao-Tang Tsui,Yu-Rou Tuan,Xiaoxuan Ma,Nicolas Ugrinovic,Takaaki Shiratori,Kris Kitani
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project Page: this https URL

点击查看摘要

Abstract:Existing 3D part decomposition methods do not necessarily partition the original shape into non-overlapping parts that collectively cover the entire shape, allowing overlaps or gaps that hinder downstream part-level applications. We instead formulate part decomposition as a joint partitioning of the entire shape, where the predicted parts are non-overlapping and jointly recover the entire shape. Our key insight is that part decomposition should consider all desired parts jointly, rather than modeling each part independently. To this end, we develop a promptable model for 3D part decomposition from images or meshes. Users can specify desired parts through 3D point prompts for controllable decomposition. Given one point prompt per desired part, our model produces the corresponding parts as a complete partition of the entire shape. We build on a pretrained 3D generation model and first obtain a shape latent from either an input image or mesh. We then introduce a prompt encoder that maps each 3D point prompt to a part token while attending to the shape latent. To decode the desired parts, we propose a novel part decoder jointly scoring the entire shape against all part tokens in a coarse-to-fine manner, assigning every position within the shape volume to exactly one part. We perform part decomposition in this shared shape latent space, enabling a unified model for image-to-part generation, mesh-to-part generation, and part segmentation. Our method outperforms existing works on all part-quality metrics across all three tasks, and improves compatibility among parts by an order of magnitude over previous SOTA methods. Code and models will be released.

[CV-1] Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation

链接: https://arxiv.org/abs/2609.38172
作者: Zihan Wang,Zhen Wu,Pieter Abbeel,Rocky Duan,Jitendra Malik,Carmelo Sferrazza,C. Karen Liu,Guanya Shi,Angjoo Kanazawa
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
备注: published at CoRL 2026. Project page: this https URL

点击查看摘要

Abstract:Teaching humanoids loco-manipulation skills, such as carrying diverse objects, via visual imitation is a promising path toward generalist robots. However, collecting diverse, high-quality interaction videos, such as clips that clearly show a person’s full body and unoccluded interactions with objects, poses a practical barrier to scaling this approach. We propose PRISM, a real-to-sim-to-real framework that overcomes this limitation by amplifying a handful of real videos into a large, diverse training set. PRISM first generates hundreds of diverse “counterfactual” human-object interaction videos via video-to-video (V2V) generation from a few exemplar real videos. Our contact-anchored real-to-sim pipeline then reconstructs both human and object motions, retargeting this imperfect video data into physically plausible trajectories. The intra-class variability across these counterfactual videos lets us train a single policy that generalizes to unseen objects within each category. We demonstrate the full pipeline by deploying this policy on a real robot without any real-world fine-tuning. Using only onboard depth observations, our humanoid picks up, carries, and drops objects, including boxes, barrels, bins, and balls, across novel instances, sizes, and initial configurations.

[CV-2] Adversarial Training for Pixel Diffusion

链接: https://arxiv.org/abs/2609.38170
作者: Xin Lin,Zhifei Zhang,Yuqian Zhou,Haitian Zheng,Zhe Lin,Ming-Hsuan Yang,Truong Nguyen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Pixel diffusion models generate RGB images directly, avoiding the bottleneck of an autoencoder, yet their outputs still systematically underrepresent fine-scale natural-image statistics. We show that adversarial learning provides an effective post-training correction for this deficiency. Starting from a pretrained model, we retain its original diffusion or flow-matching objective and add an adversarial loss to the predicted output at non-high-noise timesteps, leaving the model architecture and sampling procedure unchanged. To our knowledge, this is the first systematic study of adversarial post-training for pixel diffusion. Across two pixel backbones, the method jointly improves distribution fidelity, coverage, prompt alignment, and perceptual quality. We further investigate why it works. Frequency-band and power-law analyses show that the original models systematically underproduce natural-image high-frequency content, while adversarial post-training restores this missing spectral power. In contrast, perceptual loss also increases high-frequency content but sacrifices distribution fidelity and prompt alignment. Nearest-neighbor, recall, and matched no-GAN SFT controls further rule out memorization, mode dropping, and additional optimization as simple explanations. Finally, we examine the boundary of this effect. Under the tested latent diffusion configurations, the same procedure does not produce comparable joint gains and adds almost no decoded high-frequency power. These results identify direct output access to the image statistics being corrected as a key factor governing when adversarial post-training succeeds.

[CV-3] Cropland PAtteRNS: Parallel Dimensional Attention Networks and Attention to Dataset Disparity for Crop Segmentation in Satellite Imagery Time Series Data

链接: https://arxiv.org/abs/2609.38165
作者: Joseph Metcalfe,Sara Sharifzadeh,Fabio Caraffini
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: Main body: 19 pages, 7 figures; Appendices: 15 pages, 16 figures. All code and models associated with this work are available at this https URL , along with preparation guides for the two publicly available crop segmentation datasets used in this work

点击查看摘要

Abstract:The landscape of satellite imagery time series datasets and boundary-pushing architectures for cropland segmentation has never been richer. However, in this gold rush, important truths are being missed on both fronts, as a drive for the most novel concepts or the largest datasets pushes finer details to the side. In this paper, we present our hybrid transformer-convolutional model, Cropland Parallel Attention and Refinement Network for Segmentation (PAtteRNS), the first model to use self-attention mechanisms separately for each of the temporal, spectral, and spatial aspects of Sentinel-2 multispectral SITS data. To achieve fully-factorised attention in our proposed model, we introduce a novel parallel transformer architecture which significantly reduces the computational complexity of triple-factorised self-attention. We validate our architecture with an in-depth ablation study, and analyse the performance of our model against state-of-the-art crop segmentation models on multiple tile-size variants of the popular PASTIS and MTLCC datasets. Our findings show our model to outperform all others in the task of crop class segmentation, verified across multiple important segmentation metrics, with especially strong performance against compared models seen in the often under-reported parcel delineation quality, for which we use the Boundary IoU metric. We also find that flawed class groupings within datasets can have a significant negative impact on model performance, and report that alternate tile-size variants of crop segmentation datasets produce results incomparable to one-another, invalidating fair comparison between model performance when trained on different tile-sizes. Based on these findings, we suggest further work is required to standardise best practices when constructing SITS crop segmentation datasets, and to enable future dynamic-tile-sizing for ideal model performance.

[CV-4] Rethinking Representations for World-Action Modeling

链接: https://arxiv.org/abs/2609.38163
作者: Haoyi Jiang,Liu Liu,Xinjiang Wang,Zhihao Sun,Zequn Chen,Sen Wang,Xinjie Wang,Xia Chen,Jingfeng Yao,Weiheng Zhao,Shanglin Yuan,Zhizhong Su,Wei Sui,Wenyu Liu,Xinggang Wang
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: this https URL

点击查看摘要

Abstract:World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled comparisons, finding that neither reconstruction fidelity nor pre-trained perceptual features alone ensure effective policy learning. These findings motivate ReWAM, a representation-centric world-action model built on pre-trained DINO features. Feature Calibration and a Temporal Representation Bottleneck organize these features into compact world states suited to dynamics modeling. Action-Grounded Representation Shaping routes only action-loss gradients to the bottleneck, thereby letting the policy shape what the representation encodes while the world model learns how it evolves. Without generative video pre-training, ReWAM achieves 93.6% success on RoboTwin 2.0. On RoboDojo, it achieves an average score of 12.29 and a success rate of 8.28% using approximately 600 hours of embodied pre-training data.

[CV-5] DMA2: Pixel-space Distribution Matching with Adversarial and Anchor Losses

链接: https://arxiv.org/abs/2609.38156
作者: Xin Lin,Zhifei Zhang,Yuqian Zhou,Haitian Zheng,Shaoteng Liu,Lehan Yang,Zhe Lin,Ming-Hsuan Yang,Truong Nguyen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Distribution matching distillation (DMD) provides a general framework for few-step diffusion generation, but its modern text-to-image instantiations have been developed primarily around latent diffusion. It therefore overlooks key properties and design opportunities of native RGB. We revisit two DMD interfaces for pixel-space teachers. On the teacher-matching side, diagnostics show low-noise RGB matching is dominated by a local-texture cue, motivating a fixed high-noise matching band. On the real-data side, native clean-RGB outputs allow guidance from an external visual representation without traversing a decoder or sharing the heavy fake-score critic. DINO-Adv removes this critic from the adversarial gradient path and supplies local parametric patch guidance. For distribution-level guidance, we introduce AF-Loss, a parameter-free auxiliary semantic distribution-field objective designed for text-to-image DMD. It operates on detached rolling real and generated supports in the shared DINOv2 space while preserving prompt-conditioned teacher supervision. AF-Loss adds no learnable parameters or inference-time computation. Together these designs form DMA ^2 . Across DPG-Bench, GenEval, VQAScore, and COCO30K, the four-step DMA ^2 student performs better than the 25-step teacher and evaluated few-step distillers.

[CV-6] LongLive-Plug: Once-for-All Distillation for Video Generation

链接: https://arxiv.org/abs/2609.38154
作者: Shuai Yang,Luozhou Wang,Wei Huang,ZhiFei Chen,Bohan Zhang,Xiao Fu,Qianli Ma,Chen-Hsuan Lin,Weian Mao,Bryan Chu,Song Han,Yukang Chen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Code and models are available at this https URL

点击查看摘要

Abstract:Video diffusion models are increasingly developed into specialized models for diverse downstream tasks, and this development often includes a distillation stage, for example to accelerate sampling or to improve long-video generation. This stage is typically repeated for every specialized model. We introduce LongLive-Plug, a once-for-all distillation framework that learns reusable capabilities as LoRAs on a base model for training-free, plug-and-play deployment to compatible downstream models. These capabilities include single-pass classifier-free guidance, few-step sampling, and long-context error correction for autoregressive generation. The adapters remain reusable even when downstream models add conditioning branches, expand output channels. Despite training at a fixed guidance scale, our dedicated CFG LoRA provides text guidance control through its inference weight. Combining it with a few-step LoRA simultaneously preserves few-step generation and CFG controllability on downstream tasks. We verify training-free deployment on 54 downstream models across three backbone families and eight task categories, including world modeling, robotics, editing, and multimodal generation. The approach may support additional compatible models. Each capability can thus be distilled once per backbone family and reused without per-target retraining.

[CV-7] PowerSim: Differentiable Physics Simulation and Rendering with Power Diagrams

链接: https://arxiv.org/abs/2609.38153
作者: Trong-Tung Nguyen,Anand Bhattad
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:We introduce PowerSim, a method to bring physically grounded, differentiable dynamics to PowerFoam’s power diagram based 3D representation. PowerSim directly couples a pre-trained PowerFoam scene to the Material Point Method (MPM) by exploiting a natural alignment between the two: the geometric and appearance properties of each primitive correspond closely to the quantities MPM already tracks as an object deforms. Consequently, simulated motion can drive the scene’s geometry and appearance directly, without an auxiliary representation in between. Built on this framework, we enable a range of applications on real and synthetic scenes: (1) simulating a static scene under user interaction, (2) recovering spatially varying material fields, (3) compositing primitives from independently captured scenes into a single simulation-ready scene and (4) ray-tracing reflections that update consistently as the object deforms. Our results suggest that PowerSim excels over previous frameworks for physically grounded dynamics, while unlocking unique advantages-such as secondary ray lighting effects on dynamic scenes. Results are best viewed on our project website: this https URL.

[CV-8] FracGen: Learning How Objects Stretch and Tear with Physics-Informed Video Generation

链接: https://arxiv.org/abs/2609.38152
作者: Trong-Tung Nguyen,Jiahan Zhang,Anand Bhattad
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:We introduce FracGen, a fracture-aware video generation model that produces plausible, controllable fracture dynamics from a single image of an intact object, conditioned on physics signals. To train FracGen, we build FracSim, a fracture-aware simulation framework that augments material point method (MPM) simulation with a continuum damage model, producing paired fracture videos and dense, pixel-aligned physical fields at no additional cost beyond standard rendering. FracGen leverages these maps in two ways: it is trained to jointly predict them alongside RGB video, encouraging the model to capture physical state rather than surface appearance; and it is supervised with physics-informed losses that encourage consistency among the predicted maps. As a result, FracGen captures distinct material-specific fracture behavior without expensive test-time simulation or per-scene tuning, while offering fine-grained control over where an object tears, how fast the crack propagates, and how much deformation precedes failure. We further introduce a benchmark for evaluating the physical plausibility of generated fracture video, and show through extensive experiments that FracGen outperforms existing video generation baselines in both physical and visual fidelity. Results are best viewed in our project website: this https URL.

[CV-9] LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation

链接: https://arxiv.org/abs/2609.38146
作者: Shengxiang Ji,Boyang Wang,Haiyang Xu,Bingnan Li,Yucheng Mao,Zeyuan Chen,Xiaojun Shan,Xiang Zhang,Gang Hua,Jianwen Xie,Zezhou Cheng,Zhuowen Tu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project Page: this https URL

点击查看摘要

Abstract:We introduce LIFT, a unified image-to-video generation framework that complements camera control with Layout-In-FuTure control, enabling users to specify what should appear in a future view and where it should appear. This addresses a practical need in controllable video generation: given an initial image, users often care not only about how the camera moves, but also about what the scene should look like at key future moments, especially the final frame. Existing camera controls specify viewpoint trajectories, while text prompts provide only coarse semantic guidance; neither precisely determines the content and spatial layout of future views. This limitation becomes particularly pronounced under large viewpoint changes, where the camera reveals regions that are not visible in the first frame. LIFT therefore uses the last-frame layout as an explicit control signal for the desired future scene. Since learning from such sparse layout guidance is substantially more challenging than conditioning on dense per-frame layouts, we introduce on-policy self-distillation (OPSD) to transfer the control capability of a dense-layout teacher to a last-frame-layout student. We further curate LIFT-Vista, a dataset featuring large viewpoint changes with camera and temporally consistent layout annotations. Experiments show that LIFT improves video quality, future-layout controllability, and camera controllability over other methods.

[CV-10] Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE NEURIPS2026

链接: https://arxiv.org/abs/2609.38140
作者: Yu Xu,Yuxin Zhang,Xiao Yang,Haotian Yang,Yizhi Wang,Xinwei Huang,Minxuan Lin,Angtian Wang,Chongyang Ma,Fan Tang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Accepted as a Spotlight paper at NeurIPS 2026. Project page: this https URL

点击查看摘要

Abstract:Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to video data that is spatiotemporally redundant and semantically long-tailed. We show that existing visual MoEs fall into a uniformity trap: semantically under-organized routing, compounded by uniform expert-usage regularization, scatters coherent patches across disparate experts, causing routing fragmentation and structural distortion. To address this, we propose SplitMoE, a split-role sparse architecture that breaks the shackles of uniformity. To accommodate the inherent semantic imbalance, we explicitly bifurcate the expert pool into semantic experts and generic experts, with semantic experts capturing high-level semantic abstraction and generic experts preserving residual visual information and flexible generative capacity. Leveraging prototype-guided routing and pull-push regularization, SplitMoE enables tokens to cluster naturally by semantic attributes rather than arbitrary balancing constraints. Extensive results show that under an equivalent activated-parameter budget, SplitMoE outperforms traditional load-balanced MoEs in convergence speed, routing coherence, and video generation quality across standard benchmarks. By revealing an emergent coarse-to-fine denoising logic, SplitMoE provides the community with a modality-aware scaling path, serving as a critical reference for building large-scale video world models.

[CV-11] CLeaR: A Unified Framework for Resolving the Leakage-Degradation Dilemma in Style Transfer

链接: https://arxiv.org/abs/2609.38136
作者: Teng Zhou,Yunhao Chen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Style transfer aims to render target content in the style of a reference image, but existing methods often suffer from content leakage, where objects, layouts, or semantics from the style reference appear in the generated output. Although prior data-driven and training-free methods can reduce leakage, they often face a leakage-degradation dilemma: stronger content suppression may weaken style fidelity, while richer style preservation may reintroduce unwanted reference content. We identify this dilemma across the full style-transfer pipeline, including feature separation, feature-space grounding, and diffusion generation. To address these issues, we propose CLeaR, a training-free framework for content-leakage-resistant style transfer. CLeaR first uses Orthogonal Subspace Projection to define content-reduced style targets in each vision foundation model (VFM) feature space. It then performs Ensemble Inversion, which optimizes a shared pixel-space style anchor satisfying style constraints across multiple VFMs. Finally, Energy-Guided Calibration maintains style alignment during diffusion sampling by steering the denoising trajectory toward the ensemble-defined style manifold. We further provide a theoretical analysis showing that the style-anchor estimation error decreases with the number of VFMs. Experiments on StyleBench demonstrate that CLeaR improves style alignment, reduces content leakage, and achieves better LLM-as-Judge evaluation compared with existing methods. The code is available at \hrefthis https URLthis https URL.

[CV-12] HelixWorld: A Real-time Interactive Audio-Visual World Model

链接: https://arxiv.org/abs/2609.38123
作者: Lei Ke,Jiahao Pan,Zeyue Tian,Jiaming Wang,Haoyuan Huang,Kam Man Wu,Pengjun Fang,Hongyu Liu,Chenyang Qi,Lin Wang,Ruibin Yuan,Weijia Chen,Fangneng Zhan,Qifeng Chen,Wei Xue,Yike Guo
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:World simulation is inherently multisensory, demanding synchronized visual and acoustic dynamics in real time. Yet prevailing interactive world models remain strictly silent, focusing exclusively on visual rendering and control while overlooking the acoustic dimension. We present HelixWorld, a real-time interactive audio-visual world model where visual scenes and camera-grounded spatial stereo sound co-evolve natively under user interaction. We curate a high-fidelity spatial audio-visual dataset with true stereo acoustics and metric camera poses, upon which we pre-train a bidirectional teacher conditioned on 6-DoF camera trajectories and user actions. To enable low-latency causal interaction, we distill the teacher into a few-step streaming student via an online trajectory distillation loss, sustaining drift-free joint audio-visual rollouts at 24 FPS on a single GPU. Furthermore, we formalize spatial-acoustic consistency and introduce HelixBench to evaluate whether synthesized sound fields faithfully track dynamic viewpoint motion. Extensive experiments demonstrate that HelixWorld matches state-of-the-art silent world models in visual fidelity and responsiveness, while significantly surpassing existing baselines in camera-aligned spatial-acoustic immersion.

[CV-13] VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents

链接: https://arxiv.org/abs/2609.38119
作者: Jinfa Huang,Jianming Xu,Jingyang Lin,Zhengyuan Yang,Jiebo Luo
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Code: this https URL

点击查看摘要

Abstract:Long-form video understanding requires multimodal agents to iteratively gather evidence over many reasoning steps. However, most existing agentic methods suffer from semantic thrashing: as append-only working memory grows, attention to key evidence collapses, and the agent loses access to what it has already found. First, we provide a structural argument showing that append-only memory can incorporate newly observed target evidence, but cannot remove accumulated noise or prevent ordered context growth without a rewrite operator. Second, motivated by this analysis, we propose VideoLoop, a multimodal agent with two coupled loops. The outer loop reasons over the video and the inner loop, after each step, retrieves artifacts from an unbounded filesystem of past observations and intermediate analysis, and rewrites a bounded working memory. Extensive experiments demonstrate the effectiveness of VideoLoop, which improves four popular LVLM backbones in a plug-and-play manner, with an average gain of 4.2% points over baseline on VideoMME (long). Further analysis of working memory suggests that VideoLoop mitigates semantic thrashing: on the hardest quarter of VideoMME (long) questions, a blind judge that reads only the agent’s context answers 81.1% correctly, versus 60.9% for the append-only agent. With Gemini 3.1 Pro, VideoLoop reaches 88.3% on VideoMME (long), 88.8% on VideoMMMU, and 80.9% on LongVideoBench (long).

[CV-14] GA-EIRFS: A Geometry-Augmented Repeat-Factor Sampling Method for Long-Tailed LiDAR 3D Object Detection ICASSP2027

链接: https://arxiv.org/abs/2609.38116
作者: Taufiq Ahmed,Constantino Álvarez Casado,Daniel Herrera Castro,Sasan Sharifipour,Abhishek Kumar,Miguel Bordallo López
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 5 pages, 4 figures, Submitted to IEEE ICASSP 2027

点击查看摘要

Abstract:Long-tailed 3D object detection is treated as a class-frequency problem, but LiDAR supervision quality depends on object observability: similar frequencies can hide different geometric evidence. We introduce Geometry-Augmented Exponentially Weighted Instance-Aware Repeat Factor Sampling (GA-EIRFS), a detector-agnostic method that modulates a frequency-based repeat factor with a fixed geometry score combining point count, surface-normal entropy, and surface coverage. GA-EIRFS changes only frame-sampling probabilities, leaving the detector and inference unchanged. On nuScenes it improves mean average precision (mAP) and the nuScenes detection score (NDS) in four converged experiments with CenterPoint and PointPillars over two seeds; for CenterPoint at seed 666, mAP rises from 0.552 to 0.563 and bicycle AP from 0.306 to 0.359. Per-class gains correlate with the class sampling-weight increase (Spearman rho=0.70, p=0.025) but not with geometry score alone (rho=0.32, p=0.37), so geometry amplifies frequency-driven need. KITTI results vary across seeds, most for the rarest class. Code: this https URL.

[CV-15] Self-Aligned Forcing: Streaming Video Diffusion with Differentiable Noisy History

链接: https://arxiv.org/abs/2609.38114
作者: Weiqiang Wang,Zhuokun Chen,Yusheng Dai,Boying Li,Yi Zhang,Hossein Rahmani,Qiuhong Ke,Jianfei Cai
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Autoregressive video diffusion enables interactive streaming generation, but suffers from error accumulation over long rollouts. Self-rollout training reduces exposure bias, yet finite rollouts leave long-range drift unresolved. We observe that the noise level of the history key-value (K/V) representations trades visual quality against motion, and that restoring gradients through the history aligns causal training far more closely with bidirectional training. Motivated by these observations, we introduce Self-Aligned Forcing (SAF), a training scheme that aligns the history of each block with the noise level of the block being denoised. Specifically, the history is the K/V produced by preceding blocks at the same denoising stage, so all blocks at a stage can be denoised in a single forward pass under a causal mask. This keeps the noisy history differentiable, allowing future losses to optimize how it is encoded. SAF therefore avoids a separate no-gradient rollout and per-block timestep-zero recaching, training up to 1.8x faster than prior methods with lower memory. At inference, SAF achieves the highest single-GPU throughput among existing methods and keeps one history bank per stage for a multi-GPU pipeline, reaching 49.1 FPS on 4 GPUs. Experiments show superior long-horizon generation with a better balance between visual quality and motion. Project page: this https URL.

[CV-16] VISTA: Internalizing Collective Visual Experience via On-Policy Distillation for Active Multimodal Agents

链接: https://arxiv.org/abs/2609.38086
作者: Zheng Jiang,Houde Qian,Yiming Chen,Ling Li,Chaoyang Li,Yueqi Li,Yuxuan Liu,Lifeng Sun
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Active multimodal agents use visual tools to acquire task-relevant evidence while reasoning. Although reinforcement learning samples multiple interaction trajectories per input, outcome-based objectives primarily use the group to estimate scalar advantages, leaving complementary visual discoveries underused. We introduce VISTA, which internalizes collective visual experience through on-policy distillation by turning observations from same-input rollouts into shared supervision. Collective visual experience distillation (CVED) organizes these observations with their interaction context and aligns them with individual decisions, while heterogeneity-aware policy improvement (HAPI) reinforces successful trajectories and provides experience-guided distillation for unsuccessful attempts. An experience-conditioned teacher evaluates the student’s sampled response prefixes, allowing discoveries from one trajectory to guide learning in another without replacing the student’s original history or generating new target trajectories. The trained agent retains its visual tools and acts using its own interaction history. VISTA achieves the strongest average performance among the evaluated active multimodal agents of comparable size and consistently outperforms same-backbone training baselines across fine-grained perception and general reasoning tasks, demonstrating the value of collective experience for active multimodal learning.

[CV-17] OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

链接: https://arxiv.org/abs/2609.38079
作者: Jiaxin Ge,Yiming Qin,Ji Xie,Haozhe Jiang,Xiaochuang Han,Junyi Zhang,Andrew Dai,Yinfei Yang,Jitendra Malik,Ranjay Krishna,Sewon Min,Haiwen Feng,Le Xue,Baifeng Shi,Trevor Darrell,XuDong Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Training a model to generate visual content can encourage it to learn rich perceptual capabilities related to geometry, spatial relationships, and objectness; yet, its benefits for visual understanding remain unclear. We ask: when and how does visual generation supervision improve visual understanding? We study controlled pairs of image-to-image (I2I) generation and image-to-text (I2T) understanding tasks that express the same underlying problem in different output modalities. We find that under the correct recipe, I2I training improves downstream I2T performance, with larger gains as the amount of I2I training data increases. We next ask which generation tasks benefit which understanding capabilities. To study transfer beyond paired tasks, we introduce OmniTaskonomy, a unified taxonomy spanning 19 I2I generation tasks and 25 I2T understanding capabilities. The resulting transfer map reveals selective, task-dependent benefits. Some follow intuitive correspondences, e.g., depth prediction improving metric 3D reasoning, object pointing improving counting, and jigsaw reconstruction improving 2D ordering. Interestingly, we also uncover surprising connections: 2.5D segmentation improving category recognition and Z-depth prediction improving localization. To probe these patterns, we analyze gradient alignment between generation and understanding tasks and find that stronger alignment is associated with larger downstream transfer gains. Together, our results highlight visual generation as a rich source of supervision for visual understanding and provide a roadmap for unlocking its benefits through the right training curriculum and task selection. Project page: this https URL.

[CV-18] MUGEN: Interactive Panoramic World Exploration via Camera Control

链接: https://arxiv.org/abs/2609.38077
作者: Jiaming Tan,Zhen Li,Shuwei Shi,Minggui Teng,Siqi Yang,Yuwei Wu,Bo Zheng,Chuanhao Li,Kaipeng Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL

点击查看摘要

Abstract:Interactive panoramic video generation aims to synthesize immersive 360\textdegree videos that remain visually coherent while following user-specified camera trajectories during exploration. However, progress is limited by a coupled data-and-model gap: existing panoramic video datasets are often short, weakly annotated, or lack camera trajectories, while existing camera-controlled video generation models are designed for perspective videos and do not directly support panoramic geometry. In this paper, we introduce MUGEN and Wan360 to address these limitations. MUGEN is a large-scale real-world panoramic video dataset tailored to interactive 360-degree world exploration, comprising over 1,300 hours of at least 4K panoramic videos with rich semantic and geometric annotations. Built on MUGEN, we further present Wan360, a camera-controllable interactive panoramic video generation model. Panoramic videos are commonly represented by EquiRectangular Projection (ERP), which unfolds a spherical 360-degree view into a rectangular frame with cyclic longitude seams and pole distortions. To this end, Wan360 introduces three parameter-free ERP-aware components: periodic longitude RoPE for seam-consistent positional encoding, ERP-aware padding for reducing boundary artifacts, and random roll yaw for consistent learning. For camera control, Wan360 uses a panoramic Plücker embedding that represents camera motion with ERP rays rather than perspective pinhole rays. Experiments show that MUGEN serves as a data foundation for panoramic world exploration, and that Wan360 enables high-quality, temporally coherent, camera-controllable 360-degree video generation.

[CV-19] RS-OPSD: Reliable Privileged On-Policy-Self-Distillation for Ultra-High-Resolution Remote Sensing VQA

链接: https://arxiv.org/abs/2609.38072
作者: Chengjie Jiang,Yunqi Zhou,Jiafeng Yan,Sihang Zhao,Chun Yuan,Jing Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 16 pages, 7 figures

点击查看摘要

Abstract:Ultra-high-resolution (UHR) remote sensing visual question answering (VQA) requires models to resolve small visual evidence within extremely large images. Existing approaches typically rely on token pruning, visual search, or tool-augmented reasoning at inference time. We instead investigate whether the benefit of zoom-in visual privilege can be internalized into the model. We introduce RS-OPSD, a reliable privileged on-policy self-distillation (OPSD) framework for UHR remote sensing VQA. To provide high-quality privileged information with explicit question-relevant evidence, we construct GeoEvidence-6K, containing 6,750 VQA samples across seven task categories with evidence-region annotations, and develop Human Feedback-Guided Skill Refinement (HF-SR) for scalable annotation. To address context loss from tight crops and conflicting signals from imperfect teachers, RS-OPSD introduces Context-Preserving Visual Privilege (CPVP) and Correctness-Aligned Distillation (CAD). Without any additional visual search or tool calls at inference time, RS-OPSD achieves state-of-the-art (SOTA) performance on XLRS-Bench, MME-RealWorld-RS, and LRS-VQA, outperforming pervious SOTA models of comparable scale by an average of 4.0 percentage points. Moreover, our 2B variant, RS-OPD-Lite, surpasses most 8B-scale models while achieving the fastest measured inference speed. Our Code, GeoEvidence-6K, and the model weights for RS-OPSD and RS-OPD-Lite are publicly available.

[CV-20] WorldLine: Action-Driven Visual Simulation for Robotic Manipulation

链接: https://arxiv.org/abs/2609.38059
作者: Shenghe Zheng,Wenbo Li,Jiyao Zhang,Bin Xia,Haoyang Huang,Nan Duan,Jiaya Jia
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: A work about visual simulators for embodied AI

点击查看摘要

Abstract:Real-world robot learning is constrained by the cost of collecting experience and evaluating candidate behaviors. Video generation models offer a scalable foundation for visual simulators that predict action outcomes before physical execution. Yet they often favor visual plausibility over accurate action following and coherent robot–object dynamics, while action-conditioned simulators depend on scarce, embodiment-specific data that are difficult to share across incompatible control spaces. We introduce WorldLine, an action-driven visual simulator that decouples transferable dynamics learning from heterogeneous action grounding. WorldLine learns manipulation dynamics from more than 10,000 hours of action-free robot videos and grounds them using over 2,000 hours of action trajectories across more than ten embodiments. An image-space action representation provides a shared control interface across embodiments, while multi-view and failure-enriched training with relational regularization improves interaction-sensitive prediction. Robot-focused few-step distillation enables efficient causal rollout while preserving action-critical motion. Across held-out and out-of-domain settings, WorldLine maintains strong visual quality and robot-motion agreement; on failed trajectories, it improves robot-mask IoU by 0.1626 over the strongest baseline. It predicts trajectory success with 74% mean accuracy across RoboTwin and AgiBot, one percentage point above the strongest baseline. Without RoboTwin training or adaptation, its rollouts improve task success by up to 21.4 percentage points over direct policy execution. Together, these capabilities make WorldLine a scalable and efficient visual simulator for policy evaluation and embodied planning. More results are available at \hrefthis https URLproject page.

[CV-21] EVO-WAM: Evolving World Action Models through Video-Action Verification

链接: https://arxiv.org/abs/2609.38057
作者: Shiyang Zhou,Xionghao Wu,Wenbo Li,Shenghe Zheng,Jiyao Zhang,Songsong Yu,Yijun Yang,Jianhui Liu,Haoze Sun,Senqiao Yang,Li Jiang,Jingyong Su,Haoyang Huang,Zhuotao Tian
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Improving robot policies on new tasks without collecting additional expert demonstrations remains a central challenge in robot learning. World action models (WAMs) use broad video priors to jointly predict future videos and actions, offering a potential source of supervision for adapting to new tasks. However, generated videos may fail to depict task completion, and even visually successful videos may be paired with inconsistent actions that lead to execution failure. We propose EVO-WAM, a framework that adapts WAMs to unseen tasks by learning from their own generated video-action trajectories, without executing candidate actions in an external environment. First, we augment WAM training with state prediction and anchored multi-frame context to enable complete autoregressive rollouts without external execution feedback. Second, we identify reliable training experience by selecting task-completing prefixes with a vision-language model and verifying their video-action consistency with an inverse dynamics model. Third, we iteratively train the WAM on verified prefixes and generate new rollouts with the updated model. On seven unseen RoboTwin 2.0 tasks, EVO-WAM increases average success rates from 26.9% to 68.0% for Cosmos3 and from 28.5% to 46.4% for DreamZero, reaching approximately 2.5\times and 1.6\times their initial success rates. On three unseen long-horizon composite tasks in the real world, it improves Cosmos3’s average success rate from 20.0% to 76.7%, a gain of 56.7 percentage points. Project Page: this https URL.

[CV-22] Pow3R-SLAM: Real-Time RGB-D SLAM with 3D Reconstruction Priors

链接: https://arxiv.org/abs/2609.38054
作者: Christopher Kolios,Ishaan Mehta,Sasa Janjic,Yeganeh Bahoo,Sajad Saeedi
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: 9 pages, 4 figures, 4 tables. Project page: this https URL

点击查看摘要

Abstract:We present Pow3R-SLAM, a real-time RGB-D simultaneous localization and mapping (SLAM) system that uses Pow3R for tracking and mapping. Inspired by MASt3R-SLAM, a recent work on monocular SLAM using two-view 3D reconstruction priors, we extend the work to incorporate depth as a prior on the network’s prediction, rather than as geometry to fuse. Where traditional RGB-D SLAM systems struggle with sparsity in the depth images, Pow3R utilizes the available depth to give a better-conditioned pointmap, while inferring the depths in empty regions from the two-view photometric, depth, and intrinsic data. Evaluated against MASt3R-SLAM following its protocol on 24 sequences from TUM, 7-Scenes, and Replica, Pow3R-SLAM runs 1.6x faster in wall time, has 15% lower mean trajectory error, a 3.1x lower unscaled error, and produces denser maps, with a 30% lower Chamfer distance. We also introduce a hybrid variant that runs 2.1x faster than MASt3R-SLAM at 25.3 frames per second (FPS), while maintaining improved tracking and mapping accuracy. Against ORB-SLAM3 in RGB-D mode, Pow3R-SLAM is more accurate on TUM, 7-Scenes, and ETH3D-SLAM, and completes every TUM sequence. While Pow3R-SLAM can struggle on a small set of self-similar scenes, its overall performance shows that adding depth as a prior for two-view 3D reconstruction SLAM can be beneficial. A project webpage is available at: this https URL , and code will be made open-source upon acceptance.

[CV-23] doPlan: A Variable-Horizon Dataset for Multi-Stage Language-Conditioned Planning in Autonomous Driving

链接: https://arxiv.org/abs/2609.38028
作者: Parthib Roy,Yash Tandon,Marcus Blennemann,Giovanni Tapia Lopez,Angel Martinez-Sanchez,Mohan M. Trivedi,Ross Greer
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Autonomous vehicles interacting with passengers through natural language must reason beyond immediate commands. Passenger intent may span multiple stages of behavior, depend on future events, refer to surrounding agents or landmarks, and remain relevant as driving conditions evolve. Existing language-enabled driving datasets largely focus on short, localized interactions, leaving these longer-horizon forms of passenger intent comparatively underexplored. We introduce doPlan, to our knowledge the first publicly available, human-annotated real-world dataset designed to study passenger language as persistent task context. Built on nuPlan, doPlan contains 5,154 human-written passenger instructions spanning 169.1 hours of cumulative instruction-aligned context over 50.9 hours of unique driving, with annotation windows ranging from 30.0 to 508.8 s. The annotations capture immediate, deferred, event-conditioned, persistent, and multi-stage passenger intent. The dataset, annotation interface, and supporting resources are publicly available at this https URL. We evaluate four language-conditioned driving models and find that sensitivity to passenger language does not reliably translate into behavior consistent with the requested direction. More broadly, among 2,161 examples with a matched future maneuver, the first associated maneuver occurs a median of 24.6 s after the evaluation point, and only 9.8% occur within the models’ common 5 s prediction horizon. These findings highlight the need to connect persistent passenger intent with successive planning decisions. doPlan provides a setting for studying how unresolved goals can be retained, grounded in evolving scenes, and tracked across multiple stages, including how a planner determines when a future goal becomes relevant to the current plan.

[CV-24] Beyond Lip Sync: Reference-Grounded Oral Refinement for Audio-Driven Portrait Animation

链接: https://arxiv.org/abs/2609.38019
作者: Bangxun Tang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 19 pages, 8 figures, 5 tables. Under review

点击查看摘要

Abstract:We present RGOR (Reference-Grounded Oral Refinement), an audio-driven lip-sync framework that renders the mouth of the specific person being dubbed rather than a generic one. Existing lip-sync systems follow the audio closely and keep the face recognizable, yet the mouth they render is an average mouth: the shape and texture of the lips, the arrangement of the teeth, and how much of them shows as the mouth opens are not that person’s. The problem persists because nothing in current training or evaluation asks for the person’s own mouth: perceptual losses accept any plausible mouth, face identity is carried mostly by the skin around it, and the released inference code of inpainting systems uses the unmasked target frame as the reference, which hides the gap. To address this, RGOR conditions every generated frame on frames from separate enrollment recordings of the same person and on HD patches of the mouth that bypass the VAE, and trains the generator against a paired judge that compares each rendered mouth with the person’s reference and learns to reject a realistic mouth of someone else. We further build an evaluation protocol and use it to compare open-source and commercial lip-sync systems on held-out identities. Experiments show that RGOR achieves the best or second-best result on most metrics, and preserves the person’s own lip and dental detail while keeping synchronization and the rest of the face intact.

[CV-25] Brain-SAD: A Brain-Inspired Safe Autonomous Driving Control Framework with Dynamic Fear-Oriented Constraint on Dual-Policy

链接: https://arxiv.org/abs/2609.38016
作者: Huan Rong,Chao Yin,Anouar Imel,Yijie Xia,Tinghuai Ma
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Neural and Evolutionary Computing (cs.NE); Robotics (cs.RO)
备注: 18 pages, 11 figures

点击查看摘要

Abstract:Constrained Reinforcement Learning has recently gained increasing attention in the field of Safe Autonomous Driving, where the general mechanism is to maximize the expected reward while keeping the overall action risk bounded. In this way, the safety issues arising in AD can be mitigated through constrained actions. However, existing Constrained RL methods still lack dynamics on the imposed constraints. For instance, the action cost adopted by the existing Primal-Dual/soft-constrained methods is often defined as static state-to-cost mapping, and the safe-action projection in hard-constrained methods relies on the static projection with the fixed feasible region boundary estimated from offline demonstrations. The above drawback tightly couples the imposed constraints to the training scenarios, leaving the AD policy hard to handle different interaction scenarios, due to the improper state-level action-cost and the static projection boundary. Consequently, in this paper, we propose Brain-SAD, a brain-inspired safe autonomous driving control framework with dynamic fear-oriented constraints. By perceiving the current vehicle-interaction scene, Brain-SAD generates dynamic fear signal as fear reaction to online decide long-term policy for regular interaction or short-term policy for urgent-collision defense. In such two policy, the above fear-reaction will be constructed as the dynamic fear constraints, respectively reflecting the overall fear cost directly coupled with action-impact, and the dynamic fear boundary of the feasible region derived from different risky neighbors, both of which will in turn serve for the online policy optimization. Experimental results show that Brain-SAD outperforms existing methods, achieving higher success rate in shorter task-completion and collision-recovery time, and exhibits stronger reliability across continuous intersections of fluctuating complexity.

[CV-26] From Unity Simulation to Diffusion-Based Augmentation: Quantifying Dataset Balance for Robust Object Detection

链接: https://arxiv.org/abs/2609.38010
作者: Mohamed Benkedadra,Aissa Saoudi,Maxime Gloesener,Sidi Ahmed Mahmoudi,Matei Mancas
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Modern computer vision models achieve high accuracy when trained on large-scale annotated datasets. In critical domains such as construction safety monitoring, data collection is costly, hazardous, and ethically constrained. This paper presents a systematic study comparing two complementary data generation paradigms, (1) Unity Simulation-based rendering and (2) Controllable Diffusion-based generation (CIA), for object detection under real data-scarce conditions. A unified experimental framework enables controlled dataset mixing across real, simulated, and generative sources, while maintaining identical model and training settings. Quantitative evaluation using Precision, Recall, mAP, and custom \Delta -metrics, reveals that neither simulation nor generative augmentation alone achieves optimal transferability. Unity-only training yields an mAP@0.5 drop of -50% relative to real data, while CIA-only training shows a milder -16.5% degradation. Hybrid compositions significantly improve performance, with the 90% real + 10% Unity configuration achieving the best overall mAP@0.5 of 62.68% ( +7.64% over baseline), and the 90% real + 10% CIA configuration maximizing precision at 74.45% . Results demonstrate that limited synthetic inclusion enhances generalization, while excessive substitution induces domain drift.

[CV-27] HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents

链接: https://arxiv.org/abs/2609.38008
作者: Tongbo Chen,Junbo Niu,Zhengxi Lu,Niu Lian,Fei Tang,Yuchen Yan,Yike Hong,Yong Du,Yizhou Liu,Bofan Chen,Yongliang Shen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project Page: this https URL Code: this https URL

点击查看摘要

Abstract:Computer use agents (CUAs) have demonstrated strong capabilities in completing digital tasks. However, existing CUAs either rely solely on graphical user interface (GUI) interactions, which are often inefficient and error prone, or augment GUI interactions with application specific APIs or tools, which require substantial engineering effort and are difficult to scale across applications. We argue that the next generation of CUAs should combine GUI interactions with the command line interface (CLI), leveraging the generality of the GUI and the efficiency of shell commands. A critical challenge, however, is that current models do not know when or how to use the CLI during task execution. To address this challenge, we develop a data construction pipeline that produces three types of trajectories: GUI only, CLI only, and interleaved GUI and CLI trajectories. This pipeline results in HybridCUA-8K, containing 5K hybrid trajectories and 3K verified RLVR tasks. Building on these data, we propose a training framework with two stages: supervised fine tuning on the constructed trajectories, followed by reinforcement learning with our CLI aware rewards that encourages agents to use the CLI selectively and reliably. Experiments show that HybridCUA-9B achieves 53.6% accuracy on OSWorld, improving over the base model by 14.8 percentage points, and improves performance on WindowsAgentArena by 4.0 percentage points. These results demonstrate the effectiveness and cross platform generalizability of the hybrid GUI and CLI paradigm for computer use agents.

[CV-28] ORMA: Optimization-based Monocular 4D Reconstruction of Articulated Animals

链接: https://arxiv.org/abs/2609.37986
作者: Xuyi Hu,Francesco Palandra,Shangzhe Wu,Daniel Cremers,Riccardo Marin,Silvia Zuffi
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Recovering articulated 4D representations of animals from monocular videos remains challenging due to the large diversity of quadruped morphologies and lack of animal 4D supervision data. Existing learning-based reconstruction methods operate on individual images and rely on synthetic or model-fitted 3D supervision, which inherits the constraints of strong parametric priors and limits generalization to out-of-distribution species. When applied to out-of-distribution animals, they often recover a plausible pose while producing inaccurate geometry because the underlying shape model cannot faithfully represent the observed instance. We present ORMA, a training-free reconstruction framework that decouples articulation from shape, using the predicted pose as reference for optimization while leveraging generative 3D priors for accurate shape reconstruction. Given a reference image, we reconstruct the animal geometry and register it to the parametric model SMAL+, yielding an articulated shape adapted to the observed instance. We then combine per-frame articulated pose estimates with globally consistent camera poses to recover animal motion in a shared world coordinate frame, and further refine the reconstruction using self-supervised DINO correspondences and temporal consistency. To enable quantitative evaluation, we introduce PAW4D, a synthetic multi-species benchmark with ground-truth 3D geometry and camera motion. Experiments on PAW4D, PFERD, and challenging in-the-wild videos demonstrate that ORMA improves reconstruction accuracy while recovering globally consistend animal motion across diverse quadruped species.

[CV-29] PhysWAM: Physically Consistent World Action Model for Autonomous Driving

链接: https://arxiv.org/abs/2609.37970
作者: Dhruv Parikh,Fengcheng Yu,Quankai Gao,Jiawei Yang,Junjie Ye,Maulik Bhatt,Thang Vu,Charles Ochoa,Rowan McAllister,Igor Vasiljevic,Rajgopal Kannan,Viktor Prasanna,Vitor Guizilini,Yue Wang
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: Technical Report

点击查看摘要

Abstract:World-action models (WAMs) jointly predict how a scene will evolve and how an agent should act, however joint generation alone does not necessarily impose a shared geometric constraint on these predictions. We present PhysWAM, a unified world-action model for autonomous driving that co-denoises multiview video, metric depth, and ego motion within a single flow-matching transformer. To ground world and action generation in measured scene geometry, we introduce Coupled Point Projection (CPP) that unprojects the generated depth into 3D points, transforms them using the generated \mathrmSE(3) ego motion, and minimizes their distance to LiDAR points transformed using the recorded ego motion. This geometric constraint promotes physical consistency with the measured scene by jointly supervising generated depth and motion alongside their standard flow-matching objectives. At inference, trajectory selection relies only on a simple label-free consensus rule, with no learned scorer or simulator feedback. We evaluate PhysWAM across NAVSIM v1 and v2 planning, zero-shot closed-loop transfer, and future video and metric-depth prediction. Despite PhysWAM’s simple selection procedure, it achieves strong planning performance and transfers zero-shot to unseen driving environments. It also generates accurate metric depth and temporally coherent video, with CPP improving both planning and depth prediction. Together, these results demonstrate that the geometric relationship between scene depth and ego motion provides a direct way to couple world and action generation within a simple unified model.

[CV-30] SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video

链接: https://arxiv.org/abs/2609.37969
作者: Haozhe Liu,Tian Ye,Shuchen Xue,Yitong Li,Junsong Chen,Haopeng Li,Jincheng Yu,Duomin Wang,Ruihua Zhang,Lei Zhu,Song Han,Enze Xie
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 15 pages

点击查看摘要

Abstract:High-resolution video generation is expensive, as its cost grows rapidly with the number of spatiotemporal tokens. A practical alternative first generates a lower-resolution video and then applies a refiner, but conventional multi-step refinement introduces a second sampling bottleneck. We present SoL-Refiner, a one-step video refiner that transforms low-resolution model outputs into 4K videos with a single denoising step. Our three-stage recipe combines high-resolution continual training, reinforcement learning (RL) post-training, and a final one-step distillation. We introduce Refiner-Bench, a video refinement benchmark constructed from the outputs of different video generators, and use a shared-input protocol to compare refiners at approximately 2K output resolution. At 2K, the one-step SoL-Refiner outperforms all external refiners on the VBench and UniPercept averages, while at 3840!\times!2176 it improves both metrics over the three-step LTX-2.3 Refiner. With the complete acceleration stack, SoL-Refiner achieves an 8.91\times speedup in refinement latency over the same baseline in our 2K latency setting.

[CV-31] Video-RSI: Recursive Self-Improvement of Video Understanding Agents via Harness Evolution

链接: https://arxiv.org/abs/2609.37950
作者: Bingjun Luo,Jialin Guo,Siqi Li
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Video understanding agents acquire evidence through an executable harness that controls what they observe and how they use those observations. However, execution traces contain only the evidence acquired by the current harness, leaving competing explanations for failure unresolved and limiting the basis for self-improvement. We introduce Video-RSI, a framework for recursive self-improvement in which a video understanding agent uses its own language model to revise its harness. Through active video investigation, the model revisits the original training videos to test competing failure explanations with additional observations, grounding proposed changes in evidence beyond the existing trace. Cost-aware harness evolution turns these diagnoses into reusable revisions and determines which revisions to retain by considering both answer accuracy and visual cost. Across our evaluation settings on video understanding benchmarks, the evolved agent improves accuracy while processing fewer frames and achieves competitive accuracy-efficiency trade-offs against existing video understanding agents. These results demonstrate the potential for video understanding agents to improve their own evidence acquisition and use through harness evolution. Code is available at this https URL .

[CV-32] Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark

链接: https://arxiv.org/abs/2609.37938
作者: Yuedong Tan,Lei Qi,Yu Liu,Di Wen,Ruiping Liu,Xiaoye Wang,Yufan Chen,Junwei Zheng,Chengzhi Wu,Chen Zhang,Zhihang Chen,Haiwen Sun,Zongwei Wu,Radu Timofte,Danda Pani Paudel,Kunyu Peng
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggregate cross-video accuracy conflates failures of local perception with failures to preserve observation identity, establish correspondence, and compose evidence, obscuring whether local video understanding actually transfers. We introduce EgoGears, a complementary single- and multi-video benchmark designed to diagnose this transition. It contains 567 single-video and 1,487 multi-video questions derived from 126 human-collected egocentric recordings covering 39 outdoor routes. Repeated traversals across movement speeds and lighting conditions ground comparisons in shared physical environments; 531 questions require alignment across independent recordings. Single-video questions measure the local visual, spatial, and motion evidence available to a model, while multi-video questions test whether evidence remains bound to the correct observation and can be composed into consistent route relationships. We report 29 single-video and 31 multi-video MLLM configurations across six model families in the main leaderboard. Among the 20 configurations evaluated comparably on both splits, every model performs worse on multi-video questions, with a mean decrease of 22.5 percentage points, and the gap persists when answer format and scoring are held fixed. The gap is not explained simply by additional videos or recording boundaries. The central bottlenecks are observation–evidence binding and ordered route-state tracking. The code and benchmark are publicly available at this https URL.

[CV-33] Look Closer: Patch-wise Supervision for AI-Generated Image Detection

链接: https://arxiv.org/abs/2609.37937
作者: Zhida Zhang,Tao Wu,Siyu Liu,Jie Cao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 29 pages, 11 figures, 28 tables. Code: this https URL

点击查看摘要

Abstract:How much of an image does a detector need to see? Small RGB regions can retain useful evidence of image synthesis even when they reveal little of the full scene. Motivated by single-patch detection, we study patch-wise supervision: a shared backbone classifies explicit crops, each crop receives its own loss, and patch probabilities are averaged only at inference. The procedure requires neither handcrafted residual filtering nor a learned image-level fusion module. Experiments span single-patch selection, multiple generator collections, and four CNN and Transformer backbones. On GenImage, the reported patch-wise variants improve average accuracy over their whole-image counterparts across all four backbones. Comparisons of supervision granularity, source resolution, crop size, and inference coverage further characterize the approach, while post-processing tests and difficult-image evaluation reveal its limitations. The historical experiments include evaluation-based model selection, so their scores are not presented as a uniformly selected leaderboard comparison. Overall, the study identifies explicit local input and patch-level supervision as a simple, useful combination for investigating generalizable AI-generated image detection.

[CV-34] Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation

链接: https://arxiv.org/abs/2609.37925
作者: Chenjian Gao,Zhihao Hu,Jianqi Ma,Jun Zhang,Weidong Zhang,Tianfan Xue
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Autoregressive (AR) video diffusion enables low-latency, streamable video generation, but prediction errors often accumulate over long rollouts. Training the generator on its own rollouts exposes it to these imperfect histories. However, existing video-level distribution matching distillation (DMD) scores the whole rollout jointly. Because a chunk is evaluated together with its past and future, its correction can favor matching artifacts in the surrounding context merely to preserve temporal consistency. To provide a clearer visual-quality signal, we introduce Rollout-Marginal Distillation (RMD). RMD retains the generated history for AR prediction but scores each chunk independently against a chunk teacher, ensuring its quality correction is not compromised by an imperfect temporal context. To compensate for the lack of temporal context in independent chunk scoring, RMD subsequently applies video-level DMD to restore temporal coherence. Extensive experiments demonstrate that RMD maintains high visual quality far beyond its training horizon and outperforms video-level DMD baselines. Code and video results are available at this https URL

[CV-35] EpiCon: Collective Agent Learning through Co-Evolving Multimodal Memory

链接: https://arxiv.org/abs/2609.37923
作者: Ziyun Zeng,Hang Hua,Shaden Alshammari,Rogerio Feris,William T. Freeman,Jiebo Luo
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Preprint

点击查看摘要

Abstract:Agents can learn from past executions, but enabling different agents to reuse and build on one another’s experience remains challenging. We introduce EpiCon, a shared multimodal memory framework for agent collective learning without updating host model parameters. EpiCon links question-level memory evolution to a persistent experience bank through two independently trained 2B models: a memory controller and a tree self-organizer. The controller jointly refines textual guidance and visual evidence across attempts and selectively includes visual memory. The self-organizer consolidates lessons hierarchically and retrieves experience and rules for new problems. We evaluate EpiCon on eleven benchmarks spanning four multimodal task domains, using two harnesses and multiple backbones. A frozen bank improves other systems even with a single solving attempt. A second harness raises the original system’s macro-average score by 2.6 points across eleven benchmarks. Across four host configurations, EpiCon improves macro-average scores by 1.7 to 4.9 points over No Memory and reduces memory-operation time by 67% to 74% relative to backbone-sized memory models.

[CV-36] SYNCR: Diagnosing and Learning Cross-Video Reasoning from Simulation

链接: https://arxiv.org/abs/2609.37918
作者: Sara Ghazanfari,Siddharth Garg,Prashanth Krishnamurthy,Farshad Khorrami
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Reasoning across videos requires aligning events, matching identities, comparing motion, and integrating partial observations. Evaluating these capabilities and testing how to improve them requires both reliable labels and targeted supervision. We introduce SYNCR, a simulator-grounded framework that connects these two needs through shared task generators. Built on Habitat, Kubric, and CLEVRER, SYNCR derives answers from environment state and provides 4,000 evaluation questions and 15,960 training questions over disjoint videos, spanning eight cross-video reasoning tasks. Visual ablations and human evaluation assess dependence on the supplied evidence and answer recoverability. Evaluation of 22 multimodal large language models reveals persistent difficulties in physical comparison and scene integration that increasing model size does not consistently resolve. Supervised fine-tuning raises Qwen3-VL-8B’s average SYNCR accuracy from 32.6% to 61.6%, with gains extending to task configurations and video sources absent from training for those tasks. Transfer to real footage is most consistent for temporal ordering: accuracy improves by 9.0-20.5 percentage points on constructed Assembly101 and Panoptic ordering sets across three checkpoints spanning two model families and two model sizes, with additional gains on existing temporal reasoning benchmarks. These results establish SYNCR as a controlled setting for diagnosing cross-video reasoning failures, testing their learnability, and identifying where synthetic supervision transfers.

[CV-37] Pixels to Keys: Exploring Spatial and Motion Cues in Gameplay Inverse Dynamics ECCV2026

链接: https://arxiv.org/abs/2609.37907
作者: Abhishek Pillai,Ekta Prashnani,Joohwan Kim,Iuri Frosio
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at the Workshop on Multimodal Digital Agents (ECCV 2026): this https URL

点击查看摘要

Abstract:Video games offer scalable environments for studying perception and control in embodied this http URL online gameplay videos could supply demonstrations, but they rarely include player inputs for training. Inverse Dynamics Models (IDMs) have thus been proposed to infer inputs from frames. Large (up to 1B parameters) IDMs trained on \sim 1K-2K gameplay hours demonstrate feasibility and cross-environment generalization at this scale, but researchers do not clarify what the key components are to recover individual actions and often report only aggregate accuracy that can mask rare-action failures. We study the problem in a data-constrained scenario to evaluate how spatial motion features, model architectures, and training objectives affect an IDM’s outcome and we analyse our models on per-key and balanced metrics such as F_1^macro . Our experiments on Trackmania highlight the importance of factors like the model architecture and motion flow extraction in preprocessing, while also showing the limits of evaluation through unbalanced metrics. The application of the same architecture and training recipe to Cyberpunk 2077 reveals uneven performance across game mechanics. Our per-action evaluation and failure analysis highlight ambiguities from camera motion, delayed effects and imbalanced key-press frequencies that call for explicit modeling of 3D scene structure, long-term state and the adoption of proper losses in future implementations.

[CV-38] ReCAP: Retrieval-Guided Capability Reuse for Multimodal Continual Instruction Tuning

链接: https://arxiv.org/abs/2609.37889
作者: Tao Hu,Zhinuo Zhou,Xialiang Tong,De-Chuan Zhan,Da-Wei Zhou
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Multimodal continual instruction tuning (MCIT) aims to enable multimodal large language models to acquire new capabilities from sequential tasks while preserving previously learned knowledge. Existing methods primarily mitigate catastrophic forgetting by constraining parameter updates or separating task-specific adaptations. However, continual adaptation can also benefit from external knowledge that provides domain-specific information and reusable reasoning patterns for solving diverse instructions. For example, to answer “How many red cubes are to the left of the sphere?”, domain knowledge can provide relevant concepts about objects and spatial relations, while reasoning knowledge can specify ordered operations such as object recognition, spatial filtering, and counting. Despite this potential, how to leverage external knowledge for continual adaptation remains largely unexplored in existing MCIT methods. To this end, we propose ReCAP, a retrieval-guided framework that leverages external knowledge to guide capability reuse during continual adaptation. At each continual stage, ReCAP uses external search and an LLM to incrementally build a knowledge base of domain, reasoning, and format knowledge based on the current-stage training data. For each instruction, retrieved domain knowledge guides generation, while retrieved reasoning knowledge selects and orders capability modules to form an instance-specific capability path. As these capability modules are reused across stages, subsequent adaptation can overwrite previously learned parameters. To enable stable cross-stage reuse, ReCAP introduces adaptive subspace recycling, which parameterizes reusable capability modules with shared bases and stage-specific cores, protects historically important directions while recycling residual capacity. Extensive experiments on MCIT benchmarks show that ReCAP achieves SOTA performance.

[CV-39] Visual Branch is What You Need for CLIP-based Class-Incremental Learning

链接: https://arxiv.org/abs/2609.37888
作者: Tao Hu,Zhen-Hao Xie,Jingcai Guo,De-Chuan Zhan,Da-Wei zhou
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Class-Incremental Learning (CIL) requires models to recognize new classes over time without forgetting previously learned ones. With the rise of vision-language pre-training, CLIP has become a strong foundation for CIL. A common design in CLIP-based CIL is to construct textual classifier weights by encoding class-name templates with the CLIP text encoder, and then classify visual features by image-text cosine similarity. This design is appealing: since CLIP aligns images and text in a shared embedding space, textual weights appear to provide an off-the-shelf classifier for incremental classes. However, we show that this seemingly natural design is not always beneficial, as a modality gap can still separate the two modalities and make textual classifier weights deviate from visual class distributions. Empirically, under identical task-wise CIL training, initializing the cosine classifier with visual class centers yields lower loss and better incremental accuracy than using CLIP textual this http URL by these observations, we propose VIS, a visual-only method for CLIP-based CIL that removes the deployed textual branch and constructs the incremental classifier entirely in the visual space. To obtain stronger task-adaptive visual representations, VISuses only base-session data to enhance CLIP’s final visual representation with informative visual-layer features. Built on the enhanced visual representation, VISemploys a simple kernelized incremental least-squares SVM, whose classifier weights are solved in closed form from additive sufficient statistics. When new classes arrive, VISaccumulates their sufficient statistics and recomputes the classifier weights for all seen classes, enabling efficient incremental updates while preserving historical class knowledge. Extensive experiments show that VISachieves state-of-the-art performance without a textual branch.

[CV-40] EndoPrior-GS: Dynamic Endoscopic Reconstruction with a Joint Texture Prior ACCV2026

链接: https://arxiv.org/abs/2609.37874
作者: Jiaqi Huang,Shidong Wang,Tong Xin,Kabita Adhikari
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at ACCV 2026. Code: this https URL

点击查看摘要

Abstract:Dynamic endoscopic reconstruction is fundamental to robotic surgery and computer-assisted interventions. While 3D Gaussian Splatting (3DGS) realises real-time rendering, its application to deformable intraoperative environments remains constrained by spurious geometry and varying illuminations. To address these limitations, we introduce EndoPrior-GS, a novel pipeline that explicitly couples frame-extracted vision heuristics and estimated depth maps. EndoPrior-GS derives a joint texture prior from a tool-filtered valid tissue mask, a non-specular photometric filter, and anatomical structural salience, yielding a probability map that guides primitive initialisation and subsequent density control. The prior is further extended to the temporal domain through a texture-aware term that dynamically weighs pairwise primitive contributions during training. We conduct extensive experiments on benchmark datasets EndoNeRF and SCARED, and the obtained results show that our method EndoPrior-GS reduces Flow Error by 27.7% and 25.8% over the representative approaches while preserving competitive rendering quality and real-time rendering speed. Our project website is available at this https URL.

[CV-41] Learning from synthetic photorealistic raindrop for single image raindrop removal ICCV

链接: https://arxiv.org/abs/2609.37870
作者: Zhixiang Hao,Shaodi You,Yu Li,Kunming Li,Feng Lu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW)

点击查看摘要

Abstract:Raindrops adhered to camera lens or windshield are inevitable in rainy scenes and can become an issue for many computer vision systems such as autonomous driving. Because raindrop appearance is affected by too many parameters, therefore it is unlikely to find an effective model based solution. Learning based methods are also problematic, because traditional learning method cannot properly model the complex appearance. Whereas deep learning method lacks sufficiently large and realistic training data. To solve it, in our work, we propose the first photo-realistic dataset of synthetic adherent raindrops for training. The rendering is physics based with consideration of the water dynamic, geometric and photometry. The dataset contains various types of rainy scenes and particularly the rainy driving scenes. Based on the modeling of raindrop imagery, we introduce a detection network which has the awareness of the raindrop refraction as well as its blurring. Based on that, we propose the removal network that can well recover the image structure. Rigorous experiments demonstrate the state-of-the-art performance of our proposed framework.

[CV-42] HandAnthro: Automated Hand Anthropometry from a Single Image

链接: https://arxiv.org/abs/2609.37855
作者: Fan Zhou,Shuairan Chen,Mengying Zhang,Yulin Wu,Sadegh Jafari,Sixing Yu,Rui Li,Ali Jannesari,Guowen Song
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 21 pages, including 7 pages of main text and references and 14 pages of supplementary material

点击查看摘要

Abstract:Hand anthropometry supports protective-glove design, but existing measurement methods often require trained operators, specialized hardware, or manual landmarking. We present HandAnthro, which estimates 44 projected hand dimensions from a smartphone photograph of a palm-up hand on US letter-size paper. The pipeline reconstructs wrist-occluded paper boundaries for rectification, whitens non-hand pixels, and refines 41 anthropometry-specific landmarks from a fine-tuned You Only Look Once (YOLO) pose model using image-specific geometry and contours. Controlled evaluation comprised 720 captures from 45 held-out participants, each contributing 16 images across two smartphones, two backgrounds, two angles, and two nominal illumination settings. HandAnthro produced complete outputs for 704 captures (97.8%); among these, mean absolute error (MAE) was 3.80 mm per dimension against two trained operators’ caliper measurements. Regional MAEs were 2.48 mm for non-thumb fingers, 6.04 mm for thumbs, and 6.17 mm for palm and wrist. In a researcher-assisted mobile-app pilot, automated batch processing returned all 44 dimensions for 260 of 268 retained, researcher-screened firefighter images (97.0%). A descriptive, unpaired comparison with an independent national firefighter reference yielded a mean absolute difference of 2.40 mm across 28 sex-by-dimension group-mean contrasts. These results characterize controlled measurement performance and researcher-assisted field feasibility for future distributed hand-anthropometry studies.

[CV-43] FlowMap-OPD: Rollout–Kernel Separation for On-Policy Distillation of Few-Step Flow-Map Generators

链接: https://arxiv.org/abs/2609.37851
作者: Zhiqi Li,Bo Zhu
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 38 pages, 18 figures

点击查看摘要

Abstract:Few-step flow-map generators, including MeanFlow and consistency models, enable efficient sampling through long-range transport, yet their on-policy distillation remains underexplored. We introduce FlowMap-OPD, an on-policy distillation framework that separates student-state acquisition from teacher–student distribution comparison. A formulation based on state marginals establishes this separation, while flow–velocity consistency connects local supervision to the deployed long-range map. Within this framework, we develop flow-map, induced-velocity, and instantaneous-velocity distribution supervision, each paired with a separately specified native flow-map rollout. Cross-capacity ImageNet experiments across three teacher rewards identify instantaneous-velocity distribution supervision with independently tunable student consistency as the most effective choice. In text-to-image experiments, FlowMap-OPD demonstrates strong multi-specialist consolidation capabilities and surpasses multi-reward Flow-Map GRPO in task performance and convergence speed.

[CV-44] RelayVSR: Large-Small Model Collaboration for Efficient Real-World Video Super-Resolution

链接: https://arxiv.org/abs/2609.37850
作者: Xijun Wang,Xin Li,Zirui Lang,Suhang Yao,Haoran Li,Zhibo Chen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: The code is available at this https URL

点击查看摘要

Abstract:Large generative models can recover realistic detail in real-world video super-resolution (VSR), but processing an entire video with them is computationally expensive. In this work, we present RelayVSR, a streaming VSR framework built on the Sparse Generative Relay mechanism. A large generative model generates reference latents for sparse keyframes, while a lightweight VSR network uses these references and low-resolution video to super-resolve every frame. The lightweight VSR network, implemented as a Dual-Memory Video Transformer, reuses keyframe information across frames and updates recent video context, supporting first-keyframe conditioning and dual-endpoint conditioning with bounded lookahead. However, errors in shared keyframes can propagate and accumulate across output frames, making keyframe quality alone an insufficient optimization target. We address this collaboration gap with Video-Aware Reference Optimization (VARO), which uses reinforcement learning to update the large generative model with two reward levels: a system-level reward evaluates videos produced by the fixed lightweight VSR network, while a reference-level reward evaluates decoded keyframe quality. VARO improves final video quality over direct joint training, and its dual-level rewards outperform a system-level reward alone. At 1080p on a single NVIDIA A100 80GB, dual-endpoint RelayVSR with a 15-frame keyframe interval reaches 29.29 FPS, 13.82 GB peak GPU memory, and 0.327 s first-frame model latency, compared with 7.80 FPS, 24.447 GB, and 2.83 s for FlashVSR-Tiny. The code is available at this https URL.

[CV-45] Evaluation Choices Shape Biomedical ML Claims: A Pediatric Pneumonia Benchmark Case Study

链接: https://arxiv.org/abs/2609.37848
作者: Bhanu Prakash Vangala,Sowmya Guda,Latha Peddi,Navya Vangala
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Biomedical machine learning papers often compress model performance into one headline number. That number can look like a property of the model even when it depends strongly on how the benchmark was evaluated. We study this problem on the widely used Kermany pediatric chest radiograph dataset using nine image classifiers and a controlled evaluation protocol. Under the same protocol, the eight pretrained backbones differ by only 0.026 AUROC. In contrast, changing whether the backbone is frozen or fine-tuned changes AUROC by 0.044 on average, and changing the decision threshold changes balanced accuracy by 0.090 on average. The official test split is also measurably different from the training pool: a partition classifier distinguishes them at AUC 0.697, rising to 0.898 for normal radiographs. Most strikingly, a classifier using only file properties, with no image anatomy, reaches 0.992 balanced accuracy within the training pool but falls to 0.496 on the official test split. Validation-fitted thresholds and calibration also transfer imperfectly. These results show that a high benchmark score can support different conclusions when the split, training policy, threshold, metric, calibration, and uncertainty are not communicated with it. We end with a seven-item reporting recommendation in which each item is tied to an effect measured in the study

[CV-46] ReCaVSR: One-Step Streaming Diffusion Video Super-Resolution with Recycled Latents and Learned Cache Routing

链接: https://arxiv.org/abs/2609.37831
作者: Xijun Wang,Xin Li,Suhang Yao,Zirui Lang,Bingchen Li,Zhibo Chen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: The code is available at this https URL

点击查看摘要

Abstract:Real-time diffusion-based video super-resolution (VSR) is in high demand for online streaming, yet stringent latency requirements often compromise generative fidelity. We propose ReCaVSR, a Wan2.2-based, one-step framework for streaming VSR that builds on two observations: recycled SR latents retain local temporal context, reducing the need for full historical Key-Value (KV) caches; and individual transformer layers benefit from distinct temporal scopes. ReCaVSR combines three complementary designs: (i) layer-wise cache routing with recycled SR latents: each DiT layer learns its KV-cache temporal scope under a cache budget and exports a static inference schedule, while recycled SR latents propagate local context by conditioning each new block on the model’s own preceding predictions. (ii) Multi-Scope Query (MSQ) Discriminator: a compositional discriminator combining global, spatial-window, and temporal-tube feedback for holistic realism, local texture generation, and temporal stability. (iii) LR-conditioned adaptation of FlashDecoder: a VAE decoder that incorporates LR observations for efficient latent decoding. ReCaVSR enables streaming VSR without iterative sampling or full historical KV-cache materialization. Experiments on synthetic and real-world VSR benchmarks show better perceptual quality, temporal consistency, and streaming efficiency than representative VSR baselines. At 1080\times1920 output resolution on a single NVIDIA A100-80GB, ReCaVSR achieves 21.20 FPS with 15.16 GB peak allocated GPU memory, running 2.72 \times faster while using 38.0% less peak allocated memory than FlashVSR Tiny. The code is available at this https URL.

[CV-47] Minkowski Attractor Networks: Closed-Form Hyperbolic Flows for Visual Representations

链接: https://arxiv.org/abs/2609.37817
作者: Zhongping Ji
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 15 pages

点击查看摘要

Abstract:Geometric representation learning predominantly scaffolds representations onto flat Euclidean subspaces or compact product tori ( \mathbbT^K ). However, flat manifolds possess vanishing curvature and polynomial volume growth, inherently suffering from metric distortion when embedding multi-scale, tree-like visual hierarchies. While hyperbolic spaces ( \mathbbH^m ) circumvent this via constant negative curvature ( K0 ) and exponential volume expansion, prior hyperbolic deep architectures are hindered by computationally cumbersome Riemannian optimization, non-linear gyrovector calculus, and floating-point instabilities. In this work, we introduce \textbfMinkowski Attractor Networks (MAN), an operator-splitting-inspired framework that embeds representations within pseudo-Riemannian Minkowski spacetime ( \mathbbR^1,m ). By framing hyperbolic manifolds as quadric level sets, MAN resolves hyperbolic geometry by combining linear Lorentz group transport with non-linear cone lifting and closed-form radial rescaling, evaluating in a single forward pass without numerical ODE solvers or iterative retractions. We establish \textbfMAN-2D ( \mathbbR^1,1 \to \mathbbH^1 ) as our primary, high-throughput visual backbone, which maximizes channel factorization granularity into D/2 independent two-dimensional Minkowski blocks. We further formulate \textbfMAN-4D ( \mathbbR^1,3 \to \mathbbH^3 ) as a spacetime extension, leveraging a commuting Cartan-subalgebra parameterization of \mathrmSO^+(1,3) to evaluate 4D Lorentz isometries via two commuting 2D planar maps without matrix-exponential overhead. Comments: 15 pages Subjects: Computer Vision and Pattern Recognition (cs.CV) Cite as: arXiv:2609.37817 [cs.CV] (or arXiv:2609.37817v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2609.37817 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[CV-48] WINGS: Reference-Free Gaussian Splatting Inpainting with 3D-Native Generative Priors

链接: https://arxiv.org/abs/2609.37816
作者: Noé Lallouet,Michael Fischer,Elie Michel
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Preprint. Under review

点击查看摘要

Abstract:Inpainting 3D Gaussian Splatting scenes, a key challenge in 3D editing, requires generating plausible content within a masked region of 3D space. Prior approaches rely on 2D diffusion models to produce one or several inpainted reference views, making them susceptible to challenges associated with multi-view inconsistency and lengthy optimization times. Departing from these approaches, we introduce a reference-free Gaussian splatting inpainting method operating natively in 3D. Our method leverages the embedding space of a large, pre-trained 3D prior, combined with a structure completion network to feed a generative prior which reconstructs the missing region’s geometry and appearance. Performing content generation entirely in 3D, it avoids the need to reconcile inconsistencies of multiple inpainted reference images, and is faster than related 2D-based methods. We demonstrate the effectiveness of our method qualitatively and quantitatively, through extensive experiments and a user study. To the best of our knowledge, this work is the first Gaussian splatting inpainting method to operate in the learned representation space of a 3D-native generative prior without relying on inpainted reference views.

[CV-49] Pixel-Level Transformers in Remote Sensing: A Canopy Height Case Study

链接: https://arxiv.org/abs/2609.37809
作者: Sven Ligensa,Jan Pauls,Karsten Schrödter,Ibrahim Fayad,Fabian Gieseke
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Accepted at ACM SIGSPATIAL 2026

点击查看摘要

Abstract:Predicting canopy height from medium-resolution satellite imagery is a common and scalable approach for assessing the condition of the world’s forests, which play a crucial role in climate change mitigation. While Transformer-based architectures have shown strong performance in many domains, their straightforward application to dense (i.e., pixel-level) regression tasks often yields suboptimal results. In particular, the patch size has a crucial impact on the model performance. In this work, we consider pixel-level attention schemes and show that the resulting models generally outperform those relying on larger patch sizes. However, pixel-level attention can be a prohibitively resource-intensive operation. For this reason, we conduct an extensive experimental study using efficient attention variants to identify favorable trade-offs between prediction quality and resource requirements, facilitating the practical deployment of the proposed models. In addition, we perform a comprehensive comparison with several well-established models in the field and show that, with suitable hyperparameter choices, Transformer-based architectures can outperform competing approaches. Our findings provide practical guidance for designing models for pixel-level regression tasks on medium-resolution satellite imagery, including canopy height and biomass estimation, soil moisture mapping, and yield forecasting.

[CV-50] ByteTraX: Enhancing the ByteTrack Architecture with Optimised Thresholding

链接: https://arxiv.org/abs/2609.37801
作者: Thomas A. O’Shea-Wheller
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:The ByteTrack algorithm is a widely used and computationally efficient multi-object tracking architecture. Its core innovation lies in the combination of lenient bounding box associations with tracklet similarity matching to robustly deal with object occlusions. However, this strategy is nevertheless vulnerable to erroneous track reclassification and identity switching, as detection confidence scores dictate association priority. To address this, I present a simple enhancement of the ByteTrack architecture, named ByteTraX, that optimises track continuity via a single unified matching threshold, while penalising identity switches through stringent track initiation criteria. This approach achieves consistently improved performance across a range of diverse benchmarks including GMOT-40, LC-MOT, SportsMOT, TeamTrack, DAMUNT, and DeepSea-MOT, while simultaneously increasing processing speed by 10%. Specifically, results demonstrate a 40% reduction in identity switches, accompanied by mean increases in HOTA of 3.6, IDF1 of 5.6, and FPS of 6.3. As such, adoption of the ByteTraX algorithm has the potential to substantially enhance tracking performance over the ByteTrack baseline, while retaining the efficiency needed for real-time deployment. To facilitate usage, I provide the source code, integration functionality for the YOLO family of object detection models, and deployment instructions via an open source repository.

[CV-51] CHOQOLATE: Organizing Concept Bottleneck Latent Spaces with Choquet Integrals

链接: https://arxiv.org/abs/2609.37786
作者: Rémi Kazmierczak,Johanne Cohen,Marianne Clausel
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Concept Bottleneck Models (CBMs) built on vision-language models such as CLIP represent a latent space as human-understandable concepts. These representations are unfaithful: related concepts are entangled, so individual scores do not reflect their intended meaning. We propose CHOQOLATE, an interpretable-by-design layer based on 2-additive Choquet integrals, which merges correlated concepts into compact nodes. Across four datasets, CHOQOLATE achieves a favorable accuracy-interpretability trade-off, with weight-sparse and semantically coherent nodes. A closed-form gradient derivation, backed by experiments, explains why Choquet layers drive this organization without explicit supervision. Choquet weights also map directly to Shapley values, which enables test-time intervention. On standard bias-mitigation benchmarks, suppressing spurious concepts after training performs on par with methods that require group annotations or retraining, while needing neither.

[CV-52] Planetary Feature Fields are Scalable Earth Representations

链接: https://arxiv.org/abs/2609.37784
作者: Arjun Rao,Sebastian Loeschcke,Anthony Fuller,Isaac Corley,Nico Lang,Evan Shelhamer
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 28 pages, 16 figures, 7 tables

点击查看摘要

Abstract:Satellite observations, precomputed embeddings, and map products describe the same evolving Earth, yet are stored as independent, petabyte-scale data products. Their continued growth calls for compact representations of multiple products while preserving spatial and temporal detail. We introduce Planetary Feature Fields (PFFs), which exploit redundancy across data products by modeling them jointly as continuous functions of space and time at planetary scale. PFFs are spatially local explicit-implicit (hybrid) neural fields. Each field shares a factored feature volume—a decomposition of an explicit 3D grid with smaller factors—across products, while lightweight implicit decoders reconstruct individual products across multiple timesteps. PFFs reconstruct EO products over space and time more accurately than single-product fields at matched compression rates. At 1800\times compression relative to the uncompressed source data, reconstructed features retain approximately 90% or more of the performance achieved with the original features on pixel-level segmentation, change detection, and patch-level classification tasks. PFFs can add new timesteps by extending their factored feature volumes and add new products by attaching new decoders, while leaving existing outputs unchanged. PFFs reduce end-to-end feature access latency by an order of magnitude relative to evaluated API and cloud-storage pipelines.

[CV-53] A Benchmark Dataset for Detecting AI-Manipulated Visual Evidence in the Court System

链接: https://arxiv.org/abs/2609.37783
作者: Kelly McConvey,Sajad Ebrahimi,Nima Jamali,Jalehsadat Mahdavimoghaddam,Matina Mahdizadeh Sani,Maksym Taranukhin,Wentao Zhang,Jacquelyn Burkell,Yuntian Deng,Karen Eltis,Maura R. Grossman,Vered Shwartz,Ebrahim Bagheri
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Photographic evidence is becoming increasingly vulnerable to forms of alteration and fabrication that existing legal and technical workflows are not well equipped to evaluate. Surveillance frames, dashcam stills, and phone photographs may be used to establish presence, sequence, causation, damage, or identity, yet contemporary generative systems allow non-experts to alter or fabricate such images through ordinary prompt-based interfaces. Existing image-forensics benchmarks provide important resources for face manipulation, classical tampering, and general synthetic-image detection, but they are not organized around the forms of visual evidence submitted in courts, the localized edits that can change what an exhibit appears to prove, or the consumer-tool threat model now facing the justice system. We introduce the CIFAR Synthetic Evidence Corpus for Detecting AI-Manipulated Images, a benchmark for evidentiary image authentication in court and justice-system contexts. The corpus contains 1,505 photographic items, including 720 authentic controls and 785 manipulated or fabricated images, spanning surveillance, dashcam, and consumer-photo imagery. Manipulations are organized into scene-condition edits, localized element edits, and full fabrications produced with contemporary generative systems. Each item is released with structured metadata covering source provenance, manipulation tier, subtype, generator, prompt template, and scene attributes, enabling controlled evaluation beyond aggregate binary detection. We also establish baselines with publicly available image-manipulation detectors, showing that current systems exhibit error profiles that remain problematic for evidentiary use. The dataset, prompts, metadata, code, and baseline evaluation scripts are released to support research on visual evidence authentication, information integrity, and trustworthy AI for the justice system.

[CV-54] HiRAE: Hierarchical Representation Autoencoding with Residual Budgets

链接: https://arxiv.org/abs/2609.37775
作者: Xuanyu Zhu,Yan Bai,Yang Shi,Yihang Lou,Yuanxing Zhang,Tengfei Liu,Jing Jin,Yuan Zhou
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Pretrained visual representations support image generation, but may not fully preserve the fine-grained details needed for faithful reconstruction. Meanwhile, intermediate encoder layers contain complementary visual details, but learning to fuse them for reconstruction can produce a latent distribution that is difficult to model. Existing fusion methods require empirical tuning of layer selection or staged optimization of fusion and decoding, increasing configuration effort or training complexity. We introduce HiRAE (Hierarchical Representation Autoencoder), which learns a hierarchical fusion framework over the full encoder hierarchy to improve reconstruction fidelity while maintaining compatibility with generative modeling. HiRAE groups encoder layers by depth and learns residual corrections to the deepest representation. Group-wise norm caps bound these corrections relative to the deep anchor, with tighter budgets for shallower groups. Our HiRAE-24 preserves the latent token count and channel dimension. On ImageNet-256, HiRAE-24 reduces reconstruction FID from 0.299 to 0.209 relative to RAEv2 while maintaining competitive guided generation quality. For text-to-image generation, HiRAE-24 improves alignment over RAEv2 on GenEval, DPG-Bench, and GenAI-Bench both before and after supervised fine-tuning. Under the same generator-training and evaluation protocol, post-fine-tuning GenEval increases from 84.86 to 87.70.

[CV-55] Selective Channel Restoration for Backdoored Vision-Language Models

链接: https://arxiv.org/abs/2609.37759
作者: Shuming Liu,Zhifang Zhang,Suqin Yuan,Khin Mi Mi Aung,Zhuoyi Lin,Lei Feng
类目: Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
备注: 14 pages, 4 figures

点击查看摘要

Abstract:Vision-language models (VLMs) exhibit strong multimodal capabilities but remain vulnerable to backdoors implanted through poisoned fine-tuning data. Existing defenses often require extensive parameter updates during fine-tuning or incur per-query overhead during inference. To address these limitations, we propose Perturb-Select-Restore (PSR), a post-training defense that performs sparse updates to the projection interface and introduces no additional computation during inference. We reveal that backdoored VLM projectors are substantially more sensitive to bounded perturbations than clean VLM projectors, a phenomenon we term projection fragility. Building on this finding, PSR identifies the output channels most sensitive to perturbations in each projection layer of a backdoored VLM and restores their parameters to the corresponding pretrained values. Experiments across multiple tasks show that PSR reduces attack success rates to near zero while preserving clean-task performance.

[CV-56] Multi-Site Real-World Performance of Commercial AI for Pulmonary and Incidental Pulmonary Embolism Detection

链接: https://arxiv.org/abs/2609.37750
作者: Aawez Mansuri,Mohammadreza Chavoshi,Theodorus Dapamede,Wasif Bala,Beatrice Brown-Mulry,Rohan Isaac,Bardia Khosravi,Hanzhou Li,Frank Li,John T. Moon,Chad Robichaux,Dan I.G. Cohen-Addad,Ninad V. Salastekar,Janice Newsome,Judy W. Gichoya,Hari Trivedi
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Pulmonary embolism (PE) is a leading cause of cardiovascular mortality, yet the real-world performance of FDA-cleared AI detection models remains incompletely characterized. We retrospectively evaluated two FDA-cleared AI algorithms from a single commercial platform (Aidoc Medical BriefCase), one for PE triage on dedicated CT pulmonary angiography (CTPA; n = 30,678) and one for incidental PE (iPE) detection on routine contrast-enhanced CTs (n = 37,191), across a 17-facility academic health system. Reference-standard labels were extracted from radiology reports using a validated LLM pipeline (97% accuracy, kappa = 0.94). The PE model achieved 86.8% sensitivity and 99.1% specificity, with sensitivity declining from 99.3% for saddle emboli to 72.9% for subsegmental PE, and from 89.7% for acute to 65.3% for non-acute PE. The iPE model achieved 73.5% sensitivity and 99.8% specificity. Both models demonstrated lower sensitivity than FDA-clearance benchmarks while exceeding cleared specificity, with diminishing performance for peripheral and non-acute emboli mirroring known human reader limitations and underscoring the need for standardized post-market surveillance of AI-enabled medical devices.

[CV-57] he Camera Inside the Editor: Reading the Implicit Camera of Image Editors with Painted Calibration Patterns

链接: https://arxiv.org/abs/2609.37732
作者: Sebastian Rückerl
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 23 pages, 13 figures, 8 tables

点击查看摘要

Abstract:Instruction-based image editors insert objects, restyle scenes and render new viewpoints, but it is unknown which camera they assume when they paint into a photograph. Asked to cover the floor with a checkerboard, an editor paints projective structure from which classical vanishing-point geometry reads pitch, roll, focal length, yaw and, on renders, the principal point, without any training. Unlike a calibrator such as GeoCalib, which estimates the camera of an image, this isolates the camera under which the editor paints. On 120 rendered cameras with exact ground truth, Qwen-Image-Edit-2511 paints tile edges that meet their vanishing points within 0.26 degrees, and its implicit camera matches the true one to 0.8 degrees in pitch and 6% in focal length, more accurately than GeoCalib except in roll. Asked to draw the horizon or mark a vanishing point instead, the editor fails, so this knowledge is revealed by painting and not by the explicit tasks we tried. The implicit camera has two priors: roll is pulled towards level (slope 0.71), and telephoto perspective towards a default of about 30 mm, which roughly matches the camera the models paint without any scene. For Qwen, the priors do not grow when blur removes four fifths of the line evidence. They are stronger on real photographs, and on NYUv2 a shorter wording of the task removes the difference for roll. On photographs from a 24–240 mm zoom lens the painted perspective grows with only 0.62 of the lens’s slope, while GeoCalib and MoGe-2 saturate at about 52 and 42 mm. FLUX.1 Kontext and LongCat-Image-Edit are pulled much harder. Finally, from a level camera a camera-control LoRA executes pose commands at only 50–70% of their strength, and a board painted into its output agrees with the camera it produced.

[CV-58] CogWAM: Aligning Semantic Cognition with World Action Modeling via Event-Driven Interfaces

链接: https://arxiv.org/abs/2609.37721
作者: Sen Wang,Liu Liu,Xinjiang Wang,Zequn Chen,Haoyi Jiang,Taojun Ding,Tingyang Xiao,Zhizhong Su,Jie Wang,Sanping Zhou
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Robot policies increasingly incorporate semantic reasoning and future-world prediction, yet combining these capabilities does not guarantee that local predictions and actions remain aligned with task progress. We introduce CogWAM, a cognition-guided world-action model that establishes an explicit semantic interface between task reasoning and world-action learning through a persistent Semantic State, which stores completed task events and the active subtask. CogWAM updates this state only when observations indicate semantic transitions, allowing task-level context to persist across multiple action chunks. To bridge semantic context with physical prediction and control, CogWAM employs progress-conditioned WORLD and ACTION queries that selectively extract task-relevant information for future-world prediction and action generation. During training, the Semantic State provides shared task-progress context for both branches, while inference removes the future-prediction branch and directly generates actions from observations and the maintained state. We further introduce semantic training strategies to improve transition learning and closed-loop conditioning. Without additional robot-action pretraining, CogWAM achieves 15.56 / 11.70 % Score/SR on RoboDojo and state-of-the-art performance on BiCoord, while real-world experiments demonstrate closed-loop dual-arm manipulation with 16.4 fewer Semantic State regenerations than step-wise updating.

[CV-59] PolyOCR-Venus: Unified OCR Foundation Models for Text-Centric Visual Intelligence

链接: https://arxiv.org/abs/2609.37712
作者: GuangJian Team:Kaili Huang,Yongshuo Zhang,Bingtao Fu,Changjiang Jiang,Chenfan Qu,Chenfeng Zhang,Fangming Cui,Gaoyang Zhang,Jiangwei Xie,Jianshu Li,Jing Huang,Jingwen Bai,Mingqi Fang,Tao Fang,Weihong Zhang,Wenbo Du,Xiongfei Bai,Xuekang Zhu,Yinan Xia,Zhenming Wang,Jian Liu,Jingjing Liu,Xiang Qi,Weiqiang Wang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Technical Report

点击查看摘要

Abstract:Optical Character Recognition (OCR) is evolving from plain-text transcription toward general visual intelligence, requiring models to recognize, localize, and reason over textual information in complex visual environments. However, existing OCR systems often excel at only some tasks and struggle to balance recognition, parsing, and reasoning across scenarios. In this report, we present PolyOCR, a family of unified OCR foundation models of varying scales. PolyOCR combines a shared instruction-following framework with a large-scale data engine that converts heterogeneous visual resources into quality-verified OCR supervision. We introduce Competence-Guided Policy Optimization, which combines verifier-based Group Relative Policy Optimization with on-policy distillation through sample-wise routing based on teacher reliability and the teacher–student competence gap. We also introduce OCRBench v2.1, our revision of OCRBench v2 with manually verified annotation corrections and task-aligned scoring metrics. Extensive experiments across OCRBench v2.1, CC-OCR, in-house KIE Benchmark, OmniDocBench v1.6 and MDPBench demonstrate that PolyOCR achieves state-of-the-art or highly competitive performance.

[CV-60] VIF-Bench: Evaluating Visual Instruction Following in Multi-Reference Image Generation

链接: https://arxiv.org/abs/2609.37709
作者: Yuta Oshima,Masakazu Yoshimura,Masahiro Suzuki,Yutaka Matsuo,Hiroki Furuta
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Code: this https URL , Benchmark: this https URL

点击查看摘要

Abstract:Recent multimodal image generation models can take multiple images and textual instructions as input, enabling reference-based generation guided not only by text but also by visual instructions such as layouts, arrows, and pose cues. However, existing benchmarks do not evaluate the joint setting in which multiple references must be composed under multiple and heterogeneous visual-instruction images. To address this gap, we introduce VIF-Bench, a benchmark of 1,241 tasks designed to assess the edge of model capabilities in this joint setting by covering: (i) multi-reference generation (up to 7) under multiple heterogeneous visual instructions (up to 6), (ii) cases where reference images can potentially compete with visual instructions (e.g., a strongly posed subject vs. a target pose), and (iii) controlled comparison of visual instructions with text descriptions at different levels of specificity. Using these capabilities, we uncover three findings: (1) models face an adherence-artifact trade-off: once models reach stronger visual instruction adherence, stronger adherence tends to coincide with more instruction artifacts in generated images, (2) visual instruction adherence tends to be lower on tasks whose reference images carry a salient state of the controlled attribute (e.g., a neon-lit subject under a light-direction instruction), most consistently for light and wind, and (3) for models that can understand visual instructions, it is often better to provide visual constraints directly rather than describe them in text; when using text, a moderate level of detail works better than an exhaustive description. VIF-Bench is released as an open benchmark to establish a basis for fair comparison in controllable multi-reference image generation.

[CV-61] Honeycomb: Constant-Size Scene Memory Representation for Video World Models

链接: https://arxiv.org/abs/2609.37690
作者: Jack Wei Lun Shi,Kaichen Zhou,Haoyu Chen,Yufeng Weng,Keane Ong,Ruojin Cai,Hang Hua,Justin K. W. Yeoh,Mengyu Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project Page: this https URL Code: this https URL

点击查看摘要

Abstract:Video world models require persistent scene memory to maintain consistency during long-horizon video generation. Existing spatial memory systems accumulate RGB observations or latent features, causing storage requirements to grow as generation proceeds. We introduce Honeycomb, a video world model built on HexMemory, a compact low-rank representation that stores scene features in a fixed-size memory comprising six spatial and spatiotemporal planes. A feed-forward writer maps each newly generated video chunk to plane features. As the spatial coverage or temporal range expands, HexMemory warps the existing planes while preserving their dimensions, then integrates new features through confidence-weighted pooling and a learned residual correction. A reader retrieves latent features from HexMemory to condition subsequent video generation. Because the writer processes only observations from the latest chunk, Honeycomb avoids per-scene optimization and repeated processing of the full generation history. Experiments on WorldScore and RealEstate10K demonstrate strong video generation quality and robust consistency when revisiting previously observed regions, while maintaining constant feature-storage requirements throughout generation. Code and additional visualizations are available on our this https URL.

[CV-62] PAIQ: Patch-Aligned Semantic Injection via Residual Rotation

链接: https://arxiv.org/abs/2609.37685
作者: Pinze Ren,Yuwei Zhang,Hao Chen,Linghao Meng,Chang Li,Qiankun Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Language-aligned and self-supervised visual encoders offer complementary strengths in semantic abstraction and spatial detail. Harnessing this complementarity requires enriching local features while retaining distinctions between semantically related patches. We introduce PAIQ, a patch-aligned semantic injection framework that combines content-based cross-encoder matching with orthogonally constrained residual updates. Using DINOv3 patch features as the spatial base, PAIQ aggregates complementary SigLIP features through joint source allocation and injects the aggregate–base differences through a shared orthogonal transformation Q. This rotation adapts update directions while preserving residual norms and pairwise angles. For fixed projected features, we derive conditions for patch separability under similar semantic aggregates and show that rotation adds a nonnegative separation term over direct interpolation when the aggregate is shared. Only the projection and fusion parameters are trained; both visual encoders and the language model remain frozen, and fusion retains 196 visual tokens. Across diverse language backbones, PAIQ yields broad gains in judge-assessed correctness and reductions in hallucination severity over single-encoder interfaces on image description and visual question answering. On the 2B and 9B Qwen backbones, this compact interface outperforms the strongest evaluated fusion or token-compression baselines by about 2.9 correctness points on average.

[CV-63] Med-RADIO: Reducing All Medical Domains Into One via Multi-Teacher Distillation

链接: https://arxiv.org/abs/2609.37682
作者: Chu Zhang,Haoyu Jiang,Hongyuan Zhang,Hongbin Liu,Dong Yi
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:The rapid expansion of large-scale medical datasets and computational resources has driven significant progress in medical foundation models. Given the inherent heterogeneity of medical imaging modalities, current research mainly follows two paths: specialized models optimized for specific modalities, and generalist models designed to handle multiple modalities. However, medical generalist models suffer from both insufficient training data scale relative to natural image generalists and inadequate domain-specific depth relative to medical specialists. Empirically, generalist models establish a cross-modality performance baseline, while specialists define the performance ceiling within their respective domains. To elevate this baseline toward these ceilings, we propose Med-RADIO, a medical multi-teacher distillation framework that Reduces All Domains Into One by compressing complementary expertise from multiple domain-specific teachers into a unified medical vision foundation model. Our method curates both generalist and specialist teachers, allocates modality-aligned distillation streams to reorganize generalist pretraining data so it matches specialist domains, and uses a balanced loss to prevent any single teacher from dominating the distillation process. On internal and external classification benchmarks spanning five modalities, Med-RADIO improves over strong medical generalists under linear probing and remains competitive with representative specialists on most evaluated modalities. Code is available at this https URL.

[CV-64] MeanFlowAdvantage: Stable Reward Fine-Tuning for Few-Step Averag e-Velocity Generators

链接: https://arxiv.org/abs/2609.37670
作者: Haocheng Tang,Tianchi Xie,Xingqiao Lin
类目: Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:MeanFlow enables efficient few-step generation by predicting interval-average velocities, but this representation creates a mismatch for reward fine-tuning: existing advantage-based objectives are typically defined on instantaneous velocities or equivalent x_0 -space predictions, whereas inference directly uses the learned average-velocity map. We introduce MeanFlowAdvantage, a signed advantage-weighted least-squares objective for average-velocity generators. Our key construction uses a shared, detached MeanFlow derivative correction to express the reward objective in prediction space while making rollout and reference regularization exact penalties on the average-velocity network deployed at inference. The resulting formulation preserves MeanFlow’s native few-step sampler and provides a direct mechanism for transferring reward improvements to the deployed flow map. On SD3.5-Medium, MeanFlowAdvantage improves all eight reported metrics over the matched four-step MeanFlowNFT baseline and, with only four NFEs, matches or exceeds the 40-step DiffusionNFT baseline on six of eight metrics. The same objective also transfers to DNA promoter design, where it supports both teacher-free on-policy RL for a generator defined on a manifold and teacher-guided reward-graded distillation, with the latter yielding the lowest one-step Sei profile MSE among the compared configurations.

[CV-65] Are In-Context Images Worth 10 Dimensions?

链接: https://arxiv.org/abs/2609.37659
作者: Adhemar de Senneville,Xavier Bou,Jérémy Anger,Rafael Grompone,Gabriele Facciolo
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:There has been significant work on understanding the In-Context Learning capabilities of Large Language Models, especially on the induction circuit. For a few-shot classification task, the induction circuit leverages linear representations of each labeled example in-context in order to classify an unlabeled query. However, few works focus on how those linear representations are built in the first place. Leveraging the expressivity of the vision modality compared to text, we uncover a Shared Discriminative Geometry (SDG) inside Large Vision Language Models (LVLMs). It is a low-dimensional space, shared across all image classification tasks, in which in-context images are compressed into linearly separable representations later used to perform classification. We observe that this is the result of the model performing a dimensionality reduction of vision representations in early layers. In order to explain this phenomenon: (1) We show analytically that linear self-attention can perform a dimensionality reduction by projecting in-context data onto its principal components, with each layer implementing one gradient descent step toward this objective. (2) We provide evidence that trained LVLMs reduce the dimensionality of vision representations in early layers via a similar mechanism.

[CV-66] racing the Evidence: Faithful Token Attribution Through Vision-Language Reasoning

链接: https://arxiv.org/abs/2609.37656
作者: Bowen Yuan,Danny Wang,Ruihong Qiu,Zijian Wang,Zi Huang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Large vision-language models (LVLMs) exhibit strong reasoning capabilities, yet the visual and textual evidence supporting the generated responses remains difficult to identify. Faithful token attribution explains an LVLM’s response by assigning scores that rank image and prompt tokens by how much the model relies on them, such that removing higher-ranked tokens causes the likelihood of the generated response to drop more rapidly. However, existing token-attribution methods have been developed mainly for text-based language models, and our empirical study reveals two challenges when complex multimodal sources are involved. First, the joint image-text attribution can underrepresent visual evidence relative to text, obscuring the image regions supporting the response. Second, visual evidence may influence the generated response through multiple intermediate reasoning paths, while existing methods trace only a limited subset of these paths, causing important visual contributions to be underestimated. Motivated by these insights, we introduce VTrace, a multimodal token-attribution framework that traces input contributions through intermediate reasoning and calibrates attribution scores across modalities. VTrace constructs pairwise attributions that highlight token-specific contributions and aggregates all forward attribution paths in closed form to account for both direct and indirect contributions. Cross-modal calibration then rescales image and text attribution scores using modality contributions estimated from response-likelihood changes, enabling a unified ranking of input tokens. Evaluations against seven baselines across six visual reasoning benchmarks demonstrate the superior attribution faithfulness. Project page: this https URL.

[CV-67] Exemplar2VQA: A Scalable Exemplar-Driven Visual Question Answering Generation Framework via Multi-Agent Coding NEURIPS2026

链接: https://arxiv.org/abs/2609.37655
作者: Jiayu Ying,Qijian Tian,Ruijie Xu,Xinnan Zhu,Daoguo Dong,Jiachen Xu,Xin Tan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to NeurIPS 2026. 33 pages, 11 figures, 11 tables

点击查看摘要

Abstract:Advancing spatial intelligence in Multimodal Large Language Models (MLLMs) is bottlenecked by the scarcity of complex, scalable 3D question-answer (QA) data. While manual annotation is labor-intensive, directly utilizing LLMs to synthesize these QA pairs often fails due to their inherent deficiencies in spatial and geometric computation. We introduce Exemplar2VQA, a scalable exemplar-driven visual question answering generation framework that rapidly synthesizes large-scale spatial QA pairs in simulated environments via multi-agent coding. By equipping collaborative agents with a meticulously designed library of geometric utilities, Exemplar2VQA bypasses LLMs’ spatial reasoning flaws through deterministic code execution. Crucially, the framework exhibits remarkable versatility: taking diverse static object-centric spatial query templates as exemplars, it seamlessly and autonomously scales them into massive, high-fidelity synthetic datasets. Fine-tuning Qwen2.5-VL (3B/7B) exclusively on Exemplar2VQA-generated synthetic indoor data yields significant performance improvements across various diverse benchmarks. Furthermore, its effectiveness is not limited to in-domain indoor datasets but also robustly extends to outdoor and mixed-scene benchmarks. These results establish Exemplar2VQA as a scalable and powerful paradigm for bridging the sim-to-real gap in Embodied AI. Our code is at this https URL

[CV-68] xture Space Material Diffusion

链接: https://arxiv.org/abs/2609.37654
作者: Jacob Munkberg,Peter Kocsis,Jon Hasselgren
类目: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
备注:

点击查看摘要

Abstract:We present a method for generating high quality materials for 3D objects entirely in texture space. We finetune a video diffusion transformer for text-guided material generation, multi-view material generation, and material upscaling. Our key insight is to use the known projection from image space to texture space, enabling the diffusion process to generalize across arbitrary geometries and texture parameterizations. This approach also avoids the view consistency issues inherent in video and multi-view diffusion models. Because texture space is two dimensional, we can reuse the strong priors of pretrained video diffusion models. We apply our method to high quality material reconstruction from posed photos captured under unknown lighting, as well as to text- and image guided material generation. Our method can scale to high resolutions (8K), 100+ input views, and neural material representations. In quantitative and qualitative evaluations we show state-of-the-art results for material generation and reconstruction.

[CV-69] VoxelSage: Tool-Augmented 3D CT Analysis and Simulator-Shielded Sequential Resection Planning for Liver Tumors

链接: https://arxiv.org/abs/2609.37648
作者: Binghong Qian,Xuanhe Liu,Yifan Xing,Wenjie Deng,Jian Wu,Haochao Ying
类目: Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)
备注: 21 pages, 10 figures. Technical report. Code at this https URL

点击查看摘要

Abstract:Preoperative liver-tumor assessment requires segmentation, physical-space measurement, visual evidence, and resection planning from the same three-dimensional CT volume. Existing tools often handle these steps separately, while language models cannot reliably compute physical measurements from CT. To provide an integrated workflow, we present VoxelSage, a multi-modal system for two- and three-dimensional visualization, liver-tumor analysis, and preoperative resection planning. Its dual-port architecture separates language-model orchestration from image computation: Port A interprets requests and selects skills, while Port B applies them to CT volumes and segmentation masks and returns structured results. Keeping physical measurements in Port B prevents the LLM from computing them directly and reduces the risk of fabricated numerical results. Eight built-in skills support quantitative analysis, visual evidence generation, three-dimensional reconstruction, segmentation refinement, and sequential resection planning; user-defined skills can extend these functions. For sequence planning, a behavior-cloned neural ranker orders candidate resection targets, while a simulator-based shield checks them against predefined constraints. Across 256 unseen simulator scenes, this approach reduced mean simulated time from 34.274 to 33.388 min (0.886 min, 2.59%) and mean simulated blood loss from 300.847 to 183.852 mL (116.995 mL, 38.89%) relative to a deterministic baseline. These results demonstrate system integration and simulator-level performance, not clinical efficacy or safety. The public implementation is available at this https URL.

[CV-70] argeted Visual Counterfactual Explanations for Contrastive Vision-Language Model

链接: https://arxiv.org/abs/2609.37638
作者: Van Bach Nguyen,Jörg Schlötterer,Christin Seifer
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Current explanation methods for contrastive vision–language models such as CLIP mainly identify important regions without showing how to change the input in order to get a target prediction. We introduce \textbfMask-guided \textbfAdaptive \textbfCounterfactual \textbfExplanations (\mace), a targeted visual counterfactual method designed specifically for CLIP zero-shot classification. \mace constructs an editable region from either source attribution or source–target attribution differences and expands the mask only when needed to reach a specified target class. A latent diffusion inpainting model then modifies the selected region, while a frozen CLIP model provides modification guidance and anchors the remaining image content to the original input. We evaluate \mace on ImageNet, Food-101, Oxford Pets, and CUB-200. The source-mask variant achieves the highest target top-1 success rate across all four datasets, while the difference-mask variant produces the smallest pixel-level and perceptual changes and the best realism scores. Both variants improve proximity and realism over a Stable Diffusion-only baseline using the same generative backbone. These results show that adaptive mask-guided editing produces effective CLIP counterfactuals. They further reveal a tradeoff between counterfactual validity and source-image preservation.

[CV-71] Procedural Core: A Compact Recurrent Initialization for Vision Transformers

链接: https://arxiv.org/abs/2609.37631
作者: Zachary Shinnick,Christian Internò,Hemanth Saratchandran,Anton van den Hengel,Damien Teney
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this http URL

点击查看摘要

Abstract:Transformers are typically trained from random initialization, requiring all their capabilities to emerge from large-scale optimization. Recent work showed that a small amount of abstract procedurally generated data can help acquire generic inductive structure at low cost. However, this adds a pretraining stage that must be repeated for every target model. We propose Procedural Core, an initialization strategy that captures this generic structure into a compact set of weights that can be reused across models. We train a minimal recurrent transformer on procedural data, then expand its weights to initialize transformers of arbitrary width and depth. The resulting initialization improves performance on image classification, self-supervised visual learning (DINO), and modeling natural language (FineWeb-Edu) and code (CodeParrot). For image classification, expanding a 1M-parameter core to initialize an 85M-parameter ViT-Base improves ImageNet top-1 accuracy by 2.2 pp over standard random initialization. Our analysis identifies recurrence as essential for learning compact weights that transfer across models. In ViTs, we localize a key benefit in the suppression of high-norm tokens that produces substantial improvements in zero-shot segmentation (ImageNet-S mAP 32.3 to 42.9), object localization (VOC07 CorLoc 9.9 to 18.4), and depth estimation (NYUv2 RMSE 1.104 to 0.998). This demonstrates that transformers need not start from a blank slate, and can be initialized with generic capabilities at low cost with no domain- or task-specific data.

[CV-72] omoTransformer: Towards a Foundation Model for CT Reconstruction

链接: https://arxiv.org/abs/2609.37605
作者: AmirEhsan Khorashadizadeh,Benjamín Béjar
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Supervised deep learning has advanced sparse-view tomographic reconstruction. However, conventional models, which typically map filtered back-projection (FBP) images or sinograms to clean reconstructions, are brittle under distribution shifts. Because they require retraining whenever projection counts and angles, detector resolutions, or data distributions change, their deployment in real-world applications remains limited. To address this, we introduce TomoTransformer, a transformer-based architecture that treats each \textitlocal filtered projection as an individual token and predicts missing views via self-attention. Crucially, TomoTransformer operates in a \emphback-projection space that separates projections across spatial locations, making view interpolation geometrically well-posed and invariant to detector size. This design yields a single foundation model that can process any number of input projections, at arbitrary angular locations and detector dimensions, and query any number of target angles without retraining. Trained on a large-scale dataset spanning diverse medical CT anatomies and natural images, TomoTransformer generalizes effectively across anatomies, materials, and resolutions. Extensive evaluations on several benchmark sparse-view datasets show that TomoTransformer significantly outperforms concurrent multi-purpose models like ViewTrans and matches or exceeds strong protocol-specific baselines, while remaining fully agnostic to the number of input and target projections. Furthermore, the model demonstrates robust zero-shot generalization on real experimental nanoscale brain data collected from an X-ray synchrotron, showcasing its practical utility for real-world applications.

[CV-73] When to Adapt: Multi-Signal Domain Shift Detection for Efficient Training-Free Adaptation in Open-Vocabulary Segmentation

链接: https://arxiv.org/abs/2609.37602
作者: Michele Antonazzi,Alejandra C. Hernandez,José Araujo,Olov Andersson,Patric Jensfelt
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Robust and reliable perception is essential for autonomous robots operating in real-world environments, particularly in long-term missions where environmental conditions may change significantly over time. Although recent advances in Visual Foundation Models (VFMs) have improved open-vocabulary semantic segmentation, these models can still suffer from domain shift, which can significantly degrade performance if they are not adapted to the current environment. Training-free domain adaptation is a relevant paradigm for adaptation, consisting of adjusting the model online using lightweight adapters. Recent approaches apply this on a per-frame basis, which is impractical for deployments on resource-constrained robotic hardware. To tackle this, we propose a multi-signal domain shift detection method for training-free continual test-time adaptation (TF-CTTA) in open-vocabulary segmentation. Our method leverages temporal coherence across consecutive frames by monitoring and combining complementary aspects of domain shift (visual change, adapter mismatch, and semantic drift) to trigger adaptation only when needed. We validate our approach on a benchmark including indoor and outdoor environments and using real robotic data. We demonstrate that our approach maintains segmentation accuracy while substantially reducing adaptations, making training-free adaptation practical and feasible for long-term, real-world robotic deployments.

[CV-74] FedSocket: Recipient-Executable Knowledge Exchange for Heterogeneous Multimodal Federated Learning

链接: https://arxiv.org/abs/2609.37582
作者: Xinyuan Zhao
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 17 figures

点击查看摘要

Abstract:Federated knowledge must remain usable by recipients with different modalities, private architectures, and tasks. We present FedSocket, which makes recipient execution a design requirement of the exchanged model. A shared Q combines recipient-computable inputs, task-owned outputs, and ownership-aware aggregation, connecting heterogeneous private models through a common prediction interface. Private models teach local Q copies; the returned Q supports local learning and Joint inference, with only Q parameters and counts exchanged. Across six datasets, FedSocket improves missing-modality recipient accuracy over Local by 14.44 and 15.51 percentage points on MELD and UCF-51. Under matched inference capacity, Joint exceeds independent ensembles by 11.06 points in UCF-51 accuracy and 4.87 points in mean bidirectional Flickr30k R@1. Joint also improves over Q alone on all four heterogeneous endpoints, demonstrating the value of combining local and exchanged predictions. Teacher controls, sharing-path interventions, and component factorials identify the roles of supervision, sharing, and deployment. FedSocket makes exchanged knowledge directly usable from federated training to recipient inference.

[CV-75] ReVS: Integrating Textual Relevance and Visual Saliency for Efficient Vision-Language Model Token Pruning

链接: https://arxiv.org/abs/2609.37581
作者: Jing Wang,Zhiping Wu,Dongdong Ren,Youfang Han,Wei Zhao,Wenbin Li
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Vision-Language Models (VLMs) excel at visual understanding and reasoning but often incur substantial inference costs due to the large number of visual tokens. Recent visual token pruning methods increasingly follow a two-stage paradigm: they first remove visually redundant tokens after the vision encoder and then discard tokens irrelevant to the textual query within the Large Language Model (LLM). However, since the first stage typically relies solely on vision-encoder saliency, it may prematurely eliminate query-relevant tokens, depriving the subsequent text-guided stage of critical visual evidence. Our empirical analysis shows that incorporating query guidance into first-stage pruning better preserves task-relevant evidence and consistently improves performance over vision-only saliency-based pruning. We further find that high-variance attention heads are more sensitive to the textual query and yield more discriminative text-to-vision attention signals for second-stage pruning. Motivated by these findings, we propose TReVS, a training-free framework that combines textual relevance with vision-encoder saliency for pre-LLM pruning and leverages high-variance attention heads to remove task-irrelevant tokens at shallow-to-intermediate layers of the LLM. On LLaVA-1.5-7B, TReVS retains 92.8% of the unpruned baseline performance while pruning 94.4% of visual tokens, outperforming prior state-of-the-art methods.

[CV-76] Evaluating the Evaluators: Diagnosing Large Multimodal Models for AI-Generated Image Assessment

链接: https://arxiv.org/abs/2609.37576
作者: Yu Zhao,Jiarui Wang,Huiyu Duan,Ye Zhao,Jutao Tang,Juntong Wang,Guangtao Zhai,Xiongkuo Min
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:With the rapid advancement of text-to-image (T2I) generation, robust evaluation becomes critical yet challenging, as traditional metrics fail to capture fine-grained alignment and generative artifacts. While large multimodal models (LMMs) are increasingly adopted as evaluators, existing benchmarks typically study semantic understanding, quality perception, and authenticity identification in isolation, while largely neglecting responsibility detection. This leaves a gap in unified and comprehensive validation. To bridge this gap, we introduce SQUARE-Bench, a comprehensive benchmark that systematically evaluates LMM capabilities as evaluators of AI-generated images across four aspects: Semantics, Quality, Authenticity, and Responsibility. SQUARE-Bench introduces a granular taxonomy of 38 sub-dimensions to evaluate nearly 10K AI-generated images sampled from 22 diverse models, ranging from legacy to state-of-the-art generators, complemented by over 3K real-world images. The images are annotated with curated question-answering pairs. Extensive experiments on 23 LMMs reveal that top proprietary models, such as Gemini-3-Pro, already outperform the individual human expert baseline. However, the performance gap between models remains significant, exhibiting notable disparities in fine-grained inference and domain-specific robustness. Beyond benchmarking, we conduct a proof-of-concept study of LMM-guided iterative editing, in which dimension-specific LMMs provide diagnostic feedback to fixed image editors. The resulting guided system yields selective improvements in semantics, authenticity, and responsibility, while exhibiting a consistent visual-quality trade-off. SQUARE-Bench can serve as both a diagnostic tool for characterizing LMM evaluator capabilities and studying their use in T2I generation refinement. The benchmark and dataset will be released upon publication.

[CV-77] Decompose Radicals Then Reward: Fine-Grained Inspection for Accurate Chinese Text Rendering

链接: https://arxiv.org/abs/2609.37569
作者: Yazhen Xie,Xingsong Ye,Zhineng Chen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Rendering accurate Chinese text remains challenging for text-to-image models. Existing OCR-based reinforcement-learning rewards compare decoded transcripts with target strings. Such rewards overlook the compositional nature of Chinese writing: an ideograph consists of reusable components arranged through explicit spatial relations, yet OCR evaluates it as an atomic character. Consequently, visually different radical-level errors may receive equally coarse feedback, encouraging glyphs that merely resemble the target instead of faithfully reproducing its internal structure. We employ Ideographic Description Sequences (IDS), which comprise spatial operators and character components, and train an expert IDS recognizer to transcribe rendered Chinese text into this representation. Building on this recognizer, we introduce IDSpect, which deterministically decomposes the target text into IDS tokens and aligns crop-level visual IDS predictions with the target sequence. Globally unique token credit makes this comparison robust to the order of detected text regions. Combined with a whole-character semantic reward, IDSpect supplies fine-grained credit with component and spatial-relation without changing the image generator or adding inference-time cost. Experiments with GRPO post-training of Qwen-Image demonstrate that IDSpect achieves leading structural quality and semantic alignment on LongText and GenTextEval.

[CV-78] APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants

链接: https://arxiv.org/abs/2609.37559
作者: Jianguo Huang,Jinming Liu,Qiyao Wang,Liang Xu,Jianhang Li,Zhimian Wen,Mingda Li,Shule Lu,Zhicheng Wang,Yuhan Guo,Xin Jin,Wenjun Zeng
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 33 pages, 11 figures, 15 tables

点击查看摘要

Abstract:To serve as real-world personal assistants, streaming video models need persistent memory that retains past experiences for later use. Yet existing streaming benchmarks and methods often focus on individual continuous videos or short clips, overlooking that real-world interactions are often intermittent and require memory to persist across interruptions. To fill this gap, we introduce APM-Bench, which reformulates real-world streaming interaction as multi-session life trajectories. It contains 549 sessions, 104 trajectories, and 2,719 candidates, spanning both objective and open-ended questions. Each session is a video with fine-grained annotations, and sessions within a trajectory revolve around related activities. Models then use persistent memory to answer questions about past sessions and provide proactive responses while maintaining real-time interaction. This raises challenges: persistent memory must be storable, selectively retain information, be injected at the right time, and remain efficient. Moreover, finite storage may leave required evidence unavailable, so assistants should recognize missing evidence. Therefore, we systematically evaluate general video models under different memory protocols and diverse specialized streaming memory systems, and test whether models acknowledge insufficient evidence. Our evaluation reveals a clear utility–latency–storage trade-off: existing methods still struggle to simultaneously achieve reliable long-term recall, low overhead, and effective proactive assistance across sessions. APM-Bench provides a comprehensive testbed for developing and comparing persistent memory systems under realistic streaming conditions. We hope it encourages future work that jointly considers utility, latency, and storage toward more practical persistent memory for real-world streaming assistants.

[CV-79] Weeding Out Bad Seeds: Initial-Noise-Robust Unlearning for Text-to-Image Diffusion Models

链接: https://arxiv.org/abs/2609.37537
作者: Arian Komaei Koma,Seyed Amir Kasaei,Aida Aryafar,Matin Ghiasi,Ali Aghayari,Amirhossein Souri,Mohammad Mosayyebi,AmirMahdi Sadeghzadeh,Mohammad Hossein Rohban
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Machine unlearning has emerged as a critical post-hoc safety measure to erase sensitive concepts from Text-to-Image (T2I) models without prohibitive retraining. However, we reveal that current state-of-the-art (SOTA) approaches are brittle due to a severe lack of robustness to noise initialization. We call this phenomenon ``probabilistic forgetting’': suppressed concepts re-emerge under specific random initial noise conditions, despite appearing unlearned on other initializations. We trace this failure to the misalignment between standard Gaussian sampling during unlearning and the unlearning objective. Since the target concept manifests only in specific initial noise regions throughout the unlearning phase, uniform random sampling yields sparse, uninformative gradient updates that fail to drive robust erasure. To overcome this issue, we propose an adaptive, concept-conditioned sampling strategy that dynamically concentrates gradient updates on regions where the target concept manifests, down-weighting uninformative areas. We integrate our framework with six distinct SOTA unlearning methods across four diffusion backbones and evaluate it across safety, object, and artistic-style unlearning, as well as under black-box and white-box adversarial attacks. Our method reduces the conditional nudity re-emergence rate across random initializations by 67.2% on average over four baselines and lowers attack success rates across both adversarial evaluations. Across concept domains, Adaptive Noise Sampling strengthens adversarial robustness and non-target retention while preserving competitive generative quality and target-erasure performance.

[CV-80] RawVLA: Embodied Neural Image Signal Processor For Robotic Manipulation

链接: https://arxiv.org/abs/2609.37530
作者: Shuhong Liu,Heng Zhou,Lingfeng Qian,Yuhao Fang,Xianbao Hou,Qianyu Zhou,Lin Gu,Wei Sui,Jianfei Yang,Ziteng Cui
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Vision-language-action (VLA) models typically operate on RGB images produced by a fixed camera image signal processor (ISP), leaving the imaging pipeline outside the learning and evaluation loop. We systematically examine the consequences of this overlooked design choice across five fundamental ISP dimensions: gain, sensor noise, chromatic response, tonal response, and bit depth. Our analysis reveals that RAW-to-RGB processing materially shapes both action prediction and manipulation success, with different ISP dimensions exerting substantially different effects. Guided by these findings, we introduce RawVLA, a streaming neural ISP that adaptively renders RAW observations for frozen VLA policies while concentrating its capacity on the imaging factors relevant to embodied behavior. We further present RawVLA-Bench, a RAW-domain manipulation benchmark to expose image processing as an explicit evaluation variable across clean and adverse acquisition conditions. Experiments on RawVLA-Bench show that RawVLA preserves performance under standard conditions while substantially improving robustness under degraded imaging, establishing adaptive RAW processing as an effective interface between physical cameras and embodied policies.

[CV-81] Principled MAP estimation for inverse problems: bridging the gap between convergence and performance

链接: https://arxiv.org/abs/2609.37529
作者: Alexandre Lagier,Valentine Tosel,Anne Gagneux,Mathurin Massias,Ségolène Martin
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Pretrained denoisers provide a powerful way to incorporate image priors into restoration algorithms. Plug-and-Play and RED approaches exploit fixed-noise-level denoisers within first-order optimization schemes, with convergence guarantees, but often struggle to achieve high-quality reconstruction on severely ill-posed inverse problems. In contrast, recent state-of-the-art approaches leverage denoisers derived from flow- or diffusion-based generative models and evaluate them along a sequence of decreasing noise levels. While these methods achieve strong empirical performance, their convergence theory remains limited. In this paper, we bridge this gap by specifically designing an algorithm that combines denoisers at decreasing noise levels with a schedule tailored to ensure convergence. From a Bayesian perspective, we prove that our method converges to a \textitMaximum a Posteriori (MAP) estimate, under suitable assumptions. Subsequently, we apply our method to various ill-posed inverse problems and show that it surpasses convergent methods while competing with state-of-the-art empirical ones.

[CV-82] GeoSET: Generalist Foundation Model for SAR-to-EO Image Translation

链接: https://arxiv.org/abs/2609.37496
作者: Jeonghyeok Do,Munchurl Kim
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Please visit our project page this https URL

点击查看摘要

Abstract:Paired synthetic aperture radar (SAR) and electro-optical (EO) imagery is increasingly available across sensors, resolutions, and geographic regions. Yet existing SAR-to-EO image translation (SET) methods are typically trained on a single, limited-scale dataset, producing models specialized to particular sensing conditions. We introduce GeoSET, the first generalist model for SET, built around a single pretrained parent that is adapted to downstream datasets under a common protocol. We curate over 3 million high-quality SAR–EO pairs from a collection of more than 10 million SAR observations, spanning diverse sensors, spatial resolutions, and ground sampling distances. To bridge the modality gap between SAR observations and a pretrained image generator, we develop a speckle-robust SAR encoder and pretrain the conditional generator on this heterogeneous corpus. The resulting parent supports efficient adaptation across downstream datasets through low-rank adaptation (LoRA), updating only 0.60% of the generator parameters and requiring approximately one hour per dataset. Across six downstream benchmarks, GeoSET achieves state-of-the-art results in FID and DISTS with full fine-tuning or LoRA, demonstrating effective transfer across heterogeneous SAR-EO domains.

[CV-83] MotionMaestro: Masked Tokenization for Unified Motion Generation

链接: https://arxiv.org/abs/2609.37495
作者: Yun Chen,Munchurl Kim,Jeonghyeok Do
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Please visit our project page this https URL

点击查看摘要

Abstract:Human motion generation plays an important role in applications such as character animation, virtual environments, and embodied interaction. While existing approaches have achieved remarkable progress, many of them are developed for individual tasks, including text-to-motion, pose-conditioned generation, and trajectory control. Although these tasks involve different types of conditions, a unified framework capable of handling them within a common representation would greatly simplify motion generation systems. We observe that diverse motion conditions can be naturally formulated as different observation patterns over motion sequences, where each task corresponds to a specific masking strategy. Based on this insight, we introduce MotionMaestro, a unified motion generation framework that learns a shared representation for complete motions and heterogeneous partial observations through masked motion tokenization. MotionMaestro employs a three-stage training strategy that first learns a masked motion tokenizer, then refines its reconstruction ability on clean motions, and finally trains a conditional flow-matching generator in the learned latent space. Furthermore, we introduce an observation map and an observation loss to explicitly preserve provided motion conditions during generation. With this unified representation and conditioning mechanism, MotionMaestro supports text-guided and unconditional synthesis, pose conditioning and partial completion, temporal interpolation, trajectory control, and motion continuation. Experiments on the large-scale RoMo and MotionMillion datasets show state-of-the-art performance across diverse motion generation tasks.

[CV-84] Attention-Scoped Guidance: Training-Free Spatial Control for Image Editing

链接: https://arxiv.org/abs/2609.37492
作者: Zeyan Li,Wei Zhou,Hadi Amirpour,Minghao Zou,Panqi Yang,Jianfeng Xu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Instruction-guided image editing should change what the instruction names and leave the rest of the image untouched. In dual classifier-free guidance (CFG), an editor combines two directions at every denoising step, one that pushes toward the instructed edit and one that pulls back toward the source image, using global weights. We introduce Attention-Scoped Guidance (ASG), a sampler wrapper that makes these weights spatial. It reads a soft support map from the instruction attention that the editor already computes, then weakens text guidance where support is low and strengthens image anchoring where support is high. The wrapper requires no training, no external mask, and no additional network evaluation. On the full MagicBrush and PIE-Bench++ splits, ASG improves preservation-oriented metrics, leading three of four MagicBrush metrics and PIE-Bench++ background PSNR. A dose-matched control that removes the spatial placement loses up to 0.73 CLIP on PIE-Bench++, confirming that the spatial allocation itself carries the gain.

[CV-85] Physics-Guided Flow-Map Matching for Precipitation Nowcasting ACCV2026

链接: https://arxiv.org/abs/2609.37487
作者: Shunya Nagashima,Takumi Bannai,Makoto Misaizu,Keisuke Maeda,Takahiro Ogawa,Miki Haseyama
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: Accepted by ACCV 2026

点击查看摘要

Abstract:Precipitation nowcasting, generating future radar fields from past observations, is critical for flood warning and disaster response. It is also a demanding benchmark for spatiotemporal generative modeling, with chaotic dynamics, heavy-tailed intensities, and rare high-intensity structures that matter most. Deterministic models minimize a pixel loss and are driven toward the conditional mean, which blurs exactly those structures, while generative models that add a stochastic residual on top of a deterministic backbone inherit the same blur. We propose Physics-Guided Flow-Map Matching (PG-FMM), a conditional flow-map model that decouples predictable advection from uncertain small-scale detail. A frozen Lagrangian advection prior transports the radar field and supplies an explicit motion forecast, and a flow-map generative head, conditioned on the past frames and the prior rollout rather than summed onto it, produces sharp stochastic detail in four sampling steps. The prior serves only as guidance, so the head replaces blurred structure instead of inheriting it. Extensive experiments on four radar benchmarks show that PG-FMM outperforms state-of-the-art methods on 18 of 24 metrics, with the largest gains at heavy-rain thresholds, where the critical success index improves by up to 58.9%. The project page can be found at this https URL.

[CV-86] PoE-Fuse: Precision-Weighted Expert Fusion for Bi-Temporal Change Understanding ACCV2026

链接: https://arxiv.org/abs/2609.37485
作者: Haruki Watase,Shunya Nagashima,Takayuki Nishimura
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: Accepted by ACCV 2026

点击查看摘要

Abstract:Bi-temporal change understanding, which localizes and characterizes what changed between two satellite images, is central to disaster response and environmental monitoring, spanning change detection, building localization, and damage assessment. Strong vision-language models address these tasks, but adapting them typically requires full fine-tuning or reinforcement learning, which is costly and unstable. We propose PoE-Fuse, a parameter-efficient framework that instead composes frozen foundation experts for geometry, grounding, and language, resampling their features onto a shared spatial grid and training only a lightweight fusion trunk. PoE-Fuse treats the aligned features as Gaussian observations of a latent scene state and fuses them by learned per-cell precision. This product-of-experts estimator strictly generalizes uniform summation and scalar gating, and extends to change fields by composing the precisions of the two timestamps. A single shared trunk solves the three tasks at once, reaching a mean F1 of 59.2%, compared with 40.7% for an instruction-tuned temporal vision-language assistant, and surpassing dedicated change-detection models retrained under the same protocol and training budget.

[CV-87] Label Less Learn More: Resource-Efficient Active Semi-Supervised Learning for Onboard Satellite Image Annotation

链接: https://arxiv.org/abs/2609.37481
作者: Ahmed Abdelnaby,Mohamed Elmahallawy,Marius Bernahrndt,Tobias Hecking
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Large-scale pervasive sensing increasingly relies on high-resolution satellite imagery, yet task-specific onboard vision is constrained by costly annotation and limited computation, memory, energy, and communication resources. Existing approaches largely rely on either data-hungry supervised learning or large vision-language foundation models, limiting efficient adaptation and deployment under these constraints. We present SatLabel, a resource-aware learning framework that transforms limited satellite labels into progressively refined onboard models through adaptive sample acquisition and semi-supervised model adaptation. Rather than repeatedly training on uniformly sampled labels, SatLabel closes the loop between model uncertainty, class imbalance, and pseudo-label quality to selectively acquire informative samples while exploiting abundant unlabeled imagery. This enables a compact student to adapt to target sensing domains with reduced annotation and inference costs. We further introduce an optional Mixture-of-Experts (MoE) student with graph-based feature refinement to enhance representation capacity while retaining a lightweight footprint. We evaluate SatLabel on 11 remote-sensing datasets spanning core, extended, and unseen domains against RemoteCLIP zero-shot inference. SatLabel improves Macro-F1 on most core and extended datasets while maintaining strong cross-dataset transfer to unseen domains. More importantly, the Balanced student contains only 11.2 M parameters and occupies approximately 42.8 MB, compared with 151.3M parameters and 577 MB for RemoteCLIP, while requiring 3.65 versus 5.89 GFLOPs. Across four efficiency benchmarks, it achieves approximately 2x higher GPU-forward throughput and reduces energy per image on datasets.

[CV-88] Learning Social Navigation from Internet Videos in the Policy State Space

链接: https://arxiv.org/abs/2609.37476
作者: Jiaming Wang,Duc Thang Nguyen,Jizhuo Chen,Volodymyr Shcherbyna,Diwen Liu,Zhengcheng Shen,Harold Soh
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 9 pages, 5 figures, 6 tables

点击查看摘要

Abstract:Training robust social-navigation policies requires simulators with diverse scene layouts, terrain, and human motion, but constructing such environments and specifying pedestrian behavior is costly. We propose an efficient pipeline that converts ordinary monocular walking videos directly into closed-loop social-navigation training environments in the policy’s state space. Our key observation is that local social navigation primarily depends on two types of information: where the robot can traverse and how nearby pedestrians move. We therefore represent the static scene as a metric traversability map, which can be rigidly transformed under counterfactual robot motion, while directly replaying the pedestrian trajectories recovered from the video over time. This abstraction allows us to define the forward dynamics directly in the policy’s state space and efficiently simulate counterfactual robot states without reconstructing or rendering photorealistic observations. The resulting policy achieves 81.2% success in the independent Arena benchmark, compared with 75.0% for the strongest baseline, and succeeds in 19/20 real-robot trials without policy fine-tuning.

[CV-89] Event-Only Wingbeat Counting under Camera Motion: A Controlled MuJoCo Benchmark

链接: https://arxiv.org/abs/2609.37465
作者: Zhang Nengbo
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 10 pages, 1 figure, 6 tables. Controlled simulation study with an ideal contrast-event sensor. Follow-up to arXiv:2609.17308 using new event-only acquisitions and a motion ablation protocol

点击查看摘要

Abstract:Counting completed wingbeats requires identifying individual cycles, including during frequency changes and pauses; estimating a dominant frequency alone is insufficient. Camera motion further mixes target and background brightness changes in event observations. We present a controlled MuJoCo benchmark that separates motion training from event-only image translation compensation. The acquisition contains 324 streams from 24 independent scenes, three flapping geometries, two distances (1.5 and 3.0 m), and static, moderate-motion and stronger-motion views. Fifteen scenes are used for fitting, three for validation and six for held-out testing. A fixed causal temporal convolutional network is evaluated in a matched 2 x 2 ablation with three initialization seeds and compared with ridge, Fourier, autocorrelation and an adapted EEPPR baseline. Under moderate motion, paired motion training reduces count mean absolute error from 31.130 to 3.185 cycles at 1.5 m and from 42.019 to 5.444 at 3.0 m. Adding the tested compensation increases these errors to 4.630 and 10.185, respectively. A Fourier baseline achieves 0.944 cycles at 1.5 m under moderate motion, showing that the neural model is not uniformly best. We report exact-count accuracy and temporally matched cycle F1 alongside count error. These findings support motion-aware training in this small synthetic benchmark, while exposing limits of simple event-background stabilization. They do not establish real-sensor performance, aerodynamic flight, or generalization to unseen vehicle types.

[CV-90] LazySloth: Bounded LLM -based Lazy Tree Search for Fast Long Video Comprehension

链接: https://arxiv.org/abs/2609.37426
作者: Arka Mukherjee,Kaleen Shrestha,Larissa Zhu,Maja Matarić
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Under review at conference. Preprints allowed when under review

点击查看摘要

Abstract:Modern vision-language models (VLMs) have shown promising results in long-video understanding due to the rich semantic information they can capture. However, most methods focus on coarse captioning of extracted image frames that are computationally inefficient and require models with large context windows. While past work has explored efficient methods through multimodal retrieval-augmented generation (RAG), they rely on lossy embeddings that lose temporal context and fine-grained detail. Few works to date have investigated how VLM-based query-relevant information retrieval can be optimized. We introduce LazySloth, an efficient tree-based search method that speeds up video comprehension and retrieval tasks 2.9-8.3x (compared to existing agentic methods) through bounded captioning of portions of the video considered irrelevant by a VLM of the video. Compared to contemporary specialized video-understanding VLMs and RAG-based methods, LazySloth achieved similar or better final task accuracy across two recent open-source base VLMs–Gemma 4 31B and Qwen3.6 27B–across four benchmarks. LazySloth reduced the gap between the base open-source model and a closed-source model, GPT-4o. Ablations showed that replacing VLM scene understanding with CLIP-based retrieval cost 8.8-19.9% in accuracy, while lazy tree construction matches eager construction at a fraction of the captioning cost. With LazySloth, we demonstrate the possibility of faster long-video comprehension without substantial loss in performance.

[CV-91] Complementary Retrieval-Augmented Prompting for Consistent Long-Form Video Generation

链接: https://arxiv.org/abs/2609.37407
作者: Xianghan Wei,Xiaoda Yang,Zhi Wang,An Pan,Daoan Zhang,Huayi Zhang,Yan Zhang,Wei Xu,Zishun Liao,Jianwen Lou
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:While recent video foundation models excel at generating high-quality short videos, long-form video generation remains a critical challenge, where a major bottleneck lies in conditioning independently generated shots to preserve consistent characters, scenes, and objects throughout a story. Existing training-free approaches typically condition target shots using retrieved historical visuals. However, these references often suffer from severe informational mismatch, either introducing irrelevant contextual redundancy or failing to provide the full combination of required elements for the target shot. To resolve this, we present Complementary Retrieval-Augmented Prompting, an agentic framework that strategically aggregates a compact set of mutually supportive historical references to achieve complete and targeted conditioning for long-form video generation without retraining or modifying the underlying generator. Specifically, our framework explicitly models the visual elements required by each target shot by parsing the narrative script into a text-grounded visual element registry that tracks characters, objects, scenes, and their shot-level states. A VLM-annotated keyframe library further maps these elements to past visual observations. Guided by the required elements, our agent retrieves complementary references that maximize target-element coverage while minimizing historical noise. Finally, the retrieved references, structured element states, and grounding instructions are assembled into a unified prompt for the frozen video generator. This element-aware process provides comprehensive conditioning while remaining fully interpretable. Quantitative and qualitative evaluations on multi-shot story generation demonstrate that our method consistently outperforms recent-frame, memory-based, and entity-level retrieval baselines in cross-shot consistency and text-controllability.

[CV-92] BeatDance: Generating Beat-Consistent 3D Dance with Hierarchical Spatial-Temporal Modeling

链接: https://arxiv.org/abs/2609.37400
作者: Xiaojian Shen,Dahu Shi,Jianrong Zhang,Hai Li,Hongwei Zhao,Dawei Zhang,Yunzhi Zhuge,Zhiliang Wu,Guanghui Yue,Wei Zhou
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Published in Pattern Recognition

点击查看摘要

Abstract:Generating realistic 3D dance from music is a challenging task that requires accurate synchronization with musical rhythms while capturing the spatial complexity of human motion. Although existing methods can generate physically plausible dance motions, they often struggle to achieve precise alignment with music, such as the beat. To address this limitation, we propose a novel diffusion-based framework, BeatDance, with two components: 1) We present a Hierarchical Decoupled Attention (HDA) module, which first disentangles the learning of human pose and temporal dynamics. A hierarchical structure is then employed to capture both short-term and long-term dependencies, thereby enhancing spatial-temporal modeling. 2) We adopt cycle-consistent learning by introducing an auxiliary dance-to-music module. During training, discrepancies between the reconstructed and original music induce a stronger loss signal, effectively encouraging the consistency property between the music and dance motion. Extensive experimental results demonstrate that our proposed approach outperforms recent competitive methods on two benchmark datasets.

[CV-93] Multi-task learning for the automatic grading of enlarged perivascular space burden using MRI

链接: https://arxiv.org/abs/2609.37387
作者: Jesse Phitidis,William N. Whiteley,Joanna M. Wardlaw,Miguel O. Bernabeu,Yajun Cheng,Xiaodi Liu,Junfang Zhang,Una Clancy,Stephen Makin,Roberto Duarte Coello,Susana Muñoz Maniega,Mark E. Bastin,Simon R. Cox,Maria del C. Valdés Hernández
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Enlarged perivascular spaces (PVS) visible in brain magnetic resonance imaging (MRI) are increasingly thought to be linked to poor brain health. PVS are elongated structures of less than 3 mm in diameter and can be numerous. To reflect the incidence of PVS, radiologists visually score their burden following a clinical grading scale - a task that would benefit from automation to accelerate analyses and overcome the influence of inter-observer differences. We developed and evaluated methods for training machine learning models to score PVS incidence in the basal ganglia (BG) and centrum semiovale (CSO) leveraging the Potters/Wardlaw scale. The novelty in our work lies in the use of imperfect, semi-automatically generated “silver-standard” PVS segmentation masks during training, in addition to PVS radiological scores. We comparatively evaluated a conditional convolutional neural network (CNN) which accepts PVS masks as an extra input channel, a multi-task CNN which performs both PVS segmentation and scoring, and a logistic regression model which utilises features derived from PVS masks to predict PVS scores. Multi-task learning was the most effective method, achieving a mean average precision of 64.08% compared to 60.22% for the conditional CNN, 52.11% for a baseline CNN trained only to predict PVS scores, and 49.32% for the logistic regression model. The multi-task model showed an ability to localise individual PVS not shown by the other CNNs, and behaved in a probabilistically sensible way, predicting with lower confidence on inherently harder classes. Age, sex, hypertension status, white matter hyperintensity volume, and ischaemic stroke lesion status were shown to be associated with the multi-task model’s PVS score predictions and the ground truth in a similar way.

[CV-94] Anatomy-Aware Prediction of Bronchoscopic Accessibility from 3D CT MICCAI2026

链接: https://arxiv.org/abs/2609.37386
作者: Linkai Peng,Cuiling Sun,Bin Wang,Jamie Rowell,Catherine Gao,Oyku Ikizgul,Eminenur Sentasci,Andrea Bejar,Halil Ertugrul Aktas,Gorkem Durak,Momen Wahidi,Christopher Kapp,Ulas Bagci
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted in MICCAI 2026

点击查看摘要

Abstract:Pre-operative planning for bronchoscopy is critical for the diagnosis of lung lesions. Current accessibility assessment relies on subjective manual inspection of CT scans, which is time-consuming and prone to inter-observer variability. In this paper, we formalize bronchoscopy accessibility prediction as a novel supervised learning task and present the first end-to-end framework to address it. We propose an Anatomy-Aware Mixture-of-Experts (MoE) model that integrates specialized modules: a CT Expert for local morphological features, a Lobe Expert for anatomical priors, and a Path Geometry Expert that encodes the sequential constraints of the bronchial tree. To support this task, we curated the first clinical dataset of 438 cases with pre-operative CT scans and documented procedural outcomes. Experimental results demonstrate that our method achieves an AUROC of 0.8052, significantly outperforming both state-of-the-art baselines and experienced human experts. This work establishes a new benchmark for computer-aided interventional planning in pulmonary medicine. Our data and code will be publicly available at this https URL.

[CV-95] Do-JEPA: From Masking to Intervention in Latent World Models

链接: https://arxiv.org/abs/2609.37378
作者: Hossein Resani,Javen Qinfeng Shi
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Latent world models are trained to predict what happens next, so nothing in their objective separates what an action caused from what merely co-occurred with it. Object-masking models such as C-JEPA intervene on what the predictor can see; we intervene on what physically happens. From one saved simulator state we run the dynamics under an action a and under a reference action a_\varnothing , and train the model to predict the difference \Delta z=z^a-z^a_\varnothing between the two latent futures. The resulting objective, Do-JEPA, has an effect loss, a support loss (where the action enters), a propagation loss (where its effect travels) and invariance losses (what must not change). In a synthetic system with object-aligned variables, support supervision finds the directly intervened object in 99.95% of test cases, where a sparse action mask sends the action to a nuisance slot in every case, and response-onset supervision recovers the ring-shaped propagation graph (edge AUROC 0.975 vs. 0.624). From pixels, the effect loss beats a control trained on exactly the same data: it lowers latent effect error by 28.4% on an end-to-end LeWM model and physical effect error by 13.5% when trained and tested on natural action sequences, and on three independently generated CausalWorld benchmarks it lowers responsive effect error by about 20% under physics shifts and the latent context sensitivity of predicted effects by 66%. Trained from scratch it costs factual accuracy; fine-tuning an existing model with it removes this cost. Together, these results show that intervening on the world, rather than on what the model sees, helps latent world models predict what their actions cause.

[CV-96] MG-Thinker: Bi-Axial Self-Reflection for Multi-Image Reasoning Grounding

链接: https://arxiv.org/abs/2609.37374
作者: Heyu Huang,Chi Chen,Zonghao Guo,Yuhua Li,Maosong Sun,Ruixuan Li
类目: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
备注:

点击查看摘要

Abstract:Reinforcement learning (RL) has recently delivered substantial gains in multimodal reasoning, opening a promising route for fine-grained visual perception. Yet for multi-image reasoning grounding (MRG), reasoning over real-world multi-image contexts toward pixel-precise localization, existing RL-based approaches overlook two characteristics intrinsic to this paradigm: a coarse-to-fine hierarchical reasoning pattern, and heterogeneously distributed task–sample difficulties. In this work, we present MG-Thinker, a post-training RL framework that advances a new MRG paradigm featuring such hierarchical reasoning, supported by a curated 25K MRG dataset with task-adaptive Chain-of-Thought (CoT) annotations that elicit multi-perspective evidence before conclusion. To remedy the heterogeneous task–sample difficulties, we further propose Bi-Axial DAPO (BiA-DAPO), which decomposes rollout advantages along an intra-group signal axis and an inter-group competence axis through two complementary mechanisms, both grounded on our defined candidate pool for stable group-level statistics. Extensive experiments show that MG-Thinker achieves state-of-the-art performance on multi-image reasoning grounding while consistently improving generalization across multi-image understanding and diverse multimodal benchmarks.

[CV-97] hink Before You Score: Thinking Reward Model for Visual Generation

链接: https://arxiv.org/abs/2609.37372
作者: Xuehai Bai,Zhenchen Tang,Yang Shi,Dianyi Wang,Tengfei Liu,Wanshun Su,Xuanyu Zhu,Ruohui Wang,Haiwen Diao,Haotian Wang,Xiaoling Gu,Yuanxing Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 31 pages

点击查看摘要

Abstract:Visual reward models are essential for evaluating and improving visual generation models, yet existing approaches typically map task conditions and candidate outputs directly to scalar rewards, leaving implicit what should be evaluated for each individual case. We introduce Think Before You Score, a paradigm that explicitly determines what matters for each case before judging how well the candidate performs. Following this principle, we propose the Thinking Reward Model (TRM), which formulates case-adaptive rubrics, performs rubric-guided assessment, and produces fine-grained pointwise rewards. We further observe that conventional pairwise preference optimization can induce score polarization, and introduce Pairwise Dual-Group Relative Policy Optimization (PD-GRPO), which leverages pairwise supervision to improve reward discrimination while preserving fine-grained pointwise scoring. Extensive experiments on image generation and editing reward-modeling benchmarks demonstrate that TRM achieves state-of-the-art performance among open-source reward models while remaining highly competitive with proprietary alternatives. Moreover, using TRM as a reward for reinforcement learning consistently improves diverse visual generation models, demonstrating that its fine-grained, case-adaptive rewards translate into effective optimization signals for visual generation.

[CV-98] Visual Anomaly Synthesis for Model Selection in Data Scarcity

链接: https://arxiv.org/abs/2609.37360
作者: Daniel Pröll,Thomas Kraxner,Tobias Schaefer,Sebastian Hegenbart
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Defect detection systems for industrial condition monitoring can only be relied upon if they are validated, yet defective samples are rare and, for a specific asset, often nonexistent. We present a framework that synthesizes severity-graded defects on real non-defective images without any defect references for the target asset, that can be used for model selection and validation. A defect taxonomy for common failure modes is distilled from literature into prescriptive prompts at varying defect severities. Regions of interest are cropped from in defect-free images and edited with a pre-trained image generation model (“FLUX.2 [klein]”). Color-matching and blending are employed to improve structural coherence with the original image. Generations are filtered out by a scorer and by estimated detection difficulty. Model selection experiments on MVTecAD show image AUROC choice regret over model selection can be nearly halved compared to the best fixed model chosen with access to test data. Experiments show the need for severity-graded anomaly synthesis. A case study investigates the proposed method for in-situ monitoring of Pelton turbine runners in hydropower, where real defect images are rare and expensive to collect. A PatchCorebased anomaly detection model is fit on Pelton turbine images and selected and validated using synthetic images, showing strong detection performance (94 % correct detection at optimal threshold and AUROC 0.97). The model reliably detects moderate and advanced defects, while early-stage defects remain challenging, indicating the synthetic data meaningfully stresses detector sensitivity.

[CV-99] Encore: Few-Shot Agent ic Discovery of Manipulation Strategies

链接: https://arxiv.org/abs/2609.37359
作者: Yifan Kang,Zihan Wang,Zhiwen Fan,Bangya Liu
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Coding agents can now write, run, and debug programs with little human help. Robot tasks, however, are usually specified by a sentence that leaves out how to grasp, in what order to make contact, and what the result should look like, and an agent given only the sentence must find these details by trial and error. We introduce ENCORE, which gives the agent a few demonstrations as evidence to read rather than as training data. A deterministic builder distills each demonstration into a pack of multi-view keyframes, gripper events, frame strips, and the full trajectory. A coding agent studies the pack, writes a policy program against a fixed perception and action API, refines it iteratively over a few development rollouts, and freezes it before a sealed evaluation that never reveals the success signal. On LIBERO-PRO, the agent’s first program already succeeds in half of the perturbed tasks with demonstrations and in one task without them, and the frozen programs outperform the strongest prior agentic system run with the same language model (96.3% against 89.3%). On RoboDojo tasks whose instructions leave the goal unstated, no program succeeds without demonstrations. ENCORE also runs on a real bimanual robot, learning cube handover and cup inversion from five demonstrations each.

[CV-100] PCaPaint: Prostate Cancer Inpainting by Mitigating Shortcut Learning MICCAI2026 MICCAI

链接: https://arxiv.org/abs/2609.37350
作者: Levente Lippenszky,Hongxu Yang,Marcell Dömötör,Krisztian Koos,László Ruskó
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at the DGM4MICCAI workshop at MICCAI 2026

点击查看摘要

Abstract:The development of AI systems for tumor-specific applications is limited by the scarcity of labeled data. Synthetic tumor inpainting offers a promising approach but faces challenges for prostate cancer MRI which contains high-resolution multi-sequence data. Although methods leveraging latent diffusion models (LDMs) enable large-volume synthesis, they are prone to shortcut learning, simply reproducing the condition image created by masking the lesion region. In this work, we introduce PCaPaint, a prostate cancer inpainting method based on LDMs that explicitly addresses this failure mode. To overcome shortcut learning that compromises synthetic tumor texture, we propose a simple yet efficient conditioning strategy in which the condition image is filled with Gaussian noise, and we provide theoretical justification. In addition, we propose a novel training objective for LDM that emphasizes the error within the lesion region. Furthermore, we introduce a multi-sequence latent design, in which T2w scans and DWIADC scans are compressed using two separate autoencoders to preserve their distinct frequency characteristics. Extensive experiments demonstrate that the generated synthetic data improves downstream performance in prostate lesion segmentation, patient-level classification and lesion-level detection. Furthermore, our method significantly outperforms a recent state-of-the-art LDM-based tumor inpainting method both in downstream performance and in synthetic image quality.

[CV-101] AEC: Trajectory-Aware Evidence Coordination for Multi-Step Visual RAG

链接: https://arxiv.org/abs/2609.37349
作者: Yalun Wu,Bingzhou Wang,Boyang Wang,Peiying Wang,Shaojie He,Yunhan Wang,Shaozu Yuan,Jiawei Wang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Multi-step visual retrieval-augmented generation (RAG) answers complex questions by repeatedly retrieving visual evidence, updating an intermediate state, and deciding whether to continue searching or answer. Yet retrieving relevant evidence does not ensure its effective use throughout the reasoning trajectory. As multi-step reasoning progresses, redundant sources occupy context capacity needed for missing evidence, observations tied to resolved requirements or unproductive searches linger in context, and visual sources are revisited with insufficient detail for fine-grained reading. We term this loss of usable evidence over a reasoning trajectory trajectory-level evidence utilization degradation. To address it, we propose Trajectory-Aware Evidence Coordination (TAEC), a training-free framework that coordinates evidence use around unresolved answer requirements. TAEC tracks these requirements in a shared trajectory state to guide which evidence enters the context, how accumulated memory is retained, and at what level of detail visual evidence is examined. Under a unified evaluation protocol on ViDoSeek, SlideVQA, and MMLongBench-Doc, TAEC achieves the best overall performance against leading training-free visual RAG baselines, with the highest average accuracy across multiple proprietary vision-language models. These results demonstrate that aligning evidence with evolving reasoning needs improves evidence use throughout multi-step visual RAG.

[CV-102] When to Retrieve When to Stay: Uncertainty-Aware Temporal Evidence Allocation for Streaming Video-LLM s

链接: https://arxiv.org/abs/2609.37345
作者: Xiang Hu,Jiazuo Yu,Lu Zhang,Yunzhi Zhuge,Huchuan Lu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Streaming video understanding requires Video Large Language Models (Video-LLMs) to reason over continuous visual streams under causal constraints. As the visual history grows, a bounded visual?processing budget requires evidence selection that balances temporal recency with query relevance. Recent-only selection excludes potentially relevant historical evidence, whereas Semantic-only retrieval can displace useful recent context when relevance scores are ambiguous. We introduce WRWS (When to Retrieve, When to Stay), a training-free framework for uncertainty-adaptive evidence allocation. A lightweight external vision-language encoder scores query relevance across the observed history, while an adaptive allocation module uses the normalized entropy of the similarity distribution as a proxy for retrieval uncertainty. WRWS favors semantic retrieval when relevance cues are reliable and strengthens the recency prior under uncertainty. Following a retrieve-first, encode-later pipeline, WRWS selects evidence before target-model visual encoding, such that only the selected observations are processed by the costly target Video-LLM. Experiments across four Video-LLM families and multiple model scales demonstrate competitive accuracy on StreamingBench and OVO-Bench. In our efficiency evaluation, WRWS reduces average vision-to-answer time to 47.93% of the state-of-the-art method. Code will be released.

[CV-103] HyperSAM: A Promptable Foundation Model for Hyperspectral Remote Sensing

链接: https://arxiv.org/abs/2609.37340
作者: Li Pang,Xinqiao Wu,Jing Yao,Pedram Ghamisi,Jun Zhou,Zhengchao Chen,Deyu Meng,Xiangyong Cao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by IEEE Geoscience and Remote Sensing Magazine (GRSM)

点击查看摘要

Abstract:Hyperspectral remote sensing provides dense spectral measurements that are indispensable for material-level Earth observation, yet the construction of a general-purpose hyperspectral foundation model remains difficult. Two bottlenecks are especially limiting. First, large hyperspectral corpora rarely provide high spatial resolution together with reliable dense annotations. Second, many hyperspectral models are still trained almost from scratch, so the geometric and interactive priors learned by modern vision foundation models are not fully reused. To alleviate these issues, we \highlightpresent \textbfHyperSAM, a promptable hyperspectral foundation model that couples a data-centric hyperspectral synthesis pipeline with a spectral adaptation architecture based on Segment Anything Model 3 (SAM3). On the data side, HyperSAM synthesizes full-spectrum hyperspectral cubes from high-resolution SpaceNet multispectral imagery through a physics-informed abundance-transfer generator, while SAM3-derived pseudo-masks provide object-centric supervision. On the model side, the latest implementation uses a frozen SAM3 RGB image branch, a trainable hyperspectral side encoder initialized from the RGB vision transformer (ViT), ControlNet-style zero-initialized feature injection, and a lightweight mixture-of-experts mask refiner. To enhance training robustness against noisy pseudo-labels, Cross-modal Sample Selection (CromSS)-style confidence selection is incorporated for noisy-label weighting. Extensive experiments show that HyperSAM obtains strong generalization on diverse hyperspectral tasks (e.g., classification, anomaly detection, change detection, target detection, and airborne oil-spill mapping) and that high-quality synthetic hyperspectral data can be more effective than simply scaling noisy hyperspectral supervision.

[CV-104] UGO: Unified Architecture for General Multi-Object Tracking by Segmentation NEURIPS2026

链接: https://arxiv.org/abs/2609.37339
作者: Jer Pelhan,Alan Lukezic,Matej Kristan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to NeurIPS2026

点击查看摘要

Abstract:General multi-object tracking (GMOT) tracks all instances of a user-specified category from a single first-frame exemplar. Prior work relies on bounding boxes and surrogate training, and struggles with non-rigid objects, crowded scenes, and distractors. We introduce UGO, a unified GMOT tracker that pairs a pretrained exemplar-conditioned detection head with an instance-propagation head in a common architecture. A novel training-free, energy-minimization consolidation method converts overlapping proposals into exclusive pixel-wise masks and detections, resolving over-segmentation, duplicates, and conflicts. A hierarchical memory spanning global and instance levels improves recall and per-instance segmentation accuracy using a new memory management protocol. UGO sets a new state-of-the-art on GMOT benchmarks and video object counting, and is competitive with specialist MOT methods, establishing a strong paradigm for unified, open-category multi-object tracking.

[CV-105] OFBD: Object-Focused Background Debiasing for Long-Tailed Learning

链接: https://arxiv.org/abs/2609.37331
作者: Shenghan Chen,Yiming Liu,Zhipeng Deng,Haolin Wang,Jiale Zhou,Zhijian Wu,Xiankai Lu,Yafei Ou,Yefeng Zheng
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Balancing performance trade-offs on long-tailed data distributions remains a long-standing challenge in visual recognition. Existing methods mainly improve tail classes through re-balancing, representation learning, or data augmentation, but the underlying cause of tail class degradation is still insufficiently explored. In this paper, we find that standard long-tailed training induces background-biased representation and optimization: tail classes suffer larger background distribution shifts and become increasingly driven by background gradients. This reveals that tail degradation is not merely caused by insufficient samples, but also by the learning of irrelevant background features. To tackle this issue, we propose Object-Focused Background Debiasing (OFBD), a framework that mitigates background bias from both distribution and optimization perspectives. Specifically, Foreground-guided CutMix preserves target-related foregrounds while diversifying complementary backgrounds, and Background-guided Feature Rectification suppresses background-biased features without learnable parameters or additional training. Extensive experiments show that our method improves overall accuracy, achieves significant tail-class gains, and can serve as a plug-in for mainstream long-tailed methods without external data or pretrained recognition models. The code is available at: this https URL

[CV-106] he Domain Is a Residue: Adapting Self-Supervised Features Not Generators

链接: https://arxiv.org/abs/2609.37330
作者: Thomas Deixelberger,Markus Steinberger
类目: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Machine Learning (cs.LG)
备注: 9 pages main text, 28 pages including appendix. 12 figures, 13 tables

点击查看摘要

Abstract:Clearing fog, rain or snow from footage, or turning renders into photographs, must remove the source domain and keep the scene. Unpaired translators carry it through because their generator sees the source appearance (pixels, a near-invertible latent or a control map) and keeps it. A DINO feature map fixes what is in the scene and carries weather, lighting and rendering style as a residue of 13 to 14% of the feature norm. We propose the Representation Feature Adapter (RFA), a 2.9M-parameter network that moves this residue. We train only the adapter and its discriminators; the encoder and a feature-conditioned decoder, trained once for all conditions, stay frozen. Against CycleGAN-Turbo it is ahead on both metrics on fog and on KID on night, and level within noise on snow, rain and haze. On sim-to-real it leads REGEN and HyPER-GAN on both metrics. Only the RFA removes the rain while keeping the scene. The removal costs scene structure: CycleGAN-Turbo keeps more on every condition but fog. On VAE latents the identical adapter collapses to the identity, and decoders from other groups that never saw it render its output. The RFA has about 160 times fewer trainable parameters than CycleGAN-Turbo and under a fifth of its per-condition training time.

[CV-107] What Comes Next? Omni-StoryBench for Evaluating Story-Grounded Omnimodal Generation

链接: https://arxiv.org/abs/2609.37317
作者: Sieun Hyeon,Yejoon Lee,Mintaek Lim,Woojin Kim,Jaeik Kim,Jaeyoung Do
类目: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
备注:

点击查看摘要

Abstract:Omnimodal evaluation should go beyond independent text, image, and speech production: individually plausible outputs may not express a coherent shared event. We introduce Omni-StoryBench, a story-grounded omnimodal benchmark evaluating whether models can coherently continue stories across image, narration, and speech. Each instance provides a current storybook page and structured next-page conditions, requiring models to generate the next illustration, narration, and spoken character utterance. Omni-StoryBench contains 900 rigorously validated story transitions from openly licensed children’s books, with ground-truth next-page references and speech metadata. We evaluate systems with modality-specific metrics and consistency-centered LLM-as-a-judge rubrics for context preservation, condition following, reference consistency, and cross-modal coherence. Across 32 baseline configurations spanning orchestration, semi-orchestration, and native any-to-any paradigms, we find orchestration with strong VLM planning most reliable, while current native omnimodal models often struggle with output completeness and controllability. Our analysis shows text-side performance is associated with image and speech quality, but image generation and visual continuity form the clearest observed bottleneck among the evaluated configurations. These results position Omni-StoryBench as a system-level benchmark measuring coherent omnimodal generation beyond isolated modality quality.

[CV-108] FLASH: A “Generate Once Synthesize Many” Framework for Synthetic Anomaly Generation in Industrial Anomaly Detection WACV2027

链接: https://arxiv.org/abs/2609.37314
作者: Abhay Kumar Das,Rajesh Gangireddy,Ashwin Vaidya,Samet Akcay
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Submitted to WACV 2027

点击查看摘要

Abstract:Synthetic anomaly generation helps expand industrial anomaly datasets when real defects are scarce or unavailable. Existing approaches lie at two extremes: procedural approaches are fast but struggle to represent complex anomalies, while generative approaches produce diverse defects but require costly per-sample generation. We present FLASH, a framework that decouples defect generation from anomaly synthesis under a ``generate once, synthesize many’’ paradigm. Given only normal images, FLASH uses Vision-Language Model (VLM) guidance and an image-generation model to produce a small set of defect images, from which it extracts, validates, and banks reusable defect patches. For synthesis of anomalous images, Object Boundary Suppression (OBS) first identifies the probable foreground object-aware region of the host image, while Multi-Resolution Spectral Pyramid (MRSP) noise generates diverse, size-controllable masks that determine the defect location and spatial extent. It then composes a large and diverse synthetic anomalous image set by localizing the defect region, sampling size-controllable placement masks and seamlessly blending retrieved defects onto new defect-free images without further need for image generation. Experiments on the MVTec AD 2 dataset show that FLASH-generated anomalies nearly close the calibration gap on real defects, reaching 78.1% image-level F1 against an 83.6% real-anomaly upper bound and providing the most consistent calibration transfer across detectors among procedural and generative alternatives. Moreover, FLASH synthesizes anomalies more than 11.95x faster than per-sample generative approaches.

[CV-109] chnical note on: Zero-Training Feature-Space Alignment via Information Geometry

链接: https://arxiv.org/abs/2609.37302
作者: Behraj Khan,Tahir Qasim Syed,Syed Ahmad Chan Bukhari
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Deep vision models often degrade under distribution shift. Test-time adaptation can improve robustness but typically requires iterative optimization, hyperparameter tuning, and multiple forward-backward passes. We propose Zero-Training Fisher Geometry Alignment (ZFGA), a closed-form method that improves robustness under covariate shift without modifying model parameters. ZFGA is based on the observation that distribution shifts distort feature-space geometry. It estimates the Fisher information matrix of the predictive distribution with respect to feature embeddings and applies a linear transformation that aligns test-feature Fisher geometry with a reference geometry computed from clean data. This provides a natural-gradient-inspired preconditioning step in feature space. We evaluate ZFGA on CIFAR-10-C and ImageNet-C using ResNet-50, DINO ViT-S/16, and CLIP ViT-B/32. ZFGA consistently improves over zero-shot inference across all three models, although it is not the strongest method for every model. Covariance whitening performs better on ResNet-50, while Fisher whitening is statistically indistinguishable from ZFGA on CLIP. Across six training-free and gradient-based alternatives (covariance whitening, Fisher whitening, TENT, T3A, LAME, and AdaNPC), ZFGA is the only method that does not substantially harm any of the three model families. The Fisher geometry distortion is also positively correlated with ZFGA gain (Pearson r = 0.366, p = 0.017), providing preliminary evidence that geometric misalignment contributes to robustness degradation. ZFGA requires only forward passes and matrix operations at inference time, offering a lightweight and deterministic alternative to optimization-based test-time adaptation.

[CV-110] Scaling Full Conformal Image Classifiers NEURIPS2026

链接: https://arxiv.org/abs/2609.37298
作者: Julio Silva-Rodríguez,Ender Konukoglu
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注: NeurIPS 2026. Code: this https URL

点击查看摘要

Abstract:Conformal prediction provides set-valued predictions with distribution-free coverage guarantees, making it attractive for high-stakes image classification. However, split conformal prediction is data-inefficient, while full conformal prediction (FCP), despite its stronger statistical efficiency, is computationally prohibitive at scale because it requires candidate-specific model refits at test time. We address this limitation by leveraging zero-shot vision-language models (VLMs) to guide scalable FCP in large label spaces. We introduce Targeted Full Conformal Prediction (T-FCP), which uses a lightweight inductive conformal predictor to prune unlikely labels and applies FCP only to the remaining candidates, reducing computation while retaining the formal guarantee of the combined conformal procedure. We further propose Stabilized Online LDA (SO-LDA), an efficient VLM adaptation solver based on rank-one inverse-covariance updates. Across multiple benchmarks, including ImageNet, T-FCP enables practical full-conformal image classification with modest test-time overhead, yielding efficient prediction sets and more stable empirical coverage than split conformal alternatives.

[CV-111] Why Cross-Skeleton Retargeting Is Non-Identifiable: Structural Limits of Generative Motion Models

链接: https://arxiv.org/abs/2609.37297
作者: Zhiyuan Li,Wenyan Yang,Pekka Marttinen,Joni Pajarinen
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Cross-skeleton motion generation trains generative models to carry action structure and motion intention from one body to another. Yet a target motion that shows the right action has two explanations that the training data cannot tell apart: the model transferred the source clip, or it recovered a typical motion for the requested action. We show that this ambiguity is structural rather than incidental: under standard generative objectives, the source-conditioned retargeting map is non-identifiable in sparse heterogeneous motion domains. Unpaired distribution matching yields gauge non-identifiability: the latent spaces of different skeletons can be transformed relative to one another without changing the training evidence, so different source-conditioned maps fit it equally well. Sparse paired supervision admits the complementary failure mode, \emphconditional-mean degeneration: when clips are paired only by action, squared-error training converges to an average target motion that ignores the source clip. To make the missing evidence observable, we introduce Source-Instance Fidelity (SIF), a diagnostic that tests whether outputs differ from one another the way their source clips do, with the target skeleton and action held fixed. Under this diagnostic, methods that succeed at the standard action-level test on animal motion data often sit at the source-blind floor, while the methods that rise above it retain only a partial relational signal. Retargeting therefore needs objectives and evaluations that can identify the source-conditioned map it claims to learn. Project page: this https URL.

[CV-112] VISTA-Bench: Benchmarking Multilingual Image Translation with Image-Specific Rubrics

链接: https://arxiv.org/abs/2609.37287
作者: Bo Lv,Mao Zheng,Zheng Li,Fangxu Liu,Mingrui Sun,Tao Chen
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Image translation is a fundamental capability of multimodal models for multilingual applications, requiring visual understanding and meaning preservation across languages. However, existing benchmarks have limited language coverage and often lack explicit image-specific evaluation criteria, making it difficult to comprehensively assess this capability. To systematically evaluate this capability, we introduce VISTA-Bench, covering 22 languages and 10 domains, and develop an image-specific rubric evaluation protocol. The benchmark combines sampling for language and scenario coverage with model-assisted, human-verified annotations that group related text into coherent semantic units and provide multilingual reference translations. The rubrics specify essential content, semantic relations, and acceptable translation variants, yielding separate output-based scores for translation quality and the preservation of visual and knowledge-dependent information. We conduct extensive evaluations of 16 mainstream models, including 12 multimodal models and four text-input models, and provide systematic analyses across languages, domains, and evaluation dimensions.

[CV-113] SAM Meets VLM: Parameter-Decoupled Full-Parameter Training for Unified Medical Reasoning and Segmentation

链接: https://arxiv.org/abs/2609.37283
作者: Xuyang Cao,Enyou Liu,Jun Zhao,Zhuoyun Liu,Jintao Fei, Leo
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Medical multimodal large language models (MLLMs) are increasingly expected not only to answer clinical questions, but also to localize the visual evidence behind their predictions. A common strategy connects a vision–language model (VLM) with SAM-style segmentation through a special SEG token, yet full-parameter training of this unified architecture is difficult because image-level reasoning and pixel-level segmentation impose different requirements on the shared representation space. To address this issue, we propose a parameter-decoupled training framework for unified medical reasoning and segmentation. The framework treats the SEG hidden state as a semantic-to-spatial prompt for the mask decoder and encourages it to become separable from generic language states, reducing ambiguous segmentation prompts and potential disruption to reasoning representations. It first performs medical shallow alignment to adapt visual features to clinical language without disturbing the LLM; then controlled instruction tuning shapes separable SEG prompt states, monitored by the Davies–Bouldin Index (DBI), while scaling segmentation gradients entering the language backbone; finally, the SAM branch is specialized with the VLM frozen to improve mask precision without altering reasoning parameters. Experiments on medical referring segmentation, grounding, visual QA, and textual QA benchmarks show that our framework achieves strong language-conditioned segmentation while preserving competitive reasoning ability. Ablations show that two-phase instruction tuning, gradient scaling, and segmentation specialization all contribute to the model.

[CV-114] UniAfford: Token-Routed Multitask Learning for Generalizable 2D-3D Affordance Perception

链接: https://arxiv.org/abs/2609.37264
作者: Yuhao Liu,Yiming Zhong,Hanqing Wang,Shaocheng Yan,Yuhang Zhang,Wenzhou Lyu,Ziyang Ding,Wei Zhang,Xue Zhao,Jin Pan,Yuexin Ma,Xinge Zhu
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Affordance perception aims to localize actionable regions supporting embodied interaction, yet 2D and 3D affordance grounding have evolved as separate problems, with different task definitions, supervision formats, datasets, and evaluation protocols. This fragmentation limits the learning of transferable object-affordance semantics across visual and geometric spaces. We propose Token Router for Tasks, a multitask training paradigm for MLLM-based systems that routes contextual hidden states to task-specific branches without requiring the language head to generate predefined markers. Routed states are supervised directly by branch-specific objectives, enabling dense prediction losses to shape shared MLLM representations. We instantiate this paradigm as UniAfford, a unified framework for generalizable 2D-3D affordance perception, together with UniAfford-Data, a dataset integrating pixel-level 2D annotations, point-level 3D annotations, and language instructions under a shared object-affordance taxonomy, supporting heterogeneous supervision through semantic-level 2D-3D pairing. UniAfford adopts an MLLM as a shared semantic hub and a modality-aware token router to produce image- and point-cloud-affordance queries. These queries respectively condition a SAM-style pixel decoder and a SONATA-based point decoder, enabling flexible 2D, 3D, and joint affordance inference from image-only, point-cloud-only, or paired multimodal inputs. Experiments demonstrate strong zero-shot generalization across 2D and 3D affordance benchmarks without target-specific fine-tuning, alongside state-of-the-art branch-wise performance under modality-isolated protocols. Ablations validate token routing, joint 2D-3D supervision, and decoder coupling, while language-head diagnostics show that routed latent states carry meaningful object-affordance semantics. Project page: this https URL

[CV-115] Beyond Attention Imbalance: Mitigating Hallucinations via Spectral Surgery

链接: https://arxiv.org/abs/2609.37263
作者: Siqi Lu,Suo Wei,Yongbin Zheng,Jianhang Yao,Wanying Xu,Peng Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:While Large Vision-Language Models (LVLMs) achieve remarkable success, hallucinations remain a significant barrier to their reliable deployment. Recent studies primarily attribute these issues to cross-modal attention imbalances; most solutions therefore focus on reweighting visual tokens or suppressing language priors. However, such approaches often overlook the spectral characteristics of the visual information flow and frequently rely on Contrastive Decoding (CD), which doubles inference time. Instead of following conventional approaches, we identify two distinct hallucination patterns-Perceptual-Semantic Dissociation and Localized Fixation-and propose FLASH (Frequency-Localized Attention SHaping), a training-free and CD-free framework. FLASH utilizes a Spectral Vortex Score to detect vision heads within multi-head attention layers and applies adaptive spectral modulation to rectify the visual information flow during decoding. Empirical results demonstrate that FLASH achieves a superior balance between performance and efficiency compared to SOTA methods.

[CV-116] Collision-Aware and Observation-Aligned Object-Centric Scene Reconstruction from Point Cloud

链接: https://arxiv.org/abs/2609.37260
作者: Yuxuan Xie,Xuan Yu,Rong Xiong,Yue Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Object-centric scene reconstruction requires completing partial object observations while preserving metric alignment and avoiding collisions with the surrounding. Existing generation-based methods are often image-conditioned and suffer from scale ambiguity and insufficient geometric constraints. We propose COOL, a framework for COllision-aware and Observation-aLigned reconstruction. Based on an object generation model, COOL conditions the generation on instance and background point clouds. Instance geometry anchors generation in scene coordinates, while background geometry provides local context for scene-consistent completion. We further introduce an explicit collision loss and use joint optimization and resampling to reduce collisions during inference. Experiments on 3D-Front and Scan2CAD demonstrate strong scene-level fidelity, observation alignment, and collision reduction. Moreover, additional studies validate its robustness to mask errors and its applicability to real-world scene replicas.

[CV-117] V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents

链接: https://arxiv.org/abs/2609.37250
作者: Yang Zhang,Jiangyuan Zhao,Chenyou Fan,Jiayu Hu,Xiu Yuan,Chenjia Bai,Xiu Li
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注: 19 pages, 5 figures, 11 tables

点击查看摘要

Abstract:World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrained at scale, a prominent line of recent WAMs inherits both predictive knowledge and the models in which it was learned. We ask whether a predictive visual latent space induced by large-scale predictive pretraining can instead provide a sufficient foundation for effective WAM learning without inheriting a complete pretrained visual generative model. To answer this question, we introduce V-JEPA Policy, a simple framework that builds a WAM on the latent space of a frozen V-JEPA 2.1 encoder. An instruction-conditioned future-latent predictor and a flow-matching action expert are jointly learned from scratch in a single downstream stage, with the predictor’s future-informed context key–value states conditioning action generation. With 0.9B total parameters, of which 0.6B are trainable, V-JEPA Policy achieves competitive performance with representative WAM and vision-language-action baselines across LIBERO, LIBERO-Plus, and RoboCasa-GR1. Comparing visual foundations under the same downstream framework and training budget identifies V-JEPA latents as more effective than the discriminative, reconstructive, and video-understanding-oriented alternatives, particularly under distribution shifts. Beyond task-specific learning, pretraining the predictor on DROID video–instruction pairs without action labels and adapting it into a WAM yields substantial gains in downstream control and out-of-distribution generalization. Together, these findings establish predictive visual latents as a foundation for effective WAM learning from task-specific demonstrations and for transferring future-modeling knowledge acquired from broader in-the-wild videos. Our code is available at this https URL.

[CV-118] Codebook-Guided Cross-Modal Knowledge Distillation for Structurally Heterogeneous Features

链接: https://arxiv.org/abs/2609.37243
作者: Dae Ung Jo,Jongin Lim,YoungJoon Yoo,Daeho Um
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Cross-modal knowledge distillation transfers knowledge from a teacher modality to a student modality. Existing feature-level alignment methods typically assume that teacher and student features reside in structurally alignable representation spaces. However, this assumption does not hold when cross-modal features are structurally heterogeneous and lack clear unit-level correspondence, such as 2D spatial visual grids and 1D temporal audio sequences, thereby limiting the applicability of feature-level alignment. To address this challenge, we propose a cross-modal distillation framework that enables effective knowledge transfer across structurally heterogeneous feature spaces via a vector-quantized codebook. Specifically, teacher features are abstracted into a set of vector-form codes regardless of their original feature structure, and the selected codes serve as concept-level anchors for student learning. Code selection is guided by both task relevance and student compatibility, allowing the student to receive transferable teacher knowledge without requiring direct unit-level feature alignment. Experimental results across diverse cross-modal distillation scenarios demonstrate the effectiveness of the proposed framework on classification and semantic segmentation tasks.

[CV-119] Seeing Is Not Addressing: Auditing Linguistic Access to Frozen Visual Geometry

链接: https://arxiv.org/abs/2609.37230
作者: Woosang Jeon,Jiwon Yang,Soo Chung,Taehyeong Kim
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 27 pages, 10 figures. Code available at this https URL

点击查看摘要

Abstract:Visual distinctions are often finer than those reflected in linguistic conceptualization. Vision-language models exhibit a similar asymmetry: a distinction can remain discriminable in frozen image geometry while being weakly addressable through the native text interface. We study this gap by separating visual discriminability from linguistic addressability in text-to-image retrieval. Using FactorAtlas, a fully crossed testbed of 23,040 images spanning shape, hue, pattern, and nuisance variation, we compare both readouts on held-out images of the same distinctions. We then derive image-side contrasts that separate each value from its alternatives for matched visual grounding, and test whether this reduces the native-text access gap across factors and models. Direction-specific and visual-absence controls tie these gains to the relevant visual contrast; the gains persist after global alignment and extend to compositional retrieval and natural images. Together, these results show that visual discriminability and linguistic addressability need not coincide, and that matched visual grounding can probe and reduce the resulting access gap.

[CV-120] AESOP: Asymmetric Human-Camera Generation with Translation-Intensity Control

链接: https://arxiv.org/abs/2609.37229
作者: Jingzhong Lin,Zhanke Wang,Heng Li,Wenxiang Liu,Zhao Zhang,Kecheng Tang,Dongdong Xiang,Changbo Wang,Di Kang,Chunchao Guo,Linchao Bao,Gaoqi He
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Human motion defines an action, while a camera trajectory determines how it is presented. Camera generation for a given human motion and joint human-camera generation are usually treated as separate tasks, although both share an asymmetric dependency: human motion can be generated independently, whereas the camera responds to the realized action. We introduce AESOP, a unified framework with an independent human pathway and a shared human-conditioned camera generator. Its asymmetric architecture serves both tasks while preserving the human output during camera generation. Although human context anchors the shot to the action and camera text describes its movement, translation intensity remains underspecified. We therefore construct trajectory pairs that differ in camera translation magnitude while sharing human motion and camera text, then use these pairs to learn an explicit intensity condition. Experiments on the PulpMotion dataset demonstrate strong camera distributional and framing quality in both tasks and effective control over camera translation intensity.

[CV-121] ResComEmb: Effective and Efficient Multimodal Embedding via Residual Homogeneity Compression

链接: https://arxiv.org/abs/2609.37225
作者: Zijing Cai,Yuzhe Wang,Jingxian Zhu,Fengbin Zhu,Richang Hong
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 19 pages

点击查看摘要

Abstract:Multimodal large language models (MLLMs) have shown strong potential for universal multimodal representation learning. However, existing methods either compress each input into a single vector, limiting fine-grained expressiveness, or retain long sequences of visual-token vectors, incurring substantial storage and interaction costs. To resolve this trade-off, we propose ResComEmb, a trainable framework for effective and efficient universal multi-vector multimodal embedding. ResComEmb first encodes each input at native dynamic resolution into ordered global, intermediate, and fine-grained views. After MLLM contextualization and embedding projection, a trainable Residual Homogeneity Compression (RHC) module reduces within-granularity redundancy and cross-granularity repetition under explicit visual token budgets. Then, ResComEmb introduces a length-adaptive Bidirectional Late-Interaction Matching mechanism for robust query-document scoring, which averages the strongest token-level matches in each direction and combines the two scores using a weight based on how many valid tokens each side has. Extensive experiments on MMEB, ViDoRe V1, and ViDoRe V2 show that ResComEmb produces higher-quality universal multimodal embeddings than VLM2Vec-V2, and outperforms ColQwen2.5 in visual document retrieval using only 37.5% of its full visual token budget, demonstrating a favorable effectiveness-efficiency trade-off.

[CV-122] Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL

链接: https://arxiv.org/abs/2609.37200
作者: Songlin Yang,Xiaotong Zhao,Jiacheng Zhang,Zhe Wang,Toyota Li,Eric Liu,Alan Zhao,Anyi Rao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Multi-reward guided reinforcement learning (i.e., RL) offers a promising way to improve joint audio-video diffusion models along several complementary objectives, including modality-specific quality, cross-modal semantic alignment, and temporal synchronization. Its effectiveness, however, depends on two quantities that change during training: where reward-driven updates should act, and how competing rewards should be combined. Existing methods tend to rely on fixed routing and reward weights, failing to track evolving model functions. To address these limitations, we propose Adaptive Reward Routing to jointly adapt update locations and reward coordination during forward-process RL (i.e., DiffusionNFT) of joint audio-video diffusion models. Our method consists of two components. (i) Cross-Modal Influence-Guided Routing (Localizing Updates): We use bidirectional cross-attention responses as an efficient proxy for evolving cross-modal influence, dynamically reweighting token-aware losses and scaling gradients across cross-modal layers without additional model interventions. (ii) Preference-Preserving Modality-Aware Reweighting (Coordinating Rewards): We preserve predefined weights as preference priors and use branch-specific reward-gradient interactions as residual corrections after warm-up. This resolves evolving conflicts without letting dominant rewards suppress weak but essential objectives. Extensive experiments demonstrate consistent improvements in modality quality, semantic consistency, and audio-video synchronization over strong RL baselines. Ablations and mechanism analyses further validate the complementary benefits of adaptive update routing and reward coordination.

[CV-123] Exploring In-Context Learning for Handwritten Text Recognition

链接: https://arxiv.org/abs/2609.37195
作者: Eric Ayllon,Abel Gandia,Jorge Calvo-Zaragoza
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 19 pages, 3 figures

点击查看摘要

Abstract:Handwritten Text Recognition (HTR) systems have become an indispensable tool for the digitization of historical documents. Not only do they cut down time and cost, but they also allow democratizing access and processing of their contents by generating their transcripts. However, literature in HTR currently focuses mostly on specialized models that require large amounts of annotated samples to achieve satisfactory performance. We explore the use of In-Context Learning with pre-trained Vision-Language Models (VLMs) to create a transcription pipeline without updating the model’s parameters. We then evaluate this pipeline across multiple collections and models, and demonstrate that general-purpose VLMs can be effectively taught how to transcribe handwritten text from images. To assess how our observations may translate to practical applications, we evaluate the performance in a Cross-Domain (CD) scenario, where context examples are drawn from a different collection than the query image. Results in both the controlled In-Domain (ID) scenario and the realistic CD scenario follow the same patterns. First, as context size grows, the error range is expected to narrow towards the average performance. Thus, larger context sizes sacrifice the performance of the oracle-best sampling for lower expected error rates. The results obtained show that, without any parameter updates, this methodology has strong potential to compete with traditional HTR in the presence of domain shift. Moreover, we show and argue that some context samplings work better than others and suggest more effort should be put into finding an ideal sampling method in future work. Comments: 19 pages, 3 figures Subjects: Computer Vision and Pattern Recognition (cs.CV) ACMclasses: I.4.1; I.4.9 Cite as: arXiv:2609.37195 [cs.CV] (or arXiv:2609.37195v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2609.37195 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Eric Ayllon [view email] [v1] Tue, 29 Sep 2026 10:15:38 UTC (733 KB)

[CV-124] HaPRL: Human-Anchored Process Reinforcement Learning for Visual Search Agent

链接: https://arxiv.org/abs/2609.37190
作者: Zhangquan Chen,Yaoxin Niu,Xiang An,Mingze Sun,Zhumei Wang,Chih-Ting Liao,Hongkun Cao,Ruqi Huang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 24 pages, 9 figures. Code: this https URL

点击查看摘要

Abstract:Multi-turn visual search agents answer questions about high-resolution images by iteratively deciding where to look. Reinforcement learning for these agents rewards only the final answer, leaving the search process unsupervised. Consequently, faulty routes in which the reasoning process is erroneous yet the final result is correct arise frequently, which in turn leads to ineffective training, i.e., scaling along the wrong paths. In this paper, we introduce HaPRL, the first framework to reinforce the search process with human search behavior. We first build an annotation platform and collect 1K+ human-annotated data with fine-grained behavioral signals. During training, a carefully designed judge scores each rollout with task-adaptive weights, anchored on the distilled trace of how a human annotator actually searched the same image. Extensive experiments show that HaPRL consistently outperforms outcome-based RL, and early-stage process supervision yields 6.7x more improvement in subsequent outcome-based scaling. Our results also demonstrate the importance of aligning model behavior with human process annotation signals, which offer new insight into the training of foundation models.

[CV-125] InsightMap: Structured Spatial Modeling for Embodied Multimodal Reasoning

链接: https://arxiv.org/abs/2609.37187
作者: Hongpei Zheng,Hujun Yin
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Language-guided navigation requires connecting partial observations to a persistent spatial reference and learning how actions change that representation. We introduce InsightMap, a framework that uses top-down maps as both explicit spatial memory and action-conditioned prediction targets. Historical views are linked to labeled map locations, and a shared multimodal backbone jointly learns navigation action prediction and post-action map generation. Map prediction provides auxiliary training supervision, while navigation inference decodes actions from the observed spatial context. An aligned RGB-D data pipeline supports a common interface for navigation, visual question answering, situated reasoning, and 3D grounding. On the validation-unseen splits of R2R-CE and RxR-CE, InsightMap achieves success rates (SR) of 56.9% and 54.9%, respectively. Adding map-prediction supervision improves R2R-CE SR by 4.3 and success weighted by path length (SPL) by 3.2 percentage points. On static spatial tasks, InsightMap achieves 103.7 CIDEr on ScanQA, 60.1% exact-match accuracy on SQA3D, and 53.1% grounding accuracy at 0.5 IoU on ScanRefer with detected object proposals. On Unitree Go2, it outperforms NaVid and NaVILA in hallway, lab, and office environments.

[CV-126] Sparse cubical complexes for efficient topology-preservation in image data

链接: https://arxiv.org/abs/2609.37177
作者: Alexander H. Berger,Marco Fontana,Daniel Rueckert,Johannes C. Paetzold,Laurin Lux,Ulrich Bauer
类目: Computer Vision and Pattern Recognition (cs.CV); Computational Geometry (cs.CG); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Persistent homology (PH) is a frequently used tool for extracting and preserving topological information from image data, particularly in image segmentation, where preservation of topological structures is important. However, despite its general applicability across dimensionality, domains, and target structures, the runtime cost of PH-based methods often makes their practical use infeasible. In this work, we argue that this runtime cost is largely driven by processing information that is unimportant for downstream application (e.g. as optimization objective). We propose sparse cubical filtrations as an alternative foundation for PH computation, reducing subsequent computational costs by factors of up to 100 on real datasets. We show close agreement with the optimization signal of the dense counterpart and empirically evaluate our solution’s effectiveness as an optimization objective in realistic training regimes where other PH-based objectives can practically not operate (i.e., 3D data with large patch sizes). We show how our solution improves topological accuracy by up to 80% across six diverse datasets while maintaining pixel- and region-based accuracy.

[CV-127] End-to-End Self-Supervised RGB-T Tracking without Modality Misleading

链接: https://arxiv.org/abs/2609.37162
作者: Shenglan Li,Rui Yao,Kunyang Sun,Hong Jia,Yong Zhou,Javen Qinfeng Shi,Xinyu Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:RGB-T object tracking leverages the complementary characteristics of visible and thermal infrared modalities to improve robustness under adverse conditions. Existing supervised methods typically rely on costly modality-aligned bounding box annotations, while most self-supervised approaches follow a two-stage pseudo-labeling paradigm, making tracker training sensitive to pseudo-label quality and preventing joint end-to-end optimization. In this paper, we propose ESMTrack, a fully end-to-end self-supervised RGB-T tracking framework without offline pseudo-label generation or dense frame-level bounding box annotations. Given only the standard initial-frame annotation used in visual tracking, ESMTrack learns discriminative and temporally consistent representations through two complementary objectives: a grounding triplet loss on annotated initial frames and a cross-frame temporal triplet loss on unlabeled search frames, with reliable samples selected by forward-backward consistency. To address modality dominance bias, ESMTrack employs a three-branch architecture consisting of a fusion branch and two unimodal branches for RGB and thermal inputs. We quantify modality contributions using the Average Peak-to-Correlation Energy by measuring response discrepancies between the fusion and unimodal branches. The resulting reliability estimates guide a training-time modality decoupling mechanism that suppresses dominant-modality shortcuts and adaptively weights cross-modal contrastive learning for task-level alignment. Extensive experiments on five RGB-T tracking benchmarks show that ESMTrack achieves competitive state-of-the-art performance, strong cross-dataset generalization, and real-time inference speed. The source code is available at this https URL.

[CV-128] Improved Distributional Diffusion Models

链接: https://arxiv.org/abs/2609.37147
作者: Tommaso Martorella,Alexandre Galashov,Felix Krause,Stefan Andreas Baumann,Valentin De Bortoli,Arthur Gretton,Björn Ommer
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Distributional Diffusion Models (DDMs) replace the standard mean-prediction denoiser with a \emphdistributional denoiser trained via a scoring rule objective, learning a stochastic approximation to p(x_1 \mid x_t) rather than its conditional mean. However, scaling DDMs to modern image-generation settings faces two obstacles: (i) multi-particle training incurs overhead that scales with the number of particles, (ii) DDMs use globally fixed scoring rule hyperparameters, forcing a single trade-off across sampling budgets. We mitigate these limitations by deferring particle expansion to late transformer layers, and the hyperparameter trade-off by introducing time-dependent scoring rule schedules informed by the dynamical regimes of~\citetBiroli2024. Combined with a DiT-based latent setup, these changes make DDM training practical on class-conditional ImageNet- 256^2 , achieving 4.48 FID at 4 steps and 2.38 at 50 steps with DiT-XL/2, from a single model trained from scratch in one stage, without a teacher, self-distillation or JVPs. The result is a stochastic few-step generator whose FID does not degrade as the sampling budget grows from 4 to 50 NFE, and the same recipe transfers to text-to-image generation. Code and pre-trained models available at this https URL.

[CV-129] MSTypography: Multi-character Semantic Typography via Balancing Word Legibility and Object Recognizability

链接: https://arxiv.org/abs/2609.37141
作者: Xinye Yang,Xinding Zhu,Kai Fang,Xinyi Ren,Mengjian Li,Bin Cao,Jiazhou Chen
类目: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
备注:

点击查看摘要

Abstract:Semantic typography is a design technique where the visual representation of a word conveys its semantic meaning, while maintaining its legibility. Existing digital typography methods mainly focus on single-character scenarios. They suffer from a lack of legibility constraints and insufficient local deformation when extended to multi-character words, as the intricate structures among multiple characters are hardly preserved during the typography process. In this paper, we propose a global-to-local typography framework for multi-character scenarios. It performs mask-driven silhouette approximation at the global level, while semantic-guided refinement at the local level, with a culling step in between to improve efficiency. To preserve word legibility, we designed structural losses (including explicit collision constraints and implicit Jacobian singular value constraints) and an OCR constraint for character-level readability. To enhance the object recognizability, we leverage semantic guidance with diffusion priors, which drives the character glyph toward the target concept while preserving its structural integrity. To the best of our knowledge, this is the first multi-character semantic typography method that effectively balances word legibility and object recognizability. Evaluations on five representative languages (English, Chinese, Japanese, Korean, Arabic) demonstrate superiority over SOTA methods. Codes will be open-sourced.

[CV-130] aoFlowForge: Progressive Native Mesh Generation via Cascaded Flow Matching

链接: https://arxiv.org/abs/2609.37139
作者: Xianze Fang,Qiyuan Feng,Dongfang Sun,Yan Zhang,Xiuchao Wu,Jingnan Gao,Jiangjing Lyu,Chengfei Lyu,Gang Yu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:3D content generation technology has significantly advanced the work of designers, as well as the 3D printing and gaming industries. However, it remains difficult to produce lightweight, editable, and topologically clean artistic content that is directly production-ready. To achieve this, we present TaoFlowForge, an artistic mesh foundation model that generates production-ready meshes. Specifically, TaoFlowForge decomposes the mesh generation process into vertices generation and their connectivity prediction, i.e., edges. We formulate vertices generation as a two-stage coarse-to-fine process and incorporate several effective loss functions to further enhance its performance. In the connectivity prediction stage, we propose a simple yet effective method for estimating the connectivity affinity between vertices and additionally predict per-vertex normals, which determines the correct orientation of faces. Besides, we construct a large-scale dataset combining hand-crafted 3D assets with public high-quality topology datasets. Based on this, a carefully designed data curation pipeline is employed to filter the raw dataset, retaining only high-quality topology data for model training. Our model is trained on the combined dataset and tested on both out-of-distribution hand-crafted set of 3D assets and public datasets. Under image-conditioned generation, TaoFlowForge outperforms autoregressive methods and achieves state-of-the-art results among open-source mesh topology generators. We will release all the code and weights together with a portion of our test dataset.

[CV-131] Multi-Granularity Language-Guided Imitation Learning via Instruction Decomposition

链接: https://arxiv.org/abs/2609.37135
作者: Yi-Pei Chiu,Wei-Ta Chu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Using language instructions as conditions to guide robot policy learning has recently become an important research domain. However, existing language-guided policy learning methods typically use an overall task description to guide the entire demonstration trajectory. For manipulation tasks involving multiple execution stages, these methods assign the same language description to different subtasks, making it difficult to distinguish the behaviors required at different stages. In this work, we propose a multi-granularity language guidance method based on instruction decomposition. The proposed method decomposes an overall task description into more fine-grained, concrete subtask-level language instructions, thereby enhancing learning efficiency and improving performance. We evaluate the proposed method in the setting of multi-task imitation learning and validate its effectiveness.

[CV-132] EviViT: Evidence-Adaptive Vision Transformers for Fine-Grained Perception

链接: https://arxiv.org/abs/2609.37123
作者: Yaoxin Niu,Zhangquan Chen,Yang Zhang,Xiang An,Zhumei Wang,Chih-Ting Liao,Hongkun Cao,Ruqi Huang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 23 pages, 9 figures. Code: this https URL Data: this https URL

点击查看摘要

Abstract:Fine-grained visual perception enables vision-language models to distinguish subtle attributes and ground their answers in visual evidence. In high-resolution scenes, processing the whole image at greater resolution spends visual tokens on irrelevant content, while isolated crops can lose the context needed to interpret the selected evidence. We introduce EviViT, a lightweight attachment that learns where a pretrained vision transformer should acquire detail. Human visual-search traces supervise a question-conditioned evidence density, which guides regional re-reading from the original pixels and the allocation of visual tokens. A sparse, coordinate-aware bridge then connects the regional features to the global scene, allowing the host to interpret precise evidence in context. Learned with the host backbone frozen, the attachment serves both the base model and compatible post-trained descendants without refitting. Experiments across nine hosts show consistent gains in average fine-grained accuracy. Matched-budget comparisons further show that EviViT outperforms global-only processing at every tested token ceiling while using fewer visual tokens.

[CV-133] NRF-GS: Neural Residual Fields for Expressive and Compact Gaussian Splatting NEURIPS2026

链接: https://arxiv.org/abs/2609.37115
作者: Pratik Singh Bisht,Andreas Kolb
类目: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
备注: Accepted at NeurIPS 2026

点击查看摘要

Abstract:We revisit the role of appearance modeling in 3D Gaussian Splatting (3DGS) and show that limited expressiveness in view-dependent reflectance is a key driver of representation redundancy. In standard 3DGS, low-order spherical harmonics (SH) are used, restricting the splats’ ability to model high-frequency directional effects, which is typically compensated by increasing the number of splats. We propose \emphNRF-GS: Neural Residual Fields for Gaussian Splatting, a hybrid representation that replaces per-splat SH-bases with a shared neural residual field. Each Gaussian encodes a compact set of appearance features and a lambertian base color, while a lightweight \emphglobal scene-level MLP predicts view-dependent residuals conditioned on viewing direction, distance, and per-splat features. This formulation enhances directional reflectance modeling by combining diffuse per-splat reflectance representations with a shared global function for high-frequency details, enabling both higher expressiveness and parameter sharing across splats. Our key insight is that by accurately capturing high-frequency directional reflectance, especially in specular regions, the GS-representation becomes more expressive, reducing the need for geometrically redundant splats. As a result, NRF-GS achieves comparable or better rendering quality while reducing the number of Gaussians by up to 50%, and produces visibly improved specular and high-frequency details.

[CV-134] V2X-WAM: A Cooperative World Action Model for End-to-End Autonomous Driving

链接: https://arxiv.org/abs/2609.37098
作者: Junwei You,Weizhe Tang,Can Wang,Yan Zhao,Jun Hua,Haotian Shi,Wei Zhang,Lin Wang,Bin Ran
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Vehicle-infrastructure cooperation can complement onboard sensing with broader and more informative observations of the traffic environment, providing valuable support for end-to-end autonomous driving. However, existing cooperative driving methods mainly exploit roadside information to enhance the representation of the current scene, while the future consequences of prospective driving actions are rarely modeled explicitly. This limits the ability of the planner to anticipate how its decisions may interact with the evolving traffic environment. To address this issue, we propose V2X-WAM, a cooperative world action model that tightly couples cooperative scene understanding, action generation, and future-world reasoning. V2X-WAM constructs a reliability-aware spatiotemporal representation from vehicle- and infrastructure-side observations, while compressing infrastructure information into a compact quantized message for efficient communication. Based on the resulting cooperative representation, a multimodal planner generates prospective trajectories, which explicitly condition future occupancy and dynamic-flow prediction. The predicted world consequences are then fed back to refine the planned trajectory, forming a closed interaction between action and future-world evolution. Experiments on a large-scale real-world cooperative driving dataset demonstrate that V2X-WAM consistently improves planning accuracy and safety over representative end-to-end cooperative driving methods, while achieving stronger future-world prediction and substantially lower communication overhead. Ablation studies further validate the effectiveness of the proposed design.

[CV-135] Why MLLM s Struggle to Count: Overcoming Individuation and Aggregation Bottlenecks with ConvStack

链接: https://arxiv.org/abs/2609.37096
作者: Liwei Che,Yihao Quan,Sen Fang,Hongyi Wang,Ranjay Krishna,Ruixiang Tang,Vladimir Pavlovic
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Multimodal Large Language Models (MLLMs) consistently struggle with fine-grained visual counting, yet the underlying causes remain poorly understood. In this work, we present a mechanistic analysis of this failure mode, identifying two critical bottlenecks inherent to the global attention pipeline of MLLMs. First, we reveal an individuation bottleneck stemming from image patchification: because Vision Transformers process patches independently, they struggle to group fragmented geometric features across boundaries into distinct object representations. Second, we identify a collapse in the subsequent counting aggregation process, where representation separation rapidly diminishes as numerosity increases due to attention compression. Identifying and formalizing these twin bottlenecks constitutes our first major contribution. To overcome them, we propose ConvStack, a lightweight architecture that operates directly in the visual token space to explicitly aggregate and inject local spatial structures via zero-initialized residual connections. By explicitly addressing the individuation bottleneck, ConvStack provides unambiguous geometric evidence for downstream aggregation. Remarkably, by fine-tuning exclusively on counting tasks, the model achieves substantial improvements in dense object counting and broader spatial understanding benchmarks, without compromising on general visual capabilities.

[CV-136] ask-Oriented Visual Feature Compression via Residual Vector Quantization for Device-Edge Multimodal Inference

链接: https://arxiv.org/abs/2609.37090
作者: Luning Pang,Cheng Yuan,Jiawei Shao,Mingtao Huang,Yuan Shen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 13 pages. Submitted to IEEE Transactions on Mobile Computing

点击查看摘要

Abstract:Large multimodal models (LMMs) support diverse visual understanding and reasoning tasks but are often impractical to run entirely on resource-constrained devices. Device-edge co-inference reduces device computation, yet transmitting visual data over bandwidth-limited uplinks can introduce substantial delay. Task-oriented feature compression (TOFC) reduces the payload through feature aggregation and entropy coding. However, continuous-feature coding remains costly, and query-agnostic aggregation may discard task-relevant local evidence. We propose query-guided task-oriented feature compression (Q-TOFC) for device-edge multimodal inference. Q-TOFC employs residual vector quantization (RVQ) to encode each merged feature as a compact sequence of codebook indices, reducing its representation cost and allowing more features to be transmitted. It further incorporates query relevance into feature aggregation and uses a quantization error compensation adapter to mitigate the distortion introduced by discrete quantization. Experiments on seven multimodal benchmarks show that Q-TOFC reduces the visual payload by 53.6% relative to TOFC while maintaining comparable average normalized task performance. End-to-end latency evaluations further demonstrate lower latency under bandwidth-constrained uplinks.

[CV-137] Real2Gym: Building Gyms from Videos Bringing Skills to Robots

链接: https://arxiv.org/abs/2609.37089
作者: Kerui Ren,Yingxiang Xu,Kaiwen Song,Linning Xu,Bo Dai,Mulin Yu,Tao Lu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL

点击查看摘要

Abstract:Real-world videos provide rich demonstrations of manipulation, but turning them into reusable robot skills requires visually aligned environments, executable physical interactions, and mechanisms for learning from experience. We introduce Real2Gym, an agentic Real2Sim2Real framework that turns human and robot demonstrations into interactive simulation gyms and brings skills acquired in simulation to physical robots. The Real2Sim module reconstructs editable scenes, aligns objects and cameras with the input, validates demonstrated or retargeted actions through native physics execution, and generates task-conditioned variations with action-feasibility checks. Within these environments, the agent generates executable code for manipulation stages, observes their outcomes, and distills successful attempts and failures into reusable task procedures, object-relative motions, and recovery strategies. Through a shared perception-and-control interface, these skills guide subsequent execution in simulation and on real robots, with motions adapted to current observations and no updates to the underlying model weights. Extensive evaluations demonstrate that Real2Gym enables high-fidelity simulation environment reconstruction, outperforming GPT-6 Astra Direct Mode by 16.7% in success rate with approximately 74.9% fewer policy-execution tokens across these environments, while exceeding it by 33.3% in physical robot execution success rate across four tasks on a real Franka robot.

[CV-138] LDM-is-AE: Latent Diffusion Model is an Auto-Encoder for End-to-End Image Generation NIPS2026

链接: https://arxiv.org/abs/2609.37080
作者: Zhengqiang Zhang,Lingchen Sun,Rongyuan Wu,Qiaosi Yi,Xiangtao Kong,Chaodong Xiao,Lei Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by NIPS 2026. More info can be found in this https URL

点击查看摘要

Abstract:Latent Diffusion Models (LDMs) typically adopt a two-stage pipeline: an auto-encoder (AE) is first pre-trained to define a latent space, then a diffusion model is trained to perform denoising within it. Such a two-stage design introduces a representation mismatch, as the latent space is optimized for reconstruction rather than adapting the denoising dynamics. We reveal that the LDM itself is an AE, and consequently present LDM-is-AE, an end-to-end one-stage LDM training framework that eliminates the need for a separately trained tokenizer. Our key observation is that the LDM backbone actually performs a latent-to-feature-to-latent transformation at each denoising step, which can be interpreted as an internal decoding–encoding process. Leveraging this structure, we split the DiT backbone into two reciprocal components, DiT-E (i.e., DiT Encoding) and DiT-D (i.e., DiT Decoding), and impose image-space supervision on the intermediate features across all timesteps. Our model encourages the internal representation to align with the image domain throughout denoising, thereby establishing an explicit latent-to-image-to-latent path. At the zero-noise timestep, our model further performs an image-to-latent-to-image mapping, corresponding to an auto-encoding process. As a result, LDM-is-AE jointly learns latent representations and denoising dynamics in an end-to-end manner, yielding a diffusion-native latent space tailored to the generation process. Experiments demonstrate that LDM-is-AE exhibits highly competitive generation performance, achieving an FID of 1.80 and 1.90 on 256x256 and 512x512 class-conditional image generation, respectively.

[CV-139] Context without Commitment: Robust Dense Correspondence under Non-Rigid Deformation

链接: https://arxiv.org/abs/2609.37071
作者: Yuzhen He,Sara Homscheid
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Non-rigid point-cloud registration aims to find the corresponding target point for each point on a deforming source surface. Point-level matching keeps the full target cloud available, but correspondence becomes ambiguous when different regions have similar local geometry. Regional or coarse-to-fine methods provide broader spatial context, but an incorrect regional match can exclude the correct correspondence before dense matching. We propose CoCo-Reg, which uses regional patches to enrich dense point features without allowing patch predictions to restrict the final point-level search. CoCo-Reg constructs farthest-point-sampled patches, exchanges geometric information within and between source and target, supervises patch similarity using identity-corrected point overlap, and projects the resulting regional information back to dense point features. The final registration stage still scores the full target cloud before global point-level candidate selection. On 726 held-out ModelNet10 objects across nine deformation levels, two established learning-based baselines obtain mean correspondence errors of 0.1993 and 0.1921, whereas CoCo-Reg obtains 0.0547. Relative to its point-level baseline, this is a 72.6% reduction. CoCo-Reg achieves lower correspondence error on 92.3% of paired test objects and reduces the mean fraction of points with error above 0.1 from 47.3% to 17.3%. Chamfer distance and HD95 decrease in the same direction, and CoCo-Reg remains lower across all tested deformation levels. These results support using regional context for dense non-rigid correspondence without imposing a hard patch-level restriction on the final search. Because evaluation uses one checkpoint per method, the reported gains characterize the complete systems rather than the isolated causal contribution of an individual component. Code will be made publicly available.

[CV-140] Spatial-OPSD: Self-Improving Spatial Reasoning via Label-Free Self-Distillation

链接: https://arxiv.org/abs/2609.37055
作者: Zhenyu Liu,Zhangquan Chen,Keyi Chen,Mingze Sun,Xiang An,Haodong Jing,Ruqi Huang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Vision-language models (VLMs) increasingly operate in embodied and spatially grounded settings, where accurate understanding of depth, viewpoint, and three-dimensional relations is essential. However, improving spatial reasoning typically relies on ground-truth answers, answer-derived rewards, or other forms of task-specific supervision. We introduce Spatial-OPSD, a label-free self-improvement framework that instead exploits spatial structure naturally available from perception and reconstruction tools. During training, a privileged teacher receives automatically obtainable spatial priors, such as depth, reconstructed 3D relations, and camera geometry, while the student observes only the original visual-language input. On trajectories sampled by the student itself, the teacher provides dense token-level supervision, allowing the student to internalize spatial knowledge without ground-truth answer labels or privileged information at inference time. To extend this supervision beyond a single round, we adopt a round-wise recursive training scheme: the teacher remains frozen within each round to provide a stable learning target, and the improved student initializes both teacher and student in the next round, where privileged spatial priors re-establish an informative teacher–student asymmetry. This enables repeated self-improvement while avoiding a rapidly moving teacher during optimization. Across four VLM families, a single round of Spatial-OPSD consistently improves the five-benchmark average, while three rounds further push a strong spatially specialized model to the open-source frontier, achieving the highest average among the open models and the best results on three of five spatial reasoning benchmarks. Our code is available at this https URL. Subjects: Computer Vision and Pattern Recognition (cs.CV) Cite as: arXiv:2609.37055 [cs.CV] (or arXiv:2609.37055v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2609.37055 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[CV-141] OmniRoute: Mapping Temporal Semantic Evidence to Audio-Visual Token Budgets for Efficient Omnimodal Large Language Models

链接: https://arxiv.org/abs/2609.37052
作者: Yuchen Deng,Zidang Cai,Feidiao Yang,Yufei Wang,Jie Wang,Hai-Tao Zheng,Yuxing Han
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Omnimodal large language models (Omni-LLMs) encode audio and visual streams into temporally interleaved token sequences for multimodal reasoning. However, processing long audio-visual token sequences incurs substantial prefill costs. Existing compression methods have made progress, but often overlook temporal changes in audio-visual semantic relevance. Motivated by temporal variation and local continuity, we propose OmniRoute, a training-free, two-stage compression framework. First, Temporal Evidence-Guided Budgeting (TEGB) derives chunk-wise modality preferences and initial leading-modality budgets from semantic relevance and local content variation. Second, Budget-Constrained Semantic Compression (BCSC) compresses the leading modality and then calibrates the follower’s retention target using the actual retained fraction. For video, it combines spatiotemporal grouping with query-guided selection; for audio, it selects tokens based on encoder attention and query relevance, then merges residual tokens into context anchors under visual guidance. Experiments on four representative benchmarks demonstrate a better trade-off between inference efficiency and performance than competitive baselines. The code and interface will be released to facilitate further research.

[CV-142] NHO: A Neural Hamiltonian Operator for Anchor-based Region Localization and Dense Correspondance

链接: https://arxiv.org/abs/2609.37048
作者: Jing Li,Yawei Luo,Xiangze Meng,Ying Li,Tieru Wu,Rui Ma
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Non-rigid partial-to-full shape correspondence from sparse anchors requires identifying the corresponding region on the full surface and recovering dense correspondences between the partial shape and that region. We present NHO, which combines sparse anchors with the intrinsic geometry of the partial shape to learn a neural Hamiltonian operator whose localized eigenspace encodes both the region support and intrinsic coordinates for dense correspondence. NHO parameterizes the Hamiltonian potential as an intrinsic neural field and optimizes it using anchor evidence together with spectral and geometric constraints. To resolve the spatial ambiguity left by sparse anchors, we introduce reciprocal refinement between operator estimation and correspondence recovery. At each round, the current eigenspace provides spectral coordinates and restricts matching to its induced support, while geometrically reliable correspondences provide additional evidence for updating the potential. After refinement, aggregated eigenfunction energy yields the final localization, and the recovered map initializes dense correspondence refinement. Experiments demonstrate competitive accuracy on both tasks and robustness to uniform scaling and rotation.

[CV-143] Multi-Depth Temporal Fusion for Feedforward Locally Trained Spiking Neural Networks

链接: https://arxiv.org/abs/2609.37047
作者: Aidin Attar,Eleonora Cicciarella,Michele Rossi
类目: Neural and Evolutionary Computing (cs.NE); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 22 pages. Submitted to Neurocomputing. Code available at this https URL

点击查看摘要

Abstract:We propose a new spiking neural network (SNN) design to process static images and event streams using time-to-first-spike (TTFS) latencies. Our key research question is which architectural choices best accommodate local and online learning in multi-layer convolutional SNNs. This question is addressed via an original framework combining residual-like connections with multi-depth feature aggregation and consensus. The full SNN pipeline features an early-vision front end, to convert raw visual data into sparse spike latencies, a four-layer convolutional backbone trained layerwise with unsupervised spike-timing-dependent plasticity (STDP), a deterministic Multi-Depth Temporal Fusion (MDTF) and a final classifier trained with reward-modulated spike-timing-dependent plasticity (R-STDP). Rather than replacing early features in deeper layers, the proposed MDTF preserves early temporal evidence, adding sparse residual events from intermediate layers, and incorporating deeper features only when they agree in time with earlier representations. The resulting architecture is experimentally validated across MNIST, Fashion-MNIST, CIFAR-10, and N-MNIST, delivering strong classification performance under a fully local learning regime. Selective multi-depth fusion significantly outperforms traditional STDP/R-STDP baselines on higher-variability visual tasks (achieving +18.2 pp on Fashion-MNIST and +29.2 pp on CIFAR-10). Furthermore, activity-budget analyses show that the network retains high accuracy even when removing a large fraction of late or weak spike events, confirming its high data efficiency and reduced event-processing requirements. The codebase is publicly available at this http URL temporal-fusion-snn.

[CV-144] Speed in the Blind Spot: An Interpretability Analysis of Dynamic Perception in VLMs for Autonomous Driving

链接: https://arxiv.org/abs/2609.37046
作者: Katharina Winter,Stefan Englmeier,Fabian B. Flohr
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible

点击查看摘要

Abstract:Vision-Language Models are increasingly used in autonomous-driving systems, yet their ability to recover dynamic physical state from visual input remains insufficiently characterized. We study velocity understanding as a controlled diagnostic across three tasks: surrounding-agent speed, current ego speed, and short-horizon future ego-speed proposal. On nuScenes, we evaluate open-weight general-purpose and PhysicalAI VLMs, together with the driving-oriented Alpamayo-1.5 Vision-Language-Action model, using multiple input and output formulations. We combine verbal evaluation with temporal perturbations, counterfactual ego-speed hints and linear probes of hidden representations. The tasks exhibit distinct failure modes. Surrounding-agent speed is weakly encoded in an agent-specific form, whereas current ego speed is often internally accessible but poorly verbalized: continuous probes achieve 4.7-5.8 km/h MAE compared with 10.2-16.8 km/h MAE for verbal outputs. Multiple frames provide inconsistent verbal gains to single frame inputs, and frame order is rarely exploited. Under non-optimized QLoRA, task-specific adaptation improves both task-relevant latent speed representations and verbal readout, but continuous surrounding-agent speed estimation remains weak, while most future-speed gains survive frame shuffling, indicating limited temporal grounding. Driving specialized Alpamayo-1.5 shows stronger latent representations for surrounding-agent and future ego speed, while current ego-speed decodability is comparable and substantial probe-verbal gaps remain. Thus, driving specialization can strengthen motion representations but does not guarantee stronger encoding across both scene and ego states or reliable readout. The results show that plausible planning outputs do not necessarily imply reliable recovery or temporal grounding of the underlying dynamic state.

[CV-145] GleanVID: Complementary Token Selection for Efficient Video Large Language Models

链接: https://arxiv.org/abs/2609.37042
作者: Shuo Yang,Changbai Li,Rui Tang,Xinyu Zhao,Linlin Yang,Baochang Zhang
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Video Large Language Models (VideoLLMs) have achieved strong video understanding capabilities but incur substantial inference overhead due to the large number of visual tokens. Existing VideoLLM token compression methods largely rely on selection-independent scoring, overlooking cross-frame complementarity and consequently retaining redundant evidence across frames. Instead, we view video token selection as a progressive evidence accumulation process. It aims to retain visual evidence that is individually informative and collectively complementary under a limited token budget. Building on this insight, we introduce GleanVID, a training-free inference acceleration framework for VideoLLMs. Specifically, GleanVID first allocates the global token budget across frames according to temporal novelty and then selects tokens by jointly considering local representativeness and subspace complementarity, thereby preserving richer and less redundant visual evidence. Extensive experiments across diverse VideoLLMs and benchmarks demonstrate that GleanVID consistently achieves state-of-the-art performance. Notably, with only 25% of visual tokens, GleanVID preserves 98.6% of Qwen3-VL’s original performance while reducing its prefill latency by 44.7%. On LLaVA-OV-7B, GleanVID at a 25% retention ratio even slightly surpasses the original model.

[CV-146] NowcastDiT: Diffusion Transformers are Effective Precipitation Nowcasters

链接: https://arxiv.org/abs/2609.37038
作者: Haoran Xu,Xingzhuo Guo,Yuchen Zhang,Jincheng Zhong,Jianmin Wang,Mingsheng Long
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注: 28 pages, 11 figures

点击查看摘要

Abstract:Precipitation nowcasting demands accurate short-term forecasts under strong spatiotemporal variability. Diffusion models are well suited to modeling complex precipitation distributions, yet existing approaches often introduce increasingly specialized designs, leaving the capability of a standard diffusion architecture underexplored. We show that a standard Diffusion Transformer already provides a simple and scalable foundation for precipitation nowcasting, with domain-specific requirements accommodated naturally within its design space. Based on this principle, we develop NowcastDiT and instantiate this flexibility through two complementary adaptations: a dynamics-aware noise prior for temporally coherent forecasts, and end-to-end reinforcement learning with timestep-aware rewards for meteorological skill. Experiments on SEVIR and MRMS benchmarks show that NowcastDiT achieves state-of-the-art performance in both perceptual quality and meteorological skill. These results suggest that standard DiT can serve as an effective foundation for precipitation nowcasting.

[CV-147] UniBuild: Unified Building Mapping From Multi-Source Optical Remote Sensing Imagery With Detail Decoding and Geometry Regularization

链接: https://arxiv.org/abs/2609.37031
作者: Wei Huang,Chenying Liu,Yilei Shi,Xiao Xiang Zhu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Building extraction from optical remote sensing (RS) imagery is fundamental to urban mapping, yet existing methods are often dataset-specific and generalize poorly to unseen domains. Their practical use is also limited by insufficient detail recovery and weak geometric regularization, leading to blurred boundaries, irregular shapes, and merged adjacent buildings. To address these issues, we propose UniBuild, a unified building extraction framework for multi-source RGB optical RS imagery. First, a unified multi-dataset training scheme is constructed over heterogeneous RGB optical datasets to learn transferable building representations across sensors and resolutions. Second, a novel detail-preserving HR-DPT decoder is designed to integrate high-level semantic features with high-resolution spatial features, enhancing building detail recovery. Third, geometry-aware regularization is introduced through a structure-tensor-based direction-aware loss for boundary direction consistency and a saddle-aware loss for suppressing false activations in narrow inter-building gaps under low-resolution conditions. We train and evaluate UniBuild on multi-source RGB optical datasets, including 10 public high-resolution datasets and two self-collected low-resolution datasets. Experiments show that UniBuild consistently improves building-region accuracy, boundary sharpness, and adjacent-building separation across diverse datasets. It also generalizes well to unseen domains and supports practical building extraction from RGB optical RS imagery up to 10,m resolution. The predicted masks can be further converted into GIS-compatible building footprints through simple polygonization. The trained model and inference code are released at this https URL.

[CV-148] MotionInsight: Diagnosing Object Motion Deficiencies in Generated Videos

链接: https://arxiv.org/abs/2609.37030
作者: Jiahao Zhan,Yongrui Ma,Qunliang Xing,Xuanyu Zhang,Jingqi Tong,Junlin Li,Li zhang,Shijie Zhao,Tianfan Xue
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Despite rapid progress in video generation models, they still exhibit obvious motion deficiencies, often manifested as incorrect object motion. However, most existing video quality evaluations focus on aesthetic quality or text-video alignment. To address this gap, we study object-centric motion fidelity assessment, evaluating target objects along object consistency, motion continuity, and physical plausibility. To achieve this, we first introduce VidMotion, a diagnostic dataset of 6,879 videos with designated moving objects and fine-grained annotations including dimension-wise scores and failure causes. We further propose MotionInsight, a diagnostic evaluator that shifts assessment from implicit RGB-frame observation to explicit motion-space diagnosis. By constructing motion-aware representations, MotionInsight makes subtle motion deficiencies more observable. We also introduce motion-specific rewards during GRPO to transform observed motion into a diagnostic assessment. Experiments demonstrate that MotionInsight provides an effective basis for diagnosing object motion deficiencies, producing human-aligned scores along three dimensions and grounded explanations.

[CV-149] Back2Struct: Making Structured Images Editable Again

链接: https://arxiv.org/abs/2609.37016
作者: Pengyu Yan,Yixin Wu,Yunjie Tian,David Doermann
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Structured images, such as diagrams, charts, and flowcharts, are inherently symbolic and can be compactly represented in an editable format, yet in practice, they are often rendered as images, and therefore not graphically editable. This mismatch presents a significant challenge for researchers, engineers, and designers who wish to incorporate modified versions of existing graphic content into new materials without manually reconstructing it. In this study, we presentBack2Struct, which “makes structured images editable again” by directly recovering vector graphics code (SVG / XML) from image representations. Given an image of a structured graphic, Back2Struct predicts semantically object-level SVG / XML code that explicitly encodes text, shapes, topology, and layout, rather than performing low-level pixel vectorization. The generated code can be seamlessly imported into tools such as PowerPoint, allowing users to edit, refine, restyle, and reuse graphic content while preserving structural fidelity. Beyond supervised fine-tuning on ground-truth SVG token sequences, we further optimize Back2Struct with reward-based learning to better match deployment-time requirements: the output should be syntactically valid, properly concise, and visually faithful to the input diagram. Specifically, we design a composite reward that jointly encourages SVG / XML compilability, length consistency with the reference code, and structural or semantic similarity between the generated and ground-truth graphics. These complementary signals guide the model to produce SVGs that are not only closer to the training distribution, but also more complete, editable, and renderable in practice. Experiments show that Back2Struct improves accuracy, editability, validity, and user alignment over baselines. Dataset and code are available at: this http URL

[CV-150] RBF-GNN: Rational Basis Functions for Pseudo-Coordinate based Graph Convolutions

链接: https://arxiv.org/abs/2609.37015
作者: Paweł Batorski,Abtin Pourhadi,Paul Swoboda
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:We propose RBF-GNN, a new pseudo-coordinate based graph neural network architecture that takes into account Euclidean, spherical or angular coordinates and uses them to induce a powerful spatial inductive bias. Similar in architecture to SplineCNN, we improve upon the latter by replacing the less efficient sparse-activation based B-splines whose number grows exponentially with dimension by rational Padé basis functions. For effective training we propose a spline-subspace initialization and a variance-preserving weight rescaling. Experimentally, we evaluate on a number of popular neural network architectures that use SplineCNNs. We replace only the SplineCNNs with RBF-GNN. We achieve improved results, including on semantic keypoint matching, shape matching, event based camera computer vision tasks. We will make our implementation publicly available upon acceptance of the paper.

[CV-151] Embedded Bi-Temporal Building Damage Assessment for On-Board Data Reduction

链接: https://arxiv.org/abs/2609.37013
作者: Thomas Goudemant,Benjamin Francesconi,Marjorie Bellizzi,Adrien Dorise
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 8 pages. Accepted at OBPDC 2026 (International Workshop on On-Board Payload Data Compression), Barcelona, October 2026

点击查看摘要

Abstract:Rapid assessment of building damage after natural disasters is essential to support emergency response. Earth Observation satellites can acquire relevant imagery shortly after an event, but exploitation is limited by uplink and downlink capacity and by ground-processing latency. We address this with a bi-temporal building damage assessment pipeline built on a siamese detector derived from YOLOX, designed to compress information at both ends of the ground/space link. On the ground, pre-disaster reference images are encoded into a compact latent space – compressed by up to a factor of 64 – and uplinked to the satellite. On board, this reference is compared with a fresh post-disaster acquisition so that the downlink carries only actionable object-level products, bounding boxes and damage classes, instead of full scenes. This cuts the data exchanged in both directions, while on xBD the strongly compressed reference still preserves most of the detection performance. Because on-board acquisitions suffer from residual pre/post co-registration errors, we introduce a latent-space shift estimation and correction module that regresses the global offset from the coarse feature level and realigns the post-disaster features before fusion. It substantially improves robustness to de-registration – especially under large shifts, where fusion-only variants collapse – while also raising nominal accuracy and remaining compatible with the strongest compression. We finally port the pipeline to two embedded targets, a Xilinx Versal VCK190 and an NVIDIA Jetson AGX Orin, and report hardware performance (latency, throughput, power efficiency). The core detector and its compression port cleanly to both, but the operators needed for long-range robustness survive only on the Jetson GPU, whereas the Versal DPU does not. Comments: 8 pages. Accepted at OBPDC 2026 (International Workshop on On-Board Payload Data Compression), Barcelona, October 2026 Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) ACMclasses: I.4.8; I.2.10 Cite as: arXiv:2609.37013 [cs.CV] (or arXiv:2609.37013v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2609.37013 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[CV-152] World2Motion: Turning Video World Models into 3D Human Motion Generators

链接: https://arxiv.org/abs/2609.37004
作者: Tu Fangyuan,Xiangyue Zhang,Yiyi Cai,Yichen Peng,Kunhang Li,Bo Zheng,Zhixiang Wang,Kaipeng Zhang,Erwin Wu,Haoran Xie,Haiyang Liu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 15 pages, 6 figures

点击查看摘要

Abstract:We present World2Motion, a framework that generates scene-aware 3D human motion and corresponding video from a single image and a text prompt. While existing 3D motion generators learn from motion datasets, their generalization is constrained by limited coverage of environments. In contrast, video world models such as Cosmos 3 offer broader environmental priors but are not designed for full-body motion generation; recovering motion from their generated videos requires costly two-stage inference. To address these, we turn Cosmos 3 into a single-stage 3D motion generator. This adaptation has two challenges: the scarcity of paired video–motion data and temporal instability in the generated motion. First, we construct a training dataset combining synthetic video–motion pairs with real videos paired with estimated 3D motion. Second, we propose a shift-decoupled noise schedule that assigns different noise levels to video and motion through shared denoising progress. This design accommodates the different denoising requirements of the two modalities, reducing motion jitter. Experiments on a multi-source interaction benchmark show that World2Motion has better motion–text alignment and scene interaction compared with the evaluated 3D motion generators. It also matches the interaction success rate of the two-stage baseline while achieving approximately 3.3 \times faster inference.

[CV-153] VesselBench-800K: A Large-scale Perception Benchmark for Multimodal Vessel Detection Counting and Density Estimation

链接: https://arxiv.org/abs/2609.37003
作者: Danfeng Hong,Chenyu Li,Jocelyn Chanussot
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Vessel perception from space is crucial for a wide range of maritime applications, from traffic monitoring to environmental protection. However, most existing datasets predominantly focus on general object detection tasks in optical remote sensing (RS) images. Relying solely on single-modality optical RS images proves inadequate for effectively perceiving vessel objects in complex maritime scenarios, where ever-changing weather conditions (e.g., clouds and rain), the need for day-and-night coverage, and the inherent limitations of a single imaging modality pose significant challenges. To fill this gap, we introduce VesselBench-800K, the largest-to-date benchmark dataset on a global scale for vessel perception in multimodal RS images. As its name suggests, VesselBench-800K comprises 800,000 images, each at a resolution of 512x512 pixels, specifically curated for vessel perception tasks such as detection, counting, and density estimation. These multimodal image pairs (i.e., optical, SAR) are collected from diverse platforms, sensors, scenes, shooting heights, and synthetic sources, spanning spatial resolutions from 4.5m to 0.1m. Furthermore, we evaluate numerous state-of-the-art detection, counting, and density estimation models on VesselBench-800K through both qualitative and quantitative comparisons. By revealing previously unrecognized cues, this dataset holds immense potential to significantly advance our understanding of marine traffic. Our VesselBench dataset will be publicly available at this https URL to support and contribute to community development.

[CV-154] Visual Parallel Search: Learning to Search High-Resolution Images with Parallel Tile Inspection and Adaptive Zoom

链接: https://arxiv.org/abs/2609.37002
作者: Xijia Tao,Yihua Teng,Xinyu Fu,Cheng Gong,Ziru Liu,Xudong Xie,Rui Liu,Lingpeng Kong
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:High-resolution visual question answering often fails because a multimodal model does not acquire the small, spatially localized evidence needed to answer a question. Sequential zooming can recover detail, but it asks the main model to choose a region before obtaining a reliable overview. We introduce VPS, a visual parallel-search framework in which a main agent first invokes grid_search to inspect image tiles in parallel with question-conditioned sub-agents, and then adaptively invokes zoom_in PSisual Parallel Search improves mean accuracy over dedicated zoom-only search in 14 of 15 same-model comparisons, with gains up to 8.0 points and especially strong improvements for smaller main models. ZoomBench retains an approximately 3.2-point gain at every tested size. We further develop a supervision pipeline with hint-free verification and a paired role-specific GRPO surrogate for learning the controller and tile-reader roles. SFT improves observed accuracy on all five benchmark splits, including a 4.17-point gain on HR-Bench 4K. Role-specific RL further reshapes search behavior: main-only RL reduces mean tool use from 2.65 to 2.11 with similar pass@1 in an internal four-response evaluation, while external accuracy changes are mixed. Joint training reveals an asymmetry between local evidence reading and global search control. Together, these results support VPS as an effective inference-time scaffold and a trainable decomposition for visual evidence acquisition.

[CV-155] Parameterized Stripe Attention for Efficient Video Generation

链接: https://arxiv.org/abs/2609.37001
作者: Xingyu Jia,Baole Ai,Ang Wang,Kang Zhao,Yong Li
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Diffusion Transformers (DiTs) enable high-quality video generation but suffer from substantial inference latency, primarily attributable to the computationally expensive full spatio-temporal attention. While sparse attention methods offer potential solutions, existing approaches face an inherent flexibility–efficiency dilemma: predefined masks lack the flexibility to capture diverse attention patterns, while runtime-determined masks introduce overheads and sacrifice hardware efficiency. We identify the lack of a unified structural characterization of DiT attention as a key limitation of existing methods, and establish that video DiT attention exhibits \textbfperiodic diagonal stripe structures along both temporal and spatial dimensions. To formally encode these structured patterns within a single efficient kernel, we present \bf PSA, a parameterized stripe attention that formalizes the observed stripe regularity, unifying diverse attention patterns for efficient mask generation. This unified representation enables a single hardware-efficient CUDA kernel to process all sparse patterns, achieving FlashAttention-3-level Model FLOPs Utilization. To determine optimal sparsity configurations, we propose a training-free offline search algorithm that automatically maximizes sparsity under a specified error tolerance for each attention head. Experiments on HunyuanVideo and Wan~2.1 demonstrate that PSA achieves 1.57 \times and 1.37 \times end-to-end speedups over FlashAttention-3 baselines, with acceptable visual quality degradation.

[CV-156] Salt: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation

链接: https://arxiv.org/abs/2609.36995
作者: Xingtong Ge,Yutong Wang,Lunjie Zhu,Haitao Lin,Fangyu Lin,Yushi Huang,Xin Zhang,Yi Zhang,Yu Liu,Jun Zhang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Image and Video Processing (eess.IV)
备注: under review

点击查看摘要

Abstract:Few-step streaming audio–video generation requires both causal modeling and step distillation, yet standard training recipes face two context-related challenges. Teacher forcing pairs clean history with a noisy target, but supervises predictive contextual representations only indirectly through velocity prediction. Meanwhile, directly reusing bidirectional score models in causal Distribution Matching Distillation (DMD) creates a mismatch between generation and scoring contexts. We address these challenges with Salt++, a two-stage post-training framework comprising Causal Self-Flow (CSF) and context-aligned autoregressive DMD. CSF exploits contextual information asymmetry by varying the history while keeping the noisy target fixed: a noise-mixed-history student aligns its intermediate representations with those of a clean-history exponential-moving-average teacher. This self-supervised signal encourages the student to extract semantic information and improves cross-modal alignment. Context-aligned AR DMD shares the causal mask and prefix across generator sampling, fake-score training, and real-score evaluation to match generated and reference distributions under a block-conditional KL objective. With calibrated teacher guidance, it performs clean-prefix few-step distillation and then adapts to generated histories without switching objectives or requiring separate consistency distillation. At 480p, Salt++ improves visual and motion quality by 57% and 45% over OmniForcing on JavisBench under the same 4-step causal setting. A separate scale-wise post-training stage extends Salt++ to 4-step 1664\times960 generation, outperforming bidirectional LTX-2 on six of seven reported metrics. Project page: this https URL

[CV-157] UltraMatch: Transport Path Routing for Ultra-Fast and Memory-Efficient Image Matching

链接: https://arxiv.org/abs/2609.36980
作者: Jiajun Le,Yifan Lu,Zizhuo Li,Lei Cao,Junjun Jiang,Jiayi Ma
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 18 pages, 5 figures

点击查看摘要

Abstract:Despite recent advances in accuracy and efficiency, coarse matching remains an indispensable yet costly stage in existing semi-dense matchers due to dense token-level matching. We present UltraMatch, an ultra-efficient and scalable semi-dense matching framework that bypasses the quadratic computation and memory cost of dense token-level matching by routing only a small fraction of candidate matching paths. At its core, a lightweight Transport Path Router operates on coarse block representations to rank candidate target blocks for each source block and retain only a small set, restricting subsequent token-level matching to the selected paths and avoiding the construction of the full token-to-token matching matrix. We further design a sparse global Dual-Softmax that performs matching only over the routed block candidates while retaining global competition across the sparse matching space. Beyond matching acceleration, UltraMatch employs deployment-oriented structural reparameterization for feature extraction and a tiny fine matching head with shared parameters, further reducing inference cost and memory consumption. UltraMatch achieves competitive accuracy among semi-dense matchers, while running 1.67 \times faster than SuperPoint+LightGlue with only 0.44 GiB peak inference memory. Its scalability enables inference at up to 6K resolution on a single RTX 3090, whereas existing semi-dense matchers run out of memory before reaching 2K. Our routing strategy is also transferable, delivering about 2 \times end-to-end speedup in EDM and ELoFTR without accuracy loss. The project repository is available at this https URL.

[CV-158] A Dual-Track Curation-and-Classification Framework for Resolving Ground-Truth Label Noise in Operational Sentinel-2 Wheat Area Estimation

链接: https://arxiv.org/abs/2609.36975
作者: Kasimali Agharia,Ujjwal Kumar Gupta
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 21 pages, 4 figures, 5 tables. Preprint. Not yet peer-reviewed

点击查看摘要

Abstract:Operational estimation of wheat-cultivated area is persistently constrained by discordance between administrative record-keeping and remotely sensed classification products. We address this administrative reference discordance for the 2022 Rabi season in Patiala district, Punjab, India, using a thirteen-timestep Sentinel-2 NDVI time series. A curated 849-sample reference dataset, developed through an iterative rule-based bootstrapping procedure, underpins both a feature sensitivity analysis and an operational classifier. Feature sensitivity independently assessed via Cohen’s d and gradient-boosted information gain converges on the February-to-March grain-fill window as most discriminative. Four classifiers (1D-CNN, LSTM, hybrid CNN-LSTM, and XGBoost) were benchmarked on an identical 679/170 sample split. XGBoost achieved the highest overall accuracy (78.82%) against deep-learning baselines (64-66%), consistent with tree-based ensembles’ favourable parameter-to-sample ratio in low-sample regimes. At full-population deployment across 36.25 million valid district pixels, the operational classifier attained 86.31% precision and 71.05% recall. The predicted wheat extent deviated by only +2.99% from the official tabular target, whereas the government’s spatial reference mask exhibited a +25.11% positive area bias against the identical target. This asymmetry indicates that a classifier trained on an auditor-curated reference set reconciles more closely with the official tabular area than the spatial product conventionally used to validate it. We present this dual-track curation-and-classification framework as a methodological reference for crop-area reconciliation in label-noisy administrative settings.

[CV-159] Structured Visual Target Learning For Cross-Subject eeg-to-image retrieval

链接: https://arxiv.org/abs/2609.36971
作者: Salini Yadav,Taveena Lotey,Mickaël Coustaty,Pravendra Singh,Partha Pratim Roy
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Cross-subject EEG-to-image retrieval requires a neural represen- tation trained on source subjects to remain aligned with a visual embedding space for an unseen subject. Whereas existing methods primarily focus on the EEG side, we address this problem from the perspective of the visual target. Our approach preserves the spatial information of the Perception Encoder, converts its patch grid into a compact set of learned visual views, and aggregates them for each image with a block-structured, content-dependent router. The target is learned jointly with the EEG encoder through contrastive learning with MMD regularization across source subjects. For deployment, we propose a training-free representation refinement that aligns frozen embeddings without updating either encoder. Under leave- one-subject-out evaluation on THINGS-EEG2, the structured target achieves 35.3%/65.6% Top-1/Top-5 accuracy, the best among com- pared methods. Refinement raises this to 48.1%/77.1%, an 18.5% Top-1 gain over the strongest compared method, improving all ten held-out subjects.

[CV-160] Prior-Driven Enhancements in 3D Gaussian Splatting: Normals and Depths Regularization

链接: https://arxiv.org/abs/2609.36969
作者: Gyeonggwan Lee,Seunghwan Hong,Junghun Suh
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 7 pages, 2 figures, 1 table. Oral presentation at ISPRS Geospatial Week 2025 (Dubai). Project page: this https URL Code: this https URL

点击查看摘要

Abstract:3D Gaussian Splatting (3DGS) is a state-of-the-art technique for 3D scene rendering, offering high efficiency and excellent visual quality. However, because 3DGS relies on an initial sparse point set from Structure-from-Motion (SfM) and view-dependent properties, it can suffer from geometric inaccuracies and visual artifacts, particularly in complex scenes. To address these challenges, we propose an improved 3DGS approach that regularizes the optimization process by integrating geometric priors, including surface normals and dense depth information. Surface normal regularization improves geometric consistency by aligning Gaussian covariance with local surface structures, while dense depth priors combined with an initial points from SfM enhance per-pixel depth estimation, increasing accuracy and reducing ambiguities. These enhancements enable robust handling of diverse and complex real-world scenarios, minimizing visual distortions and improving reconstruction quality across various environments. To validate our method, we evaluate it on challenging datasets, including street-view scenes and highly reflective environments, while testing it across multiple SfM pipelines. Our results demonstrate compatibility across diverse environments and highlight the robustness of our approach. Experimental findings further show that our method enhances geometric accuracy and visual quality, establishing a reliable solution for real-time 3D scene rendering in complex environments.

[CV-161] Beyond Readability: Evaluating Task Information Recoverability

链接: https://arxiv.org/abs/2609.36957
作者: Yiwei Liu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Direct visual readability and task-information recoverability are different quantities. Failure to decode a target from a fixed observation need not eliminate access to that target through another recovery route. We develop an evaluation perspective that makes the observation, query, target, and available knowledge explicit and measures the overlap between routes’ success sets. For information available on the original visible surface under suitable imaging conditions, direct optical recovery reads the target from the image, optionally after restoration; entity-linked recovery uses residual visual evidence to identify the depicted entity and accesses its target through an entity–attribute relation in a specified knowledge resource. Such access can draw on stored knowledge or an external source. A controlled book-cover study instantiates external access with a fixed title–author catalog, comparing optical author recovery with visual title resolution and deterministic lookup under resolution degradation. Entity-linked successes persist across the tested vision–language models, revealing information access beyond the tested direct visual frontier despite substantial differences in absolute performance. A substantial optical-only region remains. These complementary outcomes show why visual degradation should be evaluated through the task information accessible along specified routes and knowledge resources, alongside direct readability.

[CV-162] DispFlow-GS: Displacement Flow Supervision with Motion Disentangling for Monocular Deformable 3D Gaussian Splatting

链接: https://arxiv.org/abs/2609.36940
作者: Thai Duy Nguyen,Haitian Zhang,Addison Lin Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Accurate dynamic scene reconstruction is important for robotic perception, where temporally consistent representations of dynamic environments are essential. Deformable 3D Gaussian Splatting (3DGS) models dynamic scenes through deformation fields, and recent methods incorporate motion supervision by aligning rendered Gaussian flow with optical flow. However, we find that such Gaussian-flow-based supervision provides only limited improvements in motion modeling. We identify a fundamental limitation of this supervision paradigm, namely a domain gap between rendered Gaussian flow and optical flow. To address this limitation, we propose a motion supervision framework built on Displacement Flow, which splats per-Gaussian 3D displacements onto the image plane to provide direct and stable optimization signals. We further disentangle scene motion from camera motion via intermediate-view rendering, enabling more reliable motion priors and targeted constraints on deformation and geometry. We also observe a discrepancy between motion fidelity and image-based evaluation, where improved motion awareness does not necessarily translate into better rendered image quality or higher image-based metric scores. Motivated by this mismatch, we introduce Deformation-Rendering Consistency (DRC), a motion-aware metric that measures the alignment between predicted deformation and rendering improvement. Experiments on dynamic scene benchmarks show substantial improvements in motion localization and motion–rendering consistency, reaching up to 39% and 6%, respectively, while image-based metrics change by only about 0.1%. These results confirm the observed mismatch between motion fidelity and image-based evaluation, demonstrating the significance of DRC for motion-aware evaluation.

[CV-163] WeLike2Party! In-Context Motion Transfer for Multi-Human Image Animation

链接: https://arxiv.org/abs/2609.36937
作者: Sangeyl Lee,Seunghyun Shin,Seungho Park,Wooseok Jeon,Hae-Gon Jeon
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Human image animation aims to transfer motion from a driving video to subjects in a reference image. Despite remarkable progress in video generation, achieving high-fidelity animation of multiple interacting subjects remains a challenge. Many existing approaches rely on explicit motion representations such as 2D skeletons or parametric body meshes and struggle to preserve identity-motion binding under inter-person occlusion. To address this limitation, we propose WeLike2Party, a multi-human animation framework built on direct in-context video conditioning without explicit pose or mesh extraction at inference. We further introduce Reference Asymmetric RoPE Conditioning to preserve fine-grained appearance details, and Identity Binding Supervision to associate each reference identity with its intended motion trajectory. To support cross-identity training, we construct MotionTwin, a large-scale synthetic dataset comprising 14.4K cross-identity video pairs with shared subject and camera motions, totaling 84.3 hours of photorealistic video. We additionally present MotionTwin-Bench, a cross-identity benchmark specifically designed to evaluate subject-level visual fidelity and identity-motion binding. Extensive experiments on MotionTwin-Bench and real-world videos demonstrate that WeLike2Party outperforms recent state-of-the-art methods in subject-level visual fidelity, identity-motion binding, and overall perceptual quality, particularly in multi-person interactions with substantial occlusion.

[CV-164] SFE-VGGT: Source-Free VGGT Distillation for Event-Based Monocular Depth Estimation

链接: https://arxiv.org/abs/2609.36929
作者: Thai Duy Nguyen,Addison Lin Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Recent event-based depth estimation methods successfully transfer geometric priors from vision foundation models via cross-modal distillation. However, their reliance on synchronized RGB-event pairs or depth annotations during training severely restricts practical deployment. To overcome this bottleneck, we propose SFE-VGGT, a novel source-free framework that distills the geometric priors of VGGT to the event domain without any paired RGB observations. Our core idea is to reconstruct surrogate frames directly from the target event stream to act as a frozen geometric teacher, entirely eliminating the need for genuine source RGB data. Crucially, as these surrogate frames inherently yield imperfect and spatially varying supervision, directly distilling from them propagates artifacts. To resolve this, we introduce a novel reliability-aware distillation strategy. This includes Density-Aware Feature Distillation to emphasize informative event regions, and Confidence-Weighted Depth Distillation to dynamically regulate supervision based on relative teacher-student prediction confidence. Meanwhile, we propose a Cross-Frame Relational Consistency loss that enforces temporal geometric stability using reliable inter-frame correspondences, bypassing the need for temporally consistent teacher’s depth. Extensive experiments demonstrate that, despite source-free, our SFE-VGGT closely matches the accuracy of RGB-dependent baselines under standard conditions and significantly surpasses them in challenging nighttime scenarios. Across MVSEC nighttime sequences, SFE-VGGT reduces the average 10 m depth error by 15.3% compared with EventVGGT. Moreover, our method exhibits robust zero-shot generalization across real-world datasets, proving that highly effective geometric priors can be transferred to event cameras using strictly source-free supervision.

[CV-165] rack-and-Complete: Learning Humanoid Skills from a Single Failed Human Video

链接: https://arxiv.org/abs/2609.36924
作者: Sarmad Idrees,Jongeun Choi
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: 8 pages, 6 figures, 6 tables. Project website: this https URL

点击查看摘要

Abstract:Learning humanoid skills from videos typically requires a successful human demonstration, which often demands custom data collection. Although failures have traditionally been treated only as negative examples in robot learning, they can still reveal a usable trajectory prefix before the task fails, as well as the intended outcome. To leverage this information from a failed-attempt video, we propose TRACC, a pipeline that imitates the useful portion of the motion trajectory and then completes the task based on the inferred task outcome. The usable motion prefix serves as prior knowledge until the failure occurs, after which the task-completion reward guides the policy to learn the intended task goal without requiring a successful task trajectory. We evaluate our method on six in-the-wild failed human tasks from the Oops! dataset. Our experimental results demonstrate the effectiveness of the proposed approach for learning from failed attempts when no successful demonstration is available. Thus, these findings establish failed human videos as a viable source of supervision for humanoid skill learning.

[CV-166] Seg3DParts: Segmentation-Grounded Controllable Part-Level 3D Generation NEURIPS2026

链接: https://arxiv.org/abs/2609.36918
作者: Jiantao Lin,Meixi Chen,Yingjie Xu,Chenbo Fu,Leyi Wu,Hao Chen,Yinchuan Li,Ying-Cong Chen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at NeurIPS 2026

点击查看摘要

Abstract:Part-level 3D assets are essential for editing, reassembly, and interaction, yet recovering such structure from a single image remains challenging due to occlusion, ambiguous boundaries, and the need for coherent multi-part reasoning. Existing approaches struggle to achieve both controllable part-level generation and coherent multi-part structure, as part identity and spatial allocation are typically inferred implicitly. We present Seg3DParts, a segmentation-grounded framework for controllable part-level 3D generation from a single image. By treating segmentation as an explicit grounding signal, our method defines part identity during generation, enabling each component to be anchored to a corresponding image region. To ensure coherent assemblies, we introduce structured cross-part interaction that allows components to exchange global context throughout the generative process. As a result, Seg3DParts directly generates well-aligned part meshes in a shared canonical space without post-hoc alignment, supporting flexible and controllable decomposition. We further introduce PartObjectNet, a large-scale dataset with over 200K objects and 1M annotated parts. Experiments demonstrate that Seg3DParts achieves superior geometry quality, cross-part coherence, and part-level controllability over existing methods.

[CV-167] Representation Dynamics Reveal Semantic Saliency and Similarity for Visual Token Pruning in MLLM s

链接: https://arxiv.org/abs/2609.36916
作者: Weixuan Li,Zikun Zhou,Xinyi Zhuang,Xinyan Guo,Rui Tian,Chuyao Zhang,Lin Gao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Preprint. 33 pages, 17 figures, 19 tables

点击查看摘要

Abstract:Multimodal large language models (MLLMs) incur high inference latency from long visual token sequences. Existing pruning methods commonly use attention maps or output features to estimate token importance or redundancy. Several recent approaches also exploit representation changes, but when and how these changes reflect foreground saliency and semantic consistency remain insufficiently understood. We analyze visual token representation dynamics across encoder depth and uncover two findings. First, the relationship between token update magnitudes and foreground saliency is layer-dependent: large token updates concentrate on foreground regions in two depth intervals, separated by several sink-dominated layers at intermediate depths. Second, similarities between token update directions better distinguish same-class from different-class tokens than those between encoder output features. Building on these findings, we propose MSDG-Prune, a training-free method that uses update magnitudes and directions to preserve salient and diverse visual information. Specifically, we group tokens by update-direction similarity and use query-weighted saliency derived from update magnitudes across a chosen depth window for group-wise token pruning. Extensive experiments across four MLLMs demonstrate the effectiveness and generalizability of MSDG-Prune. On LLaVA-NeXT, it retains 91.9% of uncompressed performance on average with only 5.6% of visual tokens, while achieving a 7.8x prefilling speedup. Code is available at this https URL.

[CV-168] SafeVantage: Vantage-Aware Memory for Reliable Embodied Decisions

链接: https://arxiv.org/abs/2609.36906
作者: Sean Hardesty Lewis,Zuyi Guo,Benwang Chen,Zirui Liu,Hongyi Lin,Heye Huang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Reliable embodied decisions under partial observability require informative observations and sufficient supporting evidence. However, semantic scores alone do not reveal which viewpoints justify a claim or where additional evidence should be acquired. We introduce SafeVantage, a vantage-aware semantic memory and active acquisition framework that retains each claim’s supporting views, camera poses, and estimated target location, keeping positive support distinct from search coverage. A learned candidate-observability model uses claim-grounded geometry to predict target visibility at reachable viewpoints. These predictions guide view selection through expected reduction in terminal decision loss, accounting for travel cost and geometrically distinct corroboration. A calibrated head then combines support, spatial consistency, and coverage to produce Yes, No, or Abstain decisions. We evaluate SafeVantage on a category-presence benchmark spanning 232 unseen ProcTHOR houses and 7,424 paired episodes per method and action budget. Compared with validation-selected equal-budget baselines, SafeVantage achieves macro-F1 gains of 24.7% and 12.0% at eight and twelve actions, respectively, with lower risk and higher answer rates at both budgets and 31.7% less travel at eight actions. Equal-input HM3D experiments show lower selective risk under fixed observations, while controlled ScanNet interventions show that restoring supporting views improves downstream VLM answers. Ablations further support the contribution of candidate observability to decision quality and acquisition efficiency. Results demonstrate the value of claim-level viewpoint evidence for connecting semantic memory, active acquisition, and reliable decision-making. Code is available at this https URL

[CV-169] DiffReID: Discriminative Diffusion Model for Object Re-Identification

链接: https://arxiv.org/abs/2609.36894
作者: Yingquan Wang,Pingping Zhang,Dong Wang,Huchuan Lu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by TIP2026. More modifications can be performed

点击查看摘要

Abstract:As a fundamental image processing task, object Re-Identification (ReID) aims to retrieve objects across non-overlapping cameras. Recently, with the development of deep learning, significant advancements have been made in object ReID. However, most existing methods suffer from generalization due to the limited size and diversity of ReID datasets. Meanwhile, current models tend to focus on extracting semantic patterns rather than learning identity-aware feature distributions. To address these issues, we propose a novel feature learning framework named \textbfDiffReID for object ReID. It leverages a discriminative diffusion model to gradually learn identity-aware distributions and generate identity-invariant features. More specifically, with the Contrastive Language-Image Pre-training (CLIP) model, we first obtain identity-aware text features by prompt tuning. Then, we propose a Vision-guided Noise Generator (VNG) to initialize probabilistic noises and gradually corrupt identity-aware text features. Afterwards, we take visual features as conditions and propose a Light Weight Denoiser (LWD) to denoise the corrupted text features step-by-step for identity-aware distribution learning. To obtain discriminative features, we further generate identity-invariant guided features from randomly sampling visual-guided noises. Finally, we propose a Mutual Enhancement Constraint (MEC) to facilitate mutual learning between visual features and guided features to enhance the representation robustness and discrimination. Extensive experiments on five object ReID benchmarks demonstrate that our method shows better results than most state-of-the-art methods. The source code is available at this https URL.

[CV-170] ProGuT: Label-Efficient Panoptic Segmentation for Forest Scenes

链接: https://arxiv.org/abs/2609.36891
作者: Pankaj Deoli,Karsten Berns
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Panoptic segmentation in forest environments is bottlenecked not by semantic quality but by instance separation; existing unsupervised panoptic approaches produce usable stuff maps but near-zero thing quality. Depth or flow-based instance discovery methods needs sensors that are not always available. We present ProGuT (Prototype Guided Training), which produces panoptic pseudo-labels without per-image training masks, needing only unlabeled images and one-time cluster-to-class mapping. ProGuT clusters CLIP patch features, then recovers trunk instances through multiscale geometric prior that falsifies non-trunk structures via structure-tensor. This is cheap compared to depth, flow or class-supervision methods to create pseudo labels. These are then used for downstream tasks which we evaluate against other unsupervised baselines. ProGuT achieves a Panoptic Quality (PQ) of 65.2 on Our-forest dataset (2.6x improvement over the initial pseudo-label quality) and reaches 65.9 mIoU on Freiburg Forest, outperforming unsupervised baselines like PiCIE (45.3 IoU) and STEGO(57.6IoU). Additionally, ProGuT outperforms existing unsupervised methods for class-agnostic trunk instance benchmark.

[CV-171] Less Supervision Better Generalization: Weakly Supervised Fake Region Localization in Diffusion-Edited Images NEURIPS2026

链接: https://arxiv.org/abs/2609.36882
作者: Junhee Lee,Donghyeon Jeon,Taeoh Kim,Beomyoung Kim,MyeongAh Cho
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to NeurIPS 2026

点击查看摘要

Abstract:Localizing AI-edited regions is essential for interpretable forensic analysis, but remains challenging due to subtle and spatially distributed artifacts that are misaligned with semantic or object boundaries. Existing approaches rely on pixel-level supervision from controlled editing pipelines, which is difficult to scale and can introduce misleading signals: artifacts frequently extend beyond annotated regions, while out-of-mask pixels are treated as authentic. This limits models’ ability to capture transferable evidence and generalize across generators and datasets. To address these issues, we propose ReGFLoW, a Reconstruction-Guided Fake Localization framework under Weak supervision, which is the first weakly supervised approach for diffusion-edited fake region localization. ReGFLoW requires only real/fake labels at the image-level and uses diffusion reconstruction errors as dense spatial guidance to inject them into both feature and score spaces. Furthermore, by artifact-centric multiple instance learning, ReGFLoW utilizes localized diffusion evidence without relying on semantic-affinity or boundary-based pseudo-mask priors. Extensive experiments demonstrate that ReGFLoW achieves stronger out-of-domain generalization than fully supervised learning baselines.

[CV-172] S4VY: Segment Anything in Feed-Forward 4D Visual Geometry

链接: https://arxiv.org/abs/2609.36875
作者: Jingdong Zhang,Xin Li,Jan Kautz,Wenping Wang,Chris Choy
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Accurate instance segmentation in dynamic scenes is important for downstream applications such as robotics and autonomous driving. Existing Segment Anything models operate primarily on 2D image or video masks and preserve identity through sequential memory, while promptable 4D instance segmentation built upon feed-forward visual geometry remains underexplored. We introduce S4VY, a Segment Anything model built on feed-forward 4D visual geometry. From a set of RGB observations, S4VY transforms shared visual-geometric features into an exhaustive set of class-agnostic 4D instance masks through a space-time query decoder, with each persistent object query binding one entity across all observations. This representation supports prompt-independent segmentation as well as point- and box- conditioned selection, without requiring a seed mask or temporal ordering. We further develop an agentic harness for natural-language grounding in the large observation space of a 4D scene. Active tree search identifies relevant frames without scanning every fixed window; a dual-stream grounder combines fine-grained VLM visual priors with geometry-consistent instance features through complementary bounding-box prediction and object-query matching; and an independent critic selects the final 4D instance mask from their predictions. Extensive experiments demonstrate state-of-the-art 4D instance segmentation and strong language-guided grounding performance under a unified evaluation spanning static and dynamic scenes.

[CV-173] S2T-Unet: A Structure-to-Style Framework for Inter-Modality MRI Translation

链接: https://arxiv.org/abs/2609.36866
作者: Yichao Liu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Inter-modality MRI translation aims to synthesize missing MRI modalities from available acquisitions, reducing the need for additional scanning while preserving clinically relevant anatomical information. However, existing image translation methods often learn intensity mappings without explicitly separating modality-invariant structural information from modality-specific appearance, which may lead to structural information loss or unrealistic image details. In this work, we propose S2T-Unet, a structure-to-style framework that explicitly models these two aspects. Specifically, vector quantization is introduced at the lower-level bottleneck to encode modality-invariant structural information using a learned discrete codebook. At higher levels, a modality transformation module uses decoder features to condition and transform encoder representations toward the target modality, thereby recovering modality-specific intensity and contrast information. Experiments on the IXI multi-contrast MRI dataset across four translation tasks demonstrate that S2T-Unet is comparable or outperform with state-of-art method.

[CV-174] Socialality Anchors: Towards Group-bounded Trajectory Prediction

链接: https://arxiv.org/abs/2609.36852
作者: Ziqian Zou,Conghao Wong,Qinmu Peng,Xinge You
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Trajectory prediction is a key component for understanding human behavior patterns in dynamic scenes. Researchers have devoted substantial efforts to modeling social interactions, especially group-wise interactions, since group membership often reflects shared intention, coordinated motion, and stable mutual adaptation, thus providing a persistent and semantically meaningful social prior for forecasting. However, existing group modeling methods may rely on a fixed threshold and infer groups mainly from agents’ relative positions within the observation window, overlooking the fact that grouping rules should be agent-specific, temporally coherent, and context-adaptive across diverse personalities, culturalities, and evolving interaction contexts. Inspired by human social perception that alternates between interpersonal distance in boundary-sensitive situations and relative speed consistency in dynamic interactions, we propose Socialality, a human-inspired trajectory prediction framework with interpretable Socialality anchors and an extended grouping window for stable, context-aware grouping inference. Concretely, Socialality introduces a duo-scalar-controlled grouping kernel Socialality that jointly leverages historical observations and short-term future trajectory previews to learn agent-specific grouping rules, and employs a group-wise perception mechanism to model in-group and out-of-group interactions in an intuitive and explainable manner. Furthermore, we conduct extensive experiments on standard benchmarks to demonstrate the performance gains of Socialality, and provide qualitative analyses and statistical studies of anchor distributions to verify the interpretability and stability of the proposed Socialality anchors.

[CV-175] RoXDrive: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving via Action-Faithful Rollouts

链接: https://arxiv.org/abs/2609.36851
作者: Hongbin Lin,Chaoda Zheng,Yiming Yang,Xiangyu Li,Shijia Chen,Jinhao Deng,Kangjie Chen,Dongbin Zhang,Jie Feng,Yu Zhang,Xianming Liu,Shuguang Cui,Boyang Wang,Zhen Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL Github: this https URL

点击查看摘要

Abstract:End-to-end autonomous driving policies are commonly trained via imitation learning on logged demonstrations without observing the consequences of their own actions, leading to causal confusion in closed-loop real-world deployment. To address this issue, reinforcement learning (RL) post-training offers a promising alternative by leveraging world models as interactive training environments to enable future scene generation for policy improvement. Nevertheless, existing approaches either rely on reconstruction-based simulators, offering limited counterfactual interaction, or adopt synthetic simulators to enable long-horizon closed-loop interaction at the cost of a substantial sim-to-real gap. Recently, video world models have exhibited the ability to generate realistic multi-step future rollouts but may not faithfully reflect action conditions, resulting in action-vision mismatch. In this paper, we introduce RoXDrive, a plug-and-play closed-loop RL framework that enables reliable policy optimization by identifying action-faithful world-model rollouts, consisting of two stages: 1) Model pre-training: In addition to imitation-based policy pre-training, we devise an Action-Vision Faithfulness Evaluator for inverse dynamics estimation with our geometry-aware auxiliary trajectory supervision, enabling long-horizon assessment of whether visual dynamics faithfully reflect the conditioning ego actions. 2) Action-faithful RL post-training: Agents iteratively interact with world models to form long-horizon scene rollouts, retaining only action-faithful ones for dense safety-aware scoring and scene-level closed-loop RL post-training. Extensive experiments on nuScenes and an in-house dataset with over 130K training scenarios demonstrate consistent gains across planners, reducing safety violations by 27.6% with DiffusionDrive on nuScenes and 33.7% with Qwen3-VL on the internal data.

[CV-176] GlassFormer: Learning Real-time Glass Segmentation using Radar-Depth Fusion IROS2026

链接: https://arxiv.org/abs/2609.36844
作者: Suhani Grover,Astik Srivastava,Viswas Dinesh,Avinash Sharma,K. Madhava Krishna
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: Accepted for presentation at IEEE IROS 2026. Code available at this https URL

点击查看摘要

Abstract:Transparent surfaces are ubiquitous in built environments, yet they remain a persistent failure case for robotic perception. RGB cameras perceive the background behind glass rather than the surface itself, while depth sensors such as LiDAR, time-of-flight, and RGB-D often return invalid or background measurements in transparent regions. As a result, systems that rely solely on optical sensing may misinterpret glass walls, doors, or mirrors as free space, compromising safe and reliable navigation. Existing glass segmentation approaches address this by learning visual cues such as reflections, boundaries, and semantic context from RGB images. While effective under favourable lighting and viewing conditions, these cues degrade in low-light environments, under glare, or when glass surfaces are featureless or partially occluded. In this work, we propose a multimodal framework that fuses millimetre-wave radar with RGB-D sensing for real-time transparent surface segmentation. Radar reflects strongly off glass surfaces, providing a geometric cue that remains reliable precisely where vision and depth fail. We exploit this cross-modal inconsistency to generate a radar-guided spatial prior, which is integrated into a lightweight transformer-based segmentation network, GlassFormer, via cross-modal attention. We report results on a mixed-condition test split covering all scene types and a dedicated low-light split designed to stress vision-only methods. GlassFormer achieves 0.88 mIoU on the mixed split, and 0.59 mIoU on the low light split, demonstrating substantial robustness gains over vision-only baselines while maintaining real-time performance on resource-constrained platforms.

[CV-177] Does the VGGT Family Need All Its Layers?

链接: https://arxiv.org/abs/2609.36842
作者: Fengyi Zhang,Holger Caesar,Xiangyu Sun,Zheng Zhang,Zi Huang,Yadan Luo
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Which layers of a feed-forward geometry model are needed to preserve both camera poses and dense 3D structure? We study layer redundancy in VGGT, \pi^3 , and VGGT- \Omega : 3,018 pruned configurations, scored on seven camera-pose and dense-geometry metrics across four indoor and outdoor datasets. Four findings follow: (i) Removable layers cluster in two redundancy regions: a dominant early region and a narrower late one, while deletions spanning the intervening layers are consistently more disruptive. This recurring pattern holds across models, datasets, and metrics, and contrasts with the middle-to-late redundancy commonly reported in the literature. (ii) Within these regions, we observe that the joint degradation from deleting two intervals is approximately the sum of their individual degradations, reducing the number of model evaluations for pruning search from O(L^4) to O(L^2) , where L is the aggregator depth. (iii) We find that CKA provides a cheaper representation-based proxy for interval degradation, offering a practical trade-off between pruning quality and calibration cost. (iv) Closed-form linear calibration recovers accuracy after pruning without end-to-end retraining. A least-squares analysis shows that using a shared map for special and patch tokens generally incurs excess reconstruction loss, motivating token-aware recovery. Recovery maps fitted on just 100 calibration scenes generalize to held-out scenes and unseen datasets. The resulting models reduce aggregator parameters by up to 44% while maintaining accuracy comparable to their intact counterparts. Code and experimental results will be available at our project page: this https URL

[CV-178] You Cannot Recover What Was Never Measured: Quantifying the Information Ceiling of Ultra-Low-Field MRI Super-Resolution

链接: https://arxiv.org/abs/2609.36837
作者: Prathamesh Pradeep Khole,Shreya Handa,Utkarsh Gupta,Razvan Marinescu
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Medical Physics (physics.med-ph)
备注: 24 pages, 10 tables, 8 figures

点击查看摘要

Abstract:Generative super-resolution models can turn portable 64 mT MRI into images that look like 3T scans, and the field evaluates them with PSNR, SSIM, and pixelwise uncertainty, most often on pairs built by synthetically degrading high-field images. Prior work acknowledges that these models hallucinate and that the problem is ill posed, but to our knowledge no study measures how much information about the individual subject the real low-field scan actually contains. We measure it. Using paired 64 mT and 3T scans of the same subjects from three public datasets, and a measurement protocol validated on tests whose correct answer is known in advance, we find that, judged over the whole brain, real 64 mT scans carry structure specific to the individual only down to approximately 3 to 4 mm half-pitch in plane, and coarser still through plane. Standard synthetic degradations preserve subject information roughly 1 mm beyond this ceiling, so models trained and benchmarked on synthetic pairs are evaluated on information that real scanners never record. We then test trained diffusion models and a publicly released external model on real paired acquisitions; 24 trained runs of five architectures (GAN, diffusion, and transformer families) give the coverage of the audit. On every subject where faithfulness can be measured, fine output detail is no more correlated with the subject’s own 3T scan than with a stranger’s, while sample-variance uncertainty does not distinguish fabricated structure from reconstruction difficulty. Because PSNR and SSIM score resemblance to a reference rather than whether detail belongs to the subject, a benchmark scored by them cannot tell recovery from fabrication. Code for the measurement protocol will be released so that recoverability claims can be tested for newer models.

[CV-179] Motion Concept Unlearning in Video Diffusion Models ACM-MM2026

链接: https://arxiv.org/abs/2609.36832
作者: Ping Liu,Chi Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to ACM MM 2026. Dr. Chi Zhang is the corresponding author

点击查看摘要

Abstract:Text-to-video (T2V) diffusion models can generate realistic depictions of actions such as kicking, stabbing, and shooting, raising safety concerns that motivate targeted concept erasure. Although concept erasure has been extensively studied for static concepts in text-to-image and T2V models, erasing motion concepts remains largely unexplored. We present a systematic study of motion concept erasure in video Diffusion Transformers (DiTs). Through causal interventions, we show that text-conditioning attention carries concept-specific motion information and supports selective intervention, whereas perturbing temporal positional encoding suppresses both target and non-target dynamics. We further find that directly adapting ESD, a representative weight-level image erasure method, to a video DiT yields modest and uneven motion suppression: reducing its erasure training loss does not by itself remove the concept signal from the difference between the conditional and unconditional predictions, which classifier-free guidance (CFG) then scales at every denoising step. From these findings, we derive three requirements for motion concept erasure: concept specificity, spatial selectivity, and temporal naturalness. Each determines one component of MUTE (Motion concept Unlearning in Text-to-video gEneration): at each denoising step, MUTE extracts a concept direction through token neutralization, derives a spatial gate from the direction’s intrinsic structure, and subtracts the resulting correction from the velocity output before CFG is applied. MUTE is training-free and requires no weight modification. Experiments on 20 motion concepts show that MUTE outperforms representative prompt-level, weight-level, and inference-time baselines on Wan2.1-T2V, and the same formulation transfers to CogVideoX, supporting its applicability across distinct T2V attention architectures.

[CV-180] Learning via Self-Consistency for Diffusion-based Video Reasoning

链接: https://arxiv.org/abs/2609.36826
作者: Zhenghao Ni,Weimin Qiu,Meng Tang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 23 pages, including 10 pages for the main body

点击查看摘要

Abstract:Video generation models have demonstrated emerging zero-shot capabilities for visual reasoning, perception, and other vision tasks. However, diffusion-based video generation is inherently stochastic, while many downstream vision tasks are deterministic. Motivated by the effectiveness of self-consistency in chain-of-thought reasoning for large language models, we investigate whether self-consistency can similarly improve diffusion-based video reasoning. We first introduce a training-free test-time scaling method that samples multiple video generations and aggregates their predictions through self-consistency. Specifically, we aggregate extracted paths, locations, or masks from multiple rollouts into a consensus prediction. To reduce the inference overhead of multi-rollout generation, we read out predictions early in the denoising trajectory, which preserves consensus quality while reducing denoising steps by more than half. We further propose Rejection Fine-Tuning (RFT) to distill consensus predictions into the video generation model. The resulting model internalizes the benefit of multi-sample consensus and requires only a single generation at inference time, while substantially outperforming the original model. Experiments on three tasks, including maze solving, visual search, and referring segmentation, show that both our self-consistency inference and consensus distillation dramatically improve video-based perception and reasoning, without requiring ground-truth videos or task-specific verification. For visual search, self-consistency raises task accuracy from 48.4% for a single generation to 99.0%. The distilled model retains much of the consensus benefit with a single rollout. For 4-by-4 maze solving, consensus-based training improves the single-generation strict success rate from 72.0% to 84.0% with the same inference latency.

[CV-181] RED: Reconstruction Evolution Dynamics for Generalizable AI-Generated Image Detection

链接: https://arxiv.org/abs/2609.36822
作者: Wenpeng Mu,Junshan Jin,Tanfeng Sun,Xinghao Jiang,Qiang Xu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:The rapid evolution of image generators calls for forensic cues that generalize beyond known generation mechanisms. Existing detectors often rely on static image representations or endpoint reconstruction discrepancies, leaving the evolution of intermediate reconstruction stages underexplored. We observe that the relative token predictability of real and generated images can reverse across reconstruction scales, suggesting that intermediate stages may expose forensic evidence overlooked by endpoint comparisons. Motivated by this observation, we propose RED (Reconstruction Evolution Dynamics), a framework that captures transferable forensic cues from coarse-to-fine reconstruction evolution. To our knowledge, RED is the first framework to use scale-wise token predictability to guide forensic evidence aggregation across intermediate reconstruction states. It represents the reconstruction trajectory produced by a frozen multiscale VQ-VAE in the shared feature space of a frozen CLIP encoder. To connect the observed predictability variations with visual evidence, RED learns image-adaptive stage weights from scale-wise token negative log-likelihoods provided by a frozen VAR model. A cross-stage evidence aggregation module then jointly models the original-image representation and the weighted reconstruction features, capturing complementary forensic cues through interactions along the reconstruction trajectory. Experiments on six diverse benchmarks demonstrate that RED achieves the highest average accuracy of 92.5% and average precision of 97.5% among the evaluated methods. Further evaluations show strong robustness to common image degradations, supporting the value of reconstruction evolution for generalizable AI-generated image detection. The code will be made publicly available upon acceptance of this paper.

[CV-182] CurvSpec: Adaptive Multi-Curvature Learning for Partial Relevant Video Retrieval

链接: https://arxiv.org/abs/2609.36815
作者: Zhen Liu,Letian Li,Jinpeng Wang,Shuzhao Xie,Yuzhi Huang,Jingyan Jiang,Zhi Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 10 pages. Accepted to ACM Multimedia 2026 (MM '26)

点击查看摘要

Abstract:Partially Relevant Video Retrieval (PRVR) seeks to retrieve untrim-med videos containing a moment that matches a text query, without temporal annotations. The relevant moment may last only seconds within a video spanning several minutes, creating an extremely low signal-to-noise ratio that makes PRVR more challenging than standard full-video retrieval. This task presents two intertwined challenges: (1) signal dilution, where coarse global representations blur the brief relevant signal into the dominant irrelevant surroundings;(2) curvature rigidity, where embedding all videos in the same fixed-geometry space distorts representations for videos that range from flat atomic events to deep compositional hierarchies. Existing PRVR methods have improved moment selection and cross-modal matching, but they still typically encode all videos in a single fixed-curvature retrieval space, limiting their ability to model diverse video structures. To address both challenges, we propose CurvSpec, a framework that learns content-adaptive curvature for video retrieval representations rather than imposing a fixed geometric prior. CurvSpec processes features through parallel Euclidean and hyperbolic attention layers, with independently learned curvatures assigned to the hyperbolic layers, and a content-aware fusion mechanism routes each input to its most suitable geometric regime. To further suppress signal dilution, CurvSpec represents each video with semantic centroids whose number is determined by the video’s content complexity, projects them onto the learned manifold, and matches each query against its nearest centroid by geodesic distance. Experiments on ActivityNet Captions, TVR, and Charades-STA demonstrate state-of-the-art retrieval performance.

[CV-183] MeteoVerse: Unified Weather-Controllable Video World Model

链接: https://arxiv.org/abs/2609.36810
作者: Renlong Wu,Guanqiao Wang,Xuan Shang,Yin Hanming,Xiaoxiao Sheng,Tianyu Huang,Hui Li,Wangmeng Zuo
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 13 pages

点击查看摘要

Abstract:Video world models aim to predict future content from an observed scene while following prescribed camera motion. Real-world scene evolution is determined not only by changes in viewpoint and object dynamics, but also by environmental conditions such as weather, which can substantially alter scene appearance and visibility. Modeling such realistic weather evolution is challenging because the required weather modification depends jointly on the observed and desired weather states. Depending on their relation, the model may need to preserve, introduce, or remove a weather effect. Existing video world models typically leave this weather transition implicit, forcing the generation backbone to infer weather evolution together with scene dynamics and camera motion, which leads to imprecise weather control. To address this limitation, we propose MeteoVerse, a unified weather-controllable video world model that generates future videos from a single sunny or adverse-weather image, conditioned on a weather-free scene description, a target-weather instruction, and a camera trajectory. Rather than conditioning only on the desired weather, MeteoVerse explicitly estimates the observed and target weather states and represents the required weather transition. A transition-aware mixture of weather experts then translates this transition into category-specific residual weather features, unifying weather preservation, introduction, and removal while enabling fine-grained control over introduced weather intensity. We further construct the MeteoVerse dataset with over 50K real-world weather video clips, generated sunny counterparts, disentangled scene and weather descriptions, weather-intensity annotations, and camera trajectories. Extensive experiments demonstrate substantially improved weather controllability while retaining competitive scene consistency and camera-control performance.

[CV-184] EGSD: Event-Grounded Self-Distillation for Streaming Video Understanding

链接: https://arxiv.org/abs/2609.36803
作者: Yuwei Miao,Xuesheng Zhang,Wenhao Zou,Jixia Zhang,Jianwei Lv,Bo Yuan,Junfeng Wang,Shiao Xie
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Real-time video understanding requires incrementally maintaining a memory of streaming content, and optimizing this requires dense process signals. On-Policy Self-Distillation (OPSD), which lets one model serve as both teacher and student with the teacher receiving additional privileged information such as the question and ground-truth (GT) answer, can supply such token-level signals. However, applying it directly to streaming video raises two problems. (1) The student cannot be optimized end-to-end, where memory is written before the question arrives, yet the teacher scores it with the question-and-GT privilege, misaligning their preferences. (2) Effective-entity memory collapses, where the question-and-GT privilege makes the teacher favor only question-relevant entities, and token-mean averaging over a memory renders its signal invariant to how many entities that memory covers, both driving memory against the streaming need for diversity. To address these issues, we propose Event-Grounded Self-Distillation (EGSD), which characterizes streaming memory as an incremental update over verifiable Events (key visual entities, actions, and details) and targets the two problems on this basis. For problem (1), we adapt the OPSD signal into a multiplicative weight combined with the outcome reward; for problem (2), we re-weight the teacher with Events as privileged information to counter its question-relevance bias, and add an entity-coverage reward to supply the coverage preference the token-mean teacher lacks. Extensive experiments on mainstream online and offline benchmarks show EGSD achieves strong performance, reaching 79.8% on StreamingBench and 73.4% on the OVO-Bench Real-Time track, while memory analysis shows effective-entity recall rises 17.4% at only 6.8% more memory length.

[CV-185] Scene Retargeting: Learning Object Placement with Analogical Transfer

链接: https://arxiv.org/abs/2609.36801
作者: Minkwan Kim,Junho Kim,Seungmin Lee,Changwoon Choi,Young Min Kim
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL

点击查看摘要

Abstract:Interactive simulations of embodied AI or spatial computing applications build on realistic 3D scenes that support daily activities. However, sparse, irregular layout structures impose scene-specific physical constraints, making it hard to define a generalizable framework for generating similar functional context. We formalize Scene Retargeting as stably transferring the semantically coherent spatial organization across layouts, rather than relying on textual descriptions or pairwise relationships. Our cluster-wise transfer flexibly handles mismatched object instances and adapts to distinctive floor plans. We optimize to preserve the rich semantic context of individual clusters by respecting the spatial distribution of foundation features. We can then impose physical constraints to refine wall contacts, pairwise alignment, or clear passageways and openings. Our framework outperforms state-of-the-art methods on layout generation on the 3D-FRONT dataset, and demonstrates downstream applications including real-to-sim transfer, analogical trajectory transfer, and multi-reference composition.

[CV-186] Decoding Affective Nuances: Enhancing MLLM s via Hierarchical Emotion Reasoning and Contrastive Discriminative Pruning

链接: https://arxiv.org/abs/2609.36782
作者: Cheng Ye,Weidong Chen,Zhaobo Qi,Beier Zhu,Zhendong Mao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 19 pages, 8 figures

点击查看摘要

Abstract:While multimodal large language models (MLLMs) have demonstrated exceptional capabilities in objective understanding tasks, their performance in affective reasoning still falls significantly short of human standards. We attribute it to a central capability gap: MLLMs are difficult to reliably distinguish semantically proximal emotions based on fine-grained visual evidence, which could be decoupled as two limitations: 1) Insufficient Attribution. The global reasoning paradigm of conventional MLLMs severely dilutes fine-grained emotion cues, where subtle emotional states are usually implicitly encoded, thereby generating emotional misjudgments in complex scenarios. 2) Insufficient Discrimination. Existing methods could only identify regions generally associated with emotions, which fails to distinguish discriminative regions between semantically similar emotions, leading to ambiguous emotion judgements. To overcome these limitations, we present a training-free inference-time optimization framework, named Decoding Affective Nuances (DAN). Specifically, we propose a Hierarchical Emotional Reasoning Chain (HERC) that enhances the insufficient attribution by harmonizing fine-grained scene/object-level cues and performing a soft-gated reasoning. Furthermore, to discriminate between semantically proximal emotions, we design a Contrastive Discriminative Visual Pruning (CDVP), which isolates discriminative visual tokens to reason the final emotion category by computing the absolute discrepancy between the attention distributions of similar emotions. Performances on several benchmarks demonstrate that DAN significantly improves discrimination for affective nuances without consuming additional training resources, especially achieving +10.47% improvements with Qwen3-VL-8B-Instruct on WebEmo25 dataset that contains 25 fine-grained emotion categories.

[CV-187] DRHeC: Differentiable Rendering for Hand-Eye Calibration with RGB-Based Gradients

链接: https://arxiv.org/abs/2609.36779
作者: Xiaotian Zhang,Yusheng Wang,Naoya Kagawa,Noritaka Takamura,Keiji Okuhara,Hiroyasu Baba,Jun Ota
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
备注:

点击查看摘要

Abstract:Accurate hand-eye calibration is crucial for precision manipulation. Traditional methods rely on markers, with their precision dependent on marker accuracy and observability. In contrast, markerless methods, such as learning-based approaches, use deep neural networks to directly extract keypoints or features from images, enabling the computation of hand-eye transformation with a single image and without the need for physical markers. Recently, differentiable rendering-based methods for hand-eye calibration have leveraged physical models to render binary masks and compare them with observations, enabling hand-eye calibration without fiducial markers in the calibration stage and providing interpretable optimization. While the state-of-the-art differentiable rendering methods achieve remarkable accuracy, the use of binary masks can result in the loss of internal profile details, reducing precision. Additionally, these methods can also suffer from unstable optimization and local minima. In this study, we propose a novel RGB-based differentiable rendering framework that provides richer geometric and appearance cues by incorporating color and mask geometric features, thereby improving calibration accuracy and optimization stability. Additionally, we propose a mask-guided image-to-image translation method to ensure explicit preservation of color and geometric consistency throughout the translation. Our approach is validated through both simulation and real-world experiments, with results demonstrating strong accuracy and robustness and clear improvements over existing differentiable rendering methods. Our method achieves a grasping success rate of 88.9% and insertion success rate of 57.4% on the UR5e real-world experiment, outperforming the state-of-the-art differentiable rendering hand-eye calibration method EasyHeC by 46.3 and 48.1 percentage points, respectively. Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR) Cite as: arXiv:2609.36779 [cs.RO] (or arXiv:2609.36779v1 [cs.RO] for this version) https://doi.org/10.48550/arXiv.2609.36779 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Journalreference: IEEE Transactions on Instrumentation and Measurement, vol. 75, Art. no. 7505816, pp. 1-16, 2026 Related DOI: https://doi.org/10.1109/TIM.2026.3712913 Focus to learn more DOI(s) linking to related resources

[CV-188] Causal-EVC: Breaking Emotional Spurious Causality via Spatiotemporal Grounding and Counterfactual Intervention

链接: https://arxiv.org/abs/2609.36776
作者: Cheng Ye,Weidong Chen,Peipei Song,Zhendong Mao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 20 pages, 7 figures

点击查看摘要

Abstract:Emotional Video Captioning aims to generate factually accurate and emotionally empathetic descriptions. While recent methods have recognized the importance of visual causes to guide emotion perception and caption generation, they fundamentally rely on simple attention matching, which inevitably suffers from causal redundancy and spurious correlations in co-occurrence bias (e.g., misclassifying sadness'' as joy’’ on a sunny beach), leading to severe shortcut learning from confusing backgrounds. Furthermore, existing evaluations fail to verify whether models have genuinely mastered causal reasoning or merely exploited background confounders. To address these limitations, we first construct EVC-CauseGround, a comprehensive benchmark with dense spatio-temporal causal annotations. Crucially, it introduces a carefully selected Causal-Faithfulness Subset to explicitly quantify genuine emotion-cause attribution. Second, we propose Causal-EVC, an emotion-grounding captioning framework, which introduces a Motion-guided Causal Spatiotemporal Localization module to precisely decouple causal triggers from background confounders. Besides, we introduce an Interpretable Sparse Emotion Routing module. By synthesizing counterfactual representations and formulating a novel counterfactual contrastive objective, we enforce the model to anchor its emotion predictions strictly on authentic causal triggers instead of confusing background. Extensive experiments show that Causal-EVC not only achieves the best performance on semantic metrics but also exhibits significant advantages in the causal-faithfulness subset, which demonstrates that our model could mine emotional cues from genuine visual causes and mitigate co-occurrence bias for interpretable multimodal emotion understanding.

[CV-189] DSPO: Diversity-aware Subjective Policy Optimization for Robust Emotional Reasoning

链接: https://arxiv.org/abs/2609.36775
作者: Cheng Ye,Weidong Chen,Bingyan Xu,Zhendong Mao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 17 pages, 5 figures

点击查看摘要

Abstract:Reinforcement Learning has significantly advanced the complex reasoning capabilities of MLLMs. However, prevailing RL algorithms suffer a severe failure in emotion reasoning tasks. These methods heavily rely on deterministic hard-label supervision and point-wise isolated evaluation, creating a fundamental gap with the inherently subjective and continuously distributed nature of human emotions. Furthermore, unlike explicit physical objects, emotional states are deeply implicit within visual cues. This abstract nature exacerbates visual hallucinations in MLLMs, leading to plausible yet ungrounded emotional evidence. To address these limitations, we propose Diversity-Aware Subjective Policy Optimization (DSPO), a reinforcement learning framework that jointly promotes subjective affective coverage and visual grounding. First, we construct a context-grounded emotional distribution prior in the VAD space by combining the lexical prior of the annotated emotion with image-specific contextual information. Based on this prior, we introduce a Distribution-Aligned Emotional Diversity Reward (DEDR), which measures the leave-one-out marginal contribution of each candidate emotion within a rollout. DEDR rewards candidates whose inclusion brings the predicted affective set closer to the context-grounded prior, thereby preserving plausible subjective interpretations without encouraging unconstrained dispersion. We further develop Counterfactual Visual Intervention Gating (CVIG), which masks the visual region highlighted in the reasoning process and uses the resulting candidate-wise probability changes to reduce the weights of interpretations unsupported by visual evidence. Extensive experiments demonstrate that DSPO achieves state-of-the-art performance across multiple public benchmarks, especially on the cross-domain performance, i.e., improving +10.8% on average cross-domain accuracy than EMO-R3.

[CV-190] Emergent Specialization in Populations of Self-Supervised Collaborative Vision Experts Without a Shared Gate or Cross-Agent Gradients

链接: https://arxiv.org/abs/2609.36770
作者: Aram Davtyan,Pablo Acuaviva,Sebastian Stapf,Paolo Favaro
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Can a population of neural networks develop a useful division of labor without a shared gate or gradients between agents? We study a setting where each network has its own weights, trains independently on the same heterogeneous data, and can ask another agent for help through a forward pass. Unlike mixtures of experts, where a jointly trained gate assigns inputs to experts, specialization here must emerge without central control. We test this in a small scale proxy for predictive visual pretraining. Initially identical agents are finetuned on an unlabeled mixture of six visual domains using masked prediction of frozen DINOv3 features. We measure specialization by asking whether the best agent for an input aligns with its latent domain, and utilization by asking whether responsibility is distributed across agents. We progressively remove central control, ending with DISCO (DIStributed COllaboration) where each agent locally selects a helper, reads its internal state through a gradient free channel, and rewards its router only for the improvement that help provides. Specialization emerges and is useful. Randomly routed populations underperform a single generalist, while semantically routed populations outperform it, showing that specialization rather than population size drives the gain. Specialization persists without a central router, and gradient free communication lets nonexperts exploit emergent expertise. In DISCO, a random agent helped by the expert matches the solo generalist, while experts surpass it, including on data outside the specialization mixture. Local routers select the emergent expert for 98% of inputs. These effects persist across population size, model capacity, data imbalance, and finetuning seeds, providing measurable evidence for the dynamics needed by decentralized predictive pretraining.

[CV-191] Dual-Mode Low-Rank Learner with Bridge-Prototype Ensemble for Vision-Language Class-Incremental Learning

链接: https://arxiv.org/abs/2609.36759
作者: Chiyuan He,Zihuan Qiu,Fanman Meng,Chao Wang,Liangjiang Chen,Linfeng Xu,Qingbo Wu,Hongliang Li
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 21 pages, 11 figures, and 12 tables, including the appendix

点击查看摘要

Abstract:Benefiting from transferable visual-textual alignment, CLIP has been widely adopted for class-incremental learning (CIL). However, existing learners either repeatedly update components shared across tasks, leading to knowledge overwriting, or overly isolate new-task updates, hindering the reuse of CLIP’s transferable knowledge and limiting plasticity. Moreover, the text-based or bimodal classifier designs still fail to effectively integrate complementary information from the visual and textual modalities. To address these challenges, we introduce DuLBE, which couples dual-mode low-rank learning with a bridge-prototype ensemble classifier for exemplar-free CIL. DuLBE allocates two visual low-rank update modes according to the gradient demand and uses gradient routing to coordinate them: a compact and rewritable shared mode is selected from historically occupied visual directions to reuse transferable knowledge, while residual modes provide low-interference channels for task-specific variations. Building on the resulting stable inter-modal structure, we further construct geodesic bridges between visual prototypes and text embeddings on the unit hypersphere, and ensemble reliable bridge prototypes to compensate for the modality-gap limitations of textual decision boundaries. Extensive experiments under multiple settings show that DuLBE achieves state-of-the-art CIL performance while retaining the high parameter efficiency of low-rank tuning.

[CV-192] FastVR: Efficient Streaming Video Restoration with One-Step Diffusion

链接: https://arxiv.org/abs/2609.36757
作者: Xiaoxu Chen,Qin Yang,Haoran Bai,Sibin Deng,Ying Chen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Diffusion-based video restoration recovers realistic details, but its practical deployment is limited by two efficiency bottlenecks: costly VAE encoding and decoding, and the quadratic cost of full self-attention in diffusion transformers (DiTs). This paper presents FastVR, a streaming video restoration framework built on a one-step diffusion model, which delivers strong restoration quality and temporal consistency while processing 1080p video at 11 FPS on a single H20 GPU. To improve inference efficiency, FastVR combines a lightweight VAE with chunk-wise causal attention, which substantially reduces the computational cost. During training, it further adopts velocity consistency regularization and continuous trajectory learning, which improve restoration quality. Extensive experiments show that FastVR is more efficient than the evaluated diffusion baselines while achieving state-of-the-art performance on synthetic and real-world benchmarks. We hope that this work supports further progress in the community.

[CV-193] NesTok: Nested Self-Aligned 1D Tokenizer for Autoregressive Image Generation

链接: https://arxiv.org/abs/2609.36756
作者: Jiawei Zhang,Shuhao Liu,Rong Huang,Yuancheng Li,Zhihui Li,Xiaojun Chang,Changlin Li
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Computer Vision, Autoregressive Model

点击查看摘要

Abstract:One-dimensional (1D) variable-length visual tokenizers enable adaptive compression by varying the number of tokens, allowing downstream autoregressive (AR) models to flexibly trade off generation quality against computational cost using a single tokenizer. However, existing approaches based on nested dropout often fail to fully exploit the representational capacity of the tokenizer, resulting in suboptimal performance in both image reconstruction and generation. In this work, we introduce NesTok, a nested self-alignment framework tailored to dynamic visual tokenizers. NesTok introduces cross-length training, which jointly optimizes reconstruction across token lengths while using the full-length sequence to guide shorter counterparts, enabling shorter token sequences to approach the reconstruction quality of full-length sequences. On ImageNet, NesTok improves substantially over standard training and achieves an rFID score of 0.98. On downstream image generation, it achieves the state-of-the-art gFID score of 1.46 on ImageNet 256 \times 256 among existing variable-length autoregressive image generation methods. Code will be available at this https URL.

[CV-194] Drag as Evidence: Motion-Grounded Latent Recomposition for Drag -Based Editing

链接: https://arxiv.org/abs/2609.36755
作者: Xinyu Pu,Hongsong Wang,Jie Gui,Pan Zhou
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Modern image editors excel at semantic manipulation and visual synthesis, yet remain limited in precise spatial control, motivating the development of drag-based editing. However, existing drag-based methods often struggle to balance drag accuracy with natural, plausible, and intent-aligned generation. We propose MoRe-Drag, a motion-grounded drag-based editing method. Our key insight is to treat pixel-space warping as coarse motion evidence, and to inject this evidence into the generative sampling trajectory. Specifically, MoRe-Drag performs region-aware latent recomposition over refinement, inpainting, and anchor regions, coupled with stage-adaptive conditioning that progressively shifts from motion-grounded structure formation to semantic refinement. We further support an instruction-free interface by adapting the MLLM-based text encoder for drag-aware instruction inference. Experiments on DragBench-SR and DragBench-DR show that MoRe-Drag substantially improves drag precision over strong base editors and achieves superior drag accuracy among SOTA drag-based methods, while delivering strong semantic consistency and visually realistic results. Code and dataset will be publicly released.

[CV-195] AESplat: Advancing Pose-Free Feed-Forward 3D Gaussian Splatting via Decoupled Appearance Modeling

链接: https://arxiv.org/abs/2609.36693
作者: Shiwei Ren,Zhiang Liu,Yongchun Fang,Hongwei Chen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Pose-free feed-forward 3D Gaussian Splatting (3DGS) has demonstrated remarkable potential for generalized novel view synthesis. However, existing methods typically predict Gaussian appearance attributes represented by spherical harmonics (SH) in the same manner, overlooking the fundamental distinction between view-independent and view-dependent appearance, which results in suboptimal rendering quality. In this paper, we present AESplat, a novel and general framework for pose-free feed-forward 3DGS that introduces an effective decoupled appearance modeling strategy based on an analysis of SH, enabling higher-quality rendering. Specifically, AESplat directly derives the zeroth-order SH coefficient, which represents the base view-independent appearance component, from the input images without training. The higher-order SH coefficients are subsequently predicted by a shallow multilayer perceptron equipped with two efficient 3D-aware inductive biases to model view-dependent appearance variations. Extensive experiments across multiple datasets demonstrate that our method significantly outperforms state-of-the-art approaches, achieving a 0.8 dB improvement in PSNR over the pose-free method NAS3R and a 1.1 dB improvement over the pose-required method DepthSplat on the RealEstate10K dataset. Project page: this https URL.

[CV-196] When Semantics Matter: Reliability-Aware Semantic-Rhythm Control for Co-Speech Gesture Generation

链接: https://arxiv.org/abs/2609.36685
作者: Zhirui Xing,Long Ye,Kaige Li,Ziyi Xu,Ming Meng
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 9 pages, 5 figures, 3 tables

点击查看摘要

Abstract:Co-speech gesture generation aims to synthesize natural gestures that are both temporally synchronized with speech and semantically consistent with the spoken content. Although recent methods can generate rhythmically plausible motions, they often rely heavily on acoustic prosody while underutilizing textual semantics, especially when semantic annotations are incomplete, noisy, or unavailable. Consequently, the generated gestures may follow speech rhythm while failing to express the intended semantics. To address this problem, we propose a reliability-aware semantic-rhythm control framework for co-speech gesture generation. We first learn a discrete motion prior that represents continuous gestures in a compact and structured motion-code space. We then introduce a dual-branch semantic contribution estimation mechanism consisting of a full multimodal branch and an audio-only branch. Their distributional discrepancy is formulated as conditional information gain to quantify how much textual semantics changes the predicted motion. Based on this estimate, a controllable semantic-rhythm objective selectively strengthens semantic guidance in content-relevant segments while limiting unnecessary semantic intervention in rhythm-dominant segments. Furthermore, we treat background noise as an acoustic reliability condition and introduce noise-conditioned feature modulation together with beneficial latent perturbation to improve generation robustness under realistic acoustic environments. Experiments on benchmark datasets demonstrate that the proposed framework achieves a favorable balance among semantic expressiveness, rhythmic synchronization, motion diversity, and robustness, enabling reliable and controllable co-speech gesture generation.

[CV-197] Reprogramming Vision-Language Models via Structured Prompt Reparameterization

链接: https://arxiv.org/abs/2609.36680
作者: Zizhao Li,Chengyi Cai,Mohammed Yaqoob Ansari,Feng Liu,Joseph West,Kourosh Khoshelham
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Visual reprogramming adapts pretrained models to downstream tasks by modifying their input and output interfaces while keeping the backbone fixed. In vision-language models, existing methods mainly rely on intra-class prompt aggregation and do not explicitly model relationships among classes. However, fine-grained categories often exhibit highly overlapping attribute descriptions and strong inter-class correlation in the text embedding space, where discriminative cues lie in subtle low-variance components. We propose Reparameterized Inter-Class Visual Reprogramming (RVP), a structured framework that aggregates multiple text prompts within each class and applies residual correction across classes. We also show that CLIP-based visual reprogramming with input-independent linear output aggregation can be expressed as a linear mapping from frozen image embeddings to downstream logits, and use this view to design a structured reparameterization that models shared semantic components and class-specific differences. RVP uses only a single visual prompt and can be reparameterized at inference into a frozen backbone followed by a linear classifier, incurring nearly zero computational overhead. Across 11 few-shot classification benchmarks and four CLIP backbones, RVP consistently improves over prior visual reprogramming methods with comparable or better inference efficiency.

[CV-198] ReWorld-Track: A Recursive Event World Model for Language-Guided Multi-Camera Tracking

链接: https://arxiv.org/abs/2609.36677
作者: Haoyang Wu,Shoudong Han,Chaoyue Li,Sijia Chen,Zhenyang Xie,Wang sihan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 41 pages, 8 figures

点击查看摘要

Abstract:Language-guided multi-camera tracking must preserve a target identity across unobserved gaps, where similar candidates and uncertain returns can make early associations unreliable. A wrong match can corrupt the history used to predict later observations and propagate identity errors across subsequent camera handoffs. We propose ReWorld-Track, a recursive event world model that carries association uncertainty into future predictions. Candidate matches and continued waiting define alternative target states, whose posterior probabilities are used to update a persistent recurrent belief. This representation preserves uncertainty about alternative trajectories through successive observations. This belief predicts the next camera, arrival time, and entry region, while appearance and language evidence guide association. By training across successive handoffs, the model learns to retain uncertainty that remains useful for later predictions and identity decisions. ReWorld-Track achieves HOTA scores of 65.19 on CityFlowV2 and 45.36 on MTMMC, with improved identity continuity across repeated handoffs. On MTMMC, its structured posterior update gains 0.50 HOTA points over a similarly sized generic updater and 0.94 points over fixed-moment soft association, raising next-camera accuracy from 86.03% to 87.41% and reducing median arrival-time error from 0.78 s to 0.71 s for subsequent target returns.

[CV-199] You Only Reprogram Once: Rethinking Prolonged Training for Visual Reprogramming

链接: https://arxiv.org/abs/2609.36661
作者: Zizhao Li,Mohammed Yaqoob Ansari,Xinyu Su,Jiayang Ao,Joseph West,Kourosh Khoshelham
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Visual reprogramming is a parameter-efficient method for adapting pretrained models, yet its training can remain computationally expensive: even with a frozen backbone, visual prompts are often optimized through the full model for hundreds of epochs. Before changing what the pretrained model sees, we ask whether we are fully using what it already tells us. We find that modeling the full source response can already yield strong downstream predictions without prompt optimization. Motivated by this observation, we introduce You Only Reprogram Once (YORO), which constructs a downstream predictor from the frozen response space in a single forward-only traversal. Its Bayesian Discriminant Mapping (BDM) derives a covariance-aware affine mapping from streaming class statistics, requiring no backpropagation, optimizer updates, or repeated visits to the training set. When further input adaptation helps, YORO-FP optionally refines the visual prompt for 20 epochs. BDM also extends naturally to CLIP by treating attribute-prompt similarities as source responses. Across three full-data settings, YORO improves average accuracy over the strongest prior gradient-free mapping by 18.4–24.4%. On 16-shot CLIP, it raises the four-backbone average from 71.4% to 77.2%. YORO-FP provides further gains on selected tasks, while validation often retains the one-pass predictor. These results suggest a different default for visual reprogramming: read out the frozen response first, and optimize the input only when needed.

[CV-200] Not Every Correction Helps: Gain-Guided Continual Test-Time Adaptation

链接: https://arxiv.org/abs/2609.36655
作者: Youjia Zhang,Huiling Liu,Soyun Choi,Jaehong Yoon,Sungeun Hong
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Continual test-time adaptation (CTTA) adapts a source model to an unlabeled test stream whose distribution may change over time. Existing TTA methods often assess prediction reliability using confidence or entropy, which primarily reflect the model’s self-certainty for the current sample. In CTTA, accumulated target observations can provide complementary evidence for correcting the source prediction, but this history may become misaligned as the target distribution changes. The key question is therefore not how much the correction differs from the source prediction, but whether and how strongly it should be applied. This paper proposes Gain-Aware INtervention (GAIN), a backpropagation-free CTTA framework guided by a simple principle: history proposes, gain decides. GAIN maintains compact target statistics to form a correction proposal and a posterior-predictive evaluator that accounts for estimation uncertainty. The resulting source-relative gain estimates the proposal’s benefit and determines a sample-specific intervention strength along a continuous path through efficient one-dimensional optimization. Gain-controlled predictions then update the target statistics online, limiting the propagation of unreliable corrections, all without backpropagation, sample storage, or replay. Across five benchmarks, our method achieves strong predictive performance, with favorable accuracy–calibration–efficiency trade-offs in continual adaptation. On ImageNet-C, for example, GAIN achieves 61.9% accuracy with near-source calibration. It remains stable under diverse and challenging continual shifts while running 15.9x faster than a representative optimization-based CTTA baseline.

[CV-201] FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution

链接: https://arxiv.org/abs/2609.36651
作者: FangZhi Zhong,Xuerui Qiu,Yuqi Pan,Ya Liu,Shaowei Gu,Bo Xu,Guoqi Li
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 23 pages, 10 figures. Code: this https URL

点击查看摘要

Abstract:Long-context reasoning in large language models incurs substantial computation and memory costs. Visual text compression (VTC) reduces input length by rendering text as images, but fixed-resolution rendering creates a compression-performance trade-off: low DPI saves tokens at the expense of legibility, whereas high DPI spends tokens on irrelevant content. We introduce FocusVTC, which breaks this trade-off through adaptive resolution while preserving general multimodal capabilities. It combines compressed low-DPI global views with selective region enhancement, integrating enhanced views into ongoing reasoning. We construct 29.4K high-quality Reasoning-Evidence Localization (REL) chain-of-thought examples (REL-CoT) that link reasoning traces to page indices and bounding boxes. Multi-resolution REL supervised fine-tuning (REL-SFT) teaches the model to localize relevant regions, and Group Relative Policy Optimization learns when to enhance resolution and how to use the resulting observations, without a separate continual-pretraining stage. At 72 DPI on RULER v1, FocusVTC scores 87.4 at 2.9\times input compression, including tool observations, versus 57.5 for Glyph at 3.0\times input compression. It surpasses its text-input backbone on LongBench (56.40 versus 55.86), improves the MRCR macro-average by 13.91 points, and achieves a 51.19 macro-average on VTCBench. The MRCR latency evaluation also shows a 2.79\times online end-to-end speedup over Text. General multimodal capabilities are preserved, with MMMU increasing from 65.12 to 66.73 and MME from 2424.02 to 2457.62.

[CV-202] VLM4Cluster: Benchmarking Deep Clustering In the Era of Vision-Language Pre-training

链接: https://arxiv.org/abs/2609.36648
作者: Yuanwei Hu,Bo Peng,Yuheng Jia,Xinting Hu,Yadan Luo,Wenjie Zhu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Vision-language pre-training has reshaped image clustering, giving rise to language-assisted image clustering (LaIC), which leverages textual semantics to complement visual representations. Despite the rapid proliferation of LaIC methods, it remains unclear how much LaIC has actually advanced image clustering, as existing studies generally suffer from major limitations, including inconsistent experimental settings, inadequate dataset selection, and limited evaluation dimensions. To address this gap, we introduce VLM4Cluster, a comprehensive benchmark for image clustering in the era of pre-trained vision-language models (VLMs). VLM4Cluster implements 17 representative methods spanning classical, deep, and language-assisted image clustering, and evaluates them on 20 datasets covering classical, challenging, fine-grained, large-scale, and out-of-distribution settings. Beyond effectiveness, VLM4Cluster systematically investigates image clustering along three complementary dimensions: robustness to adversarial perturbations, generalization under distribution shifts, and computational efficiency. Our study shows that LaIC substantially advances the clustering performance frontier on many semantically demanding benchmarks, generally exhibits stronger generalization under distribution shifts, and achieves a more favorable effectiveness-efficiency trade-off. However, its gains become less consistent on large-scale and fine-grained datasets, while language assistance does not systematically reduce sensitivity to adversarial perturbations. VLM4Cluster is released at this https URL.

[CV-203] OCA: ODE-Driven Cross-Attention for Image-to-Point-Cloud Registration ECCV’26

链接: https://arxiv.org/abs/2609.36644
作者: Pei An,Jiaqi Yang,Yulong Wang,Siwen Quan,Liangliang Nan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to ECCV’26

点击查看摘要

Abstract:Cross-attention is a crucial component in learning-based image-to-point-cloud (I2P) registration. Although existing cross-attention mechanisms have achieved promising progress, attention ambiguity remains a fundamental challenge that hinders the learning of discriminative 2D-3D correspondences. To address this problem, we revisit cross-attention and establish ordinary differential equations (ODEs) to model the ideal I2P feature interaction. Based on this formulation, we develop an ODE-driven cross-attention (OCA) module that refines feature representations and attention matrices through ODEs. In practice, OCA can be seamlessly integrated into existing I2P registration frameworks. To validate its effectiveness, we incorporate OCA into five state-of-the-art baselines and evaluate on four public benchmark datasets. Experimental results demonstrate that OCA improves registration recall by up to 5%, 9%, and 15% under the standard, fine-tuning, and zero-shot settings, respectively.

[CV-204] PE-OPSD: Internalizing Prompt Enhancement into Flow-matching Models via On-Policy Self-Distillation

链接: https://arxiv.org/abs/2609.36638
作者: Mingfeng Lin,Chengfei Cai,Lin Xu,Chengqian Ma,Yuxiang Wei,Liang Han
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Text-to-image users often provide concise and underspecified prompts, whereas generative models benefit from detailed textual conditions for reliable instruction following. Existing systems bridge this gap with Prompt Enhancers (PEs) that rewrite raw prompts at inference time, introducing additional latency and leaving prompt elaboration external to the generator. We instead view enhanced prompts as privileged training information and ask whether their benefits can be internalized. We propose Prompt-Enhanced On-Policy Self-Distillation (PE-OPSD) for text-to-image flow-matching models. During training, a raw-prompt student follows its own generation trajectory, while an enhanced-prompt teacher provides vector-field targets at the states visited by the student. This on-policy supervision distills the behavior induced by enhanced prompts into the raw-prompt student without requiring additional text–image pairs. At inference, both the PE and teacher are removed, and the student generates directly from raw prompts. Across multiple model families, PEs, and benchmarks, PE-OPSD achieves the strongest aggregate prompt fidelity among the evaluated baselines, yields positive aggregate visual appeal gains, and retains the base-model inference efficiency.

[CV-205] Beyond Binary Preferences: Graded Preference Optimization for Limb-Motion Captioning

链接: https://arxiv.org/abs/2609.36628
作者: Yanan Wang,Tingsong Li,Kaixun Jiang,Chongyang Zhong,Chenwei Xoe,Zhaohe Liao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 18 page, 8 figures

点击查看摘要

Abstract:Vision-Language Models (VLMs) can generate rich video captions, yet often misidentify which person performs an action or which limb is involved, particularly across camera cuts. Improving these details requires evaluation and training that distinguish missing information from incorrect assertions. We introduce FlexBench, a benchmark spanning 3,105 shots and 18,161 evaluation queries, with human-verified identities and systematic per-person coverage of fine-grained limb actions and states. Its reference-derived checklists support automated assessment of complete captions in their person and shot contexts. Our Graded Physical Alignment score (GPA) awards credit for correct content and deducts points for incorrect or fabricated actions, making these errors explicit in the aggregate score. Building on this rubric, we propose Graded Margin Direct Preference Optimization (GM-DPO), which assigns stronger preference margins and greater training weight to more severe action errors. Across three VLM backbones, GM-DPO achieves the highest substantive-action and GPA scores among the evaluated preference objectives, improving GPA over DPO by 2.02-3.40 points. On Qwen3-8B, it reduces the weighted hallucination rate by 21.3% relative to DPO. These gains accompany sustained long-form output, improved shot structure, and competitive performance on three additional multimodal benchmarks.

[CV-206] CrossTimeEdit: A Decade-Spanning Cross-View Dataset and Reward-Guided Editing for Historical Street-View Generation

链接: https://arxiv.org/abs/2609.36616
作者: Hanwen Lu,Jun He,Mingjia Yang,Hao Wei,Jinhao Huang,Yi Lin,Xiang Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 43 pages, 9 figures, 7 tables. Project page: this https URL

点击查看摘要

Abstract:Historical street-view imagery records urban evolution, but uneven coverage leaves substantial gaps in historical records. Generating plausible past appearances requires restoring changed structures while preserving persistent scene content. We construct VIGOR-his, a decade-spanning cross-view dataset containing 43,653 location-level quadruplets across 11 cities on three continents. Its automated pipeline performs spatial pairing, consistency screening, change classification, and the generation and validation of satellite-based change descriptions and local editing instructions. Based on VIGOR-his, we propose CrossTimeEdit, a model that reformulates historical street-view generation as editing, using recent street views to constrain viewpoint and unchanged appearance and temporal satellite differences as change evidence. Starting from FLUX.2 [Klein] 4B, we train CrossTimeEdit through supervised fine-tuning (SFT) followed by online reinforcement learning (RL). We design three street-view editing criteria, namely Instruction Alignment (IA), Background Preservation (BP), and Quality and Physical Plausibility (QP), as both RL reward dimensions and evaluation metrics. We optimize this multi-reward objective using Within Group Relative Policy Optimization for flow-matching models (Flow-GRPO) with Group reward-Decoupled Normalization Policy Optimization (GDPO), which normalizes each reward dimension before aggregation. CrossTimeEdit improves overall performance across the three editing criteria by 17.12% over the pretrained baseline and outperforms cross-view generation models in scene consistency, visual realism, and perceptual quality. The implementation code, dataset, and model weights are available at this https URL.

[CV-207] Pixel-wise Exposure for Highly Robust In-Vehicle Remote-PPG

链接: https://arxiv.org/abs/2609.36607
作者: Jieying Wang,Xinqi Cai,Caifeng Shan,Wenjin Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Remote photoplethysmography (rPPG) offers a promising non-contact solution for heart rate monitoring, yet its real-world robustness is fundamentally limited by an inherent hardware limitation: existing camera exposure control paradigms, whether fixed or auto-exposure, impose a uniform exposure time across all pixels within a frame. In high-dynamic-range scenes such as automotive cabins with strong directional sunlight, this spatially invariant exposure constraint inevitably leads to localized facial overexposure or underexposure, irreversibly corrupting the subtle pulsatile signals essential for rPPG at the point of capture, a physical degradation that no downstream algorithm can recover. To overcome this bottleneck, we propose PixExpo (Pixel-wise Exposure), a “temporal-for-spatial” framework that sequentially captures frames under a predefined cyclic exposure schedule and performs non-iterative pixel-wise fusion. At each pixel location, PixExpo selects the observation closest to an rPPG-motivated target intensity. This criterion seeks to reduce local saturation and severe underexposure rather than optimize perceptual appearance. PixExpo requires no sensor modification but assumes programmable frame-level exposure control. We validate the proposed PixExpo framework using our newly introduced MEX-Drive dataset, comprising 48 participants under real-world driving conditions. Experimental results demonstrate that PixExpo outperforms manufacture-default auto-exposure methods, reducing the mean absolute error (MAE) by 7.21 bpm (from 13.94 to 6.73 bpm) and increasing the success rate by 37.29 percentage points (from 25.95% to 63.24%) across challenging driving scenarios.

[CV-208] Scaling Video Generation for Reasoning : At What Cost?

链接: https://arxiv.org/abs/2609.36599
作者: Weihang Guo,Xiaoyu Wu,Yifei Wang,Niloofar Mireshghallah,Lydia E. Kavraki
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:We study whether scaling video generation enables models to reason about hidden information from the past frames, and at what computational cost. Our controlled benchmark requires predicting nine prescribed moves of an initially solved 2x2x2 Rubik’s Cube from a fixed view of three faces. Correct predictions require inferring how actions change hidden states, and the simulator provides exact ground truth for evaluation. Models learn plausible cube geometry early, while correct sticker configurations require substantially more training. Although validation MSE follows approximate power-law scaling, lower MSE loss does not reliably indicate downstream reasoning capabilities. Smaller autoregressive models achieve higher state accuracy with limited compute, while larger models reach higher accuracy after more training. At roughly 0.1 PF-days, the 70M-parameter model correctly predicts the visible sticker configuration in 44.6% of post-action frames, compared with 0.3% for the 1B model, which reaches 83.7% at 3.14 PF-days. Symbolic state supervision raises the 20M model’s frame accuracy from 31.1% to 67.3% at the same training-data budget, suggesting that learning representations of state changes can complement scaling.

[CV-209] Beyond Legibility: Benchmarking Visual Text Rendering and In-Place Editing in Unified Video Generation ICLR2027

链接: https://arxiv.org/abs/2609.36598
作者: Ziying Zhang,Litao Li,Junchao Liao,Tianyi Zeng,Siyu Zhu,Long Qin,Zhenghao Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: ICLR 2027 under review

点击查看摘要

Abstract:A video can exhibit convincing motion and photorealism yet fail immediately when visual text collapses. Unlike generic scene content, visual text is unforgiving in video generation: minor stroke corruption, temporal instability, or editing errors instantly break legibility and realism. Existing benchmarks overlook this challenge by treating text as incidental or using static OCR metrics that ignore temporal dynamics. We introduce VidScribe, a unified diagnostic benchmark spanning four generation regimes: writing from language (T2V), transferring text identity from a reference (R2V), sustaining text under dynamics (I2V), and localized text editing (V2V). VidScribe contains 803 human-verified samples across a 12-axis conditionally orthogonal factor space covering Intrinsic Text Properties, Physical Imaging Conditions, and Temporal Behavior. For reliable evaluation, we build a track-grounded, gated suite with 11 shared metrics and 2 task-specific probes under strict measurability conditions. Benchmarking 11 commercial and open-source systems shows that video text capability is non-monolithic, with content recognition decoupled from stroke-level glyph correctness. Performance is highly task-asymmetric: I2V sustains text most reliably, whereas V2V editing is the primary bottleneck. Counter-intuitively, degradation concentrates on a small subset of text-centric structural and temporal factors rather than adverse imaging conditions. Further probes show that visual references improve glyph and typographic fidelity rather than content accuracy, while localized editing fails to isolate target text without corrupting undeclared source text. Beyond evaluation, VidScribe also provides an actionable training signal, where benchmark-aligned preference optimization measurably improves visual text generation. this https URL.

[CV-210] AffectReveal: Event-Grounded Emotion Recognition Beyond Visual Appearances ICLR2027

链接: https://arxiv.org/abs/2609.36563
作者: Yihao Qian,Runhao Zeng,Sicheng Zhao,Feng Liang,Hongmin Cai,Mingkui Tan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 23 pages, 5 figures, 7 tables. Submitted to ICLR 2027

点击查看摘要

Abstract:Visual emotion recognition commonly assumes that all evidence required for prediction is contained in the observed image or video. Yet the same visible reaction can convey different emotions depending on events beyond the input: tears, for example, may indicate grief or joy. We formulate Event-Grounded Emotion Recognition (EGER), where emotion recognition requires recovering the affect-determining event. We construct EGER-Bench, comprising 10,052 videos and 10,734 images across 11 emotions, two source domains, and four visual settings. A study with six annotators shows that event context raises human recognition accuracy from 33.96% to 72.08%, confirming that visual evidence alone is often insufficient. Semantic relevance alone does not solve EGER: a plausible event may imply the wrong emotion if its identity, focal-person role, relationship, or outcome is misinterpreted. We therefore propose AffectReveal, a tuning-free framework that first constructs and independently verifies evidence-grounded alternatives over these affect-critical factors. It then cross-checks the recovered event against face-masked in-media facts through bidirectional atomic evidence support, while retaining the original unmasked input for final prediction. Across three downstream models and four input settings, AffectReveal yields average UAR gains of 5.26–10.53 points. For three fine-tunable models, it also enables untuned models to outperform their fine-tuned visual-only counterparts in all 12 accuracy comparisons, without updating downstream parameters.

[CV-211] hinkingGuard: Decoding Implicit Hazards via Step-by-Step Risk Attribution in Multimodal Large Language Models

链接: https://arxiv.org/abs/2609.36562
作者: Ruochen Zhang,Yao Huang,Yitong Sun,Jiahe Xie,Jin Yan,Jifan Ma,Yuanfang Guo,Xingxing Wei
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 9 pages, 4 figures, accepted by ACMMM 2026

点击查看摘要

Abstract:While Multimodal Large Language Models (MLLMs) are increasingly deployed in safety-critical domains, their reliability is threatened by multimodal implicit risks. Unlike explicit threats, these hazards emerge when individually benign text and neutral visual entities logically converge to induce unsafe outputs. Current detection methods fail to address this because they overlook the underlying risk activation mechanisms that govern cross-modal risk activation, leading to single-modality shortcut learning and hallucinated rationalizations. To bridge this gap, we first construct TriggerBench, the first dataset explicitly modeling risk compositionality (5,600 instances). By formally isolating Key Elements and Trigger Elements to build counterfactual contrastive pairs, TriggerBench eliminates risk residues and forces models to perform genuine logical deduction rather than superficial pattern matching, which provides a rigorous foundation for both large-scale training and fine-grained evaluation. Building on this, we propose a Step-Supervised Structured Reasoning training framework and employ it to train ThinkingGuard, a specialized guard model. Inspired by Situation Awareness theory, we decouple implicit risk identification into progressive cognitive stages, and utilize a step-reward Monte Carlo Tree Search algorithm to explore optimal reasoning trajectories, which are then distilled into the model through Dual-Constraint Preference Alignment. Extensive experiments across both standard and implicit safety benchmarks demonstrate that ThinkingGuard achieves strong performance. Project resources are available at this https URL.

[CV-212] FM-ReID: Selective Competitive Token Routing for Object Re-Identification

链接: https://arxiv.org/abs/2609.36560
作者: Zhiqi Li,Xiaowei Zhou,Zeyuan Sun,Feng Gao,Junyu Dong
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 12 pages

点击查看摘要

Abstract:Object re-identification (ReID) faces a recurring challenge: different identities can share highly similar global appearances, while the cues that distinguish them are localized, heterogeneous, and visible only under particular viewpoints. This challenge arises in animal ReID through markings, contours, and scars, in person ReID through subtle clothing and accessory cues, and in vehicle ReID through localized appearance details. Although visual foundation models encode such information in dense tokens, a single holistic descriptor can obscure discriminative local signals. We propose FM-ReID, an end-to-end framework that formulates local representation learning as selective competitive token routing. Its Competitive Fine-grained Mining module uses multiple mining queries and a residual query to compete for dense DINOv3 tokens. Above-prior selection retains tokens preferentially allocated to each mining query, while the residual slot receives tokens excluded from the retrieval descriptors. The resulting multi-query descriptors are jointly trained with a holistic representation for retrieval, without fixed spatial partitions or equal-area constraints. FM-ReID achieves strong results on animal, person, and vehicle ReID benchmarks, supporting competitive token routing as an effective way to augment holistic foundation-model representations.

[CV-213] How Medical VLMs Underutilize Their Vision Encoders: A Dermatology Perspective

链接: https://arxiv.org/abs/2609.36557
作者: Janet Wang,Yunbei Zhang,Xiao Wang,Jihun Hamm
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 27 pages, 16 figures

点击查看摘要

Abstract:Medical Vision-Language Models (VLMs) show significant promise for clinical image understanding, offering accurate diagnosis with interpretable reasoning. However, a critical performance gap exists between their strong vision encoders and the full multimodal model: in dermatology, the MedSigLIP encoder outperforms MedGemma by an average of 10.26 percentage points even when both use zero target-task labels; few-shot linear probing provides further evidence of strong visual representations. This gap motivates an investigation of how visual information is used in end-to-end diagnosis and why plausible-sounding predictions can lack grounding in image evidence. Using dermatology as our primary testbed, we systematically investigate three hypotheses for this phenomenon. We further provide a mechanistic analysis of the model’s internal attention patterns, showing that a simple describe-then-decide prompting strategy increases vision attention by 30-40% during generation. Task-specific fine-tuning improves dermatology classification but reduces cross-domain medical question-answering performance in our evaluation. To address these challenges, we combine label-free prompting with low-label encoder-assisted reranking while keeping the VLM frozen. We validate the interventions across five VLM backbones in dermatology and provide supporting representation and attention analyses across additional medical modalities.

[CV-214] SCCM: Spherically Consistent Coarse Matching for ERP Dense Feature Correspondence ACCV2026

链接: https://arxiv.org/abs/2609.36545
作者: Gyeonggwan Lee,Eunsoo Im,Seunghwan Hong,Junghun Suh
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to ACCV 2026. 35 pages: 16-page main paper (including references) and 19-page supplementary material. Project page: this https URL Code: this https URL

点击查看摘要

Abstract:Equirectangular projection (ERP) is the standard representation for 360 ^\circ imagery, and robust dense feature matching on ERP underpins panoramic stereo, view synthesis, and omnidirectional SLAM. Dense matchers trained on flat images degrade systematically on ERP because the chart introduces three coupled distortions – topological, metric, and area – that standard coarse matching and visibility estimation do not explicitly model. We show that correcting the three distortions at the coarse-stage interfaces where they arise – pairwise distortions in attention, per-pixel distortion in covisibility gating – improves PCK@ 1^\circ from 0.229 to 0.275 on Matterport3D under a fixed coarse scaffold, with the refiner architecture unchanged – our central result. Concretely, SCCM (Spherically Consistent Coarse Matching) augments a chart-naive cross-attention/dual-softmax coarse matcher with two sphere-derived priors: Spherical Positional Attention (SPA) pairs a yaw-periodic RoPE (topology) with a tangent-plane bias (metric), and Area-Aware Covisibility (AAC) applies a pre-sigmoid log-area correction (area). The chart-naive scaffold serves as a controlled reference, separating the scaffold-replacement effect from the spherical-prior effect. Instantiated in the RoMa V1 framework with the same frozen encoder, refiner architecture, and loss, SCCM also outperforms the ERP-native EDM (0.163) and an ERP-retrained RoMa V1 (0.198) under a unified ERP dense matching protocol, while perspective-trained matchers largely fail on ERP. It further transfers zero-shot to Stanford2D3D and, when trained on outdoor Holo360D, leads there as well.

[CV-215] Foresight at the Event Boundary: Evaluating Physical Prediction in Video World Models

链接: https://arxiv.org/abs/2609.36531
作者: Estela Monserrat Arriaga Santana(1),Julian Rosas Scull(1),Ehécatl Sacamch’en Núñez Rico(1),Hugo Jair Escalante(2) ((1) National Autonomous University of Mexico, (2) University of Texas at El Paso)
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 8 pages, 2 figures, 6 tables

点击查看摘要

Abstract:Video world models are largely regarded as predictive models of the physical world and are therefore expected to anticipate the consequences of observed events. However, evaluation has mainly focused on reference similarity, physical-law consistency, or judgment plausibility, estimating anticipation only indirectly. We address this directly: when a release or impact has just occurred but its consequence is withheld, can a world model anticipate what should happen next? We introduce an event-anchored evaluation based on 62 controlled real-world free-fall recordings and 124 clips spanning three object types, with fine-grained release and impact annotations and ground-truth trajectories. The protocol separates consequence production, temporal placement, and physical realization. Across six contemporary video generation and world models, Runway and Veo produce release and subsequent impact events at rates above 93% but often initiate them substantially late, whereas Cosmos-Predict-2.5 and MAGI-1 frequently preserve the pre-event state and produce little or no measurable consequence. Among measurable falls, plausible timing does not necessarily imply physically consistent motion. We further conduct a 15-participant, 20-condition human study in which participants describe the expected consequence from a single event-anchored frame and draw its trajectory. Human predictions favor the recorded future in aggregate while revealing genuine ambiguity among plausible continuations. Overall, physical foresight emerges as a sequence of distinct challenges: initiating a consequence, anchoring it in time, and realizing its motion.

[CV-216] Distilling Privileged Control Barrier Functions into RGB-Only Safety Filters for Dynamic Visual Navigation

链接: https://arxiv.org/abs/2609.36520
作者: Seungyeon Yoo,Gawon Lee,Seungwoo Jung,Inkyu Jang,H. Jin Kim
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Systems and Control (eess.SY)
备注: Project page: this https URL

点击查看摘要

Abstract:RGB-only end-to-end visual navigation policies remain vulnerable to collisions in real-world dynamic environments, motivating a dedicated safety layer. Existing visual Control Barrier Function (CBF) approaches seek to provide safety from RGB observations, but often rely on real-time rendering or explicit scene reconstruction and are primarily designed for static scenes, limiting their practicality for onboard deployment. We propose a teacher-student visual distillation framework that transfers the safety behavior of a privileged CBF teacher to an RGB-only student filter for dynamic environments. The student maps a short RGB history, robot velocity, and a nominal control action directly to a safe action, while the teacher uses ground-truth robot and obstacle states in a real-to-sim dynamic Gaussian Splatting environment. To reduce the teacher-student information gap, the teacher constructs safety constraints only from obstacles observable within the student’s RGB history. It also accounts for obstacle-velocity uncertainty to improve robustness to motion variations, while action augmentation exposes the student to diverse safe and unsafe nominal actions to better capture the safety boundary. At deployment, the student requires only RGB observations and robot velocity, without explicit 3D reconstruction or online rendering. Experiments show that the proposed method outperforms visual CBF baselines and improves the safety of RGB-based navigation policies under dynamic obstacle motion. Project page: this https URL.

[CV-217] Reimagine Video Dynamics

链接: https://arxiv.org/abs/2609.36496
作者: Yu Yuan,Yawen Lu,Guoxian Song,Kevin Duarte,Ratheesh Kalarot,Di Chang,Xijun Wang,Stanley H. Chan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project Page: this https URL

点击查看摘要

Abstract:Most video editing methods focus on changing the appearance of the source video, while offering limited control over its dynamics. We introduce Reimagine Video Dynamics (RVD), a framework that disentangles a compact, editable dynamics token from visual context. We learn this token through self-supervised reconstruction: given the first frame as visual context, a renderer must recover the original video from the dynamics token, encouraging it to capture how the scene evolves rather than how it looks. This disentanglement allows video dynamics to be edited directly while preserving visual context. We develop a language-guided dynamics-token editor that transforms source dynamics into target dynamics, and train it with a scalable counterfactual video-pair pipeline and a two-stage training strategy. Extensive experiments show that RVD enables effective video dynamics editing, training-free retiming, and appearance-controlled re-rendering.

[CV-218] Benchmarking Vision-Language Models on Synapse Detection and Proofreading in Connectomics

链接: https://arxiv.org/abs/2609.36492
作者: Yicong Li,Junjie Wang,Leander Lauenburg,Ella Hugie,Alexandra Irger,Wanhua Li,Donglai Wei,Hanspeter Pfister
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:We benchmarked vision-language models (VLMs) on the decisions annotators take when inspecting electron microscopy images in connectomics: synapse detection (presence and polarity) and proofreading (split errors and merge errors). For synapse detection, we evaluated 19 open and 2 closed models across various architectures and sizes under zero-shot, four-shot in-context learning and LoRA settings, against specialist models, on datasets constructed by us using public resources. For proofreading, we evaluated 3 open and 2 closed models on the ConnectomeBench2 dataset, with cross-species transfer from fly and mouse to human and zebrafish. Most models were at chance zero-shot; a few examples helped mainly the closed and largest open ones. LoRA on a few thousand labels brought open models level with specialist models. When evaluated on unseen species, the best adapted VLMs outperformed specialist models trained on the same data in identifying merge errors. The project will be publicly available upon acceptance.

[CV-219] Staircase Policy: Streaming Inference for World-Action Models with Large Action Chunks

链接: https://arxiv.org/abs/2609.36471
作者: Guoheng Sun,Chen Chen,Jin Wang,Ang Li,Teresa Lv
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:World-Action Models (WAMs) improve robotic manipulation by conditioning action generation on predicted future observations, but future prediction adds further inference overhead to already expensive iterative action generation. Action chunking can amortize this cost over multiple actions, yet performance degrades over long execution horizons because later actions remain conditioned on stale observations. We introduce STAIRCASE POLICY, a streaming inference and training framework that turns a flow-matching VLA into a JEPA-style WAM and partitions a large action chunk into sub-chunks at staggered denoising stages. Near-term actions are executed as soon as they become available, while later actions continue to be refined. At each sub-chunk boundary, the future latent is re-predicted from the latest observation and used to update all unexecuted actions, enabling long-horizon execution without repeated full policy inference. The resulting future-prediction error can further serve as a signal for adaptive chunking. S-WAM achieves 97.7% on LIBERO and 87.9% on LIBERO-Plus, and improves performance across multiple policy backbones and real-robot tasks. It reaches 292.7 executed actions per second, 3.62\times the throughput of conventional execution at comparable accuracy, while reducing time-to-first-action from 123.6 to 73.3 ms. With additional inference optimizations, throughput further increases to 642.9 actions per second.

[CV-220] DynamicHOI: Coupled Dynamics for Physics-aware HOI Reconstruction

链接: https://arxiv.org/abs/2609.36454
作者: Wenliang Guo,Zhanbo Huang,Yu Kong
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Page: this https URL

点击查看摘要

Abstract:We study hand-object interaction (HOI) reconstruction from monocular RGB videos, where partial observations can produce visually plausible yet mechanically inconsistent trajectories. Existing methods mainly enforce visual and geometric agreement, leaving the underlying interaction dynamics insufficiently constrained. We propose DynamicHOI, a physics-aware HOI reconstruction framework combining geometry-grounded diffusion refinement with coupled hand-object dynamics. Geometry spatially grounds visual evidence for trajectory refinement, while articulated inverse dynamics and Newton-Euler dynamics derive hand generalized forces and object wrenches for dynamics-level supervision. We further couple hand and object dynamics through contact-force transfer and recover active hand actuation as an interaction-level physical quantity. We formulate its empirical magnitude distribution into a probabilistic prior that penalizes unlikely actuation and suppresses mechanically implausible reconstructed motion. Experiments on three HOI datasets show consistent improvements in both hand and object reconstruction. The reconstructed trajectories further benefit downstream applications including hand world-model generation and robotic manipulation learning, demonstrating the value of physics-aware HOI modeling beyond reconstruction.

[CV-221] Online Versatile Incremental Learning: Towards Class and Domain-Agnostic Adaptation at Any Time ECCV2026

链接: https://arxiv.org/abs/2609.36442
作者: Jae-Ho Lee,Min-Yeong Park,Jun-Yeong Moon,Jung Uk Kim,Gyeong-Moon Park
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 17 pages, Accepted at ECCV 2026

点击查看摘要

Abstract:Continual learning enables vision systems to adapt to ever-changing data distributions. Despite significant advances, existing approaches fail to capture continuous and concurrent shifts in classes and domains, a critical capability for real-world deployment. This work introduces Online VIL (Online Versatile Incremental Learning), a novel scenario where class concepts and visual domains evolve simultaneously online without explicit boundaries. To better adapt to the challenges of such dynamic environments that more closely resemble real-world conditions, we propose a novel framework TopFlow, Topology preservation with Flow matching representation that contains two complementary mechanisms: Domain-agnostic Flow Matching (DFM) and Global Topology Preservation (GTP). DFM guides the model to have domain-agnostic representations by integrating the geodesic flow kernel into contrastive learning. In contrast, GTP maintains the global structure of the feature space without explicitly storing past examples. Our extensive experiments demonstrate that TopFlow effectively addresses the limitations of existing methods within the Online VIL scenario, achieving state-of-the-art performance in challenging Online VIL. The proposed methods suggest potential directions for building continual learning systems in realistic dynamic environments. Our implementation code is available at this https URL.

[CV-222] DARE to Mitigate Hallucination: Dual-path Auto-Regressive-aware Editing ECCV2026

链接: https://arxiv.org/abs/2609.36440
作者: Jae-Ho Lee,Jeong-Eun Lee,Gyeong-Moon Park
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 19 pages, Accepted at ECCV 2026

点击查看摘要

Abstract:Large vision-language models (LVLMs) have recently achieved remarkable progress across multimodal tasks, yet object hallucination remains a persistent challenge where models generate descriptions inconsistent with the visual input. Recent work mitigates hallucinations through training-free representation editing, typically by constructing hallucination-related directions from teacher-forcing (TF) contrasts between hallucinated and truthful responses. However, LVLMs operate through autoregressive (AR) decoding during generation, raising the question of whether TF-based analysis fully reflects the generation dynamics that lead to hallucinated outputs. In this paper, we analyze the relationship between TF-based editing and AR generation behavior and find that TF-based editing alone may be insufficient to capture both decoding dynamics and multimodal interactions associated with hallucinations. To address this limitation, we propose DARE (Dual-path Auto-Regressive-aware Editing), a hybrid hallucination editing framework that integrates two complementary contrast pathways: textual contrasts and image contrasts, together with autoregressive-aware representation signals. Specifically, DARE constructs hallucination editing directions from (1) TF-based textual contrasts, (2) AR-aware representation transitions during decoding, and (3) controlled visual differences between paired images. Extensive experiments on multiple LVLM hallucination benchmarks demonstrate that DARE consistently reduces object hallucinations while preserving multimodal perception capability and inference efficiency. Our implementation code is available at this https URL.

[CV-223] Merlin Plus: A Large-Scale Multi-Cancer Image-Mask-Report Dataset MICCAI2026

链接: https://arxiv.org/abs/2609.36436
作者: Pedro R. A. S. Bassi,Wenxuan Li,Szymon Plotka,Ruby Honjol,Jakub Przado,Xinze Zhou,Kang Wang,Yang Yang,Malte Jensen,Akshay S. Chaudhari,Curtis P. Langlotz,Alan L. Yuille,Zongwei Zhou
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: MICCAI 2026

点击查看摘要

Abstract:Multi-cancer segmentation in computed tomography (CT) is fundamentally limited by the scarcity of tumor masks across different organs. We present Merlin Plus, the first large-scale CT dataset with radiologist-created tumor masks across 9 organs. Merlin Plus extends the Merlin dataset by adding 1,153 per-voxel tumor masks and longitudinal metadata. To create these tumor masks, we developed a report-based active-learning framework in which radiology reports identify tumor cases for annotation and support training of a tumor segmentation model. The model generates initial masks, which radiologists review and correct to produce the final masks, reducing annotation burden while maintaining high-quality annotations. Besides tumor masks, the longitudinal metadata in Merlin Plus enables temporal modeling of cancer progression. By directly addressing the major bottleneck of limited multi-cancer segmentation masks, Merlin Plus supports scalable multi-organ cancer detection, segmentation, and longitudinal analysis in CT. Dataset is available at: this https URL

[CV-224] RA-CFGCache: From Branch-Level Criteria to Guided-Risk Control under Classifier-Free Guidance NEURIPS2026

链接: https://arxiv.org/abs/2609.36433
作者: Yiming Liu,Ben Wan,Tongxuan Liu,Ao Wang,Yuqi Xiong,Fan Zhang,Hui Chen,Guiguang Ding
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 30 pages, 17 figures, NeurIPS 2026

点击查看摘要

Abstract:Diffusion models enable high-quality visual generation, but iterative denoising remains computationally expensive, especially under classifier-free guidance (CFG), which requires both conditional and unconditional evaluations. Training-free caching reduces this cost by reuse of previously computed features or predictions. However, existing branch-local reuse criteria do not explicitly account for how cache errors combine under CFG or how local perturbations affect the final output. We identify two misalignments in cache control: a branch-guided mismatch, where guided error depends on both the magnitudes and alignment of branch errors, and a local-final mismatch, where the downstream impact of a local error varies across timesteps. We propose RA-CFGCache, a Risk-Aligned Caching framework under CFG that incorporates both factors while keeping the sampling schedule and guidance rule fixed. CFG-aware Guided-Risk Composition combines existing branch-wise proxies using CFG coefficients and offline-calibrated cross-branch alignment. Propagation-Aware Rescaling further weights the resulting guided-risk estimate with a timestep-dependent propagation prior calibrated from isolated reuse perturbations. An online threshold controller then determines when to jointly refresh or reuse both branches. Experiments on FLUX.1-dev, Wan2.1-T2V-1.3B, and CogVideoX-2B demonstrate improved efficiency–fidelity trade-offs over evaluated training-free caching baselines. Moreover, RA-CFGCache is compatible with diverse base proxy families, including TeaCache-, DiCache-, and MagCache-style estimators, and consistently improves fidelity at nearly unchanged latency. Code is available at this https URL.

[CV-225] mporal-Aware Fusion for Robust Outdoor LiDAR Localization

链接: https://arxiv.org/abs/2609.36432
作者: Minghang Zhu,Zhijing Wang,Yuxin Guo,Chen Liu,Yongshu Huang,Wen Li,Sheng Ao,Cheng Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 10 pages, 8 figures, 6 tables, conference

点击查看摘要

Abstract:LiDAR relocalization aims to estimate the global 6-DoF pose of a sensor in the environment. However, existing regression-based approaches often encounter limitations in dynamic or ambiguous scenarios, as they typically prioritize single-frame inference, leaving the potential of spatio-temporal consistency across scans not fully explored. In this paper, we propose a Temporal-aware Localization framework (TempLoc) designed to enhance the robustness of outdoor localization by effectively modeling sequential consistency. Specifically, a Global Coordinate Estimation module is first introduced to predict point-wise global coordinates and associated uncertainties for each LiDAR scan. A Prior Coordinate Generation module is then presented to estimate inter-frame point correspondences by the attention mechanism. Lastly, an Uncertainty-Guided Coordinate Fusion module is deployed to integrate both predictions of point correspondence in an end-to-end fashion, yielding a more temporally consistent and accurate global 6-DoF pose. Experimental results on the NCLT and Oxford RobotCar benchmarks show that our TempLoc outperforms state-of-the-art methods by a large margin, demonstrating the effectiveness of temporal-aware correspondence modeling in LiDAR relocalization.

[CV-226] owards Scalable Context-Aware Single-Cell Spatial Transcriptomics Prediction from Histology Images NEURIPS2026

链接: https://arxiv.org/abs/2609.36429
作者: Zijun Gao,Chunbin Gu,Jinxi Xiang,Xiangde Luo,Pheng-Ann Heng
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to NeurIPS 2026

点击查看摘要

Abstract:Predicting gene expression from HE-stained histology images offers a scalable alternative to costly spatial transcriptomics, yet most existing methods operate at the spot level, where signals from multiple cells are aggregated and critical cellular heterogeneity is obscured. Extending this paradigm to single-cell resolution is non-trivial. Naively applying pathology foundation models faces a scale mismatch: their patch-level representations mix multiple cells, whereas per-cell cropping or resizing distorts morphology and removes local context. Conversely, segmentation-based models without strong pretrained visual encoders often lack the morphological representation capacity needed for accurate molecular prediction and inherit errors from imperfect cell boundary masks. Here, we present CELLO, an efficient end-to-end framework that performs a single pathology foundation model forward pass per image and uses grid sampling to extract location-specific features for all cells simultaneously. We further introduce a distance-decay cross-attention module that refines each cell representation using spatially biased local morphological context. Using 52 public Xenium-HE pairs from HEST-1k that span 12 organs and approximately 10 million cells, CELLO improves the average predictive accuracy over the evaluated baselines while reducing the mean whole-slide inference time compared to DeepSpot2Cell, a 14.0x speed-up on average that excludes upstream cell segmentation. Our work establishes a scalable foundation for single-cell gene expression prediction from HE images.

[CV-227] Losing the name before the box: measuring and repairing what narrow fine-tuning costs a detector outside its deployment vocabulary

链接: https://arxiv.org/abs/2609.36426
作者: Trung Minh Bui,Jongsul Moon,YoungOuk Kim,Jung-Hoon Hwang,Dongin Shin
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 25 pages, 3 figures. Supplementary material (69 pages) is included as an ancillary file. Submitted to the International Journal of Computer Vision

点击查看摘要

Abstract:A detector pretrained on a broad corpus is fine-tuned on a narrow domain, its in-domain accuracy improves, and it ships. We ask what happens meanwhile to its coverage of objects the vocabulary never names, which in obstacle detection and inspection carry the risk. No in-domain test set holds an example of one. We give a longitudinal protocol: one pretrained checkpoint against its own fine-tuned descendants. It tracks held-out top- K proposal coverage C_\tau : of categories pretraining covered and the vocabulary omits, the share of boxes a detector’s top K regions still cover. The quantity is the open-world proposal literature’s; the longitudinal reading is not. C_\tau falls while in-domain accuracy rises, on four architectures and three domains, by 5.12 to 63.35 points on boxes above 1024 px ^2 . No in-domain number identifies the fall, and neither does detection average precision, which charges a missed and a misnamed box alike. On the one architecture scoring both, adaptation costs 87% of the AP against a fifth of the coverage, and the naming goes first at all six depths of its freeze ladder, every run. What breaks is structured: three architectures sharing no pretraining run agree on which categories lose coverage, and those a model never learned do not lose any. A repair follows and needs no training: mixing a quarter of the pretrained state back, normalisation statistics included, raises coverage on every cell swept for at most 2.47 points of in-domain accuracy. Seeing it costs one extra evaluation pass.

[CV-228] FineART: Fine-grained Annotated Robotic Trajectory Dataset and Vision-Language-Action Model for Bimanual Manipulation

链接: https://arxiv.org/abs/2609.36416
作者: Jade Choghari,Pepijn Kooijmans,Mansi Agarwal,Yusuf Umut Ciftci,Aseem Doriwala,Catherine Weaver,Mouli Sivapurapu,Kai Yang,Jackson Lee,Thomas Wolf,Pragna Mannam
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 26 pages. Code and model weights will be integrated into Hugging Face LeRobot this https URL

点击查看摘要

Abstract:Robots operating in real-world environments must execute complex, multi-step bimanual tasks over long horizons rather than single, isolated actions. Current manipulation datasets struggle to support this capability: although single-arm datasets reach hundreds of thousands of trajectories, they typically provide only one high-level instruction per episode while the rare bimanual effort that does label subtasks annotates only a fraction of its hours. We present FineART, a densely annotated bimanual manipulation dataset of 40,543 episodes, 1,718 hours, and 533,913 subtasks across 151 tasks. We also introduce FineART-VLA, a vision-language-action policy that predicts its own next subtask, and show that mid-training it this way yields substantial gains. Specifically, success on a spatial disambiguation task increases from 32.0% to 100.0%, and step-by-step human subtask guidance lifts success on an unseen long-horizon task from 16.0% to 76.0%. Furthermore, after minimal fine-tuning on a new robot, the policy requires one-tenth the data of baselines without mid-training and generalizes zero-shot to completely unseen tasks on the new hardware. We open-source the full dataset, model weights, and training code.

[CV-229] What Makes High-Magnification Knowledge Transferable? A Study of Cross-Resolution Distillation in Whole-Slide Imaging ICLR2027

链接: https://arxiv.org/abs/2609.36407
作者: Zhiyuan Yang,Jiahao Cheng,Mahdi S. Hosseini
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: ICLR 2027 Submission

点击查看摘要

Abstract:Cross-resolution knowledge distillation aims to improve low-magnification whole- slide analysis by transferring high-magnification representations, yet the conditions for useful transfer remain unclear. We develop a decomposition-based analysis of teacher access, representation loss, and model excess, motivating three questions: whether (a) teacher targets help the task, (b) low-magnification students can predict them, and © slide models benefit from those predictions. We investigate them through controlled experiments across ten pathology cohorts spanning classifi- cation, grading, and survival prediction. In the main comparison, providing teacher regional means alongside native low-magnification features improves downstream performance in all ten cohorts. Direct prediction achieves lower reconstruction error than residual prediction, yet the predicted features underrepresent variation in the teacher targets. Moreover, better reconstruction does not consistently improve downstream scores, and retaining native features changes performance even when the predicted teacher features are held fixed. Together, these findings expose a gap between reconstructing teacher representations and realizing their downstream value. They challenge the sufficiency of reconstruction error as a measure of cross-resolution transfer and provide a diagnostic framework for examining where that transfer breaks down. Future distillation designs must account for both what students can predict and how slide models use those predictions.

[CV-230] Stealth Is a Relation Not a Property: How Event Representations Create Blind Spots for Timing Attacks in Event-Based Perception ALT

链接: https://arxiv.org/abs/2609.36386
作者: Shoaib Ahmed Dipu,Md. Shaown Miah,Kamrul Hasan,Sayeed Shafayet Chowdhury
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 30 pages, 8 figures, 20 tables. Code: this https URL

点击查看摘要

Abstract:An event camera produces an asynchronous stream, but what is visible in that stream depends on how a downstream consumer, such as a model or detector, processes time. The same timestamp change may leave a coarse temporal representation unchanged while changing the response of a model that preserves finer timing. We characterize this dependence as observer-relative stealth. For recorded event streams, retiming an event within its protected accumulation window leaves the accumulated integer tensor exactly unchanged. We use this exact blind space to construct Null, a gradient-guided timestamp-retiming attack, and define SC-ASR_A(tau) to measure attack success while bounding the change visible to observer A. On DVS Gesture at a 10% event budget, Null reaches 81.56 +/- 5.81% ASR on ConvSNN and 98.67 +/- 0.45% on a GRU while preserving the protected tensor exactly. On DailyDVS-200, a protocol-scale Multi-View Fusion Network variant reaches 99.28 +/- 0.11% exact-null ASR, compared with 9.70 +/- 1.06% for its matched control. In a five-attack comparison, Null is the only method with nonzero attack success at exact observer equality, reaching 81.4% on DVS Gesture and 87.35% on DailyDVS-200. We also search the same exact blind space with an independently implemented constrained projected-gradient optimizer, C-PGD. At matched victim-gradient evaluations, C-PGD reaches 84.50 +/- 2.89% ASR on DVS Gesture and 89.55 +/- 4.39% on DailyDVS-200, again with exact protected equality. Perturbations that are exactly hidden from the protected observer become visible under shifted, finer, overlapping, and randomized temporal views. Adding observer constraints reduces the real-valued blind-space fraction from 87.5% to 75.0% to 62.5%, while DVS ConvSNN ASR falls from 74.9% to 61.9% to 37.2%. These results show that stealth is not a property of the perturbation alone.

[CV-231] LEGO-Anything: Coding Agents for 3D Scene Reconstruction

链接: https://arxiv.org/abs/2609.36380
作者: Xirui Li,Peng Shi,Mingwen Dong,Sheng Zhang,Zhuoyan Xu,Dongkyu Lee,Shuaichen Chang,Yi Xiang,Lin Pan,Jiarong Jiang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:A 3D scene reconstructed from a single image is most useful when represented not as a rendering or a fixed 3D output, but as an explicit scene program whose execution yields a scene that can be inspected, edited, and queried. We present LEGO-Anything, an Image-to-Code framework in which a coding agent iteratively writes and executes Blender code, inspects scenes and renderings, and revises the program. To evaluate end-to-end scene recovery, we introduce LEGO-Bench, a simulator-grounded benchmark with 208 images from 104 diverse indoor and outdoor scenes. LEGO-Bench separately scores artifact validity, visible-surface geometry, and rendered appearance. Its simulator-grounded design enables extensibility and precise automatic evaluation. Among evaluated agents, GPT-6-astra achieves the strongest overall results, with 53.4% indoor and 39.6% outdoor scores, yet substantial gaps remain between delivering valid scene artifacts and faithfully recovering scene geometry and appearance. Analysis of agent construction trajectories reveals three recurring issues: weak scene initialization, regressive edits during iteration, and unreliable self-evaluation. These findings motivate LEGO-Plugin, a training-free harness plugin for more controlled iterative scene construction, which improves all six evaluated models, with relative gains of up to 62.7% in overall score. Finally, we test whether reconstructed scenes can represent natural images and support vision tasks. In LEGO-World, we derive object detections, instance masks, and relative depth as deterministic queries on scenes reconstructed by GPT-6-astra. These readouts show non-trivial performance across all three tasks but fall well short of specialized vision models, suggesting that program-constructed scenes from current coding agents are a promising but not yet sufficiently precise representation of natural images.

[CV-232] OTT3R: Multi-View 3D Reconstruction and Fast Dataset Generation at 1% Compute ACCV2026

链接: https://arxiv.org/abs/2609.36374
作者: Brandon Leblanc,Charalambos Poullis
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at ACCV 2026

点击查看摘要

Abstract:Feed-forward 3D reconstruction models have achieved impressive performance by scaling model and dataset size, but their cost excludes most research groups and precludes edge deployment. Additionally, generating 3D supervision without sensors still relies on slow, unreliable Structure-from-Motion, as the community lacks a COLMAP-like system for neural 3D pseudo-label generation. We present OTT3R (RGB-Only Tiny Transformer for 3D Reconstruction), a knowledge distillation framework that addresses both problems on a single workstation equipped with 2 GPUs. Distilling \pi^3 (959M parameters) into a 102M-parameter student yields 9.4 \times compression and up to 7 \times faster inference, trained at 1.6% of VGGT’s training compute. An integrated pseudo-label pipeline offers a reliable, high-throughput alternative to COLMAP, generating dense per-pixel point maps and SE(3) camera poses for a 667K-image corpus in 3.5 hours on two commodity GPUs and succeeding on every sequence we tested, including those where COLMAP fails. The general student tracks the teacher on in-distribution monocular depth and, zero-shot, outperforms COLMAP on 7-Scenes and on DTU completion, but it does not replace the teacher on out-of-distribution multi-view geometry. The deployable artifact is the domain-specialized student: after specialization at 0.2% compute, it is 4 \times more accurate than COLMAP on 7-Scenes at 980 \times throughput, with near-teacher completion. Code is available at this https URL

[CV-233] AdaKerNet: Neural Kernel Decoding for Task-Adaptive Prediction with Multimodal Large Models

链接: https://arxiv.org/abs/2609.36368
作者: Konstantinos D. Polyzos,Eleni Oikonomou,Tara Javidi
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Large foundation models have been introduced with the promise of efficient adaptation to downstream tasks. Yet, under limited supervision, MLLMs, an important class of large foundation models, remain challenging to adapt to various downstream tasks. Adaptation typically relies either on MLLM parameter fine-tuning or on training neural-based decoders. Both approaches struggle under limited supervision, while fine-tuning additionally requires access to model parameters, which is often unavailable for closed-source models. We introduce AdaKerNet, a novel learnable task-adaptive neural kernel decoder. AdaKerNet is fully agnostic to the parameters of the underlying MLLM and operates solely on its (frozen) rich representations obtained from the diverse available modalities. AdaKerNet relies on (i) a set of learnable, Lipschitz-controlled multimodal features derived from these MLLM representations; (ii) a reference kernel that provides a soft structural prior on those features; and (iii) a lightweight nonlinear neural predictor that adaptively deforms that structure. Learning the kernel representation and the neural predictor jointly within a unified optimization framework allows AdaKerNet to capture features and geometric relationships relevant to the downstream task. Numerical tests across four MLLMs: BLIP-2, LLaVA-1.5, Qwen2.5-VL, and Gemini Embedding 2, and multimodal inputs spanning text, audio, images, and tabular measurements demonstrate significant and consistent improvements over direct MLP, attention-, autoencoder- and kernel-based decoders, across a range of scarce-label budgets, with average error reduction of up to 41% across baselines. These results establish AdaKerNet as an effective approach for prediction from frozen multimodal representations in the scarce label regime. Additional structural ablations highlight the complementary contributions of AdaKerNet’s components.

[CV-234] Compress to Remember: Learning Compact Memory via On-Policy Distillation for Long Video Generation

链接: https://arxiv.org/abs/2609.36364
作者: Xiaoyu Wu,Weihang Guo,Yifei Wang,Xinze Feng,Lydia E. Kavraki,Zhiwei Steven Wu
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Under Review

点击查看摘要

Abstract:Standard video generators do not natively compact historical context into reusable memory tokens. As generation continues, the growing history makes it increasingly difficult to retain information from earlier frames due to long-context degradation. Key-frame-based approaches address this challenge by retaining selected past frames, but can discard information needed for future generation. Rather than relying on frame selection alone, we study whether a frozen video generator can supply the supervision needed to learn a compact representation of the history. We propose Prediction-Aligned Context Compaction (PACC), which uses a learned compressor to aggregate information across past frames into compact memory tokens. We train the compressor through on-policy distillation, using the same frozen generator both as a student when conditioned on compressed memory and as a teacher when conditioned on the full history. The student generates continuations, while the teacher provides targets for the same noisy inputs at each denoising step. Only the compressor is updated to align the student’s predictions with these targets. We evaluate PACC on MBench, which jointly measures memory-event coverage and consistency. PACC outperforms the strongest baseline by 6.63 points on Causal-rCM and 3.19 points on Causal Forcing. Evaluation on VBench-Long using MovieGen prompts further shows that PACC produces minute-long videos with generation quality competitive with baselines. Together, these results show that learning to compact historical context can improve long-video memory without modifying the underlying generator.

[CV-235] StructRL: Online Structured Reinforcement Learning for Long-Horizon Vision-Language-Action Tasks

链接: https://arxiv.org/abs/2609.36352
作者: Ziyi Yin,Sangmin Woo,Kang Zhou,Sungyeon Kim,Aosong Feng,Haibo Ding,Jun Huan
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Vision-language-action (VLA) models perform well on shorter-horizon manipulation tasks but still struggle with long-horizon tasks that require multiple dependent manipulations from a single command. Online reinforcement learning (RL) can improve these policies through environment interaction, yet many existing methods provide reward only after the complete task succeeds. However, such terminal supervision is sparse and does not distinguish early failures from rollouts that make substantial partial progress. We propose StructRL, an online RL framework that constructs structured intermediate supervision from verifiable subtask completions. StructRL decomposes each task into verifiable subtasks, grants intermediate rewards only after the prerequisite subtasks have been completed, and scales each reward according to completion pace. Across RoboCasa365 and LIBERO-Long with GR00T-N1.5 and pi 0.5, StructRL consistently outperforms evaluated online RL baselines. These results show that verifiable, structured intermediate rewards improve long-horizon VLA post-training. Code is available at this https URL.

[CV-236] Representation by Design in Generation: Cross-View Class-Token Alignment in Diffusion Transformers

链接: https://arxiv.org/abs/2609.36348
作者: Xiaoyu Wu,Yifei Wang,Chen Wei
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Generative and representation learning remain asymmetrically connected: semantic representations are used to improve diffusion generation, whereas the models’ own representations are often treated as a by-product of synthesis. We ask whether diffusion models can instead be trained to learn substantially stronger semantic representations without sacrificing generation quality. SelfFlow takes a step in this direction by introducing self-supervised patch alignment into flow matching, but its main gains remain in faster convergence and improved generation. Inspired by DINO and iBOT, we extend this framework with cross-view class-token alignment to further strengthen semantic representations. Specifically, we form two independently noised, dual-timestep observations of each image and align each student class-token representation with the stop-gradient EMA-teacher target from the other observation. This objective is optimized jointly with the inherited flow-matching and local patch objectives. Notably, although the additional objective acts only on the class token, it strengthens both class-token and patch representations. Compared with a matched two-view baseline, ImageNet linear-probing accuracy improves by 9.4% using the class token and 10.1% using mean-pooled patch tokens, while frozen-backbone VOC2012 segmentation improves by 3.6 mIoU. These representation gains are achieved while maintaining comparable ImageNet generation FID. In text-to-image training, the same objective also improves generation FID, reducing it from 2.52 to 2.37 at matched checkpoints. Our results show that representation need not remain a by-product of generation or merely a tool for improving it: it can be directly optimized as a first-class capability of diffusion pretraining alongside generation.

[CV-237] PyroStack: A Multi-Band Spatio-Temporal Sub-Daily Dataset for Wildfires in the United States

链接: https://arxiv.org/abs/2609.36315
作者: Arya Kondur,Giosue Migliorini,Cameron Schmitt,Francesco Immorlano,Tairan Wang,Rebecca C. Scholten,Efi Foufoula-Georgiou,Gary Johnson,Chris Lautenberger,Valentin Waeselynck,J. Shane Romsos,Kasra Shamsaei,Alejandro Tejedor,Tianjia Liu,Yang Chen,Padhraic Smyth,James T. Randerson
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注: 23 pages, 6 figures, 5 tables

点击查看摘要

Abstract:Wildfires are an increasing hazard to ecosystems, air quality, and human systems, creating a growing need for datasets that support systematic development and evaluation of models for predicting fire spread across diverse landscapes. Effective prediction requires integrating meteorological conditions, fuels, vegetation, and topography at spatial and temporal resolutions suitable for both physical simulation and data-driven approaches. However, existing datasets often lack the resolution and coverage needed to capture these interacting controls. The PyroStack dataset addresses this gap by providing a harmonized, event-based collection of wildfire and environmental data across the contiguous United States and Alaska. It integrates satellite-derived fire observations with atmospheric reanalysis, vegetation, fuel characteristics, and topographic information into a unified framework spanning 6994 wildfires that occurred between 2012 and 2024 across a wide range of ecosystems and climate conditions. PyroStack offers spatial resolutions ranging from 30 m to 9 km and hourly temporal resolution, along with fire progression data at 12-hour intervals to support model initialization and evaluation. By combining broad spatial coverage with fine spatial and temporal detail, the dataset enables systematic analysis of wildfire dynamics and supports both physics-based and machine learning approaches, providing a foundation for benchmarking and improving fire spread models, with future extensions aimed at incorporating additional regions and fire suppression data streams to further advance wildfire prediction.

[CV-238] hink Before You Restore: Risk-Aware Manchu Manuscript Restoration with Stroke-Guided Attention ICASSP2027

链接: https://arxiv.org/abs/2609.36243
作者: Mingqiu Liang,Dongdong Wang,Siyang Lu,Ting Huang,Yingjun Qi
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Submitted to IEEE ICASSP 2027

点击查看摘要

Abstract:Full-page blind restoration of historical Manchu manuscripts is challenging due to scarce annotations, unknown degradation regions, and fragile connected strokes. Generic restoration models may improve visual quality but often modify intact content, leading to over-restoration. We propose SAGE-Restore (Stroke-Aware Gated rEstoration), a selective restoration framework that first assesses where restoration is needed and then uses this assessment to guide restoration candidate generation and pixel-level selection. Its encoder predicts patch-level repair probabilities from complementary appearance and stroke-structural cues to condition restoration candidate generation, while the corresponding repair logits are refined into a pixel-level soft gate that selectively controls where the restoration candidate is applied. We further introduce a fidelity-aware evaluation protocol that jointly measures degraded-region recovery, intact-content preservation, and their balance. SAGE-Restore achieves the highest R-Recovery (0.463) and RFS (0.626), while maintaining high U-Fidelity (0.968), demonstrating an effective balance between restoration and content preservation.

[CV-239] Mutually Adversarial Self-Training with Evolving Data for Unified Multimodal Models

链接: https://arxiv.org/abs/2609.36224
作者: Wentao Zhou,Weijie Gan,Jiayun Wang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Unified multimodal models (UMMs) combine image generation and visual understanding in a shared backbone. Since generation and understanding are inverse tasks, recent studies self-train UMMs by letting the two branches cooperatively supervise each other. We introduce MATE (Mutually Adversarial self-Training with Evolving data), a reinforcement-learning-based post-training framework in which the two branches instead challenge each other, and the challenges evolve as the model trains. MATE lets generation and understanding take turns to be challenger and solver. Given an image, the understanding branch proposes several candidate descriptions that the generation branch must turn back into similar images, and vice versa. The candidates are screened for consistency with the image or prompt they were proposed from, and the solver is trained on the candidate it handles worst. The adversary thus comes from the model’s own outputs, and no separate adversary is trained. Moreover, the candidates that defeat one branch become the sources of the next challenges to the other in the next epoch, which keeps the challenges evolving with the model and turns the training into self-play in data space. On Janus-Pro-1B, MATE improves GenEval by 2.4 points, DPG-Bench by 1.7 points, and the average over nine understanding benchmarks by 0.7 points, while strengthening consistency across repeated image-text cycles.

[CV-240] LeRF: Learning Reference Coordinate Frames for Perspective Taking Reasoning

链接: https://arxiv.org/abs/2609.36219
作者: Bang Xiao,Wenqi Jia,Ozgur Kara,Tiancheng Shen,Yibo Yang,Bolin Lai,Junho Kim,James Matthew Rehg
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 22 pages, 8 figures

点击查看摘要

Abstract:Perspective taking is a fundamental component of spatial intelligence, requiring models interpret spatial relations from a specified viewpoint, such as that of another entity or an imagined observer. Despite the increasing spatial reasoning capabilities of Vision-Language Models (VLMs), they still struggle with perspective taking, often defaulting to the camera viewpoint when a query requires reasoning from a different perspective. We introduce Learning Reference Coordinate Frames for Perspective Taking (LeRF), a framework that trains VLMs to construct and use explicit reference frames for viewpoint-dependent reasoning. Given an image and a query, LeRF decides whether a coordinate frame is necessary. If so, it grounds the reference entity and predicts the frame’s origin and entity-centered reference frame. A lightweight renderer overlays the frame onto the image, enabling subsequent reasoning over these visual cues without external perception models or explicit 3D reconstruction. To learn this process, we first perform supervised fine-tuning to teach selective tool invocation and reference coordinate frame prediction, followed by reinforcement learning on spatial VQA pairs to improve frame-guided reasoning. Across diverse perspective-taking benchmarks, LeRF consistently improves over its backbone and achieves strong performance against existing open-source methods. Further evaluations also show improved reference-frame grounding and orientation estimation, supporting the effectiveness of learned reference frames for viewpoint-dependent reasoning.

[CV-241] Sparse-View Interpretable 3D Animal Behavior Representations for Neural Encoding and Decoding

链接: https://arxiv.org/abs/2609.36217
作者: Xinming Dai,Qihang Jin,Tianshu Tan,Baiyuan Chen,Hanrui Lyu,Lenny Aharon,Kyle Daruwalla,Xun Helen Hou,Matthew R. Whiteway,Liam Paninski,Yizi Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:A deeper understanding of brain function requires a precise, structured characterization of this http URL, extracting behavioral representations from video in a form suitable for scientific analysis remains a fundamental challenge. Many prior studies represent behavior via pose estimation or nonlinear video embeddings. However, pose tracking discards rich information beyond predefined keypoints, while nonlinear video embeddings lack interpretability. We address this limitation with SABLE (Sparse-view Animal Behavior Latent Embeddings), a self-supervised framework that leverages a geometric inductive bias to learn behavior this http URL augmenting a multi-view transformer with priors from monocular depth and pose estimation, SABLE reconstructs 3D animal behavior from extremely sparse views while learning explicit 3D latent structure. Without ground-truth 3D labels, it reliably recovers 3D behavior from two-view videos, whereas state-of-the-art (SOTA) methods fail or yield degenerate solutions. Across the International Brain Lab and Cheese3D datasets, we demonstrate that SABLE learns 3D representations that match or exceed prior SOTA performance in neural encoding and decoding. Once pretrained across animals, SABLE serves as an off-the-shelf model that generalizes zero-shot to unseen animals without animal-specific calibration or retraining. Our method establishes 3D-aware video embeddings that capture complex behavior, opening new avenues for studying brain-behavior relationships.

[CV-242] On the spectral properties of generative denoiser Jacobians

链接: https://arxiv.org/abs/2609.36210
作者: Alexandros Graikos,Nebojsa Jojic,Dimitris Samaras
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Generative denoising models, such as diffusion and flow-matching, learn to sample from complex distributions by training a deep neural network denoiser to recover clean data from noise-corrupted samples. While such models are typically compared on the quality of their synthesized samples, these metrics provide limited insight into how the underlying denoiser, which drives generation, differs. In this work, we propose to analyze the spectrum of the denoiser Jacobian as a tool to characterize these differences. Across pre-trained denoising models, we observe that better generative performance is associated with larger Jacobian eigenvalues. Motivated by this, we introduce a regularization scheme that controls the Jacobian spectrum by training the denoiser on perturbed inputs, with perturbations suppressing or amplifying Jacobian responses. On ImageNet, we test whether directly modifying the Jacobian spectral properties leads to improved generations. Our findings suggest that denoisers benefit from both strengthening responses along data-relevant principal eigen-directions and suppressing the noisy, data-irrelevant ones. This establishes the denoiser Jacobian as a useful tool for identifying differences between generative denoising models.

[CV-243] PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents

链接: https://arxiv.org/abs/2609.36199
作者: Vighnesh Subramaniam,Boris Katz,Brian Cheung,Chun-Liang Li,Tomas Pfister,Yale Song
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 24 pages, 11 figures, 3 tables

点击查看摘要

Abstract:Diffusion models can produce striking images and videos, but they still struggle with the compositional details that make a generation faithful to a prompt, such as object counts, attribute binding, spatial relations, and temporally grounded actions. A common way to improve prompt satisfaction is to spend more compute at test time through Best-of-N sampling, but final-sample selection is fixed. Best-of-N can only choose among completed outputs and cannot repair a promising trajectory before it fails. We introduce PreviewDiff, a training-free test-time search method that turns diffusion sampling from scalar search into a multimodal critic-guided search over intermediate latents. At selected denoising checkpoints, PreviewDiff decodes a partial preview, asks a multimodal judge to score and critique it, and uses the resulting natural-language feedback to branch over semantic prompt edits and locally re-noised latent continuations. These branches are then scored and selectively rolled forward, allowing verifier compute to guide generation while the sample is still editable. Across image and video generation benchmarks, PreviewDiff consistently improves over budget-matched Best-of-N selection and strong scalar-search baselines. Ablations show that earlier interventions and increased search width provide the largest gains, while deeper search and additional semantic variants offer complementary improvements. PreviewDiff demonstrates that multimodal feedback is most useful not only as a final verifier, but as an active controller inside the denoising process.

[CV-244] FD-AA: A Lightweight Focal-Diffuse And Attenuation-Aware Head for Incidental Abdominal Abnormality Detection in Chest CT

链接: https://arxiv.org/abs/2609.36189
作者: Haoyan Ding,Kritika Iyer,Halid Yerebakan,Zhenyu Bu,Chushu Shen,Peiyu Duan,Xinyuan Zheng,Sepehr Farhand,Xueqi Guo,Chaowei Wu,Yoshihisa Shinagawa,Gerardo Hermosillo Valadez
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Submitted to a conference

点击查看摘要

Abstract:Routine chest CT captures upper-abdominal structures that may contain clinically relevant incidental abnormalities. Detecting these findings requires feature extraction from organs with different spatial extents and attenuation patterns. We propose FD-AA, a lightweight organ-aware classification head adaptable for frozen 3-D CT encoders. Within each organ, an attenuation-aware module preserves sparse focal evidence, while masked generalized-mean pooling captures diffuse anomaly patterns. In seven abdominal organs, FD-AA with Pillar-0 achieved state-of-the-art (SOTA) performance in both the CT-RATE test set (AUC = 0.798) and the external RAD-ChestCT dataset (AUC = 0.713). More specifically, FD-AA improved macro AUC/AP from 0.763/0.346 to 0.798/0.405 over direct classification using frozen Pillar-0 only (p = 0.034/0.016). Such performance gain generalizes across multiple frozen encoders (AUC improvement on MedicalNet +9.8%, CT-CLIP +14.7%, ResNet +3.7%), demonstrating the effectiveness of FD-AA across different feature representations. These results support the effectiveness of integrating focal-diffuse aggregation with explicit HU evidence for incidental abdominal abnormality detection.

[CV-245] Exploring Learning Models for Topological Relationship Recognition from Image Data

链接: https://arxiv.org/abs/2609.36172
作者: Saptak Das,Monidipa Das
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Figuring out how objects relate to each other, like whether they touch, overlap, stay completely separate or one sits inside another, matters a lot in fields like GIS, biomedical imaging, and robotics. Even though machine learning has come a long way, people haven’t really focused on spotting these topological relationships in images. The main roadblocks? Not enough good datasets and no clear way to measure results. So, we rolled up our sleeves and built a new dataset. It’s pretty sizable: over 11,000 labelled images showing all those essential relationships. We ran tests with some classic machine learning models, Naive Bayes, KNN, Random Forest, SVM, and Artificial Neural Networks, and threw in some deep learning stars like VGG16 and InceptionResNetV2. For the dataset itself, we used segmentation, contour detection, and grayscale normalization to tease out solid feature vectors. The results? Deep learning methods, especially VGG16, pulled ahead, with validation accuracy hitting 89.55%. That’s a big jump compared to the traditional models. This shows how powerful transfer learning is for analyzing topological relationships in images, and it gives researchers a new standard to aim for in future work on spatial reasoning and topological classification.

[CV-246] Boosting Metric Depth Completion via Training-Free Adaptive Response Geometry

链接: https://arxiv.org/abs/2609.36168
作者: Mia Zhang,Jizong Peng
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Depth completion aims to recover dense metric depth from sparse sensor measurements, increasingly leveraging visual foundation models as geometric priors. However, aligning these priors to true metric scale typically relies on rigid affine assumptions in predefined coordinate systems, leaving systematic calibration errors. Linearity in depth calibration depends on the response coordinate. We introduce adaptive response geometry, which makes the fixed choice of depth, log depth, or disparity an image-level unknown. A continuous response family unifies these coordinates and defines an explicit depth-dependent gain. We derive the response-gradient relation and estimate the response parameters in metric space. Hard-Dirichlet residual reconstruction completes the calibrated prior. Under deliberately incomplete metric observations, the training-free pipeline achieves macro AbsRel 0.0301 and macro NMed 14.04°, improving both aggregate measures over PriorDA, LDCM, and Any2Full. Linearity diagnostics examine how the selected response changes the depth relation and its metric error.

[CV-247] From Sharp Eyes to Expert Mind: Internalizing Expert Knowledge in MLLM s for Tampered Text Detection

链接: https://arxiv.org/abs/2609.36145
作者: Kaiqing Lin,Songze Li,Shen Chen,Yunfei Guo,Xiaoye Qiu,Haodong Li,Taiping Yao,Bo Wang,Youchang Xiao,Bin Li,Shouhong Ding
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Tampered Text Detection (TTD) is essential for safeguarding document authenticity in security-critical workflows. Existing expert models are effective at capturing subtle manipulation traces but often generalize poorly across diverse document domains, while Multimodal Large Language Models (MLLMs) offer stronger semantic understanding and transferability yet remain insensitive to fine-grained forensic artifacts. This complementarity motivates us to investigate how expert forensic perception can be internalized into an MLLM rather than merely accessed through an external module. We identify a fundamental Double Mismatch that hinders this goal: a Spatial Precision Mismatch between coarse visual tokens and tiny tampered regions, and a Perceptual Granularity Mismatch between semantics-oriented pre-training and low-level forensic perception. To address these challenges, we propose Expert Knowledge Internalization (EKI), a progressive two-stage framework that transfers forensic expertise into the MLLM itself. In Stage 1, Text-Focused and Image-Focused strategies establish precise spatial focus on small text regions. In Stage 2, the proposed Forensic-General Representation Alignment (FGRA) loss aligns shallow LLM representations with those of a pre-trained forensic expert, enabling the model to acquire fine-grained artifact perception before such cues are diluted by deeper semantic abstraction. Extensive experiments on multiple in-domain and cross-domain benchmarks demonstrate that EKI achieves state-of-the-art performance and stronger generalization than existing expert-model-based and MLLM-based methods. Moreover, the expert is required only during training, allowing the resulting MLLM to maintain inference efficiency nearly identical to the vanilla model without relying on any external expert at inference.

[CV-248] Xiaomi-OCR-0 Technical Report

链接: https://arxiv.org/abs/2609.36136
作者: Xin Chen,Anan Du,Feng Feng,Pei Fu,Jian Luan,Longwei Xu,Shaojie Zhang,Hang Li,Heng Qu,Cheng Tan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Compact OCR-specific vision-language models achieve strong document parsing performance, but often rely on costly supervision and focus primarily on visual-text reconstruction. We introduce Xiaomi-OCR-0, a unified 0.8B model for document parsing and OCR-centric understanding. We build an approximately 170M-sample OCR-centric corpus using an automated data engine that combines expert consensus, render-based verification, and targeted synthesis. Starting from Qwen3.5-0.8B, our progressive training recipe combines Q-Mask-based text anchoring, continued pretraining, and mixed-task reinforcement learning (Mix-RL). Xiaomi-OCR-0 achieves 95.24 on Real5-OmniDocBench, 96.83 on OmniDocBench v1.6, and 87.94 on Wild-OmniDocBench, while reaching an average score of 83.2 across five OCR-oriented VQA benchmarks. Ablations further show that, with sufficient parsing training, OCR-centric understanding supervision provides additional gains for document parsing. Homepage: this https URL. Subjects: Computer Vision and Pattern Recognition (cs.CV) Cite as: arXiv:2609.36136 [cs.CV] (or arXiv:2609.36136v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2609.36136 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[CV-249] Hardware-Aware Functional Kolmogorov-Arnold Networks for Efficient Medical Image Enhancement and Segmentation

链接: https://arxiv.org/abs/2609.36134
作者: Mohammad Sadegh Sirjani
类目: Computer Vision and Pattern Recognition (cs.CV); Hardware Architecture (cs.AR); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Functional Kolmogorov-Arnold Networks (FunKAN) achieve state-of-the-art accuracy on MRI Gibbs artifact removal and anatomical segmentation, but their 11.6 M parameters and 8.7 GFLOPs are too large for edge medical devices. We present FunKANLite, a two-stage, hardware-aware compression of FunKAN for point-of-care use. FunKANLite-TR reduces the spatial prior and replaces the ResBlock offset predictor with a depthwise-separable block. It has 1.9x fewer parameters than FunKAN and no loss in accuracy. We then distill FunKANLite-TR into FunKANLite-ST, which lowers the Hermite basis rank, factorizes the spatial prior into a low-rank form, and halves the filter widths. FunKANLite-ST has 5.6x fewer parameters and 3.7x fewer GFLOPs than FunKAN. It stays within 1.4 percentage points IoU of FunKAN on BUSI, GlaS, and CVC-ClinicDB, and reaches 33.95 dB PSNR on IXI. On an NVIDIA Jetson Orin Nano and a Raspberry Pi 5, FunKANLite-ST reduces energy per inference by up to 68% and raises throughput by 2.9x.

[CV-250] One Geometry Different Outcomes: Readout-Dependent Effects of the Modality Gap in Vision-Language Models

链接: https://arxiv.org/abs/2609.36101
作者: Aditya Sharma,Divya Saxena
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 9 pages, 3 figures, 3 tables

点击查看摘要

Abstract:Contrastive vision-language models learn shared embedding spaces by aligning matched image-text pairs, yet their representations remain separated by a modality gap. Prior work reports divergent effects of modifying this gap: reducing it can improve zero-shot classification and cross-modal alignment, whereas removing gap-related structure can degrade image-text retrieval. In this paper, we provide a unified geometric explanation for these task-dependent effects. Across CLIP and SigLIP encoders, we find that a single dominant direction captures 94.4-99.9% of the squared norm of the image-text mean separation, revealing that the mean-separation component is approximately rank-one. A decomposition of the similarity score then identifies three task-specific roles. In zero-shot classification, query-side fixed gap-offset subtraction is exactly equivalent to an additive class bias. In standard cross-modal retrieval, projecting out the gap direction and renormalising residuals discards candidate-specific norm information, inducing a multiplicative ranking distortion; a geometry-derived exponent tracks the grid-search optimum (Spearman rho = 0.93) and restores performance in some settings, although the gains transfer unevenly. In mixed-modal retrieval, the gap direction sorts candidates by modality; its removal can improve cross-modal ranking, unlike random or non-gap controls. Residual semantic structure after removal defines the limits of the rank-one account. Together, these results explain why gap modification can improve, degrade, or restore performance across downstream settings. By clarifying when and why gap modification changes model behavior, this account provides a principled basis for selecting gap interventions in similarity-based vision-language systems across evaluated downstream tasks.

[CV-251] AerialDojo-200K: A Large-Scale Benchmark Suite for Open-World Aerial Object-Goal Search

链接: https://arxiv.org/abs/2609.36066
作者: Tongtong Feng,Xin Wang,Haoran Hou,Ren Wang,Weiran Wang,Shaokai Zhu,Ziqi Jia,Hao Wang,Yu-Wei Zhan,Zongyuan Wu,Jinghao Cui,Wenwu Zhu
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Emerging Technologies (cs.ET); Multimedia (cs.MM); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Open-world aerial object-goal search is a foundational yet challenging task, requiring aerial agents to autonomously explore large-scale, unstructured three-dimensional environments and reach target objects specified by semantic descriptions or reference images, rather than following route-specific instructions. However, research in this task remains at a nascent stage and relies on small, environment-specific benchmarks with heterogeneous action spaces and data formats. These limitations hinder large-scale training and cross-benchmark evaluation, constraining the scalability and generalizability of aerial agents. To address this problem, we propose AerialDojo-200K, a large-scale benchmark suite for open-world aerial object-goal search, with 3 times as many scenes and 18.7 times as many task instances as the largest existing benchmark for this task. Specifically, we construct 42 simulation scenes spanning four scene families and 21 scene types, including 18 urban, 12 natural, six infrastructure, and six disaster scenes. To ensure data quality, 12 annotators spent two months manually annotating 109 landmarks, 2099 target objects, and 2099 object anchors across these scenes. We further construct 205,732 task instances, comprising over 100K semantic-goal and over 100K image-goal instances across Base, Standard, and Long-Horizon settings. Each task instance includes a collision-free reference trajectory and corresponding multi-view video recordings. We also develop a unified evaluation framework with a scene partition comprising 21 in-distribution scenes and 21 out-of-distribution scenes. Finally, our evaluation of five open-source and four closed-source multimodal large language models reveals that there is still a long way to go toward achieving general-purpose aerial agents. All can be found at this https URL.

[CV-252] CoDimRecon: Agent ic Reconstruction of Sim-Ready 3D Scenes with Deformable Curves Surfaces and Volumes

链接: https://arxiv.org/abs/2609.36024
作者: Shuzhao Xie,Lelin Wang,Guying Lin,Zhi Wang,Minchen Li
类目: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Robotics (cs.RO)
备注: this https URL

点击查看摘要

Abstract:Reconstructing simulation-ready 3D scenes from real-world observations enables robotics, gaming, and immersive applications, yet existing methods largely assume rigid objects. This leaves an important gap for deformables, whose simulation-ready geometry depends on dimensionality (curves, surfaces, or volumes) and whose behavior may require models beyond elasticity. We present CoDimRecon, an agentic framework that reconstructs editable scenes containing rigid, articulated, and deformable objects from multi-view RGB observations. Scene-level geometric priors ground scale and layout, while object-level generated meshes guide the agent toward detailed, compact geometry; articulated rigid objects are decomposed into movable parts with explicit joints. For deformables, category-wise agent sessions reconstruct curves as centerlines with radii, surfaces as manifold shells with thickness, and volumes as watertight solids for volumetric meshing. Reusable simulator skills initialize compatible physical models and parameters, while agent-guided behavioral tests expose mismatches and trigger targeted revisions of motion, geometry, numerics, or material modeling. On evaluated Replica and ScanNet++ scenes, CoDimRecon achieves competitive compositional reconstruction accuracy while additionally producing deformable assets for rod, shell, and solid simulation. We further demonstrate robot interactions across all three representations, including a controlled paper-folding case in which behavioral testing motivates plastic bending.

[CV-253] Persistence Forcing: Exploiting Feature Specialization in Pixel-Space Diffusion

链接: https://arxiv.org/abs/2609.36014
作者: Chong Wang,Zixuan Fu,Shiqi Huang,Siyuan Yang,Hao Cheng,Bihan Wen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page and code: this https URL

点击查看摘要

Abstract:Pixel-space diffusion Transformers (DiTs) directly operate on high-dimensional visual data, yet their hidden representations typically undergo uniform refinement across depth. Natural images, however, are inherently organized at different levels of granularity. Global structure can often be represented compactly, whereas local textures and fine details require richer representations. Motivated by this, we introduce heterogeneous refinement in pixel-space DiTs, assigning different feature groups distinct refinement budgets across depth. Consequently, an ordered feature specialization emerges: sparsely refined features predominantly encode global visual structure, whereas more frequently refined features increasingly specialize toward localized, high-frequency details. We refer to these two groups as persistent and active features, respectively. Building on this emergent specialization, we introduce Persistence Forcing (PerF), which explicitly exploits this persistent–active feature organization for pixel-space image generation. This enables persistent features to continuously condition actively refined features, allowing stable global information to guide the ongoing refinement of finer visual details. During generative sampling, this interaction further induces a meaningful guidance direction that promotes coherent global structure and naturally complements classifier-free guidance. On ImageNet 256\times256 , PerF-L achieves FID of 1.91 , approaching 1.86 of JiT-H with only half the parameters, while PerF-H further achieves FID of 1.63 and 1.76 on ImageNet 256\times256 and 512\times512 , respectively.

[CV-254] Making Cross-Continental Federated Learning Repeatable with FLIP: a Multi-Application Study

链接: https://arxiv.org/abs/2609.36001
作者: Rafael Garcia-Dias,Alexandre Triay Bagur,Chayanin Tangwiriyasakul,Virginia Fernandez,Parhom Esmaeili,Piyalitt Ittichaiwong,Yang Li,Lawrence Adams,Wason Buncharoen,Martin Chapman,Benjamaporn Chayanond,Sadthavud Chunrod,Tanawat Fongsri,Kass Gibson,Supat Plungprasertkul,Supawit Tangpanithandee,Kanyakorn Veerakanjana,Vicky Goh,Michela Antonelli,Joe Zhang,Kongkiat Kespechara,Sebastien Ourselin,M. Jorge Cardoso
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Distributed, Parallel, and Cluster Computing (cs.DC)
备注: 25 pages, 3 figures, 6 tables. Code and data: this https URL

点击查看摘要

Abstract:Federated learning (FL) in healthcare remains challenging, as the overhead of rebuilding governance guarantees for every collaboration stops most projects at the proof-of-concept stage. Here we present FLIP (Federated Learning Interoperability Platform), an open-source, multi-application platform that makes FL training and evaluation repeatable. FLIP implements common FL workflows as a set of composable services: cohort queries against per-site structured databases, on-demand DICOM retrieval from institutional PACS, per-site project approval, and reusable FL job types. To demonstrate FLIP, we ran two distinct use cases, federated fine-tuning and federated evaluation, on synthetic chest X-ray cohorts across two client nodes based in the United Kingdom (UK) and Thailand. In FLIP, each institution independently approves its participation in each project and operates its own node under local IT security processes. This study makes an operational rather than an algorithmic claim. It does not compare federated with centralised training; for that question, we refer the reader to existing systematic reviews and meta-analyses. The central result is evidence that such platforms enable international FL collaboration and improve repeatability, auditability, and site-specific governance. We also present a comprehensive comparison of existing platforms to help researchers and operators choose the right platform for their use case.

[CV-255] Systematic Multi-Agent Vision-and-Language Navigation: Formulation Benchmark and Method

链接: https://arxiv.org/abs/2609.35965
作者: Yunzhe Xu,Zhe Liu
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注: 39 pages, 18 figures, 16 tables

点击查看摘要

Abstract:Vision-and-Language Navigation (VLN) has largely focused on a single agent following a single instruction, yet many real-world applications require teams of robots to tackle tasks beyond the capabilities of any individual agent. We present Systematic Multi-Agent Vision-and-Language Navigation, providing, to our knowledge, the first systematic formalization of multi-agent VLN as a constrained coordination problem: each mission consists of subtasks carrying dependency and resource constraints (presence locks and holding chains). A verified four-stage crafting pipeline instantiates the task as MAVLN, comprising 11,724 episodes across 145 scenes with teams of up to four agents under three instruction regimes, accompanied by tailored constraint-aware metrics. We further present TRISS, a coordination-ready navigation system coupling an LLM-based subtask scheduler, a shared topological memory that turns each agent’s exploration into team knowledge, and a conflict-aware execution mechanism that realizes simultaneous intentions as collision-free routes. Extensive experiments establish TRISS as a comprehensive baseline and reveal substantial room for improvement across scheduling, planning, and execution, highlighting the challenges of coordinating under MAVLN task constraints. Project page: this https URL.

[CV-256] HEIR: Learning Human-Entity Interactions with Functional Roles KR

链接: https://arxiv.org/abs/2609.35955
作者: Di Wen,Wenhao Guo,Yuedong Tan,Yun Huang,Minheng Wu,Zhihang Chen,Haiwen Sun,Fei Teng,Zhiyuan Gao,Yufeng Zhang,Yuanhao Luo,Jingqi Zhang,Yufan Chen,Junwei Zheng,Ruiping Liu,Jiale Wei,Kailun Yang,Kunyu Peng
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 24 pages, 4 figures. Code and dataset: this https URL

点击查看摘要

Abstract:Understanding human-entity interactions requires recovering each person-action event’s participants, roles, and shared identities. This structure can support embodied agents by clarifying who acts on which entities and how, informing anticipation and coordination in shared environments. Standard HOI metrics score individual links, leaving complete event composition undermeasured. We introduce HEIR (Human-Entity Interactions with Functional Roles), an image benchmark for complete grounded participant-role sets across object, interpersonal, and self-directed interactions. It contains 18,730 images, six roles, 105 actions, and 437 nouns, with shared entities, role changes, and repeated fillers; 51.6% of images contain multiple actors and 62.1% contain multiple actions. HEIR pairs relation AP with complete-set AP and structural evaluation. We also introduce CoRISP (Compositional Role-aware Interaction Set Prediction), which uses shared entity identities to combine role-conditioned evidence and predict normalized participant-role sets. Cardinality and role-multiplicity potentials couple assignments through event size and role composition, with exact per-event normalization. Across 16 baselines, relation and complete-event rankings diverge even after aligning action weights. CoRISP leads the evaluated systems on repeated-role events and shared-participant images in HEIR by 2.87 and 3.82 Set mAP points, respectively. On V-COCO, CoRISP achieves 73.72/76.23 role AP and 61.06/68.59 complete-set AP on two-slot actions under Scenarios 1/2. These results show the value of learning and evaluating event composition alongside individual relations. The code and dataset are publicly available at this https URL.

[CV-257] HERO: Histology Encoder for Robust Representation in Oncology

链接: https://arxiv.org/abs/2609.35943
作者: Zhi Li(1),Eghbal Amidi(1),Yating Cheng(1),Tyson Dawson(1),Gorkem Can Ates(1),Shuzhen Kuang(1),Norsang Lama(1),Md Ashequr Rahman(1),Zhiying Lu(1),Elisabeth K. Kong(1),Milan Radovich(1),David Spetzler(1),Matthew Oberley(1),George W. Sledge(1),Ming Chen(1) ((1) Caris Life Sciences, Irving, TX, United States)
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 22 pages, 4 figures, 13 tables

点击查看摘要

Abstract:Foundation models trained on large pathology image corpora now provide strong, transferable representations for computational pathology. Over the past few years a series of such models has been released, each trained on more slides than the last; on standard classification and segmentation benchmarks, the leading models are now separated by small margins. In clinical use, however, the foundation model is applied to images from hospitals, scanners, and staining protocols outside its training data. Encoders generally embed these acquisition factors alongside biological information, which may introduce downstream errors and hinder safe clinical adoption. A pathology foundation model should therefore be robust to acquisition shift without giving up representation quality, yet robustness is seldom the axis along which models are compared. In this report, we introduce HERO (Histology Encoder for Robust Representation in Oncology), a ViT-G/14 pathology foundation model trained with the DINO and iBOT objectives and refined with high-resolution Gram anchoring on a morphology-balanced corpus of 500 million tiles from approximately 575,000 clinical whole-slide images. Across the evaluated public benchmarks, HERO shows the strongest robustness to center, scanner, and stain variation among the compared state-of-the-art foundation models, performs comparably on tile-level classification, segmentation, and gene-expression prediction, ranks first on average across 39 evaluated slide-level clinical tasks, and, under an equal-weighted framework-level analysis, has the best average rank across the six benchmark frameworks.

[CV-258] he Decision Value of Perception Compute

链接: https://arxiv.org/abs/2609.35910
作者: Hoang Pham Cong,Ho Viet Duc Luong
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Adaptive perception spends extra computation on inputs where perception is expected to improve. When perception feeds a downstream decision system, a better perception output need not produce a better decision. We define the decision value of perception compute as the change in downstream loss from escalating an input from a cheap to an expensive perception mode. Because this value can be negative, the allocation of perception compute should be judged against a budget-constrained decision oracle, with uniform full-fidelity inference as a baseline rather than an upper bound. We introduce DEEP (Decision Evaluation for Escalated Perception), a benchmark that scores pre-escalation allocators against this oracle under selection, latency and energy budgets, charging each allocator for its own computation. With deployed monocular geometry on KITTI and nuScenes, we find that 34–54% of the escalations that change downstream loss make it worse; harmful escalations also occur for the published PDM-Closed planner, evaluated open-loop on nuPlan with real detector outcomes. On nuScenes, perception-level gain frequently disagrees in sign with decision value. This mismatch has practical consequences: choosing among fixed deployable signals by missed-object perception gain rather than by decision value reduces realized test decision gain by 7.4% of the all-cheap loss on average. Learned allocators recover part of the oracle’s value by finding beneficial escalations but select nearly as much harm as random, and once their own computation is charged at a 20% latency budget, only the lightweight routers, at about 3.5% of a full detector pass, still beat random.

[CV-259] CoVLM-Bench: A Real-World Benchmark for Cooperative Driving Question Answering and Planning

链接: https://arxiv.org/abs/2609.35823
作者: Kang Yang,Shuai Liu,Hang Li,Yance Fang,Deying Li,Yongcai Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 32 pages, 11 figures, 13 tables. Main text 9 pages; appendices from page 17

点击查看摘要

Abstract:Vision-language models (VLMs) have made substantial progress in autonomous driving, but their success has primarily been studied in ego-centric scenes. Infrastructure-side observations provide views beyond the ego vehicle’s field of view, yet conventional cooperative-driving systems typically transform them into geometric representations for downstream perception and planning. Directly incorporating these views into VLMs offers an opportunity to improve cooperative scene understanding and trajectory planning. However, question answering and trajectory planning have not been jointly evaluated on the same real-world vehicle-infrastructure scenes. We present CoVLM-Bench, a benchmark for cooperative driving question answering (CDQA) and cooperative planning (CP) on vehicle-infrastructure paired scenes. CoVLM-Bench provides scene-grounded CDQA annotations, three-part rationales as auxiliary supervision, and future trajectory targets derived from recorded ego motion. It contains 2,196 paired frames with 35,136 CDQA annotations, while CP predicts six waypoints over a three-second horizon. The annotations combine model-assisted drafting, record-based computation, and human verification. Built upon CoVLM-Bench, we introduce CoVLM-Drive, a unified VLM baseline that directly uses paired views for both CDQA and CP. Experiments show that CDQA adaptation improves answer accuracy and that CoVLM-Drive reaches a lower FDE than the compared V2X planners; QA initialization and rationale supervision each reduce planning error. Together, CoVLM-Bench and CoVLM-Drive support the training and comparison of VLMs for cooperative scene understanding and planning.

[CV-260] HeadGuard: Selective Head Protection for Low-Bit VLM KV-Cache Quantization

链接: https://arxiv.org/abs/2609.35800
作者: Nenad Banfic
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Low-bit key-value (KV) cache quantization saves storage but can sharply degrade vision-language model (VLM) accuracy. We introduce HeadGuard, a composable head-protection method that augments a base KV-cache quantizer with a fixed high-precision mask. Image-sensitivity and output-sensitivity scores select physical KV heads offline, with approximately 1/8 protected in the main experiments; their image keys and optionally values remain in bfloat16 (BF16), while the base quantizes unprotected image entries. Across eight VLMs, three base quantizers, and eight benchmarks (six discriminative and two generative), HeadGuard recovers a substantial fraction of lost accuracy on weaker quantizers, with the strongest gains for Qwen and InternVL. At 2 bits, the six-task discriminative mean over eight models rises from 0.436 to 0.580 on the weakest base; protection can also improve generated answers and caption fidelity to BF16 outputs. Mean accuracy gains persist across all three quantizers with both tested calibration datasets. Keys-only protection retains substantial recovery at lower modeled storage cost. Evaluated through simulated quantization, HeadGuard offers a composable way to improve low-bit VLM accuracy without replacing the underlying quantizer.

[CV-261] PrototypeNAS: Rapid Design of Deep Neural Networks for Microcontroller Units ECML-PKDD2026

链接: https://arxiv.org/abs/2603.15106
作者: Mark Deutel,Simon Geis,Axel Plinge
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: Accepted at ECML-PKDD 2026. 18 pages, 7 figures, 4 tables. This work was funded by the European Commission as part of the MANOLO project under the Horizon Europe programme Grant Agreement No.101135782

点击查看摘要

Abstract:Enabling efficient deep neural network (DNN) inference on edge devices with different hardware constraints is a challenging task that typically requires DNN architectures to be specialized for each device separately. To avoid the huge manual effort, one can use neural architecture search (NAS). However, many existing NAS methods are resource-intensive and time-consuming because they require the training of many different DNNs from scratch. Furthermore, they do not take the resource constraints of the target system into account. To address these shortcomings, we propose PrototypeNAS, a zero-shot NAS method to accelerate and automate the selection, compression, and specialization of DNNs to different target microcontroller units (MCUs). We propose a novel three-step search method that decouples DNN design and specialization from DNN training for a given target platform. First, we present a novel search space that not only cuts out smaller DNNs from a single large architecture, but instead combines the structural optimization of multiple architecture types, as well as optimization of their pruning and quantization configurations. Second, we explore the use of an ensemble of zero-shot proxies during optimization instead of a single one. Third, we propose the use of Hypervolume subset selection to distill DNN architectures from the Pareto front of the multi-objective optimization that represent the most meaningful tradeoffs between accuracy and FLOPs. We evaluate the effectiveness of PrototypeNAS on 12 different datasets in three different tasks: image classification, time series classification, and object detection. Our results demonstrate that PrototypeNAS is able to identify DNN models within minutes that are small enough to be deployed on off-the-shelf MCUs and still achieve accuracies comparable to the performance of large DNN models.

[CV-262] Quantum Fidelity Landscape-Guided Prior Calibration for Single-Circuit QGAN Image Generation

链接: https://arxiv.org/abs/2609.36702
作者: Xue Yang,Rigui Zhou,Dax Enshan Koh,Siong Thye Goh,Yitao Tang,ShiZheng Jia,Young-Wook Cho,Hongyu Chen
类目: Quantum Physics (quant-ph); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Quantum Generative Adversarial Networks (QGANs) have emerged as representative generative models in the Noisy Intermediate-Scale Quantum (NISQ) era and have attracted increasing attention in quantum machine learning. However, most existing QGAN methods rely on patch-based decomposition strategies, which weaken the global consistency of generated images and increase quantum resource overhead. In this work, we investigate a simpler approach: pixel-level, end-to-end image generation using a single-quantum-circuit QGAN. By analyzing the structural matching relationship between the quantum prior and the target data distribution in Hilbert space, we provide a new theoretical perspective for understanding the training behavior of naive end-to-end QGANs. Specifically, we introduce the Quantum Fidelity Landscape (QFL), defined as the pairwise-fidelity structure induced by an ensemble of quantum states and preserved under shared unitary transformations of the quantum generation process. We show that, under a fixed Lipschitz readout, this invariant imposes a one-sided bound on decoded sample separation, motivating calibration of the prior-induced QFL before adversarial training. To validate this theoretical insight, we propose BasicQGAN, a QGAN framework incorporating quantum prior calibration. Before adversarial optimization, BasicQGAN aligns the prior-induced QFL with the data-induced QFL. Experimental results on small-scale grayscale image datasets show that BasicQGAN achieves stable and effective end-to-end pixel-level image generation while requiring fewer qubits and trainable parameters than representative patch-based quantum generators. Furthermore, experiments with different initial quantum-state ensembles show that QFL-calibrated ensembles achieve better generative performance.

[CV-263] CAMEO: A Class-Activation-Mapped Equitable Overlay Framework for Fair and Robust Deep Learning-based Skin Condition Diagnosis

链接: https://arxiv.org/abs/2609.36400
作者: Youssef Attia,Debasmita Mukherjee
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 21 pages, 9 figures

点击查看摘要

Abstract:Deep learning classifiers for dermoscopic skin lesions often reach high in-distribution accuracy while quietly relying on spurious background cues such as skin tone, device vignetting, and embedded rulers, rather than on lesion morphology. This undermines robustness and fairness across skin tones. This work asks whether Explainable AI (XAI), typically used only to audit a finished model, can instead be repurposed as an active training signal that corrects this shortcut without sacrificing diagnostic accuracy. We introduce CAMEO (Class Activation Mapped Equitable Overlay), a framework that improves skin-lesion classification by selecting stable model explanations and using them to separate lesions from their backgrounds. It then replaces the background with realistic synthetic skin while keeping the lesion unchanged. On HAM10000 and dark-skin ISIC images, CAMEO maintained accuracy while reducing background-driven errors by nearly four times. It also made the model’s attention more consistent when backgrounds changed. Results across multiple tests show that reducing reliance on background information improves robustness, with Fitzpatrick-based backgrounds providing a realistic and interpretable approach. Results show that XAI-guided augmentation can make dermoscopic classifiers measurably more robust and fair at no cost to accuracy. They also clarify that it is the mechanism and not the specific tone palette that matters, and that the lasting contribution of XAI here lies in stability-screened, annotation-free lesion localisation rather than in the robustness number itself.

[CV-264] One-Step Next-Latent Prediction Is Not a World Model

链接: https://arxiv.org/abs/2609.36227
作者: Shitong Wang,Zhongang Cai,Yuzhou Hong
类目: Machine Learning (stat.ML); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 24 pages, 3 figures

点击查看摘要

Abstract:Next-latent prediction fits a map from the current embedding to the next one. LeNEPA carries this objective to time series, replacing the stop-gradient of next-embedding prediction with the isotropy penalty of LeJEPA. A world model is a transition kernel that can be rolled out. The one-step regression identifies a conditional mean, and a mean is a kernel only in special cases. For a linear-Gaussian Markov latent, the mean transition and the innovation covariance are fixed by the one-step problem, and the open-loop squared error at horizon K equals the trace of the sum of the pushed-forward innovation covariances. That error grows with K after the one-step fit is exact. If the conditional mean is nonlinear, composing it is not the multi-step conditional mean. If the observation is a non-injective function of a Markov state, a memoryless one-step map does not determine future observations, while a short window can. An isotropy penalty is a function of the embedding marginal, so its partial derivative in the transition weights is zero. On a scalar autoregression with coefficient 0.9 , the one-step mean squared error is 0.998 and the 16 -step open-loop error is 5.10 . On a hidden rotation, an eight-step window reaches 16 -step error 0.056 , while the current scalar alone reaches 0.778 . Raising the isotropy weight from 0.1 to 10 leaves eight-step latent error inside [0.78,0.85] on three seeds.

人工智能

[AI-0] Skill-Space Shooting for Autonomous Robot Policy Improvement

链接: https://arxiv.org/abs/2609.38178
作者: Zihang Rui,Renhao Wang,Haoxu Huang,Yang Gao
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Robots deployed in the physical world must be able to improve beyond their initial training as they encounter new situations and failures. For this improvement to scale across tasks, it must make effective use of experience without requiring human demonstration of each correction. Recent agentic systems offer a way to reduce this reliance on human effort by using foundation models to autonomously compose learned behaviors to complete tasks. Yet completing tasks this way does not itself teach a task policy to overcome its own failures; that requires turning these behaviors into learnable corrections for the policy. Our insight is that many such corrections are familiar short behaviors, or skills: they recur across tasks and describe actions that foundation models can reason about from a scene. We introduce skill-space shooting, which uses foundation model guidance to explore corrections through these reusable skills and turn successful trials into policy improvement. Real-world experiments show repeated improvement in policies acting autonomously, while skills can also be shared to reduce the teaching needed to improve on new tasks. By making reusable skills a source of corrective supervision, skill-space shooting enables scalable and generalizable policy improvement within and across tasks. Additional results and videos at this https URL.

[AI-1] LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization

链接: https://arxiv.org/abs/2609.38166
作者: Yi Pan,Haocheng Xi,Kan Zhu,Xingyang Li,Yibo Wu,Mayank Mishra,Hongtao Zhang,William X.Zheng,Baris Kasikci,Song Han,Kurt Keutzer,Rishabh Iyer,Ion Stoica
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 17 pages, 11 figures

点击查看摘要

Abstract:Recent LLMs increasingly adopt hybrid designs that replace standard attention with linear attention, such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA). Although they compress the context into a fixed-size recurrent state and substantially reduce the cost of long-context processing, repeatedly reading and updating that state remains a major inference bottleneck. Quantization offers a natural way to reduce this cost, but can significantly degrade model quality, due to the accumulation of rounding errors and the presence of outlier rows and columns in the state. To address these challenges, we propose LeapQuant, a training-free method that achieves near-lossless performance under 8-bit recurrent-state quantization. First, to mitigate error accumulation, we propose per-window quantization, which leaps over a window of tokens and quantizes the state only once at its end. Within a window, outputs are computed from the fixed low-bit state together with high-precision buffered updates. Second, to reduce the error introduced by each quantization, LeapQuant retains the state’s largest outliers as a few high-precision Compensator Tokens, which share the update path of real tokens. We then smooth the remaining residual before quantization to further reduce the error. Comprehensive experiments across the Qwen, Kimi, and GLM model families show that LeapQuant substantially reduces memory and compute costs during inference. With accuracy comparable to the FP32 baseline, it achieves average speedups of 2.05–3.70 \times at the kernel level and 1.47 \times for end-to-end inference on NVIDIA B200, RTX PRO 6000, and RTX 5090 GPUs.

[AI-2] hinking Before Thinking: Scaling Agent ic Inference Through Meta-Reasoning

链接: https://arxiv.org/abs/2609.38147
作者: Paras Dahal,Anton Bakhtin,Taco Cohen,Zhengxing Chen,Carole-Jean Wu,Rob Fergus,Scott Yih,Gabriel Synnaeve,Ruslan Salakhutdinov,Sanjeev Arora,Jason Weston,Anirudh Goyal
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to stop. We introduce agentic meta-reasoning, an inference-time harness that makes these choices an explicit and structured reasoning process. Workers carry out the task-level computation, while a controller consolidates what the run has established, explores next options, assesses what each option is worth under the remaining budget, and dispatches the chosen work with context drawn from persistent memory. Between decisions the controller carries only a compact account of the run rather than replaying its full history. Our baselines span production coding agents and research harnesses, together with a Direct Control Agent using the same workers and compute budget allowance. On ProgramBench, which tests long-horizon agentic capability through program reconstruction, meta-reasoning achieves 71.5% with GPT-5.5 against 58.0% for Codex; with Opus 4.8 it achieves 67.2% against 65.5% for Claude Code. On the other benchmarks, spanning abstract reasoning, multi-domain long-horizon reasoning, and proof generation, it gains between 3.6 and 4.2 points over direct control, averaged across three frontier models. It keeps improving over the tested budget ranges where direct control plateaus, though its overhead can hurt at small budgets. Artifact-graph analysis reveals more reuse of earlier work, higher coverage of correct solutions in most settings, and nonuniform gains in final selection. These results indicate that spending computation on structured control becomes more important as agents scale to longer runs.

[AI-3] Stochastic World Models for Verifying Vision-Based Neural Feedback Systems

链接: https://arxiv.org/abs/2609.38120
作者: I. Samuel Akinwande,Mykel J. Kochenderfer,Clark Barrett
类目: Artificial Intelligence (cs.AI); Systems and Control (eess.SY)
备注:

点击查看摘要

Abstract:Verifying a vision-based neural feedback system requires a model of the observations its controller acts upon. Such a model must capture the variation the sensor produces, while remaining tractable for closed-loop analysis. Generative adversarial networks (GANs) have served as perception surrogates, but they are large, reproduce complex scenes poorly, and are hard to verify. We explore stochastic world models as a richer class of perception surrogates. We train a world model with physically grounded latents, built from operations that standard verifiers bound. It reproduces held-out frames more faithfully than GAN surrogates with up to 130 times as many parameters. To verify these surrogates, we develop a procedure that combines falsification, adaptive refinement, symbolic, and backward analyses. On an emergency braking benchmark with a GAN surrogate, our procedure resolves the entire state space, 38% of which the state-of-the-art verifier left unresolved. On the RGB version of the benchmark, where no verification results have previously been reported, our procedure resolves over 80% of the state space with a world model surrogate.

[AI-4] Do LLM Agents Execute the Plans They Declare? From Planning -Mode Declaration to Pattern-Specific Execution

链接: https://arxiv.org/abs/2609.38108
作者: Subba Reddy Oota,Francisco Herrera,Jordi Cabot Sagrera,Marcos López de Prado,Shadab Khan
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 51 pages, 8 figures

点击查看摘要

Abstract:Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment. However, successful planning requires two distinct capabilities: selecting an appropriate plan for the task and executing it faithfully. Existing planner–executor systems can fail at either stage, while final task success alone cannot distinguish selection from execution failures. We therefore study the Plan Declaration–Execution Gap and introduce Planning-as-Routing, where an LLM declares one of four planning modes: Predefined, Sequential, Hierarchical, or Search, and a deterministic router dispatches the task to the corresponding pattern-specific executor. Across four benchmarks and three LLMs, we find three consistent patterns. First, generic Plan+ReAct often fails to preserve declared planning structure, especially for longer plans: across three benchmarks, only (22)–(45%) of trajectories preserve it, whereas pattern-specific executors enforce the intended structure. Second, planning-mode effectiveness varies across environments and models: Search performs best on ALFWorld, Hierarchical on SWE-bench, and the strongest pattern can vary across models within the same benchmark. Third, the largest gains come from execution: pattern-specific executors improve task success from (0.48) to (0.92) on ALFWorld and from (0.36) to (0.44) on SWE-bench Verified over Plan+ReAct. Current LLMs, however, do not reliably select the strongest mode for each task, although few-shot examples improve selection in some benchmark–model combinations. Overall, reliable agent planning requires both effective mode selection and faithful execution: routing substantially closes the execution gap, while task-specific mode selection remains open.

[AI-5] NeuronEye: Query-Guided Visual Concept Activation for Vision-Language Reasoning

链接: https://arxiv.org/abs/2609.38098
作者: Ruiyu Yan,Bowen Chen,Shaowen Wan,Lin Zhao
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Current vision-language models (VLMs) encode visual information in dense hidden states where object identity, spatial layout, and local attributes are implicitly entangled rather than explicitly disentangled, limiting their ability to isolate and modulate the specific visual evidence required by a given language query. Inspired by sparse population coding and top-down modulation in biological vision, we introduce NeuronEye, a plug-in framework that constructs a sparse, concept-level neuron vocabulary from intermediate VLM representations and selectively activates query-relevant visual concepts during inference. NeuronEye decomposes vision-token states into an overcomplete sparse basis organized by concept-level clusters, uses the language query to activate relevant clusters and localize the patches where selected concepts are expressed, and injects the focused evidence back into vision tokens. A complementary suppression mechanism attenuates dominant perceptual directions to preserve weaker but relevant cues. All operations run in a single forward pass over a frozen VLM backbone. On Qwen2.5-VL-7B, NeuronEye raises CV-Bench overall accuracy by +3.1 with gains of +9.5 on Distance, and improves BLINK Multi-view by +8.3, with similar trends on LLaVA-1.6-7B. These results suggest that sparse neuron vocabularies can serve not only as post-hoc interpretability tools but also as active interfaces for concept-level visual reasoning.

[AI-6] Character Training for Risk-Averse Agents

链接: https://arxiv.org/abs/2609.38093
作者: Arav Dhoot,Punya Syon Pandey,Jamie Johnson,Daniel Tan,Elliott Thornley,David Demitri Africa
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Risk aversion in resources could prevent misaligned AI agents from causing catastrophic harm. Misaligned but risk-averse agents would tend to favor safer strategies like making deals with humans over riskier strategies like rebelling. We train agents to be risk averse through character training, finding that persona traits provide a robust mechanism for instilling risk preferences. To do this, we construct a model constitution describing constant absolute risk aversion (CARA) over an agent’s resources and instill it through on-policy distillation. Despite never seeing the benchmark’s decision format during training, character-trained models are competitive with baselines trained directly on it, and generalise better than them out of distribution on two of our four models. We also modulate different aspects of the constitution, finding that token budget and model choice are the most influential aspect of character training to instill risk aversion. We conclude from these results that character training is a promising and scalable way to instil broad dispositions, which we can use to our advantage in mitigating risk from misaligned AI agents.

[AI-7] Neural topology optimization of ship structures under propulsion machinery vibrations

链接: https://arxiv.org/abs/2609.38089
作者: Shengyu Yan,Muhammad Muztahidul Hakim Zareer,Jasmin Jelovica
类目: Computational Engineering, Finance, and Science (cs.CE); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Numerical Analysis (math.NA); Computational Physics (physics.comp-ph)
备注: 24 pages, 13 figures, 7 tables

点击查看摘要

Abstract:Ship structural vibrations contribute to noise, fatigue, and equipment damage, while dynamic-compliance topology optimization can produce pathological designs near resonance. This study extends neural-reparameterized topology optimization using a convolutional Kolmogorov-Arnold network (KATO) to forced-vibration design with active input power (AIP) as the objective. Applications include a 100 Hz engine-supporting deck panel and an 18 Hz thruster foundation frame. Helmholtz PDE filtering and Heaviside projection control feature sizes and manufacturing tolerance. Across both deck families, all eight optimized layouts reduce AIP relative to size-optimized references and, after finite-depth extrusion, also achieve lower static compliance. For unrestricted, manufacturing-aware, and stress-aware frame variants, KATO matches GCMMA in AIP within 0.5 dB while yielding 22-36x lower static compliance after matched-volume binary re-analysis. In a near-resonant 300 Hz case, both methods reduce initial AIP by more than 32 dB; KATO maintains a connected design, achieves 59x lower binary static compliance, and reduces maximum AIP over 1-500 Hz by 2.7 dB. KATO runs 6.4-10.4x faster than GCMMA for the implemented stress-aware formulations. The results demonstrate neural AIP-driven topology optimization as an efficient approach for designing connected, feature-size-controlled ship structures with improved forced-vibration performance.

[AI-8] Probability is Not Enough: Exploring and Counting Divergent Tokens for Reasoning Uncertainty Quantification in LLM s

链接: https://arxiv.org/abs/2609.38070
作者: Feiyang Li,Shengjing Liu,Qi Zhan,Sijie Cheng,Weiqing Wang,Hongwen Chen,Yuxuan Yang,Wen Wang,Yile Wang,Hui Huang
类目: Artificial Intelligence (cs.AI)
备注: 25 pages, 16 figures, 8 tables. Under peer review

点击查看摘要

Abstract:As the chain-of-thought reasoning capabilities of large language models improve, evaluating and calibrating their reasoning confidence is becoming increasingly important for quantifying the uncertainty of their answers. Current methods for estimating the confidence of large language models are generally based on probabilities of selected key tokens, but the underlying mechanism remains unclear. Our pilot study finds that replacing selected token probabilities with coarse substitutes can also improve calibration, motivating us to further explore effective signals of model confidence. We introduce Divergent Token Confidence (DTC), a framework that estimates confidence by counting tokens at which two models strongly disagree during decoding. DTC identifies these divergent tokens using the Jensen-Shannon divergence between next-token distributions evaluated along the same reasoning trajectory. We find that their count is almost negatively associated with answer accuracy, thereby serving as a simple yet effective signal for uncertainty quantification. DTC supports both white-box and black-box evaluation using auxiliary models, without explicit training and affecting the generation process. Experiments across multiple model families and six mathematical benchmarks demonstrate improved calibration over probability-based and verbalized baselines. Under white-box evaluation, the count-only estimator achieves an average expected calibration error of 13.0%, compared with 32.7%-42.4% for standard full-sequence confidence methods. In black-box settings, it also improves calibration over the original verbalized scores. For example, mean expected calibration error falls from 32.1%-40.2% to 13.7%-16.3% on DeepSeek-V3.2. These findings provide new insights for improving reasoning uncertainty quantification in large language models. The code is released at this https URL.

[AI-9] Jaxolotl: A Unified High-Performance Benchmark Suite for LTL-Based Multi-Task RL

链接: https://arxiv.org/abs/2609.38065
作者: Mathias Jackermeier,Jacques Cloete,Alessandro Abate
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Training agents to follow arbitrary instructions is an important goal of multi-task reinforcement learning (RL). Linear temporal logic (LTL) provides a precise and structured formalism for specifying instructions to agents, and has been successfully adopted for training generalist multi-task policies. However, differences in implementations, task distributions, and evaluation protocols make existing methods difficult to compare, while high computational costs limit the scale and statistical reliability of experiments. We introduce Jaxolotl, a unified high-performance benchmark suite for multi-task LTL-RL to address these concerns. Jaxolotl provides a modular, end-to-end JAX implementation of six representative algorithms and four environments, together with newly curated task suites and a standardised, statistically robust evaluation protocol. By precompiling symbolic task representations into static arrays, Jaxolotl enables fully JIT-compiled training and evaluation, achieving end-to-end speedups of up to 220\times and supporting controlled comparisons at substantially greater experimental scale. We use this framework to systematically evaluate existing approaches, revealing complementary strengths and limitations: general methods capable of non-myopic reasoning struggle as the number of propositions grows, while methods with stronger scaling rely on environment-specific assumptions and suffer from myopia.

[AI-10] UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training NEURIPS2026

链接: https://arxiv.org/abs/2609.38043
作者: Ashish Jain,Armaan Sandhu
类目: Artificial Intelligence (cs.AI)
备注: 8 pages, 4 figures. Accepted to the Agentic AI Benchmarks and Applications for Enterprise Tasks Workshop (AABA4ET) at NeurIPS 2026

点击查看摘要

Abstract:Interactive agent benchmarks and multi-turn reinforcement learning increasingly place a second language model in the role of the user. This simulated user controls what information the agent receives and when, yet current benchmarks score only the agent and do not directly measure whether the user correctly executed its assigned role. We introduce UserProxyBench, an evaluation layer over the tau-bench family, and the User Fidelity Score (UFS), which measures adherence to the benchmark’s private user instructions using task-grounded rubric criteria scored independently of agent success. Holding the agent fixed at GPT-5.5 and varying only the user proxy across 375 enterprise tasks changes mean task reward by 15.2 points, while 24.4% of successful episodes contain a user-specification violation. The dominant failure is premature disclosure: users provide information before it is requested. This behavior has little effect on task reward, yet among successful episodes it causes the agent to make 1.06 fewer tool calls on average, changing the interaction being evaluated while preserving the reward. Finally, across seven proxies we identify an empirical cost-fidelity frontier, enabling practitioners to select the least expensive simulator that satisfies a required fidelity level.

[AI-11] Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation

链接: https://arxiv.org/abs/2609.38024
作者: Jaewon Chu,Ji Soo Lee,Jihwan Park,Dohwan Ko,Jeehye Na,Seunghun Lee,Taehoon Lee,Minseo Yoon,Minseok Joo,Yunyang Xiong,Hyunwoo J. Kim
类目: Artificial Intelligence (cs.AI)
备注: 16 pages

点击查看摘要

Abstract:An agent skill is a reusable, actionable natural-language artifact that guides an agent to perform a task effectively under a given harness. Recent studies have explored the optimization of agent skills, contributing to a growing collection of publicly available skills spanning diverse tasks, domains, and harnesses. Despite millions of publicly shared skills, existing skill optimization methods largely overlook this accumulated knowledge, instead relying solely on expensive agent rollouts to iteratively refine skills for a target task. To address this, we propose \textbfRetrieval-Augmented Skill Optimization (RASO), a framework that leverages an external skill corpus as prior knowledge throughout skill optimization. RASO retrieves relevant knowledge from existing skills and adapts it to the target task and harness via Cross-Harness Adaptation, accounting for mismatches in both domain and harness. RASO comprises two complementary stages: \textbfRetrieval-Augmented Skill Initialization (RASI) constructs a knowledge-grounded initial skill without requiring agent rollouts, while \textbfRetrieval-Augmented Skill Update (RASU) iteratively refines the skill by retrieving external knowledge guided by execution feedback. Across four agent benchmarks and two models, extensive experiments show that RASO consistently outperforms baselines without retrieval-augmented skill initialization and updating.

[AI-12] PE-EK-PINN: Physics Embedding with Evolving Kernel for Scalable Physics-Informed Neural Networks

链接: https://arxiv.org/abs/2609.38023
作者: Huiwen Zhang,Feng Ye,Chu Ma
类目: Artificial Intelligence (cs.AI)
备注: 17 pages, conference submission

点击查看摘要

Abstract:Physics-Informed Neural Networks (PINNs) embed governing equations into deep learning, but enforce them only through loss residuals, leaving highly oscillatory wave behavior to be discovered by optimization. As a result, methods that achieve relative L_2 errors below 10^-3 on standard manufactured Helmholtz benchmarks can fail on practical radiation problems involving singular excitations, absorbing boundaries, and wave fields spanning tens of wavelengths. Architectural physics embedding addresses this limitation by factorizing the field into analytically derived oscillatory kernels and learnable envelopes. However, the kernel dictionary must be manually constructed and scales with the number of elementary units, growing exponentially with the depth of hierarchically structured systems such as antenna arrays and metasurfaces. We propose PE-EK-PINN (Physics Embedded with Evolving Kernels), which treats physics kernels as reusable learned representations rather than fixed analytical inputs. A converged subsystem field is frozen and promoted to an evolved kernel, whose transformed copies are reused to represent higher-level configurations without deriving new governing equations. The resulting hierarchy makes the peak number of active kernels independent of system size and reduces cumulative training cost from O(N) to O(\log N) . Experiments on dipole arrays, composite line-source geometries, and cross arrays demonstrate the dramatic training cost reduction, while achieving a reduced or comparable relative L_2 error. One notable example is PE-EK-PINN solves a 256 -dipole array more than 30 times faster than direct PE-PINN.

[AI-13] HARISSA: Inference-Time Self-Checks for Efficient and Safe Local Language Model Deployment

链接: https://arxiv.org/abs/2609.38006
作者: Kenan Alkiek,Moontae Lee,David Jurgens,V.G.Vinod Vydiswaran
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Running a language model locally offers advantages in privacy, latency, and cost, but local hardware fits only small models, which are less capable than frontier models. The usual remedy for a hard query, escalating it to a cloud model, gives up the privacy and cost advantages of running locally. A deployment that stays local faces two decisions for hard queries instead. First, it can spend more computation on a query, e.g., reasoning before answering, which raises accuracy at a cost in latency, so it must decide which queries are worth the extra computation (efficiency). Second, some queries are beyond the local model, and delivering a wrong answer is worse than deferring the query to a human in the loop, so it must decide which answers are safe to deliver (safety). We show that both decisions can be made from the model’s own hidden states. The prefill state, computed before any token is generated, predicts whether the model will answer correctly, and the answer state, at the end of the generated answer, predicts whether that answer is correct. HARISSA fine-tunes the model so that both states predict correctness, then makes both decisions with one policy that cascades through the ways of answering from cheapest to most expensive, skipping a way the prefill state predicts will fail and deferring the query when the answer it stops with is predicted wrong. On a device running a single model, HARISSA is within one accuracy point of chain-of-thought at 2.7 times lower latency. On a server holding four sizes of one model, HARISSA is more accurate than the FrugalGPT and Self-REF cascades at the same latency, and at the same deferral rate the answer state leaves fewer wrong answers than the standard confidence signals in five of six task and setting pairs.

[AI-14] Diagnosing and Improving Probabilistic Reasoning in Large Language Models

链接: https://arxiv.org/abs/2609.38005
作者: Huaman Sun,Dingcheng Wang,Jason Hartline,Jessica Hullman
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models (LLMs) are increasingly proposed as decision assistants who must reason probabilistically from available evidence under explicit decision costs. We propose a decision-theoretic framework that decomposes LLMs’ decision loss into two components: forming accurate beliefs from provided evidence and translating those beliefs into actions that optimize a provided utility function. Using a synthetic benchmark with known ground truth, we apply the decomposition to characterize probabilistic reasoning in frontier and open-sourced models. We further evaluate whether RL interventions targeting beliefs, decisions, or both improve these components across three domains, whether improvements transfer across components and elicitation formats, and whether decision performance can improve without improvement in belief formation. We find that targeting one component of probabilistic reasoning redistributes decision loss, improving the target without necessarily transferring to others, and that jointly targeting belief formation and decision-making improves both but hinges on matched formats between training and evaluation.

[AI-15] No Scale Left Behind: Multi-Scale Autoencoder with Bi-directional Attention for Time Series Anomaly Detection

链接: https://arxiv.org/abs/2609.38004
作者: Jiaheng Guo,Haochen Zhang,Yu-Chao Huang,Jinhao Duan,Nicholas Konz,Tianlong Chen
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Time series anomaly detection (TSAD) plays a crucial role in healthcare, finance, industrial monitoring, and other sectors. Within and between these settings, anomalies span vastly different temporal scales, from sub-second point spikes to multi-hour drift patterns. However, most existing TSAD methods commit to a single temporal granularity, and multi-scale designs either analyze different scales in isolation or are constrained to a predefined coarse-to-fine hierarchy, both failing to sufficiently capture multi-scale interactions. To resolve this limitation, we propose Multi-Scale Autoencoder with Cross-Scale Attention for TSAD (MSCAD), a simple yet powerful semi-supervised TSAD framework founded on parallel autoencoder branches corresponding to different patch sizes. A stack of symmetric bidirectional cross-scale attention blocks enables every pair of scales to exchange information before reconstruction without allowing any single scale to be privileged. On the comprehensive TSB-AD benchmark (40 datasets, 530 series), MSCAD achieves large performance gains against 50 baselines across multiple metrics, with VUS-PR of 0.57(+9.6%) on the univariate split and 0.47(+9.3%) on the multivariate split compared to the state-of-the-art.

[AI-16] Which Attention Heads are like the Human Head? Not the Ones that Compute

链接: https://arxiv.org/abs/2609.37991
作者: Christopher Pinier,Gustaw Opiełka,Hannes Rosenbusch,Taylor Webb,Michael D. Nunez,Claire E. Stevenson
类目: Artificial Intelligence (cs.AI); Neurons and Cognition (q-bio.NC)
备注: 25 pages, 16 figures, including appendix

点击查看摘要

Abstract:Brain-AI alignment is often interpreted as a sign that model and brain perform similar computations. Whether the aligned units are causally involved in model computation is rarely checked. On an abstract pattern-completion task (AAABAAA \rightarrow B), we compare LLM attention-head representations with human EEG and test how ablating those heads affects task performance. Alignment and causation dissociate: brain-aligned heads contribute to performance, but their removal is substantially less disruptive than removal of heads selected via attribution patching. We compare two head sets that prior interpretability work defines without reference to the brain: concept vectors (CVs), which represent abstract patterns across formats, and function vectors (FVs), selected for their contribution to correct-answer prediction. Brain alignment shows little association with FV scores, while its association with CV scores varies across models. Among brain-aligned heads, we find recurring attention profiles: one emphasizes distinctive elements (novelty heads), the other repeating elements (repetition heads). The novelty family tracks salience and attends to the same elements that humans look at, yet its removal is less damaging than random ablation on average. Repetition heads contribute modestly to performance and are associated with abstract-pattern representation (CVs). Across 17 models spanning 3B-72B parameters, FV-ranked removal is substantially more disruptive than brain-ranked removal. Brain alignment thus captures how the model reads the stimulus, and only faintly captures how it represents the pattern and solves the task.

[AI-17] KV-Kaizen: Learning Context-Adaptive Cache Compression Choices

链接: https://arxiv.org/abs/2609.37988
作者: Joao Monteiro,Louis Béthune,Anastasiia Filippova,Sonia Laguna,David Grangier,Marco Cuturi
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights. This impacts LLM throughput negatively, since decoding is memory-bound and decode cost grows with cache size. Recent work alleviates this bottleneck by discarding the least relevant tokens. Eviction introduces a tension, since a one-off decision to discard content may prove detrimental later. Instead, we focus on alternative choices that can lead to cache compression without evicting tokens. We achieve this by learning a selector that is able to produce, based on context, a per-layer cache configuration towards an overall compression budget. The selector operates along three axes: sharing one cache across layers (depth), caching at fewer bits (precision), or truncating the low-rank latent cache representations (rank). We call the resulting method KV-Kaizen, for the many small per-layer choices it compounds. We observe that these interventions taken independently and uniformly over all layers limit achievable compression because they degrade accuracy. Crucially, composing them locally and adaptively to the context can instead preserve accuracy while achieving large memory savings. At inference, the selector runs once, before pre-fill. In evaluations on instruction following and reasoning tasks, our selectors reach the Pareto frontier of accuracy against cache size, against learning-free and post-hoc baselines. On long-context tasks, KV-Kaizen improves on eviction and can be composed with it, reaching a 32x smaller decode-time cache on a 14B model while preserving accuracy. A 4x cache size reduction incurs no accuracy degradation from 7B parameters up, and a compressed model is more accurate than a smaller uncompressed one with the same cache size. Together, these findings support pre-training large models and compressing them only afterwards.

[AI-18] Dagger: Decoupling-based Model Stealing Attack against Graph Neural Networks

链接: https://arxiv.org/abs/2609.37972
作者: Ying Song,Xiaowei Jia,Balaji Palanisamy
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Under Review

点击查看摘要

Abstract:As Graph Neural Networks (GNNs) are widely deployed as Machine Learning-as-a-Service (MLaaS) APIs, model stealing attacks have emerged as a critical security threat. By querying a victim model’s black-box API, an adversary can construct a functionally equivalent surrogate model, compromising proprietary intellectual property and downstream security. Existing GNN stealing attacks, however, rely on overly permissive assumptions, such as soft-label outputs, large query budgets, full-graph query access, and prior knowledge of victim backbones that rarely hold in real-world deployments. In this work, we formalize a strictly constrained black-box, hard-label and backbone-agnostic threat model for GNN stealing attacks under a tight query budget. Given these realistic restrictions, we identify four fundamental challenges: sparse local structures and isolated nodes that degrade victim label quality, insufficient supervision signals, systematic imbalance with incomplete class coverage, and backbone mismatch. To address these interlocking barriers, we propose Dagger, a novel two-phase decoupling-based attack framework. Specifically, in Phase 1, Dagger pre-trains a surrogate using decoupled information propagation to preserve structural context over sparse local subgraphs while handling isolated nodes, combined with manifold-level node mixup to synthesize continuous supervision signals and smooth decision boundaries. In Phase 2, Dagger freezes the encoder and fine-tunes the classifier head via class-balanced sampling paired with logit adjustment to rectify severe query imbalance without requiring extra victim queries. Extensive experiments across four benchmark graphs and four GNN backbones demonstrate that Dagger consistently outperforms state-of-the-art GNN stealing attacks, achieving up to 18.16% higher fidelity while only utilizing 12.23 \times fewer queries than the strongest baseline.

[AI-19] BrainNet Studio: A Unified Toolkit for Brain Network Construction Intelligent Analysis and Visualization

链接: https://arxiv.org/abs/2609.37956
作者: Xiwei Zeng,Shengrong Li,Yiheng Liu,Chunwei Tian,Daoqiang Zhang,Qi Zhu
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Brain networks characterize structural and functional relationships among brain regions and support research on cognition, brain disorders, and brain-computer interfaces. Their time-varying topology and higher-order spatiotemporal dependencies are not adequately represented by conventional static networks. Existing tools primarily focus on static connectomes and provide limited integration of dynamic network modeling with modern graph and sequence learning methods. We present BrainNet Studio, an integrated toolkit for static and dynamic brain network analysis. It provides a unified workflow encompassing network construction, feature extraction, predictive modeling, candidate biomarker identification, visualization, and assisted interpretation. The toolkit integrates 27 algorithms, including deep learning, graph neural networks, and spatiotemporal sequence models, to support classification and the identification of discriminative brain regions and connections. A large language model generates researcher-verifiable summaries of functional connectivity, structural connectivity, and structure-function coupling at individual and group levels. Within a consistent computational framework, users can configure analytical tasks, compare methods, inspect outputs, and extend functionality without repeatedly assembling application-specific pipelines. BrainNet Studio provides a practical and extensible platform for connectome analysis in cognitive neuroscience, exploratory studies of brain disorders, and brain-computer interfaces. The toolkit is publicly available at this https URL.

[AI-20] opological Coherence for Self-evolving Multi-agent Systems

链接: https://arxiv.org/abs/2609.37953
作者: Sen Zhao,Ruiqi Kong,Zuyu Zhang,Lifeng Shen,Xinyu He,Xu Zhang,Qinghua Zhang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Complex tasks inherently couple workflow structure, agent responsibility, collaboration, and memory access: task regions delimit responsibility and tool scope, cross-region dependencies give rise to handoffs, and ownership boundaries delimit private and selectively shared memory. Existing methods can jointly optimize agent and communication structures, yet such optimization does not by itself require responsibility, handoff, and memory boundaries to remain consistent with task dependencies. We term this requirement topological coherence. We introduce TOCOMAS, a Topology-Coherent Multi-Agent System. TOCOMAS grounds a task graph in tool interfaces, organizes compatible task nodes into reusable responsibility domains, and derives dependency-induced and profile-conditioned collaboration together with boundary-regulated memory visibility. During online self-evolution, TOCOMAS proposes coupled changes to agent, collaboration, and memory policies, retaining for subsequent tasks only candidates that satisfy structural constraints and improve evaluated reward. Across BBEH, WorkBench, SWE-Bench-Verified, and CoMemBench, TOCOMAS improves task success over baselines across backbones. CoMemBench also shows gains over the self-evolving baseline in verified progress, handoffs, and memory isolation.

[AI-21] GRFBrain: Graph-Structured Rectified Flows for EEG Dynamic Modeling

链接: https://arxiv.org/abs/2609.37934
作者: Haohui Jia,Zheng Chen,Jathurshan Pradeepkumar,Xu Cao,Yasuko Matsubara,Yasushi Sakurai,Takashi Matsubara
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Forecasting time-varying functional connectivity from electroencephalography (EEG) requires modeling both history-dependent trends and structured variability across channels. Conditional flow matching provides a framework for distributional forecasting, yet it remains unclear whether graph-informed source distributions offer practical advantages over isotropic noise and strong deterministic predictors. We introduce a graph-structured residual flow framework that separates conditional mean prediction from stochastic residual transport. A history-only predictor estimates the future connectivity graph, while a graph Gaussian source encodes dependencies derived from past connectivity through a Laplacian-based covariance. A conditional velocity field transports source samples to future graph residuals, with transport time explicitly distinguished from physical EEG time. Our study identifies the conditions and controls needed to distinguish useful residual transport from improvements attributable to deterministic prediction, learned representations, and sampling effects.

[AI-22] RLX: A Unified Multi-Backend Tensor Compiler and Distributed Runtime in Rust

链接: https://arxiv.org/abs/2609.37916
作者: Eugene Hauptmann,Nataliya Kosmyna
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI)
备注: 6 pages, 4 figures, peer-reviewed and presented at 2026 IEEE High Performance Extreme Computing Conference (HPEC)

点击查看摘要

Abstract:Production machine learning (ML) stacks often split graph compilation and kernel execution across different layers and languages, making backend behavior, deployment guarantees, and performance fallbacks hard to reason about end-to-end. RLX addresses this gap with a single Rust codebase that combines compiler and runtime roles around one primitive-level, three-level intermediate representation (IR), plus a transparent dispatch contract that resolves each operator to native, common-IR, or rewritten lowering and fails compilation when legalization is not possible. The same IR targets fourteen runtime devices (cpu, metal, mlx, ane, cuda, rocm, oneapi, tpu, hexagon, gpu, vulkan, opengl, directx, webgpu) and two specialty codegen paths (Cortex-M INT8 and FPGA), ingests safetensors, GGUF, ONNX, and rten formats, supports F16/BF16/F64/C64 and quantized INT4/INT8 flows with AMP/PTQ/QAT, and scales via tensor-/pipeline-parallel collectives over TCP and RDMA transports. Beyond neural workloads, RLX also extends to scientific/physics-style domains through sparse and dense linear algebra extensions (e.g., CSR LU/CG/matvec and LAPACK- backed factorizations) and 3D Gaussian splatting operators. We evaluate RLX against PyTorch, TensorFlow, JAX, candle, burn, tch, rten, MLX, CoreML, IREE, Glow, TensorRT, and tinygrad under identical input generation and p50 measurement methodology on one host. On all-MiniLM-L6-v2, RLX-Metal is fastest at every batch (e.g., 16.6 ms at batch 32 vs. PyTorch-MPS 26.7 ms). In the MNIST training table, RLX also has the top-throughput entry (graph-fused MLP: 946,487 img/s), above NumPy+BLAS (787,349 img/s), while retaining 100% top-1 parity on reference checks (e.g., Qwen3).

[AI-23] Beyond Interaction Capacity: Estimator Scaling with Recursive Models for CTR Prediction

链接: https://arxiv.org/abs/2609.37905
作者: Shivang Chopra,Fotis Iliopoulos,Zsolt Kira,Gaurav Menghani
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Click-Through Rate prediction, a core task in recommendation and advertising systems, relies on modeling interactions among sparse categorical features. Explicit cross networks are a central paradigm for CTR prediction, and recent progress has largely come from increasing the interaction capacity of a single predictor through deeper cross networks and more expressive cross operators. We revisit whether continually increasing interaction capacity remains the most effective way to improve predictive performance, and find that its benefits quickly exhibit diminishing returns even as capacity continues to grow. This motivates a complementary scaling direction that we call estimator scaling, where additional resources are used to incorporate multiple related estimators rather than only enlarging a single predictor. Through theoretical analysis, we show that the gains from estimator scaling are governed by the amount of non-shared predictive variation available across estimators. However, exploiting this variation naively can be expensive: independently trained models provide substantial estimator diversity but require deployment cost to grow with ensemble size. This motivates a parameter-efficient realization of estimator scaling that can incorporate diversity from multiple estimator sources without maintaining multiple full models. Building on this view, we introduce RECursive Averaged Predictor (RECAP), a parameter-efficient recursive CTR model that operationalizes estimator scaling at three levels: distillation across independently trained models, exponential moving averaging over training trajectories, and aggregation over inference-time routes within a weight-shared recursive backbone. Experiments across multiple benchmarks establish new state-of-the-art predictive performance on standard benchmarks, while placing the RECAP on a favorable performance-parameter Pareto frontier.

[AI-24] You Cannot Pick a Provider From the Price List: Market-Aware Routing for Open-Weight LLM Inference

链接: https://arxiv.org/abs/2609.37902
作者: Liang He,Jingbo Wen,Yixiong Chen,Yue Yang,Qizhen Lan,Kangning Cui,Xilu Wang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Existing LLM routers choose among models using static per-model costs. We show that open-weight inference markets introduce a second, largely ignored decision axis: after choosing a model, a client must still choose which provider serves it. Measuring live endpoints across [nummodels] open models, competing providers, multiple task types, and three measurement waves, we find that provider choice cannot be inferred from the price list. The same model can vary sharply in quality, latency, availability, and price across providers; higher-priced providers are consistently faster, but price does not reliably predict quality or availability; and provider feasibility is task-selective, with one deployment nearly normal on knowledge tasks but catastrophically degraded on multi-step reasoning. We formulate same-model provider selection as a price-taker market-aware routing problem. A simple measured-map policy routes to the cheapest provider that is both quality-equivalent and healthy, yielding matched-quality savings while avoiding degraded endpoints. Because the map drifts, we introduce FACET, an online provider router that certifies per-(provider x task) feasibility facets and fails safe to an anchor before serving uncertified arms. Across relaxed deployment assumptions, FACET tolerates imperfect task assignment and sparse feedback, while systematic evaluator bias exposes a quality-signal trust boundary that can be mitigated with ground-truth probes or audits. Live provider runs further confirm that certification can move real traffic from a premium anchor to a substantially cheaper certified endpoint. Our results suggest that market-aware LLM routing must measure not only which model to use, but also who serves it.

[AI-25] Guide Then Let Go: Gap-Adaptive Teacher Scheduling for Sparse-Reward Agent ic RL

链接: https://arxiv.org/abs/2609.37898
作者: Youling Huang,Tiankuo Xu,Jiaji Liu,Tong Zheng,Shuo Zhou,Shaotong Qi,Junchi Yao,Shiyang Liu,Hao Xu,Pengcheng Xu,Bo Huang,Hongyi Fu,Lin Lin
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Reinforcement learning for long-horizon agents typically relies on sparse outcome-based rewards. This leads to a severe cold-start problem, as early-stage policies often fail to solve sampled tasks, leaving little useful reward signal for learning. To mitigate this problem, we use on-policy distillation (OPD) to provide token-level guidance on the student’s own rollouts. We find that the benefit of this guidance depends on the performance gap between the teacher and the student. When the teacher substantially outperforms the student, distillation helps guide the student through the early training stage where outcome rewards provide little learning signal. As the gap narrows and eventually reverses, however, continued distillation becomes less beneficial and may hinder further improvement. Motivated by this observation, we propose Gap-Adaptive Teacher Scheduling (GATS), which augments the student’s RL objective with an OPD term whose weight adapts to the teacher-student performance gap. Specifically, GATS gradually reduces teacher guidance as the student approaches the teacher’s reference performance and withdraws it once that reference is reached. This enables GATS to leverage task-trained teachers smaller than the student, since teacher guidance is primarily needed during early training. Across ALFWorld, WebShop, and ScienceWorld with three Qwen2.5 teacher-student configurations, GATS achieves the highest average success rate among the compared methods in all three configurations, improving over reward-only GRPO by 4.37%-11.87% under matched student rollout budgets. Code is available at this https URL.

[AI-26] Co-PiLOT: Constrained Physics-Informed Latent Optimization for Target-Driven Inverse Design

链接: https://arxiv.org/abs/2609.37875
作者: Mahish K. Guru,Mayank Nagar,Ayush vyas,Jan Bohlen,Roland Aydin,Noomane Ben Khalifa
类目: Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE)
备注:

点击查看摘要

Abstract:Inverse design of physical systems (molecules, devices, microstructures) often reduces to optimizing a high-dimensional structure against an expensive black-box simulator. Direct search is difficult because the space is non-Euclidean, feasibility is hard to encode, and each evaluation is expensive. We present Co-PiLOT, a latent optimization approach that maps candidates through a generative encoder-decoder, uses the decoder as a learned validity prior, and searches the latent space with physics-informed black-box optimization. The framework is applied on the inverse design of magnesium alloy microstructure/texture. We develop a vision transformer based-encoder; paired with latent diffusion, diffusion transformer and rectified-flow transformer-based decoders on \sim80,000 EBSD-derived microstructure dataset to learn a minimal bottleneck, z . The ViT-FMDiT model ( z = 768 ) reconstructs high-fidelity microstructure images (FID 27.86 , MS-SSIM 0.178 ), which our self-segmenting orientation codec converts into input grids for crystal plasticity solver. Finally, we introduce MERIDIAN, an active latent optimizer driven by deep-kernel Gaussian-process uncertainty, failure-aware feasibility prediction, manifold-aware trust regions, and target-aware acquisition. Within a budget of 160 simulations, the ViT-FMDiT and MERIDIAN combination yields the best target-driven objective score, reducing the relative target error by 3 – 22% against seven baselines (DANTE, TuRBO, BAxUS, CMA-ES, DDOM, SEIKO, DDPO) on the same decoder.

[AI-27] ExceptionDrive: A Planning -Oriented Counterfactual Corner-Case Benchmark for Autonomous Driving

链接: https://arxiv.org/abs/2609.37871
作者: Ziyi Luo,Zhe Sun,Yehao Lu,Lei Zhou,Lisheng Wu,Xuewei Li,Zequn Qin,Xi Li
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: 13 pages, 5 figures, 3 tables; supplementary material included

点击查看摘要

Abstract:Average performance on routine driving benchmarks does not establish planner reliability under rare, safety-critical hazards. We proposed ExceptionDrive, a counterfactual planning benchmark that uses VLM-assisted screening, localized multi-view editing, and quality auditing to insert hazards into real nuScenes scenes while preserving their context. Its 21 tasks span six safety families and define hazard or conflict regions, local safety constraints, and acceptable responses. Because hazard insertion can invalidate the recorded human trajectory, our reference-free protocol evaluates edited predictions using Unsafe Rate (UR), Hazard Clearance Compliance (HCC), Hazard Proximity Response (HPR), and Counterfactual Trajectory Shift (CTS), which measure core-region intrusion, clearance compliance, clearance relative to a prescribed margin, and counterfactual trajectory change. Seven representative planners frequently intrude into hazard regions or provide insufficient clearance. We also develop a Reminder Agent that, without sample-specific task labels, converts visual evidence and the shared taxonomy into structured records of hazard presence, type, and a recommended high-level strategy. The agent neither predicts trajectories nor controls the vehicle; its records guide a VLM-based decision agent. In zero-shot experiments, the reminders improve strategy accuracy and reduce under-warning.

[AI-28] Agent Bug-Smith: Automatically Reproducing Real-World Harness Bugs in Agent ic Systems

链接: https://arxiv.org/abs/2609.37864
作者: Yiming Cheng(1),Alfin Wijaya Rahardja(2),Mengshi Zhang(3),Zihao Chen(3),Zhenpeng Chen(4),Yiling Lou(5) ((1) The University of Chicago, (2) Fudan University, (3) TensorBlock, Inc., (4) Tsinghua University, (5) University of Illinois Urbana-Champaign)
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注: 20 pages, 8 figures. Code: this https URL Data: this https URL

点击查看摘要

Abstract:Agent harness bugs exhibit unique characteristics and remain challenging for state-of-the-art software agents to repair. Progress in this area is further hindered by existing benchmarks, which contain only a small and fixed number of executable harness bugs while requiring hundreds of human hours to construct. This work presents AgentBug-Smith, an automated harness bug reproduction approach that continuously discovers and reproduces real-world harness bugs from open-source agentic systems. Across different backbone LLMs, AgentBug-Smith consistently outperforms existing bug reproduction techniques designed for general software, achieving 10.67% - 27.56% higher success rates of reproducing harness bugs. By applying AgentBug-Smith to open-source agentic systems in the wild, we construct Live-Harness-Bench, a live and extensible benchmark that currently contains 200 reproducible harness bugs. We further demonstrate the utility of Live-Harness-Bench through two downstream applications. First, we use Live-Harness-Bench as the evaluation benchmark to systematically evaluate state-of-the-art software agents, revealing their limited capabilities in repairing real-world harness bugs. Second, we use Live-Harness-Bench as a knowledge base of real-world harness bug fixes, from which reusable repair skills can be distilled to improve existing software agents, increasing their harness-bug repair rates by 6.32%. Together, AgentBug-Smith and Live-Harness-Bench establish a scalable foundation for continuously evaluating and improving software agents on harness bug repair, turning real-world agent failures into executable evaluation instances and reusable knowledge for harness improvement, thus contributing to the ultimate goal of recursively self-improving agents.

[AI-29] Active Budget Can Kill Sensitivity: Diagnosing and Repairing TopK Sparse Autoencoder Reliability

链接: https://arxiv.org/abs/2609.37857
作者: Zhenting Huang,Junnan Liu,Qianren Mao,Zhixing Tan,Bo Jiang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Sparse autoencoders (SAEs) are increasingly scaled to wider dictionaries to recover fine-grained structure from large language model activations. However, a feature is useful for interpretation only if it remains a stable unit of analysis when the same meaning is expressed in different surface forms. We study this reliability question for TopK SAEs via feature sensitivity. Experiments demonstrate that scaling selectively reduces the sensitivity of rare features, while common features remain comparatively stable. A controlled width(\times k) factorial experiment identifies the active budget k as the root cause: the degradation arises from the selection boundary rather than dictionary width alone. We attribute this failure to the geometry of TopK selection. The active margin, the distance to the cutoff, predicts feature loss without thresholds. Guided by this margin diagnosis, we introduce pairwise rank stabilization. Our method targets ordering failures at the cutoff and improves rare-feature sensitivity by (8.83) percentage points, while keeping reconstruction and alive-feature coverage near the baseline. Overall, our results suggest that wide TopK SAEs should be evaluated not only by reconstruction, sparsity, and feature count, but also by feature reliability under semantic variation and boundary geometry for stable interpretability.

[AI-30] Is manual software optimization a thing of the past?

链接: https://arxiv.org/abs/2609.37849
作者: Pavlin G. Poličar,Martin Špendl,Tomaž Hočevar
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Scientific software is increasingly required to process larger datasets while maintaining acceptable execution times. Software optimization traditionally requires substantial expertise in programming, algorithms, and numerical methods. Recent advances in large language models (LLMs) offer the possibility of automating much of this process. We investigate whether LLM-based agents can autonomously achieve substantial performance improvements in scientific software, including mature implementations that have already been extensively optimized by human developers. We tasked an LLM-based agent with optimizing software for three computational problems: t-SNE, single-sample gene set enrichment analysis (ssGSEA), and graphlet counting. Humans defined the scope, correctness criteria, and a verification mechanism, after which the agent worked autonomously, in some cases for several hours. Code maintainers reviewed each resulting implementation and verified its correctness. The optimized implementations were faster in all tested configurations, by up to two orders of magnitude over the fastest existing tools. The improvements included low-level code optimizations, mathematical reformulations, and an entirely new algorithm for graphlet counting. Software optimization can increasingly be delegated to autonomous agents, with the human role shifting from implementing optimizations to deciding which software to optimize, defining objectives, providing verification mechanisms, and ensuring the correctness of the final software. For well-scoped, verifiable problems, we argue that manual software optimization may be a thing of the past.

[AI-31] Scaling Influence Functions in LLM s through Eigenbasis-Corrected One-Bit Gradient Projection

链接: https://arxiv.org/abs/2609.37842
作者: Jaeseung Heo,J Rosser,Dongwoo Kim
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Influence functions estimate how individual training examples affect the behavior of large language models (LLMs). Analyzing how training data influence different behaviors of an LLM involves repeated influence computation. Reusing stored training gradients reduces the computational cost, but storing full gradients is prohibitively expensive at LLM scale. We study how to compress these gradients while preserving influence estimates for future queries that are unknown at storage time. Through a worst-case analysis, we characterize the optimal fixed-dimensional linear representation and propose eigenbasis-corrected one-bit gradient projection (EOGP) to approximate it at scale. Specifically, EOGP uses EK-FAC to reduce gradient dimensionality, then applies PCA within the retained subspace to learn compression directions from the training gradients. We then apply one-bit quantization to the resulting coordinates, allowing more coordinates to be retained within a fixed storage budget. On GPT-2, EOGP predicts retraining outcomes more accurately than the evaluated compression baselines while using one-sixteenth of their per-example storage. On OLMo 2 SFT models from 1B to 32B parameters, EOGP remains competitive with the baselines allocated over 100 times as much storage per example.

[AI-32] Mixture of Self-Improving Branches For Agent Harness Optimization

链接: https://arxiv.org/abs/2609.37834
作者: Haoyu Dong,Yuhang Zhou,Zihao Lin,Yifan Wu,Bo Peng,Mingyi Wang,Xiangjun Fan,Lizhu Zhang,Zhuokai Zhao
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Harness optimization provides a practical setting for recursive self-improvement (RSI), where agent-generated modifications inform subsequent changes through execution feedback. Recent work such as Meta-Harness implements this process through iterative code generation and evaluation, but retains a fixed development set and proposal policy. These constraints channel evolution along a single search trajectory, increasing the risk of converging to a local optimum. We make the improvement process itself adaptive by organizing search into branches with evolving development subsets and proposal policies. Each branch retains development cases solved by more of its leading harnesses than by those of other branches, drops cases solved by every leading harness across all branches, and revises its proposal policy using its own search history. To deploy the resulting complementary harnesses, we propose a router to select one development-selected branch head for each new input before execution. Across mathematical reasoning and agentic coding benchmarks, our system achieves relative improvements over Meta-Harness of 34.8% on Olympiad-level mathematical reasoning, 11.6% on Terminal-Bench 2.0, and 3.8% on SWE-bench Lite, with harness selection and router configuration based solely on development data. These results show that evolving branch objectives and proposal policies can yield complementary harnesses whose strengths a router combines without access to test outcomes.

[AI-33] DIET: Deletion-response Expert Trimming for Video Diffusion Transformers

链接: https://arxiv.org/abs/2609.37829
作者: Jiachang Zhang,Teng Hu,Bohao Feng,Songhang Shen,Wenqiang Wang,Hongqian Deng,Ran Yi
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Video diffusion transformers (DiTs) increasingly adopt mixture-of-experts (MoE) architectures to reduce active computation, but their full expert storage remains costly. Existing one-shot pruning criteria mainly rely on static activation or routing statistics and cannot capture layer-level re-routing after expert deletion. We introduce DIET, a training-free expert pruning framework based on deletion responses. A single all-expert calibration pass records expert outputs and router states for matched conditional and unconditional tokens. Candidate deletions are then replayed from cached tensors, requiring no additional model forward passes. The resulting deletion-response signatures characterize each expert by the changes induced when it is removed. DIET selects retained experts by minimizing Overall Diversity Loss (ODL), which preserves directional coverage in signature space, and combines intra-layer local search with an inter-layer regression-guided budget search to allocate experts across layers. On LingBot-Video 30B-A3B, pruning 50% of experts (6,144 to 3,072) reduces the checkpoint from 57 GB to 30 GB and enables single-card deployment on a 48 GB GPU without fine-tuning. Under a fixed 284-case VBench protocol, the VBench Total increases from 0.7941 to 0.8115. Across tested retention budgets, DIET consistently outperforms competitive pruning baselines adapted from large language models.

[AI-34] Privy to the Foil: Recasting Value Estimation with a Self-Privileged Critic for RLVR

链接: https://arxiv.org/abs/2609.37825
作者: Kun Liang,Chenming Tang,Clive Bai,Weijie Liu,Zeyuan Liu,Qingyang Zhang,Saiyong Yang,Yunfang Wu
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Assigning credit to intermediate steps remains a central challenge in training Large Language Models (LLMs) on multi-step reasoning tasks with sparse terminal rewards, and actor-critic methods such as PPO address this by learning value functions to construct token-level advantages. Their effectiveness, however, hinges on reliable value estimation, a difficult task requiring the critic to both assess progress toward a correct solution and anticipate an evolving policy’s future behavior; errors in either can compromise credit assignment and destabilize online training. In this paper, we revisit the standard state-only formulation of value estimation and propose \pi PPO, a self-privileged actor-critic framework. By reusing verified same-prompt rollouts as contrastive evidence, \pi PPO helps the critic assess intermediate reasoning against successful and failed attempts, while preserving standard policy optimization and the deployment interface. Experiments show that \pi PPO consistently improves value-estimation quality by a substantial margin and outperforms representative actor-critic and critic-free RLVR baselines on challenging mathematical reasoning benchmarks, while remaining effective even when paired with substantially smaller asymmetric critics.

[AI-35] Making Duplicate Reimbursement Unrepresentable: A Verified Ethereum E-Invoice System for Humans and AI Agents

链接: https://arxiv.org/abs/2609.37819
作者: Jia Cai
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE)
备注:

点击查看摘要

Abstract:Electronic invoices are replacing paper invoices worldwide, but today’s centralized architectures leave three problems unsolved on the consumption side: an invoice can be submitted for reimbursement repeatedly, authenticity is difficult for recipients to verify, and data is siloed at a central authority that forms both a performance bottleneck and a single point of failure. This paper presents the design, formal analysis, and implementation of a complete blockchain-based electronic invoice system on Ethereum. We formalize the invoice lifecycle as a guarded labeled transition system and prove, under standard cryptographic and consensus assumptions, that the system guarantees: (i) reimbursement uniqueness–an invoice is reimbursed at most once, even across mutually distrusting organizations; (ii) face integrity–any verified invoice matches the recorded one unless keccak256 second-preimage resistance is broken; and (iii) authorization soundness for every lifecycle operation. The core invariants are machine-checked using Solidity SMTChecker, proving inductive validity across all reachable transaction sequences. The architecture models each invoice as a non-fungible, non-tradable token whose state transitions through five guarded subsystems, employing a lock-based protocol that makes duplicate reimbursement unrepresentable rather than merely detectable. We implement the design as a Solidity 0.8 contract with a four-role web application and evaluate it on a private Ethereum network: issuing costs 646,773 gas, full reimbursement costs under 135,000 gas, all operations run in O(1) time, and a single node sustains 137 issuances/s. Finally, the verified contract serves as a safety envelope for LLM-based reimbursement agents, provably rejecting unsafe actions (duplicate, over-limit, or forged-receipt claims) even when the agent’s internal policy fails. All code and benchmarks are open-source.

[AI-36] Explore Execute Evolve: A Skill Acquisition and Reuse Loop for Embodied Agents

链接: https://arxiv.org/abs/2609.37810
作者: Sicheng Xie,Yitong Chen,Haidong Cao,Shunlin Lu,Zuxuan Wu,Yu-Gang Jiang
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains challenging. More recently, general-purpose multimodal agents have shown great potential for zero-shot robotic task solving. However, they often incur high execution costs by reasoning and exploring the physical world from scratch. To reduce these costs, we introduce RoboSkill, a framework that connects skill acquisition and reuse through an Explore, Execute, Evolve loop. Within this loop, the agent explores to gather task-relevant information, executes tasks while adapting to feedback, and evolves its skill library based on execution records. It then reuses these skills to guide exploration and execution in the next cycle, closing the loop. To improve loop efficiency, we complement vision with tactile feedback to reduce uncertainty during physical interaction. We further augment textual guidance with reusable code to reduce reasoning overhead during skill reuse. On LIBERO-10, RoboSkill improves first-episode success rates by 12.5–25.0 percentage points and reduces average runtime by 7.6–72.4% across four agents. On real robots, it improves success rates by 8.3 percentage points and reduces average runtime for successful trials by at least 14.4%.

[AI-37] Challenges and Solutions for Bandits in the Wild: Warm-Started Mixture Bandits for Cross-Cohort Slate Recommendation

链接: https://arxiv.org/abs/2609.37800
作者: Serafima Lebedeva,Sumantrak Mukherjee,Ali Arshad Sadal,Ilias Ekşi,Rahul Sharma,Julia Mueller,Theresa Dombrowski,Jakob Karolus,Viktor Bengs,Eyke Hüllermeier,Sebastian Vollmer
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 11 pages, 3 figures, preprint

点击查看摘要

Abstract:Many recommender services repeatedly encounter cold-start cohorts, where new users arrive with little or no interaction history. This creates two challenges: learning user preferences quickly from limited feedback and sustaining useful recommendations when each user has a finite catalog that can become repetitive or depleted over time. We propose CohortMix-TS, a warm-started mixture bandit that learns latent user groups from earlier cohorts and uses available metadata to construct group-informed priors for new users. Starting from these fixed priors, the model personalizes independently as feedback from each user becomes available. Session slates combine Thompson sampling with diversity and inventory-depletion controls. We evaluate CohortMix-TS through simulation, semi-synthetic experiments, and a 25-day randomized in-the-wild deployment with 713 registered participants in a Campus Games quiz application. Our evaluations show that cross-cohort transfer improves early recommendation quality and user-level regret, while inventory-aware slate construction helps prevent premature exhaustion of preferred items. In the field deployment, treatment users also showed a larger early-to-late change in correctness than users receiving random recommendations. Together, these results show how warm-start transfer and inventory-aware recommendations can support personalization for short-lived, repeatedly cold-starting cohorts.

[AI-38] A neural network that maintains and retrieves memories based on context

链接: https://arxiv.org/abs/2609.37791
作者: Hayoung Song,JeongJun Park,Qihong Lu,Giacomo Vedovati,Monica D. Rosenberg,Zachariah M. Reagh,ShiNung Ching
类目: Artificial Intelligence (cs.AI); Neural and Evolutionary Computing (cs.NE)
备注:

点击查看摘要

Abstract:Every day, people continuously infer situational context and adjust the way they understand and remember the world. Context, signaled by the prefrontal cortex, is known to modulate working memory and episodic memory, but the algorithmic understanding of this modulation remains limited. Here, we train a recurrent neural network (RNN), augmented with an episodic memory buffer, to infer context using Bayesian inference as it continuously makes predictions of upcoming scenes while watching naturalistic movies. When the inferred context modulates the RNN’s recurrent connectivity (the basis of working memory) in a low-rank manner, the model’s activity patterns best match neural responses in human participants who watched the same movies during fMRI. Context also modulates episodic memory retrieval, such that the model retrieves memories based on not only content similarity but also context similarity. This is implemented as a key-value system with self-attention, designed to additionally encode context and retrieve context-congruent memories. The resulting model not only better resembles human brain representations but also learns to retrieve memories like humans much faster than a model without context modulation. Together, our findings suggest a computational mechanism by which context modulates information maintenance and long-term memory retrieval in naturalistic environments.

[AI-39] Predictive Self-Supervised Learning Provably Identifies Stochastic Signals under Nuisance

链接: https://arxiv.org/abs/2609.37789
作者: Fabian A. Mikulasch,Friedemann Zenke
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Self-supervised learning (SSL) by predicting in latent space, without generating the input data itself, learns highly abstract, useful representations. Intuitively, this success is often attributed to its ability to discard nuisance information that is irrelevant to prediction. However, this poses a conundrum: both stochastic variation in a prediction-relevant latent signal and true nuisance make observations partly unpredictable; how could they be distinguished? Surprisingly, we prove that common SSL methods can achieve exactly this, by implicitly instantiating a latent-variable model with stochastic dynamics and observation-private nuisance. We trace their ability to recover the stochastic signal to two complementary principles: Predictive mutual information maximization ensures that representations retain the information needed for prediction, while latent distribution matching constrains how this information is encoded, thereby making the retained signal identifiable. We confirm this identifiability result in simulations for Gaussian predictors, which recover the true signal up to an affine transformation even in dynamic, nuisance-laden environments.

[AI-40] Adam under Generalized Smoothness with Second-Moment-Type Stochastic Gradients NEURIPS2026

链接: https://arxiv.org/abs/2609.37787
作者: Ruinan Jin,Difei Cheng,Ling Chen,Jun Luo,Hao Zhou,Youzhi Zhang
类目: Artificial Intelligence (cs.AI)
备注: 37 pages, 4 figures. Accepted at NeurIPS 2026

点击查看摘要

Abstract:Adam is widely observed to remain stable even when the objective deviates significantly from global smoothness. Under the generalized smoothness framework, however, existing analyses rely on strong tail assumptions on the stochastic gradients, such as almost-sure boundedness or sub-Gaussianity. Whether Adam converges on generalized smooth objectives under only second moment information on the stochastic gradients, without such concentration assumptions, was identified as an important open direction by Li et al. (2023). This paper gives an affirmative answer under fairly general conditions: such tail assumptions are not necessary. Building on the Adam self-normalization framework of Jin et al. (2026), developed for classical smoothness and bounded variance, we extend the stopping-time and de-preconditioning strategy to the L_0 - L_p generalized smoothness condition and a generalized second moment ABC condition. Even when the stochastic-gradient condition provides only second moment information that may grow along the trajectory, the stochastic trajectory of Adam remains in a locally well-behaved smoothness region, with stretched-exponential tail decay under bounded variance and global smoothness. Consequently, we establish high-probability convergence rate guarantees over the full range p2 , with confidence dependence of order \delta^-1/2 , while the stepsize prefactor depends on \delta only through a single logarithmic factor. We further construct a hard instance showing that, under only second-moment information, this \delta^-1/2 -type confidence dependence is sharp. Finally, in the regime p1 , we combine the trajectory control with polynomial-growth estimates on rare events to obtain convergence rate guarantees in expectation.

[AI-41] OmniVCBench: Benchmarking Evidence-Grounded Multimodal Reasoning Towards AI Virtual Cells

链接: https://arxiv.org/abs/2609.37773
作者: Manyu Li,Xunkai Li,Yongfu Xiong,Yi Liu,Rong-Hua Li,Guoren Wang
类目: Artificial Intelligence (cs.AI); Quantitative Methods (q-bio.QM)
备注: 45 pages, 16 figures;

点击查看摘要

Abstract:Artificial Intelligence Virtual Cells (AIVCs) are envisioned as scientific agents that simulate cellular responses, explain underlying mechanisms, and support hypothesis-driven discovery. Existing AIVC benchmarks, however, operate primarily at the simulation layer, motivating complementary evaluation of how models interpret experimental evidence and formulate biological hypotheses. We introduce OmniVCBench, a figure-centric, source-traceable benchmark for the interpretation component of an AIVC. It contains 6,077 curated single- and multi-subfigure question–answer pairs derived from figures and experimental contexts in the scientific literature. Guided by Bloom’s taxonomy, we instantiate interpretation-layer counterparts of the AIVC Predict–Explain–Discover agenda through three scientific reasoning tasks. We further introduce AIVC-Judge, a task-conditioned MLLM-as-a-judge framework with category-specific, reference-aware rubrics for evaluating open-ended responses. A complementary Model-Derived Hard-Negative Mining (MDHNM) strategy converts plausible errors observed during model inference into MCQ distractors for lower-cost evaluation. Within the evaluated heterogeneous model pool, MCQ accuracy correlates positively with AIVC-Judge scores, providing a complementary view of performance alongside open-response evaluation. Code and data demo are available at this https URL.

[AI-42] Cross-Entropy Guided Routing in Mixture-of-Experts Large Language Models

链接: https://arxiv.org/abs/2609.37751
作者: Yury Nahshan,Nati Daniel,Jacob Goldberger,Yoli Shavit
类目: Artificial Intelligence (cs.AI)
备注: 25 pages, 6 figures, 12 tables

点击查看摘要

Abstract:Sparse mixture-of-experts (MoE) large language models scale model capacity by routing each token to a small subset of experts. Their routers are regularized with load balancing terms and learn affinity scores through the language-model objective. However, these objectives do not provide direct alignment between routing affinities and token-level error. We introduce token-error supervision for sparse routing in two forms. The first form predicts an error score per expert. The affinity-weighted aggregate of these scores is aligned to the next-token cross-entropy loss, while the individual scores attenuate affinity before top- K selection. The second directly aligns the router’s affinities to the model’s objective without requiring an additional head or inference-time modification. Both formulations use the Itakura–Saito divergence or an exponential negative log-likelihood for aligning affinities and token errors. Across two sparse MoE backbones and four multiple-choice question-answering benchmarks, we evaluate both supervision mechanisms. On Granite, our method improves accuracy by approximately 2.3 percentage points on average over a parameter-matched routing baseline. With stronger supervision, the gain on ARC-Challenge reaches 2.94 points. Both mechanisms preserve the native sparse execution budget and aggregation policy. Our code is available in the supplementary materials.

[AI-43] ContextRender: From Execution Dependencies to Agent Context

链接: https://arxiv.org/abs/2609.37743
作者: Savini Kashmira,Jayanaka L. Dantanarayana,Lingjia Tang,Jason Mars
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:LLM agents performing long-horizon tasks accumulate tool results that later steps may need. Passing the full history to every invocation is costly even when it fits within the context window, while reducing it risks omitting needed information. Existing context management methods can overlook how earlier tool results are used in subsequent execution, leaving needed information out of context. We introduce ContextRender, which manages context through a persistent graph of execution dependencies. We develop Tool-Flow Analysis to track how later operations reuse information from earlier tool results, providing a signal called observed reuse. A renderer combines this signal with recency and semantic relevance to select results within a fixed history budget, retaining omitted results for later use. Across AppWorld and 8-objective QA with three execution models, ContextRender outperforms the evaluated context management baselines using a 6K history budget, well below the models’ maximum context windows. Within this budget, it achieves task performance close to or above that of passing the full history while reducing mean inference cost by 10.2%-32.2% relative to Full history. Ablations show that observed reuse improves task performance and retention of results reused later.

[AI-44] Spatiotemporal Hyperedges for EEG Seizure Detection and Prediction CIKM2026

链接: https://arxiv.org/abs/2609.37730
作者: Hyunju Kim,Sheo Yon Jhin,Noseong Park,Nabil Imam
类目: Artificial Intelligence (cs.AI)
备注: Accepted at CIKM 2026. 7 figures, 6 tables

点击查看摘要

Abstract:Seizure detection and prediction from EEG are clinically important but challenging because seizures are rare, temporally localized, and propagate as coordinated events across multiple channels. Recent dynamic graph neural networks model this by running a temporal model over a sequence of per-time-step pairwise channel edges. However, this pairwise construction misses the spatiotemporal coupling that constitutes a seizure, at substantial training cost. We propose HyBrain, which summarizes spatiotemporal EEG evidence through a small set of soft hyperedges rather than pairwise edges. A per-channel Mamba backbone produces one token per (channel, second), and a spatiotemporal hyperedge block pools these tokens into E_h shared group embeddings through soft memberships and broadcasts them back. The same encoder serves three downstream tasks: window-based detection, one-second point-wise detection, and preictal seizure prediction. On TUSZ and CHB-MIT, HyBrain achieves the best AUROC on every reported setting against ten baselines, with the largest gap on long-clip preictal prediction. It also matches the most efficient baselines in training time and peak GPU memory. A qualitative analysis shows that even a single learned hyperedge cleanly captures the preictal - ictal - postictal trajectory on a real seizure clip.

[AI-45] Generative Interactions: Weaving Multiparty Human Motion with Bilevel Latent Dynamics

链接: https://arxiv.org/abs/2609.37708
作者: Ojas Shirekar,Yash Surange,Agustinas Jučas,Chirag Raman
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Human social behaviour is not a collection of independent motions, but a jointly organised process in which group dynamics and individual variation continuously shape one another. Yet existing social motion models often prioritise plausible trajectories while leaving interaction state implicit, limiting their ability to transfer across groups, tasks, and partial-observation regimes. To address this gap, we introduce Bilevel Representations for Agent Interaction Dynamics (BRAID), a hierarchical sequential latent-variable model for generative multi-person interaction. BRAID explicitly formulates social motion generation as a meta-transfer learning problem: shared interaction priors are learned across datasets and adapted through arbitrary context sets of observed people and joints. The model represents each scene through a group-level latent state that captures shared interaction dynamics and person-level latent states that capture individual behaviour conditioned on the evolving group context. This modelling choice enables coherent generation under full, sparse, or partial observations while exposing compact social-state vectors that can serve as an interface for downstream embodied-agent systems. We evaluate BRAID under a unified SMPL-based representation on social forecasting, tracking and in-filling, and response generation, using metrics that assess not only reconstruction accuracy but also realism, diversity, temporal alignment, and interpersonal coordination. We further analyse the hierarchical latent space, showing that it captures separable group- and individual-level structure.

[AI-46] Width Expansion as a Method for Class Incremental Learning

链接: https://arxiv.org/abs/2609.37702
作者: A. L. S. Conde,Y. Elkhatib,C. M. Ranieri
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Class Incremental Learning (Class-IL) requires models to learn new classes over time while preserving previously acquired knowledge without access to past data or task identity. This setting intensifies the stability-plasticity dilemma and makes catastrophic forgetting a central challenge. Existing approaches include regularization, knowledge distillation, replay, and architectural expansion. However, many expansion methods rely on explicit task identifiers or predefined growth strategies, limiting their applicability when task boundaries are unavailable at inference time. This work proposes a dynamic width expansion method that increases the number of neurons within existing layers according to a normalized loss criterion, without requiring task-specific information. An attention mechanism with persistent key-value memory is also incorporated to stabilize feature representations and reduce interference between previously learned and newly introduced classes. The approach is evaluated on Split MNIST and Split CIFAR-100 under the standard Class-IL protocol. Experiments compare fixed-capacity and dynamically expanding architectures, both with and without attention, combined with established continual learning methods including EWC, LwF, and A-GEM. Results show that progressive width expansion consistently improves performance over fixed architectures, particularly when combined with functional methods and A-GEM. The combination of width expansion and attention provides the most consistent gains. Overall, dynamic width expansion based on representational demand provides an effective and flexible strategy for Class-IL, although uncontrolled growth may increase overfitting and computational cost.

[AI-47] Locating Answer-Correctness Signals in Frozen Large Language Models

链接: https://arxiv.org/abs/2609.37700
作者: Yuansen Liu,Yixuan Tang,Anthony Kum Hoe Tung
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Language models expose internal signals that predict whether an answer is correct, readable from a single forward pass of a frozen model without additional generations. Yet existing probes often commit to one signal family or layer and can be brittle under distribution shift; in retrieval-augmented settings, many specialized detectors instead target passage faithfulness, which can diverge from correctness when retrieved evidence is unhelpful or conflicting. We therefore ask where answer correctness is readable, which internal signal families carry it, and how they should be combined. We search over hidden states, token probabilities, residual-stream features, attention, and their fusion, treating the selected readouts as a predictive measurement rather than a mechanistic localization. We run this analysis separately in closed-book and with-context settings, since context can change which readouts are informative. A consistent anatomy emerges: correctness concentrates in the answer span, recovered from the answer tokens even under retrieval, and the families carry it complementarily, so fusing them helps most out of distribution, where a single signal is weakest. The protocol is effective across two backbones and gates a retrieval controller as one downstream use.

[AI-48] GARDiff: Graph-Aligned Residual Diffusion for Probabilistic Multivariate Time-Series Forecasting

链接: https://arxiv.org/abs/2609.37694
作者: Rui Han,Min Yang,Xu Zhang,Xinghao Yang,Wei Liu,Yongshun Gong
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Diffusion models have recently shown strong potential for probabilistic multivariate time-series forecasting by modeling complex conditional distributions. Recent decoupled diffusion frameworks further separate forecasting into deterministic prediction and stochastic residual generation, making it natural to derive dependency graphs from deterministic representations and use them to guide residual diffusion. However, we show that this direct structural transfer is unreliable. Although deterministic-derived graphs encode useful global dependency priors, they exhibit substantial edge-level misalignment with residual dependency structures, introducing inaccurate or redundant conditions during residual generation. This reveals a previously overlooked deterministic-to-residual structural alignment problem in decoupled diffusion forecasting. To address this problem, we propose GARDiff, a Graph-Aligned Residual Diffusion framework for probabilistic multivariate time-series forecasting. Instead of treating deterministic-derived graphs as fixed diffusion conditions, GARDiff progressively adapts them to residual generation. Specifically, GARDiff estimates residual uncertainty to distinguish high- and low-uncertainty regions, enabling uncertainty-aware structural refinement, and further performs timestep-aware edge sparsification during reverse diffusion to evolve graph conditions from broad dependency aggregation to localized residual refinement. Extensive experiments on six real-world benchmarks demonstrate that GARDiff consistently improves probabilistic forecasting performance and uncertainty calibration over strong baselines.

[AI-49] WISE-ATTA: When to Ask for Labels in Budgeted Active Test-Time Adaptation

链接: https://arxiv.org/abs/2609.37687
作者: Muhammad Huzaifa,Lea Schönherr,Thorsten Eisenhofer
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Active test-time adaptation (ATTA) improves robustness under distribution shift by updating a deployed model during inference while selectively querying supervision. However, most existing ATTA methods implicitly assume that supervision can be requested for every incoming test batch, which can incur substantial annotation cost over long test streams. In this work, we introduce \emphbudgeted ATTA in which labels are available for only a fraction of test batches. This formulation shifts the central challenge from deciding \emphwhat to label within a batch to deciding \emphwhen supervision should be applied over time. To address this challenge, we propose a budget-aware approach \emphWISE-ATTA that allocates supervision over the test stream based on lightweight signals computed online, prioritizing periods where supervision is likely to be most useful. When a batch is selected for supervision, we further employ a drift-based sample selection criterion that targets samples exhibiting ongoing, unconverged adaptation dynamics, enabling effective updates from a single labeled example. We evaluate this approach on synthetic corruptions (ImageNet-C) and natural distribution shifts (ImageNet-R/K/A). Across settings, WISE-ATTA achieves competitive or improved performance compared to recent ATTA methods while requiring substantially fewer labels. Overall, we find that the timing of supervision is a key, yet underexplored, aspect of active test-time adaptation. Code: this https URL

[AI-50] Learning from Shared-Control Overrides: Context-Driven Acceleration Profile Prediction for Personalized Overtaking

链接: https://arxiv.org/abs/2609.37684
作者: Ruizheng Xu(Heudiasyc),Lounis Adouane(Heudiasyc),Javier Ibañez-Guzmán,Clément Zinoune
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Adaptive Cruise Control (ACC) systems are typically calibrated for an average driver, often resulting in a mismatch between vehicle behavior and individual expectations during time-critical maneuvers such as highway overtaking. When the ACC is perceived as too conservative and inconsistent, drivers intervene through throttle overrides, providing implicit feedback on the system’s behavior. This paper reframes these override actions as human-in-theloop supervisory signals and proposes a data-driven framework for personalized vehicle adaptation, termed Context-driven Personalized ACC (CoP-ACC). Rather than relying solely on end-to-end regression, which tends to over-smooth dynamic responses, we introduce a hybrid pipeline combining: (i) unsupervised hierarchical clustering to extract representative acceleration profiles from override events; (ii) a context classifier that maps pre-maneuver driving conditions to the appropriate profile; and (iii) a residual regressor that refines the selected profile into a smooth, personalized acceleration profile tailored to the immediate context. Evaluated on real-world public-road data against a withheld forced-ACC baseline, the approach demonstrates high reconstruction fidelity and generates acceleration profiles that tend toward the driver’s expected behavior in potential override contexts. The results highlight the potential of learning from shared-control overrides to enable anticipatory, personalized ACC behavior, reducing manual interventions and improving ride comfort.

[AI-51] Retrieve Reproduce Reveal: Dissecting Retrieval-Augmented Software Vulnerability Detection

链接: https://arxiv.org/abs/2609.37669
作者: Sabrina Kaniewski,Tim Krämer,Julius Bächle,Markus Enzweiler,Michael Menth,Tobias Heer
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Retrieval-Augmented Generation (RAG) is increasingly used to enhance Large Language Model (LLM)-based software vulnerability detection by grounding predictions in retrieved vulnerability knowledge, such as vulnerability reports. However, existing RAG-based software vulnerability detection (RAG4SVD) systems are often evaluated using proprietary models, which challenges open science and reproducibility. Further, studies use different datasets, custom knowledge bases, different backbone models, and diverse metrics, which hinders meaningful cross-system comparison. In this work, we study six open-source RAG4SVD systems and address these reproducibility and comparability challenges through (i) reproduction of their experimental settings under an open-weight setting, and (ii) a unified benchmark using a common dataset, metric suite, and pool of open-weight models. Further, RAG4SVD systems typically consist of multiple components, yet are often evaluated only as a whole system, i.e., end-to-end. Therefore, we perform (iii) a component-level analysis that decomposes representative RAG4SVD pipelines into input abstraction, knowledge retrieval, and detection. Our results demonstrate that reproducibility varies substantially across systems. Under the presented unified benchmark, published RAG4SVD performance does not transfer under a controlled open-weight evaluation and depends strongly on the used model. The component analysis shows that effective RAG4SVD depends on the alignment between pipeline stages. For example, oracle knowledge raises retrieval to near-optimal, yet performance remains low (0.51 pairwise accuracy), demonstrating that retrieval effectiveness alone is insufficient for reliable detection. These findings motivate evaluating RAG4SVD not only end-to-end, but at the level of pipeline components, and provide a basis for more standardized, RAG-aware evaluation practices.

[AI-52] Semantic Map Sharing and Capability-Aware Coverag e Planning for AI-Native 6G Robotic Coordination ALT

链接: https://arxiv.org/abs/2609.37666
作者: Abdulqader Dhafer,Qi Wang,Zhou Daniel Hao
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: An alternative version of this work was accepted for presentation at IEEE CSCN 2026

点击查看摘要

Abstract:Search and Rescue (SAR) operations increasingly deploy heterogeneous teams of aerial and ground robots. However, conventional coverage methods typically do not translate perceived terrain into platform-specific reachability, while continuous image exchange imposes a high communication cost. We propose an edge-centric, semantic-aware coverage planning framework that integrates aerial terrain perception, robot-specific traversability reasoning, and payload-efficient semantic state sharing. Aerial observations are converted into compact semantic grid maps, enabling reachability-constrained area decomposition and capability-aware coverage paths that assign only regions admitted by each robot’s capability profile. The resulting perception-sharing-planning loop feeds semantic corrections into traversability reasoning and replanning, forming an application-level mechanism motivated by AI-enabled goal-oriented communication envisioned for AI-native 6G networks. For the high-update case, transmitting semantic corrections reduces the application payload by a factor of approximately 82 relative to periodic full-map sharing. Across matched benchmark scenarios, the proposed method achieved 91.5% coverage with no capability-infeasible allocations, compared with 78.8% coverage and a 21.5% capability-infeasible allocation rate for LS-MCPP. Semantic corrections update the shared planning state without requiring repeated transmission of the complete map.

[AI-53] EnterpriseBench: Benchmarking LLM Agents on Enterprise-Level Strategic Reasoning and Decision-Making

链接: https://arxiv.org/abs/2609.37658
作者: Min Yang,Yichen Pan,Jinghua Piao,Dandan Song,Yongshun Gong,Yong Li
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:LLM agents are increasingly expected to support enterprise workflows, where tasks often involve missing information, uncertainty, feedback, and long-term trade-offs. However, existing enterprise and financial benchmarks mainly test static capabilities such as information extraction, numerical calculation, domain knowledge, and financial QA, leaving interactive and long-horizon decision-making underexplored. To bridge this gap, we introduce EnterpriseBench, a benchmark that evaluates LLM agents across this spectrum, from static question answering to dynamic decision-making. Specifically, EnterpriseBench reorganizes existing enterprise and financial QA datasets into a unified foundational suite annotated by capability and difficulty, and introduces three professional interactive settings: Consulting, based on management-consulting-style business cases for client problem diagnosis through multi-turn information seeking; the Beer Game, adapted from a classic supply-chain management simulation for inventory control under delayed feedback; and Enterprise Digital Twin, a project-based business simulator for workforce, risk, and project planning. Experiments with nine agent methods under four backbone models show that current agents have not yet achieved stable, comprehensive, and cross-task reliability in enterprise scenarios. These results show that EnterpriseBench provides a practical benchmark for evaluating LLM agents in realistic enterprise strategic reasoning and decision-making.

[AI-54] Beyond a single latent space: a dual-latent world model for long-horizon planning

链接: https://arxiv.org/abs/2609.37644
作者: Delin Zhao,Zhengrong Yue,Shaobin Zhuang,Junlin He,Xiaoyu Chen,Zikang Wang,Yuxin Liu,Limin Wang,Yali Wang
类目: Artificial Intelligence (cs.AI)
备注: 31 pages, 22 figures, 9 tables. Main text: 9 pages

点击查看摘要

Abstract:Latent world models often struggle with long-horizon planning despite accurate short-term predictions. Recursive rollouts accumulate errors, while distance concentration in high-dimensional latent spaces can weaken goal discrimination. We introduce the Dual-Latent World Model (Dual-WM), which separates local execution and long-range planning through distinct state representations and dynamics models. The low-level model predicts action-conditioned transitions, while the high-level model uses learned macro-actions to plan over longer temporal spans. We also propose Long-Horizon Representation Learning with Weighted Rollout (LoRe), which supervises self-generated predictions at both levels. An analysis of recursive error propagation motivates exponential horizon weights with separate decay rates for the two temporal scales. During planning, the high-level model generates latent subgoals that the low-level model refines into actions for precise execution. We evaluate from-scratch Dual-WM on five goal-conditioned visual control tasks against the task-wise strongest baselines without actor-guided proposals. At goal offsets of 50 and 100 environment steps, mean success increases from 75.9% to 84.4% and from 61.4% to 69.5%, respectively. At offset 100, Dual-WM outperforms these baselines on all five tasks and improves mean success over LeWM by 30.8 percentage points. Ablations and supporting analyses provide evidence of more informative representations for goal evaluation and greater consistency under recursive prediction. These results highlight the value of separating temporal roles and training across multiple horizons for reliable latent planning. Our core implementation is available at this https URL.

[AI-55] Flattening the Connectome Spectrum: A Spectral Filter for FC Induces a Pretraining Target for fMRI Encoders

链接: https://arxiv.org/abs/2609.37642
作者: Giovanni Marraffini,Victoria Shevchenko,Carlo Alberto Barbano(UNITO),Demian Wassermann(MIND)
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Neurons and Cognition (q-bio.NC)
备注:

点击查看摘要

Abstract:Self-supervised pretraining reshaped prediction in language and vision, and brain foundation models (BFMs) inherited its promise. Representations learned from large unlabelled corpora should capture individual functional dynamics and generalise across cohorts. However, kernel ridge regression (KRR) fitted on functional connectivity (FC) matrices still predicts individual phenotypes more accurately than any BFM we tested. In this paper, we show that KRR is weighted by the eigenvalues of the FC which are miscalibrated for phenotype prediction. We apply an efficient spectral filter to recalibrate the eigenvalues of each subject’s FC matrix, enabling the model to exploit more inter-individual variance. Across the 5 datasets, 11 parcellations and 6 prediction targets we tested, we match or exceed the KRR baseline. Based on this finding, we then pretrain a small encoder model on about 4,000 hours of fMRI from 162 open datasets, whereby we align the pairwise similarities between the embeddings of recording snippets with those between the recalibrated connectomes. Our model performs on par with the best of the 6 published BFMs we tested while having an order of magnitude fewer parameters. Our encoder performs better than FC on short scans and in smaller cohorts, especially in fingerprinting. We release the pretrained model weights, the code and the pretraining data, preprocessed and parcellated.

[AI-56] ProCTI: Prototype-Refined Global Conditioning for Diffusion-Based Time Series Imputation

链接: https://arxiv.org/abs/2609.37632
作者: Fariza Rashid,Duc Van Le,Rahat Masood,Gustavo Batista,Aruna Seneviratne,Suranga Seneviratne
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Time series imputation has progressed from statistical and deep learning approaches to diffusion-based models, which have shown strong recent performance. Existing diffusion-based methods typically condition the reverse process using local contextual information from the current or neighbouring windows. Meanwhile, global dataset-level structure often remains implicit, limiting performance when local observations are sparse, noisy, or unrepresentative. To address this issue, we propose ProCTI, a diffusion-imputation framework that augments local conditioning with retrieved global dataset-level priors through learned prototypes. A hybrid conditioning mechanism integrates this global context with local signals during reverse diffusion, enabling more accurate reconstruction under varying missingness scenarios. Experiments across multiple benchmark datasets show that ProCTI outperforms strong baselines overall under random missingness, while remaining competitive under attribute-wise missingness. Furthermore, we use a latent-regime data model to characterise the precise conditions under which prototype-derived global conditioning provably improves imputation. We support this with a general theoretical analysis of local-global conditioning.

[AI-57] SPLASH: Switching Parallel Layouts of Attention with Seamless Handoff for LLM Serving

链接: https://arxiv.org/abs/2609.37626
作者: Chuan Liu,Shuoming Zhang,Zhicheng Li,Qianqi Sun,Ruiyuan Xu,Qiuchu Yu,Xiyu Shi,Huimin Cui,Jiacheng Zhao
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI)
备注: 28 pages, 11 figures, 8 tables. Code: this https URL

点击查看摘要

Abstract:No single way of parallelizing attention serves large language models well under all loads. Low concurrency favors tensor parallelism, many independent requests favor data-parallel attention, and long prompts favor context parallelism. Reasoning, agentic, and RL-rollout workloads make a fixed choice untenable: a batch that begins as many short requests ends as a few very long ones, so the best layout changes while the same requests run. Serving engines nevertheless fix one layout at launch, because changing it has meant draining requests and restarting workers. We present SPLASH, a serving system that switches the parallel layout of attention while requests are running. It builds on one observation: modern attention, with few or no KV heads, decouples where a request’s KV cache lives from how attention weights are sharded. This has two consequences. First, layouts differ only in who owns the weights and the cache, and most of that state already sits where the next layout needs it; SPLASH reuses it, moves the rest in the background of ongoing inference, and hands off at a batch boundary, making a switch nearly free: its median overhead is under 0.51% of the step it runs in. Second, the decoupling exposes a layout that existing engines lack: Decoupled Ownership Parallelism (DOP) shards attention weights as tensor parallelism does while keeping each request’s cache on a single owner as data-parallel attention does. DOP replicates neither, offers 27-60% more KV capacity than data-parallel attention, and gives the scheduler a choice when KV memory limits admission. A transition-aware scheduler follows the best of the four layouts as load changes. On B200 GPUs serving GLM-5.3, SPLASH improves end-to-end serving throughput by 1.3-1.73x over fixed-layout deployments, and the same layout regimes appear with DeepSeek-V3.2 on H200 and GLM-5.3-Flash on DCU.

[AI-58] AS2D: Accelerating On-Demand Audio Understanding on Mobile Devices

链接: https://arxiv.org/abs/2609.37617
作者: Yunzhe Li,Kyoungjun Park,Hongzi Zhu,Lili Qiu
类目: ound (cs.SD); Artificial Intelligence (cs.AI)
备注: 43 pages, 9 figures, 16 tables

点击查看摘要

Abstract:Speculative decoding accelerates autoregressive generation by using a smaller drafter to propose tokens for batched verification by a larger target. However, conventional speculative decoding couples drafting to the target’s evolving verified prefix, serializing drafting and verification. We ask whether this dependency is necessary for source-conditioned generation. Our key observation is that, for audio language models, the input audio and user request can provide useful speculative candidates without following the target’s evolving text prefix. We propose AS ^2 D (Audio Speculative Speculative Decoding), which enables target-decoupled drafting: an audio-conditioned drafter follows its own generation history while the target independently verifies and corrects ready candidates. Without usable candidates, the target advances alone. Thus, target feedback determines which candidates are committed but no longer determines when the drafter can make progress, enabling drafting and verification to proceed concurrently while retaining target-side verification and correction. We implement AS ^2 D in MNN for Android and evaluate two target models across four phones, seven datasets, and three tasks covering 12.2 hours of audio. Across four phones, AS ^2 D improves pooled ASR throughput by 42-76% over target-only decoding, while only 5.7% of evaluation windows are slower than target-only, compared with 58.1-63.0% for speculative baselines. For ASR, AS ^2 D reaches 97.33-98.20% of a hindsight per-window oracle’s pooled throughput over the evaluated drafter/budget catalog. Native on-demand execution with a 7B target achieves up to 78% higher throughput than target-only. These results show that source-conditioned audio generation can relax the conventional dependence of speculative drafting on the target’s evolving output prefix, exposing substantial parallelism for efficient inference.

[AI-59] Independent Verification Paths Are Not Independent: A Case Study of Common-Mode Failure in a Satellite Catalogue Pipeline NEURIPS2026

链接: https://arxiv.org/abs/2609.37603
作者: Fabio Rovai
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Databases (cs.DB)
备注: Accepted at the AI for Science workshop (NeurIPS 2026). Code, prompts, raw model outputs and per-trial records: this https URL (paper/gates/), archived at this https URL

点击查看摘要

Abstract:A common safeguard for a data pipeline is redundant computation: derive each published number by two routes built on different technology and refuse to exit when they disagree. We report one such gate failing, in a cross-catalogue integrity study of two open registers of Earth-orbiting objects. A gate comparing a set-based Python path with SPARQL queries over the emitted RDF graph printed ALL CROSS-CHECKS AGREE on seven counts. Three were wrong, one overstated more than fourfold (932 against 220). Both paths imported the same constants, which encoded a misreading of the source’s status vocabulary, so the error was common-mode and the gate could not see it. We give the mechanism, an object-level ledger reconciling every figure, and three checks that go back to the source’s documentation, measured on the defective code and on its correction. We then checked that correction against each object’s phase history, held in a source file the pipeline never read. The correction was also wrong: 42 of its 261 disagreements are artefacts, and none of our three checks flagged them. Finally, in a controlled replication with three pinned models and tools disabled, 72 of 75 paths generated on request as independent checks computed the defective count, 29 of 30 even when the prompt carried the source’s own definitions of the codes. The evidence is one pipeline and one defect family. Within it, redundancy verified implementation, and the errors that reached publication were errors of meaning.

[AI-60] XU-RS: Explaining Credal Width in Random-Set Language Models

链接: https://arxiv.org/abs/2609.37594
作者: David Achara,Maryam Sultana,Alexander D. Rast,Fabio Cuzzolin
类目: Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
备注: 32 pages, 1 figure, 10 tables

点击查看摘要

Abstract:Uncertainty estimates tell us how unsure a model is, but not why. Without knowing which parts of an input influences a model’s uncertainty, we cannot tell whether that uncertainty score depends on input features that are relevant for the task. We study this problem in randomset classifiers built using pretrained language models. These classifiers assign probability to individual answers and to groups of answers, producing lower and upper probabilities for each answer; The difference between these probabilities, called credal width, is used to represent epistemic uncertainty about an answer arising from limited training data. We propose XU-RS, a framework that attributes an answer’s credal width to the input tokens (words or word pieces) supplied to a language model. XU-RS uses Expected Gradients (a standard feature attribution method) to estimate how input tokens contribute to credal width. The proposed framework is evaluated on a MedQA dataset using SmolLM3-3B and Llama-2-7B models, demonstrating that setting the embedding of a token ranked highly by XU-RS to zero (zero-masking) causes larger changes in credal width than zero-masking randomly selected tokens. In addition, we show that normalisation can cause other answer groups to influence an answer’s width, reveal how token attribution can mask numerical errors, and provide diagnostic checks to verify whether a token ranked highly by XU-RS meaningfully explains model uncertainty.

[AI-61] Credit-Guided Policy Improvement for Test-time Adaptive Vision-Language Navigation

链接: https://arxiv.org/abs/2609.37591
作者: Yang Li,Sijia Zhang,Yihan Li,Aming WU,Zihao Zhang,Ziju Han,Yahong Han
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Test-time adaptation for vision-language navigation (TTA-VLN) enables pretrained policies to adapt online to unseen environments using only test-time observations and interaction history. However, distribution shifts can distort local action preferences and lead to off-course decisions. Existing methods rely on predictive uncertainty, trajectory-level feedback, or accumulated adaptation experience to correct such deviations. These signals, however, do not directly reveal whether an executed action supports instruction-guided progress toward the goal. Moreover, a plausible corrective signal does not guarantee a reliable policy update. The key challenge is thus twofold: identifying interactions that support goal-directed improvement and determining whether the resulting updates are worth retaining. We observe that each executed action induces an immediate observation transition, providing evidence of its local consequences. Based on this insight, we propose Credit-Guided Policy Improvement (CGPI), which recovers signed, reference-relative decision credit from action-induced observation transitions without external outcome feedback. With the pretrained navigation policy frozen, CGPI uses this credit to propose lightweight adaptation updates and verifies them against prior credit-supported interactions. Updates are retained only when supported and rolled back otherwise. CGPI achieves consistent gains across the evaluated VLN benchmarks and navigation backbones, while qualitative robot trials further illustrate the feasibility of zero-shot sim-to-real transfer.

[AI-62] ReLMem: Learning Recurrent Memory for Longitudinal EHR Modeling

链接: https://arxiv.org/abs/2609.37587
作者: Zijie Meng,Xiwei Dai,Yingying Zhang,Jian Wu,Xian Wu,Zuozhu Liu
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Longitudinal electronic health record (EHR) modeling requires integrating new visits with an expanding patient history. Yet the continual accumulation of clinical information imposes increasing computational and memory costs on large language models (LLMs) when they process and retain complete patient histories. A practical alternative is visit-wise recurrent compression, which incorporates each incoming visit into a compact, continually updated patient memory. However, under a fixed memory budget, successive updates must integrate new information without progressively losing critical historical evidence needed to subsequent tasks. To address this challenge, we introduce Recurrent Longitudinal Memory (ReLMem), a framework that learns to maintain fixed-capacity patient memory for efficient downstream prediction with a frozen LLM. ReLMem equips this LLM with lightweight compression adapters to recurrently update the memory from its previous state and each incoming visit, without rereading earlier records. Specifically, we develop a multi-granularity optimization strategy to preserve task-relevant information throughout recurrent updates and support downstream prediction from the final memory. The intermediate supervision aligns attention outputs from compressed memory and the full history under identical queries, while prediction supervision minimizes cross-entropy with ground truth answers conditioned on the final memory. On EHR-based medication prediction, ReLMem approaches the F1 scores of full-history baseline while reducing average retained historical storage by 97.1%. Under the same memory budget, it improves macro- and micro-F1 over the strongest compressed-memory baseline by 4.66 and 4.75 percentage points, respectively. These results highlight the value of learning recurrent patient memory for efficient longitudinal EHR modeling.

[AI-63] Learning as Deepfakes Evolve: RF-Prompt for Continual Audio Deepfake Detection

链接: https://arxiv.org/abs/2609.37586
作者: Yuankun Xie,Xiaoxuan Guo,Xiaopeng Wang,Siqing Qin,Shaole Li,Kong Aik Lee
类目: ound (cs.SD); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Continual audio deepfake detection requires learning newly emerging deepfake methods while retaining discrimination of previously encountered speech. Existing dataset-incremental evaluation changes both real-speech domains and deepfake mechanisms, making their effects difficult to distinguish. We construct five task organizations over identical training, development, and evaluation pools to study these factors under a controlled sample budget. Our proposed Real-Anchored Mechanism-Incremental (RAMI) protocol reflects the practical setting in which available real speech provides a recurring mixed-domain reference while new deepfake mechanisms arrive incrementally. We further propose RF-Prompt, an asymmetric continual prompt-learning method that preserves reusable real-speech knowledge through a shared real prompt and expands mechanism-specific knowledge through inherited fake experts with orthogonal residuals. Input-adaptive soft fusion combines the accumulated experts into a fixed number of injected tokens without requiring task identity at inference. On RAMI, RF-Prompt achieves 10.110% average EER and 10.370% pooled EER, outperforming all evaluated continual-learning baselines. Across the five controlled protocols, RAMI yields the lowest common-average and pooled EER. Component ablations, limited-data experiments, and cross-backbone evaluations further validate the proposed design.

[AI-64] Concealing LLM -Based Multi-Agent Topology via Phantom Structure Injection

链接: https://arxiv.org/abs/2609.37567
作者: Longzhu He,Zelang Wen,Xinfeng Li,Sen Su,XiaoFeng Wang
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Driven by the rapid advancement of large language models (LLMs), LLM-based multi-agent systems (MAS) have emerged as a powerful paradigm for collaborative reasoning over complex tasks. A key design element of MAS is the communication topology, which governs information flow among agents and often encodes proprietary knowledge about the system architecture. However, recent work has shown that such topologies can be inferred even in black-box settings by exploiting semantic dependencies in observable reasoning traces, posing significant risks of intellectual property leakage and exposure of system vulnerabilities. To address this threat, we propose MIRAGE, a topology-concealment framework that preserves the genuine communication topology for task execution while shaping adversary-facing semantic evidence toward a carefully constructed phantom topology. Specifically, MIRAGE operates in three stages: (1) phantom topology synthesis, (2) semantic edge realization, and (3) protected MAS execution. It constructs a phantom topology structurally distinct from the genuine one, materializes phantom edges as plausible semantic dependencies, and suppresses source-specific cues that could reveal genuine edges absent from the phantom topology. Extensive experiments across three topology optimization frameworks and four benchmark datasets demonstrate that MIRAGE substantially reduces the effectiveness of topology inference attacks while largely preserving the task utility of the protected MAS.

[AI-65] Risk-Aware Semantic Grounding for Trustworthy LLM -Based Robot Planning

链接: https://arxiv.org/abs/2609.37554
作者: Łukasz Sobczak,Nur Keleşoğlu,Sławomir Piotr Nowak
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models (LLMs) are increasingly used as high-level planners in robot navigation, but their outputs may become unreliable when instructions are ambiguous, unsupported by the environment, or semantically inconsistent. This paper presents a Risk-Aware Semantic Grounding framework for trustworthy LLM-based robot planning. Unlike existing LLM-based planners that primarily optimize plan generation, we formulate semantic grounding reliability as a multi-dimensional risk estimation problem. The proposed architecture explicitly models grounding uncertainty through ambiguity, hallucination and semantic-conflict risks before planning occurs, enabling the system to decide whether to execute the instruction, request clarification, or reject it. To evaluate the approach, we introduce TRUST-NAV, a benchmark containing both standard navigation tasks and risk-inducing instruction scenarios. Experimental results show that while conventional LLM planners achieve strong performance on valid navigation tasks, the proposed framework substantially improves ambiguity detection and semantic conflict rejection. These findings suggest that trustworthy robot planning should be evaluated not only by task completion, but also by the ability to recognize when execution should not occur.

[AI-66] How Can Recommendation Feedback Evolve Agent Memory?

链接: https://arxiv.org/abs/2609.37544
作者: Shanwen Mao,Mingming Li,Hao Zhang,Zhiheng Li,Yige Wang,Penghua Yu,Junxiong Zhu
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Content-generation agents continuously receive impressions, clicks, conversions, and negative feedback from recommendation systems, providing real-world outcome signals for memory evolution. However, these signals are delayed and noisy, confounded by audience composition, placement, and recommendation policies, and may result from the combined influence of multiple memories, making accurate attribution difficult. Existing methods rely primarily on immediate feedback or semantic retrieval and therefore struggle to reliably translate recommendation outcomes into memory fitness. To address this challenge, we propose TIDE (Trajectory-Informed Directed Memory Evolution), an external memory evolution framework driven by delayed recommendation feedback. We further introduce Memory Evolution Gain (MEG), which measures the utility improvement of evolved memory over a no memory baseline on strictly future tasks. TIDE treats memory as a capacity-constrained population of experiences: temporal and semantic credit assignment estimates contextual fitness, while responsibility credit distributes outcome signals according to the memories referenced during generation. These signals are then used to reinforce, crossover, mutate, or evict memories. On an e-commerce membership marketing content-generation agent, TIDE achieves a +7.75-percentage-point MEG in offline temporal replay and significantly improves both unique click-through rate (UCTR) and activation rate in an online A/B test. On a delayed-label benchmark, TIDE achieves the lowest mean absolute error (MAE) and root mean squared error (RMSE) and the highest MEG among the compared methods, demonstrating its effectiveness.

[AI-67] SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation

链接: https://arxiv.org/abs/2609.37539
作者: Renxi Wang,Mingshan Hee,Fajri Koto,Timothy Baldwin,Haonan Li
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Skills equip LLM agents with professional knowledge and guidance to complete long-horizon and complex tasks. Although skills have been widely adopted in recent agent paradigms and harnesses, how to synthesize reliable training data and how to train agents for skill use remain underexplored. In this work, we propose SkillGym, an automatic pipeline to build verifiable environments, collect trajectories, and train skill-use agents. SkillGym first crawls a large volume of skills from the internet, then keeps those whose workflows can run reproducibly offline. A builder-reviewer pipeline is used to construct difficulty-controlled tasks, spanning four task types, each with a reference solution and an executable verifier. With this pipeline, we build 6.8k environments and collect 19k verified successful trajectories for supervised finetuning. Finetuning on these trajectories improves LLMs of different families and sizes, from 2B to 122B parameters across four skill-use benchmarks; Our Qwen3.5-9B SFT model outperforms the 397B untrained model on two of them. Further analysis shows that training teaches agents to invoke skills, raising the rate of reading the relevant skill from 28% to 96%, and that the gains hold across reasoning structures, extending to task types that form a minority of the training data and to skills held out from training

[AI-68] DScale: Scaling Block-Diffusion Speculative Decoding with Adaptive Verification

链接: https://arxiv.org/abs/2609.37532
作者: Rongjian Chen,Minxian Xu,Zhengxin Fang,Kejiang Ye,Chengzhong Xu
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI)
备注: 12 pages

点击查看摘要

Abstract:Growing large language model applications demand efficient inference. At high concurrency, block-diffusion speculative decoding suffers from verification padding, rejected candidates, and incompatibility between variable prefixes and fixed-shape graphs. Uniform truncation sacrifices acceptable tokens. We present DScale, preserving drafter architecture, weights, and full draft length. A separate 112K-parameter predictor requires neither confidence calibration nor hardware speed-curve preparation. Path-aware tiles reduce padding. Dynamic verify-length (DVL) allocation packs scored prefixes into half the native verification capacity. Fixed-address workspaces propagate changing boundaries through verification and acceptance while reusing captured graphs. On A100-40GB with tensor parallelism 1, Qwen3-8B and Qwen3-4B cover four datasets and concurrency 8-32, reusing each target’s frozen predictor. Geometric-mean throughput gains across these configurations are respectively 43.9% and 48.8% over DFlash, 22.2% and 37.7% over DSpark, and 24.4% and 32.0% over Domino, with lower request latency. Cumulative ablations show that adding the three mechanisms successively increases geometric-mean throughput, while budget adjustment improves accepted-token retention. GPU profiling shows that complete decode-step time on GSM8K decreases by 30.8-52.5% relative to DFlash

[AI-69] Video2STL: Grounding VLM-Generated Temporal Specifications for Robot Learning

链接: https://arxiv.org/abs/2609.37519
作者: Merve Atasever,Keyan Azbijari,Cagan Bakirci,Bo-Ruei Huang,Tolga Izdas,Zahra Shahrooei,Richard Yang,Erdem Biyik,Jyotirmoy V. Deshmukh
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Video-based policy learning is particularly promising, as it illustrates target behaviors without requiring action annotations or embodiment-matched demonstrations. A central challenge is deciding what information should be transferred from the video to the robot. Existing approaches commonly convert visual observations into scalar similarity or value signals, or ask foundation models to directly generate reward code. These approaches can make the temporal structure of a task difficult to inspect, ground, and reuse. We present Video2STL, a framework that converts observation-only videos into parametric Signal Temporal Logic (STL) specifications and uses the resulting formal representation for robot learning. A vision-language model extracts an embodiment-independent semantic event trace and constructs a bank of symbolic temporal specifications. The model determines the task structure, while numerical predicate thresholds and temporal bounds are grounded from successful robot trajectories. For policy learning, we separate short- and long-timescale temporal information: short-horizon specifications provide dense rewards through rolling-window quantitative robustness, while a causal monitor over a retained long-horizon specification provides one-time progress rewards for valid temporal prefixes. The same representation supports cross-embodiment transfer from human or animal videos to robot control. Across four manipulation tasks, Video2STL achieves 85.8% average success-once and 67.0% success-at-end, compared with 81.5%/59.5% for native dense PPO and 65.0%/42.3% for Text2Reward; in quadruped locomotion, Qwen-3.8 and GPT-5.6-based Video2STL policies achieve 100% success across velocities from 0.3 to 2.1,\mathrmm/s while remaining competitive in high-speed energy efficiency. Project webpage: \hrefthis https URLvideo2stl.

[AI-70] Boundary-State Control for Tool-Using Language-Model Agents : Commit-Time Consistency under State Drift

链接: https://arxiv.org/abs/2609.37475
作者: Wesley Shu
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Tool-using language-model agents can decide that an action is permissible and execute it only after security-relevant state has changed. We study this proposal-to-commit gap and introduce BSC-R, a deterministic effect-boundary mechanism that binds a single-use commit authorization to the exact action and to a semantic projection of the authorization state that justified it. On 2,847 attacked AgentDojo episodes, the boundary kernel preserves the unprotected agent’s behavior exactly (80.576% utility; 1.616% attack success). On 10,302 frozen proposals, it accepts every unchanged commit and rejects every instance of ten prospectively specified changed or replayed classes. In an independently generated boundary-drift experiment, full joint binding commits 0/4,403 invalid contexts while retaining 5,899/5,899 valid contexts. A prospective external evaluation on the 3,460-scenario CONTINUITY suite retains all 700 benign cases, prevents 1,200/1,200 represented non-replay invalid effects, handles 160/160 replay lifecycles correctly, and withholds 200/200 ambiguous no-release cases. The broader external suite also exposes the method’s limit: across all 2,560 attacks, BSC-R has a 25% invalid-effect commit rate versus 0% for CONTINUITY. The result is therefore a scoped commit-time consistency mechanism, not a universal agent-safety claim.

[AI-71] Authority Before Utility: Non-Compensatory Control for Persistent LLM Memory

链接: https://arxiv.org/abs/2609.37474
作者: Wesley Shu
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Persistent memory creates a control problem that retrieval relevance alone does not solve: a memory can remain highly useful after an update, deletion, or revocation makes it inadmissible for the current answer. We formalize this as a separation between utility and authority. A fixed finite penalty applied to an unnormalized utility score cannot guarantee exclusion under arbitrary positive-affine reparameterization of that score; by contrast, rank-normalized compensation is scale-invariant and therefore forms a stronger empirical comparator. Our prospectively frozen TIDE/LongMemEval primary was quarantined before a valid HELDOUT comparison because the materialized TIDE adapter conflated historical age with query-relative inadmissibility and the aligned LongMemEval split left no DEV set for the predeclared penalty selection. We therefore report a post-primary replacement diagnostic on Memora Remembering, where update/delete operations provide item-level forgetting state. On Qwen3-8B, DEV selected lambda = 0.6 from a ten-point normalized SOFT family. Across 185 HELDOUT units in 28 dependency clusters, HARD exclusion yields 4.04% balanced construct error versus 19.66% for locked SOFT, a paired difference of 15.61 points with a 20,000-replicate cluster-bootstrap 95% interval of [13.07, 18.76]. The effect is driven primarily by forgotten-value leakage while current-value recall is preserved. This is same-Q operator-comparison evidence, not a universal claim that scalar control fails, not an evaluation of learned authority inference, and not an independent downstream-harm endpoint. Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2609.37474 [cs.AI] (or arXiv:2609.37474v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2609.37474 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Wesley Shu [view email] [v1] Sat, 26 Sep 2026 08:45:34 UTC (16 KB)

[AI-72] CRASM-Gate: Deterministic-First Constraint- and Role-Aware Semantic Mapping with Selective Model Assistance Across Heterogeneous Industrial Standards

链接: https://arxiv.org/abs/2609.37458
作者: Kabeh Mohsenzadegan,Vahid Tavakkoli,Kyandoghere Kyamakya
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Industrial standards often encode the same engineering concept through incompatible hierarchies, identifiers, roles, and structural constraints, so the nearest lexical or embedding match can still be technically inadmissible. This article presents CRASM, a deterministic constraint- and role-aware semantic mapping method, and CRASM-Gate, its selectively model-assisted extension. The framework separates standard-specific canonicalization, bounded retrieval, deterministic rules, destination-versus-origin role interpretation, semantic and structural ranking, ambiguity refusal, and target validation. CRASM-Gate adds a confidence/disagreement gate that may invoke a candidate-constrained large language model, while final authority remains with deterministic validation. A controlled artifact covers six directed industrial-standard pairs, three difficulty levels, and ten configurations, yielding 14,400 sample-level decisions. With a fixed local model endpoint, CRASM-Gate reaches mean F1 0.9938 and an implemented structural-validity rate of 1.0000; deterministic CRASM reaches 0.9931 without generative-model calls; and the model-only baseline reaches 0.5347. Relative to the model-only baseline, CRASM reduces top-1 errors from 670 to 10 while exhibiting 0.0688 s/sample rather than 24.5263 s/sample observed latency. CRASM-Gate improves CRASM by one additional correct decision out of 1,440, with observed latency increasing to 15.7666 s/sample. The results support a deterministic-first interoperability architecture in which model assistance is optional, measurable, candidate-bounded, and unable to bypass structural validation.

[AI-73] Commitment Hierarchies under Intent Revision: A Belief-Revision Account of Salvage in Tool-Use Agents

链接: https://arxiv.org/abs/2609.37453
作者: Spandan Ghose Chowdhury
类目: Artificial Intelligence (cs.AI)
备注: Accepted in the 27th International Conference on Principles and Practice of Multi-Agent Systems

点击查看摘要

Abstract:When a user changes their mind partway through a task, an agent that has already split the task into sub-goals and paid for tool calls must decide, per cached sub-result, whether to keep, patch, or discard it (salvage), restarting wastes valid work and continuing unchanged answers the old question. Our main finding is that salvage quality is a matter of role design rather than model capability: a language model asked the keep/patch/discard question one node at a time is unreliable, but asked to classify the revision once, with a deterministic layer propagating the decision, it reaches the cost-optimal oracle on all three models tested, from two vendors. Modeling the plan as a commitment hierarchy and the intent change as a belief-revision operator with AGM style postulates, we prove that no policy observing only a node’s local view can be both safe and cost optimal, while the single classification design is both. Across three environments the policy recovers the full achievable savings, 43% cheaper than restart, at 100% correctness.

[AI-74] Engineering Efficient Self-Play Chess: Search Replay and Throughput Under Limited Compute

链接: https://arxiv.org/abs/2609.37447
作者: Bertil Braun
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 26 pages, 17 figures, including appendices. Code and experimental evidence: this https URL ; model artifacts: this https URL ; interactive demo: this https URL

点击查看摘要

Abstract:How strong can an AlphaZero-style chess system become under limited training compute when its entire learning loop is engineered for efficiency? We train from random initialization through searched self-play on a single eight-GPU node for 2.5 days. The resulting 6.32-million-parameter model reaches 3,251 benchmark Elo [3,206, 3,297] at 100,000 searches per move (estimated at under five seconds of thinking time) against a fixed-node Stockfish 13 ladder. The run ingests 3.25 million completed games, involves an estimated 100 billion search simulations, and makes 836.6 million training presentations. We investigate search allocation, replay and restart-state selection, policy representation, progressive model sizing, quantized inference, and throughput engineering. Alongside the retained design, we document plausible alternatives that failed to improve the complete learning loop or did not justify their cost. The reported strength is a result of the integrated system, not an isolated Elo gain attributable to any single choice.

[AI-75] Demistifying Data and Simulator Assumptions in Supervised Causal Discovery

链接: https://arxiv.org/abs/2609.37446
作者: Pingchuan Ma,Rui Ding,Bojun Huang,Shuai Wang
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Supervised causal discovery learns to infer causal structure for a new dataset from training datasets paired with structural labels. These training pairs are typically simulated, making the simulator both a source of supervision and a carrier of assumptions about causal graphs, mechanisms, and noise. Understanding the resulting predictions therefore requires examining how these assumptions supplement the information available in observational data, which may be compatible with multiple causal graphs. This paper examines that relationship across representative methods available through June 2026. We organize these methods by prediction target, prediction granularity, encoder, structural decoder, and training regime to relate what each method predicts to how it uses data and simulator-based supervision. Using this framework, we distinguish two questions: whether the target is identifiable under the assumed model class, and whether a trained predictor generalizes beyond its training distribution. Restrictions on mechanisms and noise can make otherwise ambiguous causal directions identifiable, but predictive accuracy under those restrictions does not establish transfer when they change. This distinction motivates evaluation that matches metrics to the identifiable graph target and tests changes in graphs, mechanisms, and noise between training and deployment. Extending such evaluation to real data also requires documenting the external causal evidence and uncertainty behind benchmark reference graphs. Together, these analyses guide method comparison and identify open questions in transfer, test-time adaptation, and uncertainty assessment.

[AI-76] Simultaneous Neural Optimal Transport

链接: https://arxiv.org/abs/2609.37424
作者: Milena Gazdieva,Kirill Sokolov,Jiawei Chen,Evgeny Burnaev,Alexander Korotin
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Optimal Transport (OT) provides a principled framework for learning transformations between probability distributions from unpaired samples. In many applications, however, a single transformation must map several source distributions to a common target distribution. For example, image restoration might require handling different types of degradation without knowing the degradation of each input at inference time. Simple approaches of pooling the source distributions only encourage alignment with the target at the aggregate level and may leave individual sources misaligned. In our paper, we consider the simultaneous OT problem which formalizes the task of learning a shared transport map that minimizes the average transport cost while aligning each source distribution with a prescribed target. We propose a neural method for solving the simultaneous OT problem by learning a shared transport map that minimizes the average transport cost while aligning each source distribution with a prescribed target. We derive a max-min formulation for learning this map. We illustrate its application to image restoration, where a single model handles multiple degradation types using a common collection of clean target images.

[AI-77] Complexity-Aware Evaluation of LLM Comprehension

链接: https://arxiv.org/abs/2609.37405
作者: Ali Mohammadi Esfahani,Nafiseh Kahani,Samuel A.Ajila
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models (LLMs) are increasingly used for software engineering tasks that require understanding existing source code, including behavior prediction, function explanation, debugging, and code review. However, aggregate benchmark accuracy can conceal how model reliability changes as source code becomes structurally more complex. This paper presents a complexity-aware framework for evaluating LLM code comprehension using cyclomatic complexity, nesting depth, branching factor, and Halstead volume. We evaluate DeepSeek-Coder-V2 and Llama through two complementary tasks: automatic input-output prediction over 300 Python functions and manually assessed semantic comprehension over a balanced subset of 60 functions. The functions are grouped into Low-, Medium-, and High-complexity bands. DeepSeek-Coder-V2 achieves an overall automatic accuracy of 78.33%, compared with 70.33% for Llama. However, accuracy decreases substantially from Low to High complexity, from 93.52% to 52.78% for DeepSeek-Coder-V2 and from 87.04% to 47.22% for Llama. Incorrect predictions are consistently associated with higher values of all four complexity metrics, and correlation and logistic-regression analyses confirm broadly comparable negative associations between structural complexity and correctness. Manual semantic comprehension shows the same degradation pattern, with accuracy decreasing from 100.00% to 75.00% for DeepSeek-Coder-V2 and from 90.00% to 60.00% for Llama. These findings demonstrate that complexity-aware evaluation provides a more diagnostic assessment of LLM code-comprehension reliability than aggregate accuracy alone.

[AI-78] Routing Should Pay for Itself: Sparse Supervision for Economical LLM Routing

链接: https://arxiv.org/abs/2609.37402
作者: Guannan Lai,Gelin Bian,Hao-Xuan Ma,Jun-Peng Jiang,Long Chen,Jian-Dong Liu,Zhi-Hao Tan,Han-Jia Ye
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language model (LLM) routing reduces serving cost by assigning each query to an appropriate model while preserving response quality. Learning such a router, however, often requires executing multiple candidate models on historical queries to collect query–model quality feedback, creating a nontrivial supervision cost before deployment. Existing work largely focuses on serving-time efficiency, overlooking whether the resulting savings are sufficient to recover this upfront expenditure. We further observe that routing quality often saturates well before all query–model feedback is collected, suggesting that dense supervision can be economically over-provisioned. We propose SaveRouter, a sparse-supervision routing framework that selectively acquires informative model feedback and shares capability information across related queries, while retaining query-level refinement for fine-grained routing. We evaluate routing by jointly accounting for supervision expenditure and subsequent serving-time savings. Across four routing benchmarks, the main setting uses only about 33–41% of available training feedback while maintaining competitive or better routing quality, and reduces the break-even deployment volume by approximately 1.9–9.5 times compared with the fastest conventional router. Further analysis shows that acquiring more supervision is not always economically preferable: the supervision level that minimizes serving cost can differ from the one that achieves the earliest payback. Our code is publicly available at this https URL.

[AI-79] Direct Experience World-Model Optimization: Learning the World Beyond Action Imitation

链接: https://arxiv.org/abs/2609.37398
作者: Xiangcheng Zhan,Zirui Chen,Yicheng Zhao,Ziteng Gao,Shuo Yang
类目: Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注: World Action Model; post-deployment training; dexterous manipulation

点击查看摘要

Abstract:World-Action Models (WAMs) couple action generation with predictions of how physical interactions unfold. However, current post-deployment learning paradigms typically improve behavior without requiring better world predictions. Especially in dexterous manipulation, small execution errors can compound in high-dimensional action spaces, hindering policy improvement and pushing interactions beyond the world model’s training distribution. Motivated by this, we propose Direct Experience World-Model Optimization (DEWO), a post-deployment learning paradigm for WAMs that, alongside action imitation, refines world representations through visual experience to better condition action generation. Specifically, it identifies interaction turning points and learns from successful and failed futures to support classifier-free guidance. An additional value head estimates task progress from video representations and activates guidance when progress stalls during inference. Across five DexJoCo tasks, DEWO improves average success across all three WAM formulations. Ablations show that visual supervision from successful and failed continuations improves both prediction and control beyond action supervision alone. On four real-world tasks across Wuji and Sharpa, 3 x 3 grid evaluations show that two rounds of deployment learning increase success from 51.0% to 71.7% in cells with at least one initial success, a gain of 20.7 percentage points. These findings support continued predictive learning for improving control through deployment experience, making world modeling an active part of WAM adaptation.

[AI-80] Beyond Prompt Count: How Data Shapes Transfer in On-Policy Distillation

链接: https://arxiv.org/abs/2609.37377
作者: Jiaxuan Wang,Jiafei Lyu,Yuchen Cai,Siye Wu,Pengyuan Wang,Jiashun Liu,Xiang Cheng,Kai Yang,Yangkun Chen,Saiyong Yang,Lan-Zhe Guo
类目: Artificial Intelligence (cs.AI)
备注: 39 pages. Code: this https URL

点击查看摘要

Abstract:On-policy distillation (OPD) trains students using teacher feedback on their own sampled responses, yet how prompt choice shapes transfer across teacher-student pairs remains poorly understood. We systematically study prompt quantity, source, and selection across RL- and SFT-continuation pairs and cross-model settings. We find that OPD can be highly prompt-efficient: a few prompts can approach large-pool performance, with four DAPO prompts matching the observed mathematics score of 3,840 DeepMath prompts. However, prompt utility is relational rather than intrinsic: changing only the teacher can reverse the relative effectiveness of mathematics and code prompts. To characterize these transfer differences, we analyze parameter and functional changes across prompt supports and model pairs. Functional alignment with the teacher varies across supports and target tasks; in continuation pairs, teacher-aligned prediction changes can coexist with weak parameter alignment. Continued OPD on effective supports can restore performance after unfavorable transfer. Finally, targeted selection does not consistently outperform uniform random sampling, and filtering out a source that performs poorly alone yields no consistent gain across three paired support draws. Overall, our results distinguish prompt efficiency from prompt interchangeability and show that effective data choice depends on the teacher-student pair and target capability, with random sampling providing a competitive baseline in the studied settings.

[AI-81] Pretrain Once Route Anywhere: Towards a Foundation Model for LLM Routing

链接: https://arxiv.org/abs/2609.37362
作者: Guannan Lai,Han-Jia Ye
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language model (LLM) routing aims to assign each query to the most suitable model from a heterogeneous candidate pool, improving the quality–efficiency trade-off of LLM inference. Existing routers are typically learned through local fitting: a router is optimized for a particular query workload and candidate pool, and often requires additional supervision or retraining as the routing environment changes. We ask whether LLM routing can instead be approached from a foundation-model perspective, learning a reusable routing capability that generalizes across tasks, candidate models, and deployment conditions. To this end, we introduce RouteFM, which learns to characterize anonymous candidate models from behavioral context and infer their target-specific capabilities, rather than binding routing decisions to fixed model identities or a single environment. Through episodic pretraining across heterogeneous routing environments, this capability can be reused by a frozen router and adapted to new environments through context alone. Experiments demonstrate transfer across changes in domains, modalities, candidate pools, and context budgets, with the largest gains when behavioral evidence is limited. On MMR-Bench, which is excluded from pretraining, RouteFM outperforms the strongest baseline by 2.23 quality points with only eight observations per candidate. These results support moving LLM routing from repeated local fitting toward a pretrain once, route anywhere paradigm. Our code is publicly available at this https URL.

[AI-82] aching LLM s to Generate Challenging MILP Instances via Solver Feedback

链接: https://arxiv.org/abs/2609.37356
作者: Jitin Singla,Parikshit Pareek,Pratik Jawanpuria,Parag Singla
类目: Artificial Intelligence (cs.AI); Optimization and Control (math.OC)
备注:

点击查看摘要

Abstract:Generating optimization instances that are both feasible and computationally challenging is crucial for benchmarking solvers and training learning-based optimization algorithms. Existing non-LLM generators rely on seed instances or parameter tuning, resulting in high test-time computational cost, while existing LLM generators lack explicit hardness measures. Recent reinforcement learning methods with verifier feedback evaluate only binary correctness, which is misaligned with generating challenging problems. We note that an optimization solver reports the cost of solving at several stages of its pipeline, and leverage this to design a reward that scores both the solvability and the hardness of generated problems, measured by branch-and-bound nodes and post-cut relaxation gaps. Our key idea is a challenger-solver asymmetric self-play approach, where an LLM challenger generates progressively harder instances and the solver verifies feasibility and hardness, so no seed or training MILP instances are required. We fine-tune Gemma-4-12B and Qwen3.5-4B with GRPO and a size curriculum into OptiScribe-12B and OptiScribe-4B, which generate feasible yet challenging MILP problems from natural language instructions. On capacitated facility location and max-cut, OptiScribe-12B raises median SCIP search nodes by 1.7-5x and post-cut gaps by 1.1-1.7x over its base model and improves the feasibility rate on facility location by 9-19 points, while OptiScribe-4B raises median nodes by up to 15.6x. The problems cover a wider difficulty range than public benchmarks of the same size, follow instructions on density and difficulty, and can tune solver settings for families that public libraries lack. These results indicate that optimization-specific rewards, used in self-play mode, can teach LLMs to generate high-difficulty optimization benchmarks. We will release our code and models publicly on acceptance.

[AI-83] Seek Before You Move: Evidence Seeking for Progress Grounding in Vision-Language Navigation

链接: https://arxiv.org/abs/2609.37353
作者: Zhimin Wang,Meiyuan Zhu,Duo Wu,Linjia Kang,Yajun Wang,Yuan Ni,Xiaohang Wang,Tianlu Pan,Jingyan Jiang,Yaowei Wang,Zhi Wang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Vision-Language Navigation (VLN) requires agents to continuously ground task progress from long-horizon instructions and partial egocentric observations. Existing VLM-based navigation agents typically reason only over available observations and may remain confident even when task-relevant evidence is missing. For example, an agent may confidently proceed forward and get lost even though the landmark indicating the next turn lies outside its current field of view. We term this failure mode Progress Myopia: the agent fails to recognize unreliable progress grounding and continues acting on insufficient evidence. To address it, we propose SeekVLN, an evidence-seeking framework that couples semantic progress reasoning with active acquisition of task-relevant observations. SeekVLN is trained in two stages: First, Future-guided Reverse Generation (FRG) uses future expert actions to augment offline expert trajectories with supplementary views and evidence annotations. Supervised fine-tuning on these trajectories establishes a prior for evidence seeking and progress reasoning without additional expert interaction. However, imitation alone does not reveal whether seeking improves subsequent navigation. We therefore introduce Counterfactual Contrastive Policy Optimization (C2PO) for reinforcement fine-tuning. By comparing each evidence-seeking branch with a counterfactual direct-navigation branch from the same state, C2PO uses a contrastive reward to assign credit to seeking decisions based on subsequent navigation benefit. Experiments on simulated benchmarks show that SeekVLN achieves state-of-the-art performance, improving success rate by 12.7% and 7.5% over the base model on R2R-CE and RxR-CE, respectively. Both simulated and real-world evaluations exhibit human-like evidence-seeking behaviors for more reliable progress grounding.

[AI-84] A Sharp Transition in Data Reconstruction under Differential Privacy

链接: https://arxiv.org/abs/2609.37344
作者: Max Cairney-Leeming,Simone Bombari,Marco Mondelli
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
备注:

点击查看摘要

Abstract:Data reconstruction attacks have empirically been successful in recovering training samples from learned models, raising privacy concerns and motivating defenses with guarantees that remain valid against future threats. While differential privacy (DP) provides formal protection, choosing the privacy budget remains a challenge: small budgets severely reduce utility, but it is hard to quantify how large the budget can be without allowing accurate reconstruction. In this work, we study informed attackers who aim to reconstruct a single d -dimensional training sample from a \rho -zero-concentrated DP model, knowing all other training data. Our main contribution is to establish a sharp transition at \rho \asymp d for data reconstruction: on the one hand, we derive entropy-based lower bounds for any private mechanism and any attack, characterizing a set of target priors for which reconstruction is information-theoretically impossible for \rho \ll d ; on the other hand, we analyze a simple attack on private linear regression with output perturbation, showing that reconstruction is practically feasible for \rho \gg d . Remarkably, the transition moves to \rho \asymp s for data lying in an s -dimensional subspace, demonstrating that the privacy budget guaranteeing adequate protection must be assessed in terms of the effective dimension of the data. We validate our findings via experiments on synthetic data and natural images (CIFAR-10, ImageNet).

[AI-85] VISTA: Value-Informed Event Appraisal for Multimodal Emotion Conflict

链接: https://arxiv.org/abs/2609.37324
作者: Jiale Dai,Liuxian Ma,Xiaoke Niu,Wenjing Zhang,Huiying Zhao,Zhaoxiang Liu,Shiguo Lian,Guojie Song
类目: Artificial Intelligence (cs.AI)
备注: 46 pages, 13 figures, 37 tables, including appendices

点击查看摘要

Abstract:Conflicting emotional cues can be individually valid: a subdued voice may reflect a blocked goal while a smile satisfies a social obligation. Their interpretation depends on what the event means to the person. We introduce VISTA (Value-Informed Semantic Trust Arbitration), a learned seven-field appraisal interface that conditions modality arbitration on concerns, event relations, and expression conditions while retaining a joint-evidence residual. A log-odds decomposition separates emotion expectation from cue diagnosticity, motivating an interface that lets appraisal change how evidence is interpreted. With a shared Qwen2.5-Omni-7B backbone and matched training examples and steps, VISTA reaches 64.5% conflict accuracy on CA-MER, improving on modality gating by 2.5 percentage points on conflict and 0.2 on consistency. Shuffling appraisal across scenes or removing its decision connection reduces this benefit. A common frozen-backbone probe reaches 0.600 macro CCC for appraisal readout, compared with 0.505 for emotion-only fine-tuning. Evaluations across five benchmarks connect recognition under increasing conflict with appraisal readout and downstream decision use. Together, the analyses and experiments support scene-specific appraisal as an intermediate representation that helps interpret conflicting emotional evidence.

[AI-86] Mubric: Mutation Testing-Guided Rubric Generation for LLM Evaluation

链接: https://arxiv.org/abs/2609.37322
作者: Jiayuxuan Yang,Jie M. Zhang,Yiling Lou,Zhenpeng Chen
类目: Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注:

点击查看摘要

Abstract:Rubric-based evaluation is widely used to assess LLM-based systems by decomposing response quality into task-specific scoring criteria. However, automatically generating rubrics that reliably capture task-specific quality requirements remains challenging. We introduce Mubric, a mutation testing-guided approach to rubric generation. Mutation testing, a classic software testing methodology, evaluates a test suite by injecting faults into programs and checking whether the tests detect them. We draw an analogy between test suites and rubrics: if a rubric captures an important quality requirement, introducing a corresponding defect into an otherwise high-quality response should reduce its score. Mubric first mines common defects from real pairs of preferred and dispreferred responses and abstracts these defects into reusable mutation operators, each specifying how to introduce a particular type of response defect. For a new task, it applies relevant operators to a reference response, checks whether the injected defects reduce response quality, and uses insufficiently penalized defects to refine the rubric. We evaluate Mubric on 703 tasks across four representative domains against six advanced rubric generation methods. Mubric achieves the highest overall evaluation accuracy, outperforming the strongest baseline by 7.48 percentage points.

[AI-87] Do Agent Benchmarks Do What They Say? An Executable-Contract Audit of Tool-Using Agent Environments

链接: https://arxiv.org/abs/2609.37315
作者: Rohith Reddy Bellibatlu,Zichong Wang,Wenbin Zhang
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 10 pages. Submitted to IEEE BigData 2026, Intelligent Data Mining special session. Code, contracts and data: this https URL

点击查看摘要

Abstract:Tool-using agents are entering settings where a wrong action carries real cost, and the benchmarks certifying them grade what each simulated tool call reports having done, assuming the tool did what its interface advertises. The audit taxonomies we survey publish no category for that assumption, and a defect beneath a score is present on every rerun. We treat a tool’s advertised surfaces as an executable contract, check the implementation against it, and trace each score’s provenance through the task files and evaluator code to the verdicts that derive from state a defective tool should have written. Across 34 audited mutating tools in four benchmarks we confirm seven tool defects and one evaluator property at pinned commits. On injected defects the checker raised no false positive in 25 flags, flagged 2 of 5 negative controls, and missed most: in 29 of 33 scored misses a clause covered the defect but no probe revealed it. The checker’s own static half, run alone, flags 14 of 17 confirmed sites, so on these findings the dynamic half confirms and traces rather than discovers. Twelve further AgentDojo tools, with six held-out tools and the seven audited first, complete its 25-tool mutating surface, on which at least 5 tools diverge from their advertised surface as our contracts read it, a rate for AgentDojo alone. No gold trajectory reaches either tau2-bench defect; on 1,120 paths built to isolate the telecom defect, a number fixed by construction, the evaluator rewards a refuel of a suspended line and fails the repaired tool. The clearest case is a clinical benchmark whose tool tells the agent each write executed under a documented no-write design its interface does not disclose; its grader takes that message as evidence, so its action success rate records whether a request carried the expected payload, not whether any record changed.

[AI-88] MetaCtrl: Your Large Language Models Can Reason Better and More Concisely with a Metacognitive Controller

链接: https://arxiv.org/abs/2609.37304
作者: Zhibin Wen,Tao Han,Lei Bai,Can Li,Yang Xu
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Large reasoning models improve performance on challenging problems by allocating additional computation before answering, but longer reasoning does not always lead to better results and can introduce substantial redundant reasoning on simple problems. Conversely, aggressively shortening reasoning can degrade performance on difficult ones. Effective reasoning therefore requires dynamically deciding when additional computation is useful based on the reasoner’s capabilities and evolving solution state. Existing approaches often rely on predefined budgets or intervention rules, retrain the target reasoner, or require additional supervision. We introduce MetaCtrl, a lightweight controller that adaptively regulates a frozen reasoner without predefined token budgets or reasoner retraining. We formulate reasoning regulation as a sequential metacognitive control problem: MetaCtrl observes the evolving reasoning trace and decides whether to continue, simplify, skip redundant steps, or conclude reasoning. It is trained directly with reinforcement learning using a reward that prioritizes correctness while favoring shorter trajectories among correct solutions, requiring neither supervised intervention trajectories nor problem-specific budgets. Across seven benchmarks spanning mathematics, science, and code, MetaCtrl consistently improves the accuracy of LRMs while reducing their reasoning length. On DeepSeek-R1-Distill-Qwen-7B, it improves average accuracy by 4.7 points while reducing generation length by 53.3%. Without further training, the same controller transfers to an unseen reasoner (e.g., Qwen3-14B), improving average accuracy by 2.9 points and reducing generation length by 50.3%. These results establish MetaCtrl as a plug-and-play controller for improving reasoning accuracy while substantially reducing inference-time generation. The code is available at this https URL.

[AI-89] AssayRouter: Historical Utility Priors for Frozen Molecular Predictor Routing

链接: https://arxiv.org/abs/2609.37285
作者: Dong Xu,Zhangfan Yang,Jiantao Wu,Shipeng Zhang,Zexuan Zhu,Jiangqiang Li,Jun Zhang,Junkai Ji
类目: Artificial Intelligence (cs.AI); Biomolecules (q-bio.BM); Quantitative Methods (q-bio.QM)
备注: 26 pages, 6 figures

点击查看摘要

Abstract:Laboratories often face a new molecular assay with 16-64 labels and a bank of predictors whose training data and parameters are unavailable. The practical question is which frozen outputs to include in a small local model. AssayRouter treats completed assays as pseudo-targets and labels each candidate by its post-fit utility: the reduction in held-out discovery loss when the candidate is added to the local target predictor. A shared regressor learns to predict this utility from candidate behavior on the support set, without source identity; on a new assay, one frozen ranking selects four sources and separate labels fit a convex combiner. We train only on completed ChEMBL-MT assays and evaluate 24 external regression assays across six frozen interface families. AssayRouter-C lowers strict four-call negative log-likelihood (NLL) by 0.0409 relative to Support-CV@4. Frozen candidate-label permutations confirm that candidate-utility correspondence carries the transferred information, and leave-one-interface-out training shows that the mapping generalizes to unseen predictor families. Completed assays therefore provide transferable supervision for scarce-label routing through frozen prediction interfaces.

[AI-90] ransolver-σ: Joint Spectral-Physical Subspace Modeling for Neural PDE Solving

链接: https://arxiv.org/abs/2609.37279
作者: Haonan Shangguan,Hang Zhou,Haixu Wu,Yuezhou Ma,Jianmin Wang,Mingsheng Long
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Neural solvers offer efficient surrogates for numerical simulation of partial differential equations (PDEs). For time-dependent problems, strong one-step accuracy does not necessarily translate into reliable autoregressive rollout. We observe that a solver based only on physical-state modeling can achieve lower one-step error, whereas its spectral-only counterpart can become more accurate at later rollout steps. Motivated by this observation, we present Transolver- \sigma , a neural PDE solver based on joint spectral–physical subspace modeling. Within each block, adaptive physical-state interactions and spectral transformations are modeled in dedicated latent subspaces, whose responses are recomposed to enable information exchange between the two representations. Within the physical subspace, we introduce Slice-Residual Physics-Attention (SRPA), which preserves an explicit slice-space identity path while retaining learnable cross-slice interaction. In parallel, an axis-factorized Fourier operator captures global spectral structure. Across five well-established PDE benchmarks spanning steady-state prediction and time-dependent dynamics, Transolver- \sigma achieves state-of-the-art with a benchmark-averaged relative error reduction of 33.4% over the strongest baseline for each metric, while consistently improving autoregressive rollout over single-operator counterparts. Transolver- \sigma further delivers strong gains on coupled multiphysics systems and real-world fluid and combustion measurements from RealPDEBench, demonstrating its effectiveness beyond standard simulation benchmarks.

[AI-91] ask-Relevant Null-Space Residuals for Non-Injective Neural Mappings

链接: https://arxiv.org/abs/2609.37272
作者: Bizu Feng,Zhimu Yang,Shuming Wang,Yuan Cheng,Shaode Yu,Xiaojun Qian,Zixin Hu
类目: Artificial Intelligence (cs.AI)
备注: 25 pages

点击查看摘要

Abstract:Non-injective mappings in neural networks map distinct inputs to the same representation, thereby implicitly inducing equivalence relations in the input space. However, the input differences eliminated by these mappings may still be required by downstream tasks, creating a mismatch between operator-induced indistinguishability and task-required distinctions. For non-injective linear operators realized in the current forward pass, their null spaces exactly characterize these invisible input variations. We propose Task-Relevant Null-Space Residuals (NSR), a general residual framework for non-injective linear mappings. NSR combines null-space component extraction from pre-mapping representations, member-level encoding and gating, and application-specific integration to exploit potentially task-relevant information under downstream supervision while preserving the original aggregation or merging rules. We evaluate NSR in two structurally different settings: token merging and graph aggregation. In token merging, NSR achieves higher semantic segmentation performance than the corresponding compressed baselines in 34 out of 36 evaluated configurations, with a maximum observed gain of 31.51 mIoU points under strong compression. In graph aggregation, NSR achieves 100% training accuracy on Tree-NeighborsMatch at depths d=2–6 across three backbones, alongside gains on heterophilic node classification and molecular graph regression. Together, these results support null-space residuals as a practical complement to non-injective linear mappings, enabling downstream models to learn from input distinctions invisible in the original operator’s output.

[AI-92] Foundations of Proactive Agents : Principles Technical Layers and Proactivity-Gym

链接: https://arxiv.org/abs/2609.37267
作者: Jio Oh,Seunghyun Do,Young-Jun Lee,Steven Euijong Whang,Dongyeop Kang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Proactive LLM agents can turn idle compute into useful support before users ask. Yet even correct work can misread user context, impose review costs, or undermine trust. This work proposes foundations for designing, realizing, and evaluating proactive LLM agents around three joint principles (3T): Task Capability, anticipating relevant needs and correctly performing useful work; Temporal Allocation, allocating compute according to resource availability and when results are needed; and Trust, sustaining users’ confidence and appropriate reliance on the agent. We connect these objectives to a design space organized around five dimensions: task scope, anticipation horizon, activation trigger, processing timing, and intervention depth, and specify the situation and system modeling needed to support its choices, including user and environment representations, backbone LLMs, and agent harnesses. Lastly, we propose PROACTIVITY-GYM, a simulation-based evaluation testbed including multi-day scenarios, stateful environments, and persona-conditioned simulated users that can evaluate the consequences of proactive assistance across interactions. Evaluations across 23 model-harness configurations uncover substantial performance gaps across 3T and reveal that LLM judges often conflate task capability and trust. A human study with 30 participants demonstrates the importance of the joint 3T optimization: participants show sharp trust declines after intervention misalignment despite correct outcomes, and prefer sleep-time assistance, even when imperfect, to preserve ongoing focus. Together, these findings support designing and evaluating proactive agents through the joint consideration of useful work, compute allocation, and evolving user trust.

[AI-93] Loss-Guided Pretraining Data Selection for Time-Series Foundation Models

链接: https://arxiv.org/abs/2609.37255
作者: Yike Li,Shaoxu Song,Jianmin Wang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 15 pages, 5 figures

点击查看摘要

Abstract:Time series foundation models (TSFMs) are pretrained on heterogeneous collections containing billions of observations, yet their training windows are typically sampled without estimating whether they provide useful learning signal. We introduce a static data-selection framework that scores each window with a reference forecaster and retains an intermediate interval within every source dataset. Specifically, we connect forecasting loss to optimization difficulty by showing that normalized squared loss controls the per-sample gradient norm under a local Jacobian condition. We then define a reference loss score and apply dataset-stratified selection to preserve the diversity of samples. Across various TSFM architectures, our method outperforms random selection by an absolute margin and even improves both relative MASE and CRPS over full-data pretraining by retaining fewer candidate pretraining windows. Further analyses show strong cross-scale and cross-architecture score correlations, indicating that a small reference model can often select data for larger targets, provided that the reference and target share compatible difficulty orderings.

[AI-94] DatalogBench: Evaluating Large Language Models on Text-to-Datalog Synthesis

链接: https://arxiv.org/abs/2609.37233
作者: Yuan Li,Hanyun Jiang,Guowei Tian,Chengpeng Wang,Peisen Yao
类目: Programming Languages (cs.PL); Artificial Intelligence (cs.AI)
备注: 33 pages

点击查看摘要

Abstract:Datalog underpins reasoning tasks such as program analysis, but its programs are hard to write. Existing synthesizers automate this task but require users to state their intent as input-output examples. Large language models (LLMs) suggest a more natural route, text-to-Datalog synthesis from a natural-language question, yet how well they do so has not been systematically evaluated. We present DatalogBench, a benchmark of 136 text-to-Datalog synthesis tasks curated from existing Datalog-based artifacts. Synthesized programs are graded by execution on held-out inputs against an oracle validated by mutation analysis. Across six LLMs and four prompting configurations, exact match peaks at 68.4%, and relation descriptions or an input-output example have only modest, model-dependent effects. Under direct prompting, most failures occur at compile time, typically because a model invents auxiliary predicates that it never declares or types consistently. Two coding agents reach up to 83.8% and eliminate nearly all such failures, leaving mostly semantic errors concentrated in recursive tasks. DatalogBench thus identifies recursive reasoning and decomposition as open challenges for current LLMs and agents, and offers a reliable, execution-grounded measure of both.

[AI-95] OptiCom : A Unified Framework for State-Conditioned Composition in LLM -Driven Optimization

链接: https://arxiv.org/abs/2609.37221
作者: Chenxing Wei,Sichen Liu,Lizhao Liu,Ningyuan Sun,Chen Bingzhou,Ying He,Bo Jiang,Fei Yu,Yao Shu
类目: Artificial Intelligence (cs.AI)
备注: 43 pages, 11 figures

点击查看摘要

Abstract:Large language models (LLMs) are increasingly deployed to solve complex scientific and practical problems via iterative optimization. However, dynamically coordinating diverse search mechanisms as candidate quality, failure modes, and resource budgets evolve remains a critical open challenge. Targeted empirical diagnostics reveal that mechanism effectiveness is highly state-dependent. Motivated by this, we analyze how individual decisions drive final outcomes, decomposing the expected terminal improvement under a shared budget into cumulative decision opportunities minus cumulative selection losses. Guided by this opportunity-loss theoretical foundation, we propose OptiCom, a unified framework that represents LLM-driven optimizers within a shared configuration space: C=(A,Q,O,E,M,S), corresponding to artifact, query, operator, evaluation, memory, and strategy. Operating within this space, a fast LLM-based Optimization Controller dynamically composes immediate mechanisms through structured Action Packages, while a slower Strategy Adapter refines long-term selection preferences, operator weights, and templates based on accumulated trajectory feedback. Comprehensive evaluations across 32 benchmark groups demonstrate the superiority of framework: OptiCom achieves an average Max-score rank of 1.72 among 14 evaluated configurations, securing the top score in 23 groups. Ultimately, these results highlight the broad applicability and high extensibility of OptiCom as a general-purpose paradigm for robust LLM test-time scaling.

[AI-96] Information Bottleneck-Guided Adaptive Hypergraph Transformer for Brain Disease Diagnosis NEURIPS2026

链接: https://arxiv.org/abs/2609.37220
作者: Jingxi Feng,Xudong Chen,Yifan Zhang,Heming Xu,Hongcheng Han,Xijing Wang,Dong Zhang,Shaoyi Du
类目: Artificial Intelligence (cs.AI)
备注: Accepted by Neurips 2026

点击查看摘要

Abstract:Exploring high-order correlations and long-range dependencies in brain networks holds significant value for both neuroscience research and clinical diagnosis. However, previous studies have lacked a unified integration of high-order and long-range dependency information in brain networks, and there is substantial redundancy behind various types of information. These issues limit their effectiveness in the diagnosis of brain diseases. To address this, we propose an Information Bottleneck-Guided Adaptive HyperGraph Transformer (IBAHGT). By incorporating the information bottleneck (IB) principle, this approach enables adaptive learning of high-order correlations and both short- and long-range dependencies within a unified framework for brain network analysis, achieving high-precision brain disease diagnosis. IBAHGT consists of three key components: an information bottleneck-guided adaptive hypergraph convolution, which introduces a novel hypergraph information bottleneck (HIB) principle to adaptively learn hypergraph message-passing weights between nodes and hyperedges, optimizes information flow and captures high-order information in brain networks that is maximally informative and minimally redundant (MIMR). The Transformer encoder captures global information within brain networks through the attention mechanism, specifically modeling short- and long-range dependencies. An information bottleneck-guided node-level adaptive fusion employs the IB principle to learn independent weights for each node, facilitating the fine-grained integration of high-order information and global information to obtain an efficient representation for downstream tasks. Extensive experiments demonstrate that the proposed method outperforms current state-of-the-art methods and can identify biomarkers for clinical applications.

[AI-97] Beyond Semantic Narrowing: Robust and Efficient LLM Watermarking with Hamming Neighborhoods

链接: https://arxiv.org/abs/2609.37218
作者: Zewen Sun,Tongyang Zhao,Liyao Xiang,Mingxuan Ma,Lingzhe Wang,Zhiyuan Li
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 30pages

点击查看摘要

Abstract:Semantic watermarking improves robustness against watermark removal attacks by embedding detectable signals into sentence-level representations. However, existing watermarking methods typically impose watermark-specific semantic preferences on generated sentences without explicitly accounting for the highly non-uniform and context-dependent semantic preference of LLM generation. When these two preferences are poorly aligned, many natural continuations become incompatible with the watermark, causing semantic narrowing: reduced semantic freedom, increased resampling cost, and potential degradation on tasks with strict semantic requirements. To alleviate this problem, we propose HammingMark, which uses the semantic hash of the preceding sentence as a dynamic center and accepts candidates whose hashes fall within its Hamming neighborhood. Defining watermark validity over a Hamming neighborhood in compact hash space retains a larger fraction of naturally likely semantic continuations. The coarse many-to-one hash mapping further allows diverse semantic realizations to remain watermark-valid. Experiments on C4 and BookSum show that HammingMark achieves strong robustness, high detectability, and near-unwatermarked generation quality, requiring only 2.2 sampled candidates per accepted sentence,a 72.8% reduction compared with the most sampling-efficient existing method. On more complex tasks with strict semantic constraints, HammingMark achieves the highest detection rates with the highest or tied-highest ROUGE-L scores, demonstrating its effectiveness in balancing watermark detectability and generation quality under constrained generation settings.

[AI-98] CRJudgeBench: Can AI Detect Plausible but Invalid Code Reviews? ICLR2027

链接: https://arxiv.org/abs/2609.37216
作者: Yue Pan,Jiawei Li,Ziyuan Zhang,Xiangxin Zhao,He Ye
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注: 26 pages, 5 figures, and 14 tables. Under review at ICLR 2027. Dataset available at this https URL

点击查看摘要

Abstract:Large language models can generate plausible code-review comments, but such comments may contain technically incorrect claims that mislead developers. We study technical trustworthiness judgment: determining whether a review comment’s core technical claims are correct and applicable to the reviewed code in its repository context. Existing code-review benchmarks primarily evaluate review generation, issue discovery, or general comment quality, but do not directly assess whether an agent can determine the technical trustworthy of an individual review comment. To fill this gap, we introduce CRJudgeBench, a benchmark of 1199 instances constructed from real pull requests and expert-verified perturbations, covering both trustworthy and plausible but untrustworthy comments. We further present Sentinel, a repository-grounded agentic judge that actively gathers code evidence to verify review comments before making judgments. Starting from Qwen3-Coder-30B-A3B-Instruct, Sentinel is trained on the CRJudgeBench training split through iterative action-level learning from a privileged teacher. On the 359-instance CRJudgeBench test set, Sentinel achieves 76.60% accuracy, outperforming GLM-5.3 by 6.13 percentage points and its base model by 19.78 points. These results show that even state-of-the-art general-purpose LLMs struggle to identify untrustworthy comments, while iterative action-level learning substantially improves the accuracy of repository-grounded trustworthiness judgments. Our dataset is available at this https URL

[AI-99] Learning to Prove Not Just to Answer: Reinforcement Learning from Formal Verification for Natural-Language Logical Reasoning

链接: https://arxiv.org/abs/2609.37203
作者: Qili Zhang,Qianren Mao,Hanze Cai,Kaiming Zhao,Yuening He,Xihan Lei,Yashuo Luo,Hanwen Hao,Yutong Gu,Likang Xiao,Zhijun Chen,Weifeng Jiang,Haoyi Zhou,Jianxin Li
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models (LLMs) are increasingly deployed for natural-language logical reasoning, where the final answer is easy to check but the proof behind it is not. In natural-language logical reasoning, an intermediate conclusion should follow from its premises, and the resulting derivation should support the final answer. Existing methods lack machine-checkable verification of intermediate conclusions and answer-supporting proof dependencies, so they may assign credit to invalid or answer-irrelevant steps. We propose Proof-R1, an RL framework from formal verification that trains LLMs to construct verifiable proofs for natural-language logical reasoning. Proof-R1 admits a generated conclusion into the verified proof state only when the corresponding reasoning action satisfies the proof obligations through UNSAT-based machine-checkable formal verification. Proof-R1 also recovers the answer-supporting dependency closure to trace the proof structure of the final answer and align outcome credit with the proof dependencies. Experiments demonstrate that Proof-R1 improves answer accuracy across three logical reasoning benchmarks and four backbone models and outperforms training-free agents and training-based methods in terms of reasoning-process verifiability.

[AI-100] V-Engram: Trigger-Indexed External Memory for Modular Text-to-Image Personalization

链接: https://arxiv.org/abs/2609.37198
作者: Haoran He,Runyuan Cai,Yiming Wang,Lin Yu,Xiaodong Zeng
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Pretrained text-to-image models contain broad visual knowledge, yet they cannot reliably acquire or refine a specific visual identity from only a few references while preserving compositional control. Token-embedding methods are compact but often underfit identity, whereas adapter-based methods improve fidelity through persistent weight updates that can be costly to store and interfere when concepts are composed. We introduce V-Engram, a trigger-indexed external memory mechanism for Stable Diffusion 3.5. Each concept is assigned an explicit trigger that retrieves concept-specific memory, whose gated directions enter frozen text-encoder and MMDiT context states as relative residuals. Separating this memory from backbone adaptation enables prompt-selective and multi-concept access without merging model updates. Experiments show that V-Engram broadly matches DreamBooth-LoRA in overall subject fidelity while showing advantages in settings such as contextual subject preservation. Prompt-matched loading retrieves only matched entries, reducing most additional adaptation-state loading for a single-concept query. Qualitative results further demonstrate paired-trigger composition and same-class separation, while prompts without registered entries retain the frozen model’s base behavior. Together, these results establish trigger-indexed memory as a modular interface for adding targeted visual evidence without rewriting the generator.

[AI-101] oolFence: Fine-Grained Authorization for Secure Tool-Using LLM Agents

链接: https://arxiv.org/abs/2609.37196
作者: Yanjie Li,Xiangyu He,Xuelong Dai,Bin Xiao
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 9 pages, 4 figures

点击查看摘要

Abstract:Tool-using LLM agents remain vulnerable to indirect prompt injection because trusted instructions and untrusted observations share one context, allowing malicious content to steer consequential input-filtering defenses. Multi-path consensus defenses still leave a high attack success rate because they examine content or aggregated outputs rather than authorizing effects, especially for the within-tool attack, which preserves the intended tool but manipulates its arguments. Data-Flow Control such as CaMeL provides stronger guarantees, but incurs substantial time latency that limits practical deployment. We introduce ToolFence, which compiles a typed authorization blueprint before execution, enforces it through a deterministic monitor, and when the blueprint is incomplete asks a judge to grant new capabilities rather than adjudicate each concrete call. ToolFence provides two key advantages. First, its fine-grained provenance-aware authorization enables the system to distinguish user-authorized values from untrusted observations, effectively addressing the within-tool attack. Second, its deterministic fast path and capability-level runtime grants substantially reduce the frequency of expensive judge calls, improving runtime efficiency. On AgentDojo with Qwen3-max, ToolFence reduces overall ASR to near zero with only a 3.80 percentage-point clean-utility drop and practical runtime overhead.

[AI-102] Accelerated surrogate dynamics for dynamical stochastic system evolution

链接: https://arxiv.org/abs/2609.37184
作者: Marco Jochum,Ioannis Kouroudis,Gohar Ali Siddiqui,Taher Amine Hamzaoui,Manuel Gößwein,Alessio Gagliardi
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Dynamic simulations are an entrenched way of gaining insight into the evolution of system dynamics. Their computational cost however is often prohibitively high, especially in cases of stochastic frameworks. Machine learning algorithms are especially suited as simulation surrogates. Nevertheless, they face some very distinct limitations. Firstly, the sheer dimensionality of these systems, however, precludes the use of traditional time series models who struggle with high dimensional feature spaces. Additionally, traditional time series focus exclusively on either long or short range effects, causing local or global drift given enough time. In this paper, we propose a framework that addresses those limitations. Our framework combines a Variational Autoencoder, with a convolutional or graph basis that reduces the dimensionality of the system. This latent vector is propagated in time using a Temporal Fusion Transformer model, which includes both long range and short range effect encoding, as well as static covariate support. We test our framework on three distinct cases, to prove its robustness and in all three we have achieved practically identical to the simulation results at a fraction of the time. Further, our framework is flexible enough to be adapted to any new system and provides an inbuilt uncertainty quantification for targeted experiment design.

[AI-103] EgoHumanoid-V2: Human-to-Humanoid Transfer of Coordinated Whole-Body Skills for Loco-Manipulation

链接: https://arxiv.org/abs/2609.37181
作者: Jin Chen,Yiming Jiang,Chongyang Xu,Modi Shi,Shijia Peng,Li Chen,Tianyu Li,Mu Xu,Yilun Chen,Steven Hoi,Hongyang Li
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Human demonstrations capture diverse scenes and rich whole-body skills without requiring robot teleoperation. Prior work on egocentric transfer has emphasized scene generalization in loco-manipulation under decoupled control, leaving direct transfer of coordinated whole-body skills less explored. We present EgoHumanoid-V2, the first egocentric human-to-humanoid skill transfer framework for coordinated whole-body loco-manipulation. At its core, coarse-to-fine action alignment combines kinematic reference correction with dynamics-aware refinement. It improves end-effector pose accuracy while preserving whole-body coordination. We also use robot-arm rendering and training-time image augmentation to reduce the visual embodiment gap and improve viewpoint robustness. On four real-world tasks, vision-language-action (VLA) policies trained on aligned human data show zero-shot skill transfer without target-task robot demonstrations. Task scores are comparable to those of policies trained on teleoperation data at a lower collection cost. These results support human data as direct skill supervision.

[AI-104] Absorbed in Inertia: Activation Analysis for Computer-Use Agents

链接: https://arxiv.org/abs/2609.37176
作者: Giulio Segalini,Zhi Wen Soi,Jérémie Decouchant,Lydia Chen
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Computer-use agents have become increasingly capable of executing tasks on live desktops through natural-language instructions, based on trajectories of screenshots, actions, and reasoning. We discover that they can stealthily exhibit inertia, in which they repeat fruitless actions despite recognizing that these actions are ineffective. We hypothesize that inertia is reflected in the agent’s internal state, i.e., the activation values of the agent’s underlying model, and propose a protocol to measure the relationship between the two. Extensive analysis of high-dimensional activation states shows that inertia corresponds to an absorbing region of activation space, where activation values become stale across actions and even after attempts to steer them. We conjecture that drastically changing the agents’ activations by re-initializing them is necessary to escape inertia. Specifically, we propose R ^3 (Reset, Reroute, Restore), which temporarily resets the agent’s context trajectory to escape the absorbing region and then restores the historical context to effectively complete the task. Our approach yields 17-55% lower measured inertia across models relative to unmodified agents. These results suggest that changing the context can interrupt recurrence more effectively than directly steering the resulting activations. Our code is available at this https URL

[AI-105] SimpleEvol: An Agent -Loop Framework for LLM -Driven Automated Heuristic Design with Minimal Human Priors NEURIPS2026

链接: https://arxiv.org/abs/2609.37172
作者: Jianghan Zhu,Cong Zhang,Rongjie Zhu,Chi Zhang,Zhiguang Cao
类目: Artificial Intelligence (cs.AI)
备注: Accepted at NeurIPS 2026. 48 pages, 13 figures

点击查看摘要

Abstract:Large language models (LLMs) have emerged as powerful tools for automated heuristic design (AHD), enabling iterative generation and refinement of heuristics. However, the dominant paradigm embeds LLMs as narrow, fixed components, such as crossover or mutation, within heavily hand-engineered evolutionary frameworks. We argue this misapprehends LLMs. It treats them as specialized tools rather than general reasoners, constrains them to low-level operations, and underutilizes their autonomy. Moreover, the extensive human priors in these frameworks violate the bitter lesson principle that general methods scaling with computation surpass hand-crafted solutions. This raises a key question: which AHD framework designs best convert stronger LLM capabilities into better heuristics? To address this, we propose metrics for LLM-driven AHD framework handcraftedness (AHI) and intelligence conversion efficiency (ICE). Evaluating ten LLMs across three challenging combinatorial optimization problems, we obtain a notable finding that frameworks with fewer human priors consistently yield higher ICE. Based on this finding, we propose SimpleEvol, an agent-loop framework for AHD which removes nearly all human priors and allows the LLM to operate autonomously. SimpleEvol consistently achieves the highest ICE, often by a large margin. Our results challenge the trend toward complex AHD pipelines and point to a lighter and more model-centric alternative, suggesting that reducing human priors is a more effective strategy to scale up with model intelligence. The source code is available at this https URL.

[AI-106] Interpolated Policy Distillation: A Controllable Continuum Between Off-Policy and On-Policy Distillation

链接: https://arxiv.org/abs/2609.37170
作者: Youxu Shi,Yifan Sun,Dacheng Yin,Haomiao Tang,Guangting Wang,Fengyun Rao,Jing Lyu,Dong Liu
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Off-policy and on-policy distillation have traditionally been formulated as separate paradigms, each favoring a different property of distillation trajectories. Teacher-generated (off-policy) traces are typically high-quality but lie far from the student’s distribution, whereas student-generated (on-policy) rollouts are more learnable but often contain erroneous reasoning. We view these paradigms as the endpoints of a policy continuum and posit that a more effective rollout policy may lie in between. We introduce \textbfInterpolated Policy Distillation (IPD), which defines the next-token distribution at every decoding step as an explicit linear interpolation between the student and teacher distributions. The interpolation operates at the distribution level, token by token, and its coefficient provides direct control over the balance between trajectory quality and student learnability. Naively sampling from this policy would require sequentially querying the teacher at every token and is thus expensive. To make IPD practical, we accelerate it with a new speculative-decoding rule while exactly preserving the interpolated next-token this http URL the trajectory level, the resulting rollouts naturally interleave student- and teacher-generated segments. Unlike recent heuristic segment-interleaving methods, however, this interleaving is induced by an exactly realized token-level interpolated policy rather than by hand-designed switching rules. Across text-only and multimodal reasoning benchmarks, IPD consistently outperforms both endpoint policies (SFT and OPD), their conventional two-stage combination (SFT-then-OPD), and recent heuristic segment-interleaving methods, demonstrating that token-level policy interpolation better balances trajectory quality and student learnability.

[AI-107] From Learner Behavior to Reusable Skills for Effective and Efficient Learner Simulation

链接: https://arxiv.org/abs/2609.37157
作者: Zijian Chen,Zheng Zhang,Miao Jia,Xingchen Hu,Weibo Gao,Linan Yue
类目: Artificial Intelligence (cs.AI)
备注: 16 pages

点击查看摘要

Abstract:Learner simulation aims to reproduce how a particular learner behaves on new tasks. Although Large Language Models (LLMs) can generate increasingly fine-grained learning behaviors, existing approaches often need to repeatedly process a growing interaction history to reconstruct the learner. This introduces additional context and inference costs and makes the acquired learner-specific simulation capability difficult to reuse across different LLMs. We therefore propose Learner2Skill, which externalizes the simulation capability acquired from historical interactions into a persistent and reusable Simulation Skill. The Skill captures the learner’s current learning state and recurring response patterns, evolves as new real interactions arrive, and can be adapted to a new LLM through lightweight executor calibration without reconstructing the learner from scratch. Experiments show that Learner2Skill more faithfully reproduces fine-grained learner behavior while reducing overall token cost, and that the same constructed Skills can be effectively reused across different LLM executors.

[AI-108] Lucid Dreaming for World Models: Learning to Doubt Imagination and Decide by Trust

链接: https://arxiv.org/abs/2609.37156
作者: Ziqi Wen,Ting Xu,Lianyu Wang,Xian Lin,Yanda Meng,Huazhu Fu,Meng Wang,Ching-Yu Cheng
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 30 pages, 22 figures, 8 tables. Project page: this https URL

点击查看摘要

Abstract:World models enable agents to learn and plan in imagination, but predictions beyond their experience can become unreliable and mislead decisions. Existing uncertainty estimates derived from predictions can remain overconfident on unfamiliar state-action pairs. We propose the Lucid World Model (LucidWM), which learns doubt from experience and propagates trust through imagination. By integrating Subjective Logic into categorical latent transitions, LucidWM distinguishes predicted outcomes from their evidential support and assigns each transition a degree of doubt. The complement of this doubt defines transition-level trust, which accumulates multiplicatively along imagined trajectories to reweight returns for policy learning and guide action selection. Uncertainty estimation requires no additional parameters or forward passes. Evaluated on four base world models against seventeen uncertainty readouts, LucidWM detects environmental changes and signals uncertainty during action-corrupted rollouts. In a controlled navigation case study, acting on trust reduces the number of steps required to reach the goal from 362 to 190. Fifteen demonstration videos show how LucidWM doubts its dreams and acts on that doubt. Videos are available at this https URL.

[AI-109] When Tools Silently Lie: Evaluating and Mitigating Blind Compliance in Tool-Augmented Data Agents

链接: https://arxiv.org/abs/2609.37153
作者: Zifu Tao,Changqing Yin
类目: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
备注: 28 pages, 8 figures. Code and benchmark materials: this https URL

点击查看摘要

Abstract:Tool-augmented data agents rely on tool outputs for analytical decisions. Yet successful execution can return plausible but incorrect evidence, requiring agents to decide whether to trust or verify it. Understanding this failure requires examining both the evidence obtained through checking and the answer ultimately adopted. We introduce ToxicBench to measure checking and adoption under numerical, label, schema, and retrieval errors, pairing clean and poisoned observations over fixed source data. In the 118-task GPT evaluation across three adapters, poisoning lowers task success by 26 to 39 percentage points. Ordinary retries help under one-shot poisoning, whereas repeated poisoning reveals wrong-answer adoption after checking. Controls on three public tables isolate how supplied evidence affects recovery. After freezing the scorer, we compare its judgments with human annotations on 200 trajectories, finding 96% task-success agreement. Human judgments support retry gains over Base and confirm adoption after checking on audited tasks. We release trajectories, versioned scoring, and reference and delivery audits. These findings highlight evidence availability and answer selection as complementary dimensions of agent reliability.

[AI-110] From Judgment Quality to Downstream Utility: Rethinking LLM -as-a-Judge for Open-Ended Tasks

链接: https://arxiv.org/abs/2609.37145
作者: Zheng Zhang,Lufei Li,Xinyue Tan,Yuanhao Zeng,Ziwei Shan,Yexin Li,Kan Ren
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:LLM-as-a-Judge is increasingly used to evaluate policy responses on open-ended tasks that lack ground-truth answers. Existing work often directly converts the resulting judgments into reward signals for policy training, paying limited attention to intrinsic judgment quality and largely restricting the use of Judges to training-time supervision. We systematically investigate judgment quality and downstream utility by examining both how judgments are elicited and how they are used. For judgment elicitation, we vary the Judge protocol along three dimensions: verdict granularity, critique usage, and evaluation batching. For judgment usage, beyond policy training, we extend Judge to test-time inference through Best-of-N selection, Judge-guided revision, and beam search. We find that, (i) Surprisingly, judgment quality and downstream utility do not always align. (ii) Judge protocol design substantially affects both intrinsic judgment quality and downstream utility. (iii) Judge guidance effectively converts test-time compute into performance gains, with benefits varying across inference strategies. Our results call for a multifaceted evaluation of LLM Judges on open-ended tasks, encompassing intrinsic judgment quality, and downstream utility.

[AI-111] rain Ahead Distill Back: Bootstrapping On-Policy Self-Distillation for Large Language Models

链接: https://arxiv.org/abs/2609.37132
作者: Zheng Zhang,Xinyue Tan,Lufei Li,Xinyi Zhang,Yexin Li,Kan Ren
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:On-policy self-distillation (OPSD) improves large language models by letting a self-teacher with privileged information provide dense token-level supervision on the model’s own trajectories. Yet existing methods typically construct the self-teacher from the current, initial, or slowly averaged policy state, leaving the quality of supervision constrained by the teacher’s ability to exploit privileged information. We ask whether the model’s own optimization progress can instead be recycled into a stronger self-teacher. In this paper, we introduce Bootstrapped On-Policy Self-Distillation (B-OPSD), which temporarily trains the policy ahead to obtain a future teacher, restores the student to the original policy state, and then uses the future teacher to supervise the restarted student. The future teacher improves supervision in two complementary ways, it can generate more reliable privileged trajectories and, conditioned on them, provide more informative token-level targets along the restarted student’s on-policy trajectories. Experiments on mathematical reasoning with Qwen3-4B and Qwen3-8B show consistent improvements over standard OPSD in both settings, including gains from 27.50 to 41.30 and from 48.80 to 64.44 in the rollout-privileged setting. Our findings point to a broader principle for self-improving models that future learning progress can be distilled backward, preserving acquired knowledge while bootstrapping beyond the optimization state that produced it.

[AI-112] SkillCome: Group Contrast Skill Optimization with Dual Memory

链接: https://arxiv.org/abs/2609.37128
作者: Haolin Li,Feng Hong,Ang Li,Chilin Fu,Weichang Wu,Ya Zhang,Yanfeng Wang,Xiaolu Zhang,Jiangchao Yao
类目: Artificial Intelligence (cs.AI)
备注: Preprint

点击查看摘要

Abstract:Skill evolution improves the capabilities of large language models by analyzing trajectories generated under a given skill and modifying the skill accordingly. Existing approaches typically generate a single trajectory per question. However, this provides insufficient optimization signals since it requires inferring effective skill edits from a solitary path. It is difficult to pinpoint which actions caused the failure in a failed trajectory, or to determine which actions in a successful one should be incorporated into the skill. Furthermore, they rely on a local batch of trajectories for analysis, making the optimization direction susceptible to noisy evidence. To address these, we propose SkillCome, a Skill-evolution method based on group Contrast optimization with dual memory. For each question, SkillCome generates trajectories and performs group contrast analysis to precisely identify key behavioral divergences between successful and failed trajectories, offering reliable optimization signals. The dual memory system further accumulates evidence from historical steps to track patterns shared across different groups, leading to more generalized optimization directions. Together, SkillCome builds a systematic optimization process that transforms experience from observed successful trajectories into reusable skills. Extensive experiments on six benchmarks spanning question answering, reasoning, and agentic tasks demonstrate the effectiveness of our method. SkillCome consistently outperforms baselines across five models of varying families and scales, with gains up to +5.69 points.

[AI-113] When Should Agents Check External State? Budgeting Observations for Stored Intentions

链接: https://arxiv.org/abs/2609.37125
作者: Zhengkun Di,Bin Shi,Kai Sun,Yiming Xu,Bo Dong
类目: Artificial Intelligence (cs.AI)
备注: 23 pages, 2 figures

点击查看摘要

Abstract:Prospective memory allows an agent to retain an intention tied to a future condition, but the stored intention does not reveal whether that condition currently holds. Checking it may require web access, multi-step tool use, and paid calls. Existing systems decide when intentions require attention, but do not allocate the resulting observations under a shared budget. We introduce the first resource-allocation formulation for the external observations required by stored intentions under a shared episode budget. BudgetPM offers two policy variants that share a hard-budget executor. BudgetPM-Static uses a lightweight Logistic scorer to learn whether a check improves the current decision. BudgetPM-Sequential distills full-episode hindsight schedules into a lightweight policy that decides when to spend or reserve capacity using only pre-query information at deployment. We evaluate BudgetPM against two public memory-agent systems, five matched controls, and four hand-designed monitoring or budget-adaptation rules. Across two benchmarks and three backbones, BudgetPM-Static outperforms adapted Mem0 and PMA workflows. On PM-Bench, its Logistic scorer reaches competitive quality–cost operating points alongside higher-capacity scorers and retains 99.9–100% of unconstrained quality with 42–54% fewer observations. Under severe scarcity and the same hard caps, BudgetPM-Sequential exceeds the strongest tested natural monitoring schedule by 1.92–2.58 Set F1 points. It reaches the same Set F1 and on-time recall with 16–33% fewer observations. Matched attribution, exact-cost analysis, and a fixed-budget load intervention link this gain to competition between present and future opportunities. These results yield a demand–capacity design rule: local gating works when capacity covers demand, while future-aware supervision adds value when observations compete across time.

[AI-114] Multichannel Audio Quality Assessment: Extending Pretrained Perceptual Models to Spatial Audio

链接: https://arxiv.org/abs/2609.37116
作者: Gouthaman KV,Shiv Gehlot,Vishnu Raj,Lars Villemoes,Arijit Biswas
类目: ound (cs.SD); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Accurate perceptual quality assessment is essential for evaluating and optimizing spatial audio, where perceived quality depends on both signal fidelity and inter-channel spatial relationships. However, subjective evaluation is costly, while existing perceptual models are often trained for limited channel configurations and cannot be directly applied to higher-channel-count audio. This raises the question: how can pretrained perceptual knowledge be effectively reused for multichannel spatial audio? Using 5.1-channel audio, we study four levels of multichannel integration: signal, prediction, latent, and feature and propose two learned approaches: latent-level aggregation of spatial-group representations and the feature-level Feature-Band Group Attention (FGAtt), which adaptively fuses spatial groups at the feature level before perceptual processing. Across five 5.1-channel test sets, FGAtt achieves the strongest over- all performance, demonstrating the effectiveness of feature-level adaptation for reusing pretrained perceptual knowledge

[AI-115] Learning from Viable Failure Prefixes: Milestone Viability Potential Policy Optimization for Long-Horizon LLM Agents

链接: https://arxiv.org/abs/2609.37111
作者: Qi Zhou,Yuanfan Li
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Long-horizon LLM agents require reinforcement learning methods that can assign credit to intermediate decisions under sparse and delayed rewards. Existing group-based methods such as GRPO and GiGPO alleviate this issue by comparing rollout returns or repeated anchor states, but they still fail when the compared returns have no variation. We identify this failure mode as zero-credit failure: during early training, many failed rollouts contain useful prefixes, yet existing methods assign them no task-discriminative advantage. To address this issue, we propose Milestone Viability Potential Policy Optimization (MVPO), a potential-routed policy optimization algorithm that learns from viable failure prefixes. MVPO estimates prefix potential over Union-Find viability regions, repairs zero-credit groups with potential-difference advantages, and attenuates the potential branch according to relative performance progress. Experiments with Qwen2.5-1.5B-Instruct show that MVPO outperforms eight strong baselines, including GRPO and GiGPO. Under the same training length, MVPO improves over the GiGPO baseline by +4.4 success points on ALFWorld and +5.3 on WebShop, while adding only 0.16%-0.20% advantage-construction overhead.

[AI-116] Breaking the Illusion of Review Reliability under Static Evaluation: SCOPE Fuzzing for LLM -based Scientific Reviewers

链接: https://arxiv.org/abs/2609.37097
作者: Zhuo Chen,Hao Zeng,Jiawei Liu,Guoxiu He,Le Cai,Liu Haotan,Li Wenbo,Yong Huang,Wei Lu
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:The rapid growth of submissions and reviewing workload has accelerated the use of large language models (LLMs) in peer review. Prior studies suggest that LLM-based reviewers can penalize content perturbations, such as overclaiming, indicating a certain degree of reliability. Yet these conclusions are largely based on a narrow set of perturbation strategies instantiated with static templates, providing limited evidence of actual reliability. In this paper, we construct a three-level evaluation framework covering perturbations to surface presentation, argumentative logic, and value judgment. Experiments on representative LLM-based reviewers reveal two limitations of static evaluation: stratified vulnerability, where perturbation effects depend on whether the paper’s original review score is high or low, and perturbation undercoverage, where a single template misses vulnerabilities exposed by diverse realizations. To address these limitations, we propose SCOPE-Fuzzer, a strategy-aware fuzzer that combines feedback-driven strategy selection with adaptive mutation of paper content. By iteratively probing reviewers with dynamic perturbations, SCOPE-Fuzzer consistently uncovers vulnerabilities overlooked by static evaluation and other baselines.

[AI-117] ARGOS: Reinforcement Learning-Driven Multidimensional Elasticity for Service Orchestration in the Computing Continuum

链接: https://arxiv.org/abs/2609.37085
作者: Javier Mateos-Bravo,Sergio Laso,Juan Luis Herrera,Ilir Murturi,Pantelis Frangoudis,Schahram Dustdar
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Data-intensive services in the Computing Continuum must balance analytics quality, resource usage, and cost across heterogeneous nodes with limited and uneven capacity. This balance becomes especially difficult when resource scaling reaches capacity limits, because changes in demand and cluster pressure must then be absorbed without violating client-defined quality ranges. Existing orchestrators mainly adapt resources, placements, or replicas, while analytics requirements such as coverage, sample, and freshness remain fixed. This article presents ARGOS, the Adaptive Reinforcement Learning-Driven Governance for Orchestrated Services, an end-to-end controller that formulates multidimensional elasticity as a per-request Markov decision process over analytics quality and cluster pressure, supported by capacity-aware admission. ARGOS is evaluated under controlled workloads and time-varying multi-tenant arrivals on a heterogeneous cluster. Across the controlled scenarios, the deep reinforcement learning policies consistently outperform the non-learning baselines and approach the independently tuned best-fixed reference. A separate live evaluation reports improvements over the static midpoint under realistic and saturated arrivals, with no recorded CPU or memory violations but remaining coverage violations. These results support deep reinforcement learning as an adaptive mechanism for multidimensional elasticity when resource scaling alone is insufficient.

[AI-118] Identifying ODEs from Unstructured Data with Causal Representation Learning

链接: https://arxiv.org/abs/2609.37083
作者: Alessandro Trenta,Riccardo Massidda,Davide Bacciu,Sara Magliacane
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
备注:

点击查看摘要

Abstract:We study the problem of recovering the governing ODE of a dynamical system from unstructured, high-dimensional observations such as images. Existing methods for ODE discovery typically assume direct measurements of the variables, or do not provide theoretical guarantees on the learned variables and equations. While Causal Representation Learning (CRL) methods provide guarantees on identifying variables from high-dimensional observations up to component-wise diffeomorphisms, we show that in general these variables cannot be used directly as input to equation discovery methods, which typically assume that the variables will lead to sparse equations. So we introduce SParse Equivalent Equation Discovery AutoEncoder (SPEED-AE), a framework that combines a pretrained CRL method with a component-wise autoencoder that learns transformations of variables that are amenable to sparse ODE discovery. We show that for polynomial ODEs, this additional step allows us to restrict the identifiability of each variable from polynomial to monomial diffeomorphisms. Experiments on Lotka-Volterra, Lorenz, and a two-pendulum system show that SPEED-AE improves on the disentanglement of the CRL methods and that it recovers ODEs that are closest to the ground truth, while achieving state-of-the-art forecasting performance.

[AI-119] Predictive Safety Curricula for Robust Legged Locomotion

链接: https://arxiv.org/abs/2609.37070
作者: Ivan Ovinnikov,Pascal Sutter,Christian Gehring,Jordis Herrmann
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Rare but consequential failures can persist in learned locomotion policies for legged robots even when average task performance is high, in part because standard curricula primarily adapt task difficulty rather than the distribution of safety-critical experience. We introduce Predictive Safety Curricula (PSC), a framework for allocating locomotion training experience using learned predictions of future safety cost. PSC trains a distributional safety critic from policy rollouts and uses its predictions to prioritize both terrain contexts and previously encountered randomized events. The resulting curriculum modifies the training distribution while leaving the task reward and policy-optimization loss unchanged. We evaluate PSC in controlled rough-terrain locomotion and in production locomotion systems. PSC improves reliability relative to standard terrain progression, advantage-based replay, and learning-progress curricula, with the largest gains on difficult terrain and under degraded observations. The same allocation principle transfers to two production locomotion stacks. On ANYmal-D hardware, PSC reduces shank-collision incidence by 63% relative to the learning-progress curriculum across three matched training seeds, with a reduction in every seed. On a production stair-climbing platform, PSC eliminates observed shank collisions in the evaluated hardware trials. These results show that learned predictions of future safety cost can provide an effective signal for allocating training experience toward rare failure modes and improving locomotion reliability.

[AI-120] FACT: Fidelity-Aware Construction of Articulated Twins

链接: https://arxiv.org/abs/2609.37067
作者: Kuixiang Shao,Chuansen Nie,Yinuo Bai,Jiayuan Gu,Jingyi Yu
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Visually plausible articulated assets may still fail during contact interactions or exhibit inaccurate motion. We present FACT (Fidelity-Aware Construction of Articulated Twins), an agentic framework that progressively constructs articulated twins to improve geometry, contact, and dynamic fidelity. The agent drives an evidence–diagnosis–revision loop on a shared editable representation, selecting measurements and model edits using quantitative feedback, while numerical tools execute and validate the updates. It reconstructs editable articulated geometry from images through feature planning, targeted measurements, and diagnostic refinement. On this reference, it repairs collision proxies through task-aware local repartitioning before fidelity-constrained compression. Finally, it constructs response models from passive-response videos, using simulation residuals to guide model revision and constrained physical parameter fitting. Experiments show that FACT improves geometric reconstruction over baselines, enables more reliable interaction with simpler collision proxies, and better reproduces held-out physical responses than direct parameter inference.

[AI-121] Beyond Compression: Diagnosing How Post-Training Changes Mathematical Reasoning

链接: https://arxiv.org/abs/2609.37066
作者: Hongyang Li,Yiming Zhu,Xiao Li,Caesar Wu,Said Mammar,Pascal Bouvry
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 18 pages, 17 figures, 10 tables. Includes technical appendix. Under review

点击查看摘要

Abstract:Post-training is central to mathematical reasoning in modern large language models (LLMs), but endpoint pass@1 alone underidentifies what has changed. Gains may reflect newly reachable solutions, cheaper sampling of latent solutions, surface robustness, or memorisation. We compare three post-training paths under a common diagnostic readout: our sufficiently trained off-policy distillation trajectories, released Qwen3 off-policy-plus-on-policy distillation endpoints, and a released DeepSeek-Math endpoint trained with Group Relative Policy Optimisation (GRPO). Our probe uses cross-surface pass@K over verbatim prompts, paraphrases, numerical isomorphisms, and translations, plus consistency, distribution-shape, and verified supervised-fine-tuning (SFT) membership analyses. We find two regimes. On easier AMC problems, large-K ceilings are near saturation, so post-training mainly compresses sample cost. On harder AIME problems, post-training expands the large-K ceiling over the base model: sufficient off-policy distillation already raises this ceiling, Qwen3 released endpoints raise it further, and DeepSeek-Math GRPO does not dominate sufficient off-policy distillation at large K. English-dominant distillation improves non-English reasoning but preserves language-tier gaps. A controlled-overfit audit finds limited sensitivity in current SFT-membership probes. Compression is one regime of post-training, not a universal explanation.

[AI-122] A Comprehensive View of Fairness through Distributional Stability

链接: https://arxiv.org/abs/2609.37061
作者: Gayane Taturyan,Charlotte Laclau,Stephan Clémencon
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:We view fairness as a property of distributional stability. Rather than assessing a predictor under a fixed data distribution, we study how its predictions change under perturbations that modify the composition of protected groups. A predictor is fair if it remains stable under such shifts. Under this perspective, several classical notions of fairness arise as stability with respect to specific perturbations, with the associated unfairness gap given by a Lipschitz constant of a prediction-rate functional. This formulation also yields guarantees that hold uniformly over a range of demographic compositions at test time, without requiring knowledge of the deployment distribution. It leads to a learning procedure based on convex combinations of reweighted predictors, formulated as a second-order cone program, for which we establish generalization bounds. Experiments on standard benchmarks illustrate the approach.

[AI-123] Evolving Towards Better Codes: LLM -Guided Search for High-Distance Binary Linear Codes

链接: https://arxiv.org/abs/2609.37056
作者: Amal Seddas,Vladyslav Shashkov,Maryna Viazovska,Emmanuel Abbe
类目: Information Theory (cs.IT); Artificial Intelligence (cs.AI); Neural and Evolutionary Computing (cs.NE)
备注: 17 Pages, 3 Figures, 3 Tables

点击查看摘要

Abstract:Evolutionary program search driven by large language models (LLMs) has produced record-breaking constructions for open problems in combinatorics and beyond. We apply this approach to the longstanding problem of improving the best-known bounds for binary linear codes. Building on the EvoTune evolutionary framework and the ShinkaEvolve codebase, we introduce LinCodeEvolve, which evolves code-construction programs against an exact minimum-distance evaluator. A strategy loop combines diversity-driven search and expert supervision: when progress plateaus, new strategies are used to redirect the search. LinCodeEvolve discovers seven record-breaking codes, [172,21,66] , [173,20,68] , [176,21,68] , [181,21,70] , [184,21,72] , [189,22,72] and [200,21,77] , six of which have concise quasi-cyclic descriptions. With standard code modification techniques, they improve 22 entries of the tables. Every code is verified by exhaustive enumeration. These results suggest that LLM-guided search can help find improved codes and complement existing methods in coding theory.

[AI-124] actr: aligning thoughts and responses for multilingual safety in reasoning llm s

链接: https://arxiv.org/abs/2609.37054
作者: Xianhui Zhang,Jian Yu,Chengyu Xie,Chenhang Cui,Shuyi Miao,Pengyang Shao,Yu Zheng,Fei Shen,Tat-Seng Chua
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Ensuring the safety of reasoning large language models (LLMs) across languages is essential for their reliable deployment. However, when exposed to jailbreak attacks in non-high-resource languages, these models may generate unsafe responses even when their reasoning traces identify safety risks. To address this issue, we propose aligning cross-lingual thoughts and responses (ACTR), a framework that improves multilingual safety alignment by strengthening the use of existing safety reasoning. Specifically, we first present the think gap score (TGS) to compare the normalized contributions of reasoning traces to attention outputs during response generation across languages, and use reasoning- trace substitution to measure the cross-lingual safety gap. Next, using a corpus of jailbreak queries, we assess neuron importance through changes in response representations caused by neuron masking and compare the high-importance neuron sets obtained with reasoning enabled and disabled to identify safety think neurons that support the use of safety reasoning. Finally, we devise neuron-selective consistency optimization (NSCO), which uses a frozen judge model to reward agreement between the safety categories of reasoning traces and responses while updating only the parameters associated with the selected neurons, requiring no human-annotated responses or preference data. Across two reasoning models, ACTR achieves lower average attack success rates than the evaluated state-of-the-art methods on AdvBench-X and MultiJail, with safety gains extending to unseen languages, while preserving or improving average performance on multilingual knowledge and mathematical reasoning tasks and limiting false refusals of benign requests. Warning: this paper contains examples with unsafe content.

[AI-125] MatToolBench: Benchmarking Multimodal Agents in Real-World Materials Science Workflows

链接: https://arxiv.org/abs/2609.37053
作者: Mei Wu,Rui Xie,Runyu Zhang,Yuqiang Li,Tianfan Fu,Bo Chen,Kai Yu,Xin Chen,Lu Chen
类目: Artificial Intelligence (cs.AI)
备注: 25 pages, 15 figures. Mei Wu and Rui Xie contributed equally. Bo Chen and Lu Chen are corresponding authors. Project page: this https URL ; code: this https URL

点击查看摘要

Abstract:Multimodal GUI agents have achieved impressive results on general software benchmarks, yet their ability to operate professional scientific software remains largely unexplored. In materials science, sparse domain-specific web data, specialized interfaces, and tacit workflow conventions create blind spots that general-purpose pretraining cannot readily bridge. We present MatToolBench, the first real-environment benchmark for evaluating multimodal GUI agents on professional materials science software, comprising 204 tasks across 10 tools in three modalities: GUI operation, OriginPro scripting, and code-based database queries, all executed inside a Windows 11 VM. Each task is decomposed into fine-grained sub-criteria by domain experts, enabling interpretable partial-credit scoring; the GUI component of our multi-level evaluation pipeline achieves an average F1 of 0.98. For OriginPro figure-generation tasks, we further conduct a human-LLM agreement study to validate the use of a multimodal judge for secondary aesthetic assessment. Our experiments show that strong performance on general benchmarks does not transfer to professional scientific workflows, and that this gap is not a visual-grounding problem alone: failures arise from domain-specific operational knowledge, sparse pretraining coverage of scientific software, weak cross-tool artifact handoff, and critical states exposed only visually. Even the best model reaches only 25% success rate on GUI tasks and 45% on code tasks. MatToolBench therefore serves as a challenging diagnostic benchmark and real-environment testbed for data-scarce, knowledge-intensive scientific workflows.

[AI-126] Beam Search as Test-Time Self-Distillation via Counterfactual Contexts NEURIPS2026

链接: https://arxiv.org/abs/2609.37041
作者: Su Ee Tan,Xiaotong Ji,Rasul Tutunov,Haitham Bou-Ammar,Matthieu Zimmer
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Accepted at NeurIPS 2026 Workshop on Towards Test-Time Continual Learning Agents

点击查看摘要

Abstract:Self-Distillation Fine-Tuning (SDFT) enables a language model to act as its own teacher: by conditioning on a demonstration, the model produces an implicit reward via pointwise mutual information, which guides on-policy learning without external supervision. However, SDFT operates at training time: it requires gradient updates and access to expert demonstrations, making it inapplicable at inference. We propose test-time self-distillation, a decoding-time method that extracts a steering signal from the self-distillation framework without any parameter updates, reward models, or training data. Our key insight is that counterfactual contexts, i.e. fixed textual templates that hypothetically prime the model for excellent versus poor reasoning, can substitute for the demonstration. The log-odds ratio of a candidate answer under these two counterfactual conditions defines a new reward signal. We derive the optimal KL-regularized policy under this reward, which takes the form of a Gibbs reweighting of the base distribution. Crucially, this reweighting is global: it cannot be decomposed into independent per-token operations without ignoring future trajectory quality. We therefore approximate the target distribution via beam search. Experiments on mathematical reasoning (MATH500), code generation (HumanEval), and graduate-level science QA (GPQA) across multiple model scales show that test-time self-distillation improves over standard sampling, low temperature, beam search and power sampling baselines on average, demonstrating that the self-distillation principle can be operationalized at inference time.

[AI-127] Watch-Think-Interact: Bootstrapping Long-Horizon Multi-Turn Streaming Video Reasoning with Reinforcement Learning

链接: https://arxiv.org/abs/2609.37035
作者: Ziheng Huang,Yicheng Bao,Xueheng Li,Zhenkun Gao,Bangwei Liu,Kunquan Li,Yuxiang Shen,Bangyan Li,Xuejiao Wang,Changbo Wang,Gaoqi He
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Streaming video assistance requires models to answer asynchronous questions from an observed prefix under a fixed context budget. Existing approaches model response timing or compress history, but an online state formed before future questions are known can omit visual details before later questions reveal their relevance; the retained state alone cannot recover them. We introduce Watch-Think-Interact (WTI), a closed-loop framework for multi-question streaming video reasoning. WTI maintains compact natural-language memory entries tagged with source-video time ranges; these entries support direct reasoning when sufficient and otherwise anchor selective recall of finer visual evidence. For each question, WTI answers when current context and memory suffice, continues watching when required evidence has not appeared, or recalls a relevant past interval and decides again after incorporating the returned chunks, without replaying the full observed history. To train this behavior, we construct WTI-82K, comprising 82,335 timed questions across 4,812 causally aligned trajectories, and develop Stream-GDPO to optimize complete multi-question streaming rollouts using trajectory-level feedback for response timing, source-video recall, and memory updates. WTI achieves state-of-the-art aggregate performance among the compared open-source streaming baselines, reaching 83.3% on StreamingBench and 73.6% weighted overall accuracy on OVO-Bench.

[AI-128] FedLAFP: Low-Rank Aggregation Meets Full-Rank Personalization in Federated Fine-Tuning

链接: https://arxiv.org/abs/2609.37033
作者: Mengjun Yi,Huaian Gu,Yinghao Ai,Furao Shen,Jian Zhao
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Federated parameter-efficient fine-tuning enables clients to adapt pre-trained models without sharing raw data or communicating the full model, but statistical heterogeneity makes a single global adapter insufficient for personalized prediction. Existing personalized methods typically use the same low-rank structure for both shared and private adaptation, overlooking their distinct requirements for aggregation and personalization. We propose FedLAFP, a role-aware framework that couples a compact, globally aggregated LoRA branch with a client-private, full-rank-capable RandLoRA branch. The shared branch provides an efficient interface for transferring common knowledge, whereas the private branch combines fixed random low-rank bases with learned scaling coefficients to provide expressive client-specific adaptation without additional communication. Client- and layer-specific mixing coefficients jointly fuse the two branches, and only the shared LoRA parameters are exchanged. A controlled linear study supports this role assignment: LoRA yields more aligned client updates and lower aggregation error, while RandLoRA more accurately recovers client-specific residuals. Experiments across four visual recognition benchmarks show that FedLAFP consistently outperforms local-only and federated LoRA baselines, achieving an average personalized accuracy of 86.93% and exceeding the best baseline average by 1.30 percentage points.

[AI-129] LongSpark: Efficient speculative decoding with a fixed-cost parallel drafter

链接: https://arxiv.org/abs/2609.37029
作者: Hao-Yuan He,Peng-Fei Liu,Si Shen,Ming Li
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Speculative decoding accelerates autoregressive inference by verifying multiple draft tokens in a single target forward pass. However, as the context grows, existing state-of-the-art drafters become increasingly expensive, eroding the very efficiency advantage they are designed to provide. We argue that this scaling is unnecessary. A standalone language model must grow with its prefix because it is solely responsible for every token it produces. A drafter, by contrast, only proposes candidates; the target catches and corrects every error before any token is committed. The drafter’s decoding cost can therefore be made entirely independent of the prefix length. We introduce LongSpark, a block-diffusion drafter that achieves this by extracting fixed-size, multiscale views from the target’s verification pass, thereby eliminating the need for a growing persistent state. Extensive evaluations demonstrate that LongSpark achieves state-of-the-art end-to-end efficiency across multiple model scales and realistic serving conditions. Notably, it delivers the lowest time-per-output-token on long-context tasks while reducing the drafter’s context state by several orders of magnitude.

[AI-130] Beyond Low-Rank Parameterization: Narrowing the Gap Between LoRA and Full Fine-Tuning via Gradient Decomposition

链接: https://arxiv.org/abs/2609.37027
作者: Yihao Ouyang,Shiwei Li,Haozhao Wang,Xiandi Luo,Zhuoqi Hu,Jinglun Yu,Yichen Li,Ruixuan Li
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Low-Rank Adaptation (LoRA) is a widely used approach to parameter-efficient fine-tuning (PEFT), yet a performance gap can remain relative to full fine-tuning (FFT). Many LoRA variants improve the initialization or optimization of low-rank factors. At each training step, however, their first-order weight-space directions are constrained by the current parameterization. We characterize the corresponding LoRA-accessible gradient space and show that it coincides with the tangent space induced by the current LoRA parameterization. This characterization yields an orthogonal decomposition of the full weight gradient at the current model parameters. We term the component orthogonal to this space the normal gradient. Based on this decomposition, we propose GDLoRA (Gradient-Decomposed Low-Rank Adaptation). GDLoRA reconstructs the full weight gradient from forward activations and backward signals, extracts its normal component, and directly updates the base weights with this component, while retaining standard AdamW optimization for the LoRA factors. GDLoRA incorporates complementary normal gradients without increasing standard LoRA’s optimizer-state memory budget under matched adapter and optimizer configurations. Experiments on natural language understanding, mathematical reasoning, commonsense reasoning, and image classification show that GDLoRA consistently improves over LoRA and narrows the performance gap to FFT. The code is available at this https URL.

[AI-131] AnyAct: Universal Action for Self-Evolving Agents

链接: https://arxiv.org/abs/2609.37025
作者: Lingrui Xu,Yangqin Jiang,Jiachang Zhang,Xubin Ren,Chao Huang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:As large language models (LLMs) advance, AI agents are increasingly deployed in open-world environments to tackle complex sequential tasks (e.g., document processing, cross-application collaboration), relying heavily on actions ranging from GUI operations to semantic APIs. However, three core challenges persist: the “scale dilemma” of massive tool ecosystems exceeding LLM context windows, the “non-stationarity” of tool quality due to updates or outages, and the “heterogeneity” of feedback formats (pixels, text, structured data) creating information silos. To address these, we propose AnyAct, a universal action layer that unifies available capabilities into a self-evolving action space, enabling agents to operate efficiently and reliably in large-scale, dynamic tool ecosystems. AnyAct’s core design focuses on two objectives: constructing this action space via hierarchical progressive retrieval (filtering task-relevant actions) and test-time reliability evolution (pruning unreliable actions), and enabling reliability-aware action orchestration through a heterogeneous observation grounding module that unifies multi-modal feedback. Additionally, it defines a hybrid action space (primitive + semantic actions) and optimizes for a balance between task success rate and execution cost. Evaluations on LiveMCPBench and OSMCP (a new benchmark we developed for multi-granularity action collaboration) demonstrate state-of-the-art performance. AnyAct delivers substantial performance gains over baseline methods across various LLM base models on LiveMCPBench and improvements are particularly notable for models with constrained native capabilities. On OSMCP, it achieves 77.27% overall success with only 50 steps, which is half the steps required by most competitors.

[AI-132] Language as the Interface: Foundation-Model Contrastive Learning Links Transcriptomes and Electrophysiology ATC

链接: https://arxiv.org/abs/2609.37024
作者: Junbo Shen,Jinying Gao,Bo Lei
类目: Artificial Intelligence (cs.AI)
备注: Code: this https URL

点击查看摘要

Abstract:Integrating transcriptomic and electrophysiological data is essential for building multimodal foundation models for neuroscience. Patch-seq provides paired measurements of gene expression and intrinsic electrophysiology from the same neuron, establishing a basis for training cross-modal models. Here we introduce LangPatch, a foundation-model-based contrastive learning framework that uses paired Patch-seq data to align pretrained GenePT representations with electrophysiological phenotypes through a language-based interface. Gene descriptions and verbalized electrophysiological profiles are embedded by the same frozen text encoder. A context adapter and projection modules connect the modalities through paired contrastive learning. Across mouse visual, mouse motor, and human cortical cohorts, LangPatch achieves the highest mean transcriptome-to-electrophysiology prediction correlation among the evaluated foundation-model and representation-learning methods. It also improves held-out cross-modal alignment in the two mouse cohorts (FOSCTTM 0.107/0.135 vs. 0.208/0.222 for JAMIE, an existing cross-modal Patch-seq imputation method). It predicts transcriptomic family, type, cortical layer, and marker-gene expression from electrophysiology, exceeding other baselines on most endpoints. More importantly, the method transfers across brain areas and species: a model trained on mouse visual cortex predicts electrophysiology in motor cortex with approximately 70% correlation retention and in human cortex with 47% (58% on acute-slice recordings). Together, these results demonstrate alignment between molecular and functional representations of neurons, providing a building block for multimodal foundation models in neuroscience.

[AI-133] CADOC: Cache-Aware Dynamic Object Context for Long-Horizon Agents

链接: https://arxiv.org/abs/2609.37012
作者: Junjie Yao,Zhangchen Zhou,Zhi-Qin John Xu
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:For a long-horizon agent, context is the bottleneck: the history is resent with every request, the window caps task length, and reasoning degrades as the history grows. Replacing structured objects with compact retrieval Cards shortens the prompt and keeps the exact originals retrievable, but editing the history can break prefix-cache reuse, and prior recoverable methods time their edits by forecasts of future reuse or by preset intervals. We propose CADOC (Cache-Aware Dynamic Object Context), an online algorithm that replaces structured objects with compact Cards while preserving exact, on-demand retrieval of their original contents. CADOC schedules replacements in batches by balancing accumulated waiting cost against shared cache-reconstruction cost. Its scheduling rule follows from an economic order quantity trade-off, recovers the optimal integer batch under stationary assumptions. Across evaluation, CADOC consistently achieves the lowest aggregate input cost among the compared configurations, which reduces input cost by approximately 40% on average while maintaining task performance close to full context. CADOC thus provides a cost-derived approach to compressible context management, demonstrating that efficient compression depends not only on shortening prompts but also on scheduling edits to preserve cache reuse.

[AI-134] OPFL: Optimistic Verification of Federated Learning via Empirical Boundary

链接: https://arxiv.org/abs/2609.37011
作者: Hongxu Su,Jianzhu Yao,Xuechao Wang,Pramod Viswanath
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 15 pages, 3 figures, 9 tables

点击查看摘要

Abstract:Federated learning enables multiple clients to collaboratively train models without sharing their private data. However, the lack of visibility into local training makes it difficult to verify whether clients follow the prescribed training procedure or submit malicious updates, such as model poisoning. A natural approach is to replay client training for verification. However, privacy-preserving replay produces numerical results that cannot be directly matched with local client execution because the two run in different environments. We present OPFL, an optimistic verification framework for privacy-preserving federated learning. To protect data privacy, OPFL performs replay inside secure multi-party computation (MPC). Although gradients computed on MPC and local GPUs are not bitwise identical, we observe that their absolute differences are stable and bounded. OPFL therefore calibrates an empirical boundary offline and uses it to distinguish benign numerical deviations from malicious manipulation. To reduce the cost of expensive MPC replay, OPFL adopts optimistic verification by post auditing only sampled training steps. Experiments on LeNet, BERT, and Qwen show that the boundary generalizes across datasets, input lengths, and GPUs, while achieving 0 % ASR against model poisoning and PGD-based attacks. On a LeNet workload, at p=0.01 , OPFL is approximately 98.6\times faster than full MPC-based FL and 625.5\times faster than ZK-based approach.

[AI-135] Cross-Organizational SysML Model Integration: A Survey of Challenges and AI-Supported Tasks

链接: https://arxiv.org/abs/2609.37000
作者: Zirui Li,Torsten Brix,Stephan Husung
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注: IEEE ISSE 2026

点击查看摘要

Abstract:Cross-organizational collaboration is widely regarded as a key promise of SysML-based Model-Based Systems Engineering (MBSE), yet practitioners still face persistent challenges when exchanging and integrating system models. In parallel, Large Language Models (LLMs) raise expectations for AI-assisted model understanding and integration, while reliability and required human oversight continue to pose challenges. This paper reports the results of an online questionnaire survey with 29 MBSE stakeholders involved in cross-organizational collaboration. Respondents rated eight predefined integration challenge categories and six AI-supported task types on five-point Likert scales. The results indicate that stakeholders perceive model integration as a multi-dimensional alignment problem across semantics, behavior, traceability, and exchange interoperability. These perceptions vary by organizational role and frequency of integration involvement. AI is rated highly useful for analysis tasks such as semantic structure analysis and inconsistency detection, and respondents predominantly prefer human-in-the-loop use with mandatory verification. These findings motivate AI support that enhances, rather than replaces, engineering responsibility in SysML-based integration.

[AI-136] ImbalancE: Inference-Time Latent Search Against Degree Imbalance in Link Prediction

链接: https://arxiv.org/abs/2609.36996
作者: Alberto Bernardi,Luca Costabello,Christophe Gueret
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Knowledge Graph Embedding models have been extensively used to learn representations of entities and relations in Knowledge Graphs for predicting missing links. However, the quality of the learned representations varies a lot across different areas of the graph. If previous research has loosely linked the problem to relation types or degree bias, we show that it is more widespread and it correlates with the degree imbalance of the entities in test triples. In particular, the prediction of a target entity that has a degree much smaller than the degree of the anchor entity is extremely problematic. This is critical in recommender systems and other use cases, where these triples represent important corner cases. To address this issue, we propose an inference-time latent search optimization method capable of significantly improving model predictions on the most imbalanced triples. Built on top of a pre-trained model, it explores the embedding space at evaluation time, blending known and out-of-band information to mitigate the degree imbalance bias. We show the value of our approach on imbalanced triples from common benchmark datasets, where we outperform conventional methods, opening the door to the successful adoption of Knowledge Graph Embedding models on these critical corner cases.

[AI-137] CF-LoRA: Decoupled Factor Aggregation and Adaptation-Aware Client Clustering for Federated LoRA Fine-Tuning

链接: https://arxiv.org/abs/2609.36986
作者: Mengjun Yi,Langxing Yang,Suhan Guo,Furao Shen,Jian Zhao
类目: Artificial Intelligence (cs.AI); Distributed, Parallel, and Cluster Computing (cs.DC)
备注:

点击查看摘要

Abstract:Federated LoRA fine-tuning enables parameter-efficient adaptation of pre-trained models without sharing private data, but suffers from two fundamental mismatches under heterogeneous client data: a structural aggregation mismatch caused by independently averaging LoRA factors, and a statistical collaboration mismatch caused by enforcing a single global adapter across divergent clients. To address these issues, we propose CF-LoRA, a clustered federated LoRA fine-tuning framework that combines decoupled factor aggregation with adaptation-aware client clustering. CF-LoRA first learns a globally shared A factor while retaining personalized B_i factors, then identifies clients with similar adaptation patterns based on the cosine similarity of their learned B_i factors, and finally performs intra-cluster B -factor aggregation with a frozen A factor. By decoupling LoRA factor aggregation, CF-LoRA preserves the low-rank structure and mitigates the structural aggregation mismatch, while adaptation-aware clustering promotes collaboration among clients with similar adaptation patterns and reduces negative transfer caused by statistical heterogeneity. Experiments on four language tasks and four vision datasets with RoBERTa and ViT show that CF-LoRA achieves the highest average accuracy in both modalities while communicating only one LoRA factor per optimization round.

[AI-138] Abductive World Modeling via Causal Representation Learning

链接: https://arxiv.org/abs/2609.36985
作者: Ziqi Liu,Songhan Yang,Linfan Zhou,Jiatong Liu,Lijun Peng,Long Wan,Yinqi Bai
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 21 pages, 4 figures

点击查看摘要

Abstract:The central challenge of world modeling is to learn representations that capture how the world evolves. However, existing world models predominantly represent future states without explicitly capturing the latent causes underlying their evolution, limiting their ability to reason about why and how the world changes. To address this limitation, we propose Abductive World Modeling (AWM), a framework that learns structured causal representations by abductively inferring latent causes from predicted futures. Specifically, we realize AWM through the Hierarchical Abductive State Pyramid (HASP), which organizes the inferred world state into three complementary components - Entity, Dynamic, and Relation - capturing what exists, how it changes, and how entities interact, respectively. By jointly reasoning over the current observation and its predicted future, HASP abductively infers these latent factors and integrates them into a structured state representation for downstream reasoning. To the best of our knowledge, AWM is the first framework to introduce abductive state inference into latent-space world modeling for learning structured representations of world dynamics. Experiments across physical prediction, causal reasoning, and action understanding demonstrate the effectiveness of our approach. Compared with V-JEPA, a state-of-the-art latent-space world model, AWM improves physical prediction AUROC by 10.7%, causal reasoning accuracy by 16.8%, and action Top-1 accuracy by 68.0%.

[AI-139] REALHOP: Rethinking Multi-Hop Reasoning Evaluation via Behavioral Auditing

链接: https://arxiv.org/abs/2609.36984
作者: Jiawen Tao,Xiaokun Yuan,Yaoming Li,Chenxu Liu,Mengzhou Wu,Tong Yang,Maxm Pan
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Complex questions often require multi-hop reasoning that connects facts distributed across sources or distant regions of a long context through intermediate steps. Benchmarks commonly evaluate this ability with questions built around predefined reasoning chains, treating a correct answer as evidence that the intended composition was used. Yet answer correctness alone leaves open whether success depends on the evidence associated with each intended step: models may instead rely on memorized associations, shorter paths, or partial evidence. We examine this dependence using the Behavioral Necessity Rate (BNR), which measures how often targeted evidence removal prevents answer recovery on initially correct instances. Across five existing benchmarks, panel-mean BNR ranges from 16.6% to 48.9%, exposing a substantial gap between annotated structure and observed dependence. Guided by this diagnosis, we introduce REALHOP, a diagnose-construct-verify framework that rebinds entities, factorizes selected relations, adds complete competing paths, and places evidence at traceable locations. Structural and semantic checks precede freezing; behavioral interventions follow. On 790 paired MuSiQue questions, REALHOP raises panel-mean BNR from 27.4% to 94.4% while retaining high Full accuracy. It also yields high BNR on REALHOP-FRAMES and REALHOP-LONGBENCH. On 216 long-context questions, the matched multiple-choice spread across 16 models grows from 13.9 to 59.2 points and persists under repeated open-ended evaluation. Together, these results show that a conceptually coherent chain and a correct final answer do not by themselves establish multi-hop reasoning. Verifying that success depends on every intended hop is therefore as fundamental to multi-hop evaluation as measuring answer accuracy itself.

[AI-140] JudgeCast: Time Series Forecasting with Experience-Informed Covariate Judgements

链接: https://arxiv.org/abs/2609.36966
作者: Donguk Kwon,Wooseok Jeong,Dongha Lee
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Covariate effects vary across contexts and shift over time, requiring forecasters to assess how to use them for each forecasting context. As forecasting proceeds, observations for earlier forecasts become available, providing feedback on past covariate use for subsequent forecasts. However, when multiple covariates act together, the forecast error reveals the numerical discrepancy from the observation but not how the covariates should have been used. We introduce JudgeCast, an experience-based framework for time series forecasting with covariates. Following the judgmental adjustment practice, a frozen TSFM provides the base forecast, while a frozen LLM uses the current context and relevant experience to adjust it. Within the adjustment, assessing covariate effects and determining the numerical adjustment serve distinct roles, so JudgeCast first forms explicit covariate-wise judgments and then determines the adjustment. After observation, JudgeCast uses the observed residual of the base forecast to reconstruct alternative judgments and evaluates the original and alternatives through their resulting adjustments. The best-performing decision is selected and retained as validated experience for subsequent forecasts. Across diverse real-world datasets, JudgeCast outperforms strong baselines. Ablations show that explicit covariate-wise judgment can improve forecast-time adjustment, while residual-guided experience construction yields more reliable forecasting gains than retaining raw decisions as experience.

[AI-141] Controlled Decoding Attacks on Black-Box LLM s

链接: https://arxiv.org/abs/2609.36956
作者: Jesson Wang,Shawn Li,Wei Yang,Franck Dernoncourt,Ryan A. Rossi,Charith Peris,Yue Zhao
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Manipulating next-token probabilities during generation can bypass the safety alignment of large language models. Existing approaches, however, rely on access to model weights or numerical token probabilities and therefore do not apply to interfaces that return only sampled text. Reconstructing probabilities from sampled outputs offers a possible alternative, but finite sampling produces sparse and noisy estimates, while repeating this process at every generation step incurs substantial query costs. Our empirical observations suggest that large distributional changes along successful jailbreak trajectories are concentrated at a small subset of positions, motivating selective control. We introduce \method, a framework for jailbreaking through text-only continuation interfaces that permit repeated sampling and assistant-prefix continuation. Sample-Based Distribution Reconstruction combines sampled outputs with a prior over unobserved actions to obtain a usable control signal. Risk-Gated Residual Control uses the evolving response prefix to decide when to reconstruct and modify the distribution, concentrating sampling costs at selected positions. Speculative Multi-Token Execution further amortizes target calls by verifying and accepting draft prefixes that require no intervention. Across four target endpoints and three benchmarks, \method achieves the highest mean score most comparisons against baselines.

[AI-142] Purlin: Separating Orchestration from the Datapath of Collectives

链接: https://arxiv.org/abs/2609.36954
作者: Osayamen Jonathan Aimuyo,Swapnil Gandhi,Christos Kozyrakis
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Distributed inference depends on GPU collective communication that must keep pace with evolving hardware and specialized workloads. However, existing collective implementations often couple semantics, orchestration (where and when data moves), and the datapath (how data moves). This coupling makes it costly to adopt new hardware mechanisms and customize communication for applications. We present Purlin, a scale-up communication framework that separates these concerns. At the top of Purlin, we specify collectives as a naming of an input and output layout and a copy or reduction operation. In the middle, we introduce a shared orchestration protocol, Stage, Notify, And Consume (SNAC), which derives coordination from these specifications. Below SNAC sits a hardware-specific datapath we call Atom, which implements two key data movement primitives for collectives: copy and reduce. This separation lets us customize collectives and adopt new hardware mechanisms while reusing orchestration via SNAC. We evaluate Purlin on A100, H200, and B200 GPUs. Across seven collectives, Purlin achieves latency speedups of up to 5.14x and bandwidth improvements of up to 4.50x over baselines. Integrated into SGLang, Purlin improves offline LLM serving throughput and interactivity by 1.13x on average and up to 1.37x over baselines. For online LLM inference, Purlin improves interactivity by 1.26x on average and up to 2.85x, with the largest gain occurring under overload. For diffusion image generation, Purlin reduces end-to-end latency by up to 1.13x.

[AI-143] Dual-Channel Robust Group-Relative Policy Optimization via Advantage and Sequence-Weight Estimation

链接: https://arxiv.org/abs/2609.36944
作者: Zhongyi Li,Wan Tian,Xiang Xu,Yutian Xiao,Yikun Ban,Yijie Peng,Fuzhen Zhuang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Group-relative policy optimization relies on reward-derived advantages and sequence-level likelihood weights, both of which can be sensitive to localized outliers. Extreme rewards can collapse the contrast among clean responses after group normalization, while token-level log-ratio perturbations can alter sequence weights and clipping decisions. We introduce RoVR-GSPO, a dual-channel robust optimizer that addresses these failure modes separately. Its reward channel combines robust reference estimation with bounded residual credit, while its ratio channel uses differentiable SoftRoVR aggregation to construct robust sequence weights. We provide stability and efficiency analyses for both channels. Experiments on mathematical reasoning, long-context summarization, and tool-call annotation show consistent improvements over GSPO, while controlled perturbation studies demonstrate stronger robustness to reward contamination and token-ratio anomalies.

[AI-144] Safe-by-Design Learning via Energy-based Neural Networks

链接: https://arxiv.org/abs/2609.36942
作者: Simone Betteti,Morteza Lahijanian,Luca Laurenti
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Systems and Control (eess.SY)
备注:

点击查看摘要

Abstract:Learning neural-network models of dynamical systems with safety guarantees is a fundamental requirement for their deployment in safety-critical settings. Safety is commonly established by proving the invariance of a desired subset in state-space, ensuring that every trajectory initialized in this subset remains confined to it for all time under admissible inputs. Existing frameworks, however, either rely on computationally expensive post-hoc verification or employ safety-enforcing mechanisms without formal correctness guarantees. In this paper, we introduce a novel neural architecture grounded in energy-based modern Hopfield networks to guarantee safety-by-design while retaining sufficient expressiveness to model complex nonlinear dynamics. Specifically, we integrate modern Hopfield networks with a port-Hamiltonian neural ODE, enabling by design the construction of barrier functions yielding explicit admissible-input sets and quantitative robustness radii. Across several benchmarks, including an 12-dimensional nanodrone model, our framework achieves state-of-the-art performance while producing certified invariant sets that are more robust to external solicitations than comparable existing approaches.

[AI-145] SCA: Spatial Credit Assignment for Reinforcement Learning of GUI Agents

链接: https://arxiv.org/abs/2609.36939
作者: Shengtian Yang,Ziyu Xiong,Kaibing Yang,Guangfeng Cai,Yewen Li,Peng Jiang,Gai Kun,Qingpeng Cai,Lei Feng
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:GUI agents automate tasks on digital devices by grounding language instructions in visual interfaces. Existing group-relative reinforcement learning improves GUI action prediction by comparing the rewards of multiple responses sampled from the same GUI state. However, binary evaluation treats spatially different failed clicks as identical and provides no relative signal when all sampled clicks fail. To address these limitations, we propose Spatial Credit Assignment (SCA), which uses the screen coordinates of sampled clicks to refine group-relative credit. Specifically, SCA predicts each held-out response’s reward from the other responses in groups containing both successes and failures, then uses the prediction residual to adjust credit. When all sampled clicks fail, SCA instead orders them by distance to the annotated target. These spatial references are used only to construct the training update; the deployed policy remains unchanged. We evaluate whether this correction improves the policy update itself by comparing its error and directional alignment with the exact return gradient in a controlled synthetic study. Across GUI grounding and offline action-prediction benchmarks, SCA improves grounding across professional domains and achieves the strongest results among reinforcement-fine-tuned models on most action-prediction metrics, with consistent gains across the reported GUI suites.

[AI-146] VLALight: A Vision-Language-Action Model for Traffic Signal Control

链接: https://arxiv.org/abs/2609.36934
作者: Pan Zhang,Siqi Lai,Kemu Dong,Hao Liu
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Traffic signal control (TSC) is essential for improving urban mobility and reducing congestion. Although roadside cameras are widely deployed at signalized intersections and provide rich visual observations of evolving traffic, existing TSC methods typically rely on manually engineered traffic states or separate perception modules, creating a gap between physical observations and control decisions. We present VLALight, the first vision-language-action (VLA) model for end-to-end traffic signal control from multi-view roadside videos. VLALight directly maps visual observations to coordinated signal actions through multi-target spatiotemporal traffic reasoning and topology-aware cooperative perception across intersections. To establish this capability, we develop a two-stage supervised cold-start training strategy for visual traffic understanding and signal decision-making, followed by cooperative agentic reinforcement learning that jointly optimizes local control and network-wide traffic efficiency. Furthermore, VLALight introduces adaptive fast and slow reasoning modes, enabling the policy to allocate deeper reasoning only when additional deliberation provides sufficient control benefits. Through balanced mode-aware rollouts and relative advantage optimization, VLALight learns to trade off decision quality and inference cost. Extensive experiments on seven real-world traffic-flow datasets across three urban networks demonstrate that VLALight consistently outperforms transportation-based, RL-based, and LLM/VLM-based baselines. Ablation studies validate the effectiveness of cooperative perception, network-level optimization, and adaptive reasoning. These results demonstrate the potential of VLA models for real-world physical traffic control. Our project is available at this https URL.

[AI-147] Learn from the Gap: Differential-Aware Advantage Pruning with Adaptive Rollout Sampling for GRPO

链接: https://arxiv.org/abs/2609.36932
作者: Jiahua Yang,Zhiwei Yang,Xianpeng Zhang,Dongyu Chen,Xing Chen,Tianhuang Su,Haonan Lu,Quanlong Guan,Kai Tang,Chuangchuang Wang
类目: Artificial Intelligence (cs.AI)
备注: 18 pages,6 figures

点击查看摘要

Abstract:Recently, Group Relative Policy Optimization (GRPO) and its variants have been developed for policy optimization and demonstrated notable performance gains. However, these methods usually incur substantial computational overhead due to per-question multi-rollout sampling and repeated per-token probability evaluation across rollouts. Furthermore, low-information or highly homogeneous trajectories can degrade downstream learning signal efficiency, hindering model optimization and limiting final performance. To address these issues, we propose FastRL, a novel plug-and-play reinforcement learning framework that simultaneously improves training efficiency and the effectiveness of policy learning. Specifically, 1) We introduce an advantage-aware pruning strategy to selectively preserve high-advantage trajectories while maximizing inter-trajectory gradient diversity. 2) Then, we design an adaptive rollout sampling mechanism to dynamically adjust the sampling scale across different training stages based on historical pruning distributions, balancing exploration adequacy and computational efficiency. Experiments demonstrate that FastRL can be seamlessly integrated into GRPO, DAPO, and GSPO variants, achieving an average 2.07 \times training speedup on Geometry3K and GeoQA8K-R1V, along with an approximately 1.64% improvement in average accuracy on visual reasoning benchmarks. Source codes will be available at this https URL.

[AI-148] Neuro-Symbolic Computer Use: Learning Reusable Policies for Reliable and Efficient Execution

链接: https://arxiv.org/abs/2609.36927
作者: Hyewon Suh,Thanh Minh Nguyen,Chih-Lun Lee,Darrow Hartman,Lizhao Liu,Xin Eric Wang,Ang Li,Jiachen Yang
类目: Artificial Intelligence (cs.AI)
备注: 26 pages, 8 figures, 10 tables

点击查看摘要

Abstract:Many computer tasks recur: the same workflow runs many times, with new inputs and from different starting states. Current computer-use agents re-plan every step of every run, which makes them costly and unreliable on such tasks. We introduce neuro-symbolic computer use, in which a recurring workflow is executed by a learned policy rather than re-derived by an agent on each run. The policy fixes the decisions that are stable across runs (ordering, variables, loops, and branches) in executable code, and delegates observation-dependent decisions, such as grounding and state checks, to neural models. We learn these policies with neuro-symbolic policy iteration: starting from one agent trajectory, it executes the policy, diagnoses failures with task-completion and step-level judges, and revises the code with a coding model informed by an agent’s continuation from the point of failure, without access to the benchmark evaluator. Iterating on generated parameter and initial-state variants makes the policy reusable, and a pre-action verifier guards each state-mutating step at deployment. On OSWorld-Verified and ScienceBoard, the learned policies achieve the highest Pass^3 of all methods in all four settings, 3.6-15.8 points above the base agent, while cutting per-run cost by 15-217 \times and latency by 3.4-5.1 \times . On OSWorld-Verified, policies built only on variants transfer to the held-out original tasks, exceeding AutoRPA by 8.6-17.5 points in Pass^3.

[AI-149] State Transport Routing for Short-horizon Adaptation in Multi-horizon Photovoltaic Forecasting

链接: https://arxiv.org/abs/2609.36926
作者: Xu Yuqing,Zhou Liguo,Sun Ze,Yu Lei,Jiang Mingming
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Recent power measurements provide valuable information for photovoltaic(PV) power forecasting, but directly extrapolating short-term trends can introduce substantial errors over longer forecast horizons. To address this challenge, we propose state transport routing (STR), a lightweight adapter that refines the predictions of a frozen forecasting model. STR combines the original forecast with two complementary trajectories derived from the latest measured power level and its recent trend. A horizon-conditioned router adjusts their contributions over the first 120 min, while leaving subsequent predictions unchanged. Experiments on four public PV datasets show that STR consistently outperforms a parameter-matched residual adapter. On PVDAQ, the same approach improves five neural forecasting backbones, reducing all-horizon normalized mean absolute error by 0.0201-0.2364 percentage points, with paired 95% confidence intervals excluding zero. No reliable improvement is observed for LightGBM. These findings demonstrate the potential of structured state adaptation to improve short-term forecasting across different neural architectures without retraining the underlying models or altering their longer-horizon predictions.

[AI-150] PrecogUI: Proactive GUI Agents via Pre-cognitive Simulation and Experience Retrieval

链接: https://arxiv.org/abs/2609.36923
作者: Bin Kang,Jiarui Ouyang,Li Jiang,Bin Chen,Zhuotao Tian
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Existing reactive Graphical User Interface (GUI) agents often fail in long-horizon, dynamic scenarios, where unexpected disturbances trigger attention-diverting and cascading failures. To address this, we propose PrecogUI, a pre-cognitive architecture that shifts the paradigm from reactive execution to proactive decision-making. Specifically, we design a Proactive Experience Pool (PEP), which caches recurring anomaly and success patterns as “state-action-result” tuples in a dual-memory repository. Furthermore, we introduce a Proactive Simulation Executor (PSE) that learns to forecast the next symbolic UI layout given a candidate action, enabling early anomaly avoidance and ranking candidate actions by predicted reliability. Finally, a Pre-cognitive Execution Controller (PEC) fuses these priors and predictions, prioritizes handling of foreseen anomalies, and ensures execution robustness through a closed-loop error correction mechanism. For robust evaluation, we develop AutoTraj, an automatic data-generation engine, to construct InterfereBench, a benchmark for long-horizon tasks with strong disturbances. Experiments demonstrate that PrecogUI surpasses state-of-the-art methods on InterfereBench while maintaining competitive performance on public benchmarks. The code will be publicly available.

[AI-151] AeroManip-VLA: Scalable Vision-Language-Action Learning for Aerial Manipulation with RL-Generated Demonstrations

链接: https://arxiv.org/abs/2609.36915
作者: Rui Huang,Yanlin Mu,Lidong Li,Yucong Wang,Zichen Yan,Lin Zhao
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: this https URL

点击查看摘要

Abstract:Aerial manipulators extend robotic manipulation into 3D workspaces that are difficult for ground-based robots to access, creating new opportunities for general-purpose manipulation. However, extending Vision-Language-Action (VLA) models to aerial robots introduces distinct challenges due to the tight coupling between manipulation and flight, continuously changing observations, and safety-critical physical interactions. These challenges demand diverse training data and systematic policy evaluation, yet collecting demonstrations and evaluating policies directly on physical aerial platforms are costly, difficult to scale, and hard to repeat under controlled conditions. We present AeroManip-VLA, a scalable benchmark for aerial VLA data generation and policy evaluation. AeroManip-VLA provides a GPU-accelerated simulation framework with low-level payload-aware flight and manipulation control in massively parallel environments. Building on this framework, we combine reusable reinforcement learning policies with expert task rules to automatically generate demonstrations without human teleoperation across diverse objects, environments, and randomized initial conditions. The generated data include basic skills such as grasping and placing, as well as long-horizon tasks that require both navigation and manipulation. We further introduce automated event labeling and trajectory categorization to filter demonstrations. These mechanisms enable fine-grained analysis of task progress, behavioral outcomes, and safety-related failures. Finally, we evaluate a range of imitation learning and VLA baselines across different task settings, revealing their performance characteristics and failure modes. Together, AeroManip-VLA enables scalable aerial manipulation data generation, structured trajectory analysis, and systematic VLA evaluation in simulation prior to real-world deployment.

[AI-152] STAR-GRPO: Canonical Anchoring and Reliability-First Advantages against Representation-Dependent Reward Hacking

链接: https://arxiv.org/abs/2609.36900
作者: Wan Tian,Zhongyi Li,Xiang Xu,Minhao Zou,Yijie Peng,Fuzhen Zhuang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Reward hacking occurs when policy optimization exploits a brittle reward interface or an overly permissive proxy objective, improving the training score without improving the underlying response quality. This phenomenon is amplified in group-relative policy optimization: an unsupported reward can shift the group baseline and alter the updates of other rollouts, while post-hoc or purely relative weighting cannot represent group-wide uncertainty. We propose \emphSelf-Tuned Anchored Reliability Group-Relative Policy Optimization (STAR-GRPO), a reliability-first advantage estimator based on paired assessments of the same rollout. STAR separates the quality signal from its learning influence: score disagreement determines rollout reliability, relative reliability enters a self-tuned robust location–scale fit before group normalization, and absolute group reliability attenuates the resulting bounded advantage. The analysis establishes coordinate and second-moment bounds, characterizes exact centering through the weighted location equation, and gives reliability-dependent attenuation guarantees for outlying rewards. We evaluate STAR-GRPO in two complementary reward-hacking regimes. In token-interface exploitation, STAR prevents runaway optimization of the deployed-interface score while improving the canonical quality signal. In rubric-proxy overoptimization for medical reasoning, STAR improves independent semantic evaluation, narrows the proxy–judge discrepancy, and reduces overclaim while optimizing the same task proxy. Together, these results show that reliability-first normalization offers a principled way to limit unsupported reward influence on both group baselines and policy updates, while retaining the task reward as the optimization target.

[AI-153] HorizonFlow: Variable-Length Planning for Offline Goal-Conditioned RL

链接: https://arxiv.org/abs/2609.36896
作者: JunHyeok Oh,Zian Jang,Byung-Jun Lee
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Recent advances in generative planning have made trajectory inpainting a promising approach to offline goal-conditioned reinforcement learning. However, these methods typically specify the planning horizon before generating plan content, even though the appropriate horizon depends on the route itself. A horizon that is too short can force infeasible transitions, whereas one that is too long can introduce redundant motion. We introduce HorizonFlow, a hierarchical planner that treats plan length as an output of generation rather than a prescribed input. Its subgoal route planner guides its action-prefix controller through a sequence of latent subgoals. Both components combine insertion-based generation with flow matching to jointly generate continuous plan content and length, using the partially generated plan to guide token insertion. HorizonFlow reuses the resulting length information to select candidates and steer generation toward shorter plans without a separate learned value model. Across Maze2D, Multi2D, and OGBench navigation and visual manipulation benchmarks, HorizonFlow achieves the highest average performance among the compared methods.

[AI-154] Beyond Sub-Gaussian Detector Scores: Robust Weighted Profile-Loss Change Point Detection for Human-LLM Text Segmentation

链接: https://arxiv.org/abs/2609.36888
作者: Wan Tian,Zhongyi Li,Yawen Li,Rui Zhang,Yijie Peng,Fuzhen Zhuang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Mixed human-LLM documents require locating authorship transitions from detector scores whose reliability varies across text units. Existing weighted mean contrasts are vulnerable to extreme scores, while directly replacing means with robust centers obscures how a misplaced boundary changes the population objective. We propose Robust Weighted Profile-Loss Change Point Detection (RWCP), which combines capped reliability weights, Huber profile gains, and narrowest-over-threshold search in reliability coordinates. Our key analysis expresses the population gap between a true and a displaced split as a merge cost, avoiding a closed-form solution for the nonlinear center of a mixed segment. Under explicit curvature, spacing, and dependence conditions, core RWCP recovers the number of changes and localizes their boundaries; its quadratic-loss limit recovers squared weighted CUSUM. We also study RWCP-R, a separately evaluated decoder that shares source centers across nonadjacent passages. Across five retrospective cached-score benchmark families, core RWCP reduces family-macro WindowDiff by 17.6% relative to weighted change-point detection, and RWCP-R lowers it further. Boundary recovery improves most clearly for isolated changes, while both fixed configurations miss changes in collaborative and densely alternating text.

[AI-155] WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents

链接: https://arxiv.org/abs/2609.36887
作者: Bo Mao,Hang He,Linting Wang,Lizhi Lin,Maosen Zhou,Guanming Liu,Jinxiu Liu,Tianyu Huai,Chaoyun Zhang,Bingxuan Li,Kepeng Lei,Guanting Dong,Zhou Shao,Rui Zheng,Hang Yan,Jie Zhou,Chengcheng Wan,Tao Gui,Liang He,Xipeng Qiu
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Recent efforts to scale tool-use post-training have largely centered on the synthesis of executable environments, which constitute only one component of a broader agentic interaction system comprising the environment, task, agent harness, and evaluator. Scaling environments in isolation, however, does not guarantee commensurate gains in model performance, because reliable learning signals depend on coherent interactions among all components of the agentic interaction system. To address this problem, we introduce WEFT (Whole-system Evolution For Tool-use Post-training), which couples scalable agentic interaction system construction, execution-driven self-evolution, and stable post-training. WEFT scales agentic interaction system construction across environment breadth, task complexity, and interaction diversity. Execution-driven self-evolution iteratively uses execution traces and state evidence to attribute failures and revise the responsible components, with fresh rollouts evaluating the changes and providing evidence for subsequent evolution rounds. For stable post-training at scale, WEFT addresses both optimization and execution reliability: prefix-preserving sampling retains verified progress and atomic-turn credit assignment localizes learning signals, while MegaMCP maintains isolated, recoverable state across concurrent rollouts over shared tool services. Extensive experiments across various models and benchmarks demonstrate the effectiveness of WEFT for tool-use post-training. WEFT-8B and WEFT-14B outperform all evaluated matched-size environment-scaling baselines on BFCL V4, \tau^2 -Bench, and Claw-Eval. In particular, WEFT-14B improves over Agent-World-14B by 6.41, 2.23, and 12.27 percentage points. WEFT-35B-A3B further extends these gains to more challenging long-horizon workflow benchmarks, including Toolathlon-Verified and AutomationBench.

[AI-156] SKILLLITE: Evidence-Guided Malicious Skill Auditing with Compact LLM s

链接: https://arxiv.org/abs/2609.36879
作者: Haoran Ou,Gelei Deng,Xuanye Zhang,Wenbo Guo,Tianwei Zhang,Kwok-Yan Lam
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:As LLM-based agents perform increasingly complex tasks, Agent Skills have emerged as a flexible mechanism for extending their capabilities. An Agent Skill packages task-specific instructions with executable components and auxiliary resources to provide specialized functionalities. However, the growing adoption of third-party Skills introduces a new supply-chain attack surface. Malicious Skills can embed harmful behaviors that abuse agent privileges and compromise the agent execution environment or accessible resources. Although recent LLM-based malicious Skill auditing approaches have achieved promising performance, they often rely on capable commercial LLMs. How to achieve effective auditing with compact, locally deployable LLMs in security-sensitive and resource-constrained settings remains largely unexplored. Our investigation reveals that compact LLMs struggle to identify malicious behaviors hidden in complex Skill packages. This difficulty arises from both the implicit nature of such behaviors and the limited reasoning capacity of compact LLMs. To address these challenges, we propose SKILLLITE, an evidence-guided agentic framework for malicious Skill detection. SKILLLITE effectively extracts security-relevant behaviors and infers the intended functionality from complex Skill packages. It then employs a compact LLM to assess the maliciousness of the Skill based on the observed behaviors and their functional context. Experiments show that SKILLLITE improves malicious Skill detection across different compact LLM backbones and outperforms existing representative auditing baselines. Its effectiveness generalizes to behaviorally confirmed in-the-wild malicious Skills. Meanwhile, SKILLLITE maintains a low inference latency, supporting its practical deployment.

[AI-157] State Trace Rationale As Auxiliary Task in Reinforcement Learning ICLR2027

链接: https://arxiv.org/abs/2609.36867
作者: Muhammad U. Nasir,Alex Vogt,Steven D. James,Julian Togelius
类目: Artificial Intelligence (cs.AI)
备注: Under review as a conference paper at ICLR 2027

点击查看摘要

Abstract:We propose STRAT, an auxiliary task that trains deep reinforcement learning (RL) agents to predict a short textual trace of their own state. Inspired by human spatial navigation, the description combines landmark, route, and survey knowledge, tracking the agent’s position, inventory, goals, and immediate progress. Environment rules generate this text online without human labelling. Our method adds a single auxiliary head to a standard policy. Across 60 sparse-reward XLand-MiniGrid tasks, STRAT solves complex environments where standard RL fails outright, while compacting state representations and preventing rank collapse. Beyond performance gains, the predicted trace provides a readable account of agent beliefs at every step for no extra cost.

[AI-158] IronLLM : Forging Compact Edge-Native Language Models for Real-Time Embodied Intelligence

链接: https://arxiv.org/abs/2609.36860
作者: Changdi Yang,Fengquan Jiao,Haochih Lin,Haoran Yang,Jing Xiao,Liangyu Huo,Suxin Lu,Tiance Chen,Wei Liu,Yinggan Xu,Yunxiang Lu,Zai Zheng,Zhirui Xie,Zhongyang Che,Ziyan Tang,Zuoxiang Zhao,Jian Yao
类目: Artificial Intelligence (cs.AI)
备注: Technical report

点击查看摘要

Abstract:We present IronLLM-0.6B, a 654M-parameter language model designed for efficient on-device inference. IronLLM-0.6B combines a hybrid attention architecture with X-MTP, a lightweight shared-KV multi-token prediction design that eliminates per-depth KV-cache replay and employs a lightweight verification head for rollback-free drafting, achieving a 1.48x decoding speedup. The model is pretrained on approximately 6.2 trillion tokens using a quality-oriented data pipeline and is further post-trained with Multi-Domain On-Policy Distillation to integrate capabilities from domain-specialized teachers. To better meet the low-latency requirements of on-device scenarios, IronLLM-0.6B adopts an Instruct-Only design. Evaluations show that IronLLM-0.6B achieves competitive performance relative to larger models such as Qwen3.5-0.8B and MiniCPM5-1B, while producing more concise responses on many tasks. We further present IronLLM-0.6B-Light, which replaces RMSNorm with Dynamic Tanh and simplifies several computationally expensive components to improve inference and quantization efficiency. Together, the IronLLM models provide an effective performance-efficiency trade-off for resource-constrained deployment.

[AI-159] When Upstream Messages Override Correct Answers: A Controlled Study of Multi-Agent LLM Collaboration

链接: https://arxiv.org/abs/2609.36855
作者: Yaxin Gong,Gangyi Zhang,Chongming Gao,Leyang Shen,Chenxiao Fan,Jiakai Wang,Dong Wang,Yang Liu,Wenjie Wang,Xiangnan He
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Multi-agent LLM systems rely on message passing among specialized agents to accomplish complex tasks. However, an upstream agent may provide useful information or an incorrect answer that causes a downstream agent to override a correct answer supported by its own evidence. Prior work has not clearly separated the benefits of communication from the damage caused by incorrect messages. We study this problem with controlled experiments across five benchmarks and five receivers, keeping the downstream task and evidence fixed while comparing answers under three conditions: no message, the upstream agent’s original message, or a message with the opposite conclusion. Our experiments reveal three key findings. First, messages often help when the downstream agent would otherwise answer incorrectly. Second, messages can also hurt: when the downstream agent would answer correctly without a message, an incorrect upstream message changes the answer in up to 32% of cases. Third, in 94% of audited harmful cases, the downstream agent copies the upstream’s specific wrong answer–a pattern we term answer substitution. Removing unreliable messages recovers part of the lost accuracy, suggesting that communication should be selective based on upstream reliability and the evidence already available to the downstream agent.

[AI-160] Automated Screw Planning for Reduced Pelvic Fractures Based on Statistical Shape Models and Deep Learning

链接: https://arxiv.org/abs/2609.36847
作者: Yang Gao,Sutuke Yibulayimu,Yanzhen Liu,Zian Zhao,Yudi Sang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Percutaneous iliosacral screw fixation is an important minimally invasive treatment for unstable pelvic fractures. Because the sacroiliac region has complex anatomy and narrow screw corridors, the accuracy and safety of screw placement directly affect surgical outcomes. Accurate and reliable preoperative screw planning is therefore essential to improve surgical success and reduce intraoperative risks. Conventional preoperative planning typically requires surgeons to determine screw trajectories through manual measurements, a labor-intensive process that depends on subjective clinical experience. To address these challenges, we propose a fully automated pipeline for preoperative iliosacral screw planning in patients with pelvic fractures. Using patient-specific three-dimensional anatomy, the pipeline automatically identifies safe screw corridors and generates individualized insertion trajectories to support clinical preoperative planning. We evaluated the proposed pipeline on 200 clinical cases of pelvic fractures. Compared with conventional manual measurements, the safety margin of the safe insertion corridors increased by 2% across the four screw types, the mean planning time decreased by more than 90%, and the clinical acceptance rate reached 95%.

[AI-161] DSWM: Decomposed Spatio-Temporal World Model for Demand-Driven UAV Base Station Repositioning

链接: https://arxiv.org/abs/2609.36845
作者: Shengjie Zhong,Zhongliang Zhao,Jingxuan Chen,Xianbin Cao,Xinmei Qiang,Dapeng O. Wu,Tony Q. S. Quek
类目: Networking and Internet Architecture (cs.NI); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 13 pages, 13 figures, 4 tables, 2 algorithms. Submitted to IEEE Journal on Selected Areas in Communications (Special Issue on Agentic AI for Intelligent Networks)

点击查看摘要

Abstract:Uncrewed aerial vehicle base stations (UAV-BSs) are expected to cover traffic demand that shifts across space and time, yet most repositioning schemes either re-solve an optimization problem per slot or learn reactive policies without an explicit demand model. We cast demand-driven fleet repositioning as latent-space decision-time planning and propose DSWM, a decomposed spatio-temporal world model: an agentic controller that perceives the demand field through a rolling observation window, retains operational context in a latent recurrent state, reasons about candidate motions by imagined rollouts under an uncertainty penalty, and coordinates the fleet through replanned first actions. DSWM learns a recurrent state-space model shaped by an exponential-moving-average (EMA) based latent predictive objective with variance regularization. It attaches a differentiable service simulator that replays the association, probabilistic line-of-sight channel, and Shannon rate chain inside latent rollouts. Planning uses a cross-entropy method whose imagined demand is anchored on the current observation window with mixing coefficient \rho=0.95 . On a unified pipeline over three real datasets (Milan CDR (call detail record), Shanghai Telecom, YJMob100K) and 14 methods including five reproduced IEEE baselines, DSWM attains weekday served ratios of 0.889, 0.908, and 0.898, ranking first among non-ablated configurations on every dataset. On Milan it improves over the strongest non-learning baseline (Greedy, 0.780) by 0.109, a margin that comes from decision-time use of observations rather than prediction accuracy.

[AI-162] ARC-KV: Amortizing Anchor Search for Reconstruction-Based KV Cache Compaction

链接: https://arxiv.org/abs/2609.36835
作者: Zheyu Shen,Guanhua Wang,Dezhan Tu,Mengchi Zhang,Yanjia Li,Adnan Aziz,Chunqiang Tang,Ang Li
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Long-context large language model inference is bottlenecked by KV caches that grow linearly with sequence length. This burden is especially severe for long, reusable context prefixes, whose cache must serve many downstream queries. Reconstruction-based methods such as Attention Matching achieve strong downstream task performance with compact KV caches. However, iterative anchor search dominates the compaction cost of OMP-based Attention Matching. This motivates our selective amortization principle of learning a reusable anchor-selection policy across contexts while retaining context-specific reconstruction. In this work, we propose ARC-KV, a novel reconstruction-based KV cache compaction method that follows this principle. To this end, we first train a value-aware indexer to select real-key anchors in a single scoring pass. ARC-KV then applies convex-hull-constrained key merging and fits an attention-mass bias and compact values against the full cache. At inference time, ARC-KV builds the compact cache once per context using the frozen indexer and reuses it for all subsequent queries. Extensive experiments demonstrate that ARC-KV outperforms reported compaction methods in most settings across QuALITY, RULER, and LongBench on Llama-3.1-8B-Instruct. In particular, at 10% KV retention on QuALITY, ARC-KV improves accuracy from 0.6409 to 0.6474 over Attention Matching while reducing compaction time by a factor of 25.73, from 959.8 s to 37.3 s.

[AI-163] Where Does Staleness Accumulate? Pool Aware Effective Staleness Control for Asynchronous RL in LLM Post-Training

链接: https://arxiv.org/abs/2609.36830
作者: Chenliang Li,Neiwen Ling,Zijun Wei,Alfredo Garcia
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Fully asynchronous reinforcement learning (RL) improves resource utilization in large language model post-training by overlapping rollout generation with policy optimization, but it also introduces policy lag as trajectories are generated and queued while the trainer continues to update. We study how this lag accumulates over a trajectory’s lifetime and how it can be controlled without sacrificing the wall-clock benefits of asynchronous execution. We decompose trajectory staleness into Generation Staleness, accumulated before rollout completion, and Waiting Staleness, accumulated after a completed trajectory enters the pool. Motivated by this decomposition, we introduce PACE (Pool-Aware Control of Effective Staleness). PACE converts excess pool occupancy into an adaptive rejection budget and ranks completed trajectories using an effective-staleness score that combines Waiting Staleness with prefix-aware Generation Staleness. This avoids penalizing long or interrupted rollouts solely because they span multiple policy versions. In single-turn mathematical reasoning, PACE improves the six-benchmark average validation accuracy by 18.7% over unfiltered asynchronous RL at the same wall-clock budget and matches synchronous RL performance with 47.1% less GPU time. PACE also improves validation performance in multi-turn tool-integrated reasoning, outperforming both synchronous and unfiltered asynchronous RL. Further experiments with the mixture-of-experts model and an alternative RL algorithm support its applicability across model architectures and training algorithms.

[AI-164] he Default Trap: Rethinking Plan Evaluation in Tool-Using LLM Agents

链接: https://arxiv.org/abs/2609.36829
作者: Xueqi Li,Jingjie Ning,Yibo Kong
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:An executor can respond strongly to a change in a supplied plan’s priority while showing a small change in the same information-selection probability when a default-aligned whole plan is removed. We call the risk of interpreting the latter as weak responsiveness to alternative priorities the default trap. We compare paired plans that prioritize different information targets with a shared no-plan reference. An accounting identity relates these distinct behavioral contrasts. Across 3,200 decision windows on 160 selected Retail, Airline, and AgentDojo tasks, switching priorities strongly redirects two models’ choices, while the two plan-versus-default contrasts differ. In 2,160 additional windows, reversing account-list order shifts default target selection by 63.3-98.3 percentage points; priority-switching effects remain 96.7-100.0 points in either order. A separate 3,240-window component study finds strong control under single priority sentences, with effects of additional text varying by group and direction. Finally, 1,080 full-task episodes yield observed success differences of -19.4 to +8.3 points relative to no plan. All Retail and Airline success intervals include zero; AgentDojo results describe four fixed application worlds. These findings support joint reporting of priority responsiveness, presentation-dependent defaults, and task success and cost.

[AI-165] Calibrate the Decisions That Change the Future: On-Policy Post-Training Quantization for Multimodal Large Language Models

链接: https://arxiv.org/abs/2609.36828
作者: Wenxiao Fan,Jingling Fu,Lichen Ma,Yu He,Luohang Liu,Jinbao Xue,Ke Zhang,Junshi Huang,Kan Li
类目: Artificial Intelligence (cs.AI)
备注: preprint

点击查看摘要

Abstract:Post-training quantization (PTQ) lowers deployment cost for multimodal large language models, but calibration typically reconstructs fixed sequences with local objectives. This overlooks autoregressive feedback: a quantization-induced token change redirects the prefix and changes future states. Yet on-policy coverage alone is insufficient because many decision mismatches barely affect future generation. We propose OnPTQ, an on-policy framework that calibrates on trajectories visited by the current quantized policy. On shared prefixes, OnPTQ identifies quantization-eroded boundaries, evaluates competing tokens through short counterfactual rollouts, and combines current discrepancy with branch consequence into a Decision–Consequence risk. The risk prioritizes critical states, while context anchoring and trajectory refresh preserve multimodal behavior and keep calibration aligned with the updated policy. We further derive a Decision–Consequence bound linking behavioral deviation to current policy discrepancy and action-conditioned future-value span. Across vision–language and omni-modal Qwen models under multiple low-bit settings, OnPTQ improves downstream performance and yields fewer correctness flips against the corresponding Dense/FP16 references, without changing the deployed inference graph.

[AI-166] Diffusion Policy Improvement with Proposal-Conditioned Refinement Flows

链接: https://arxiv.org/abs/2609.36812
作者: Junhyun Ha,Juho Lee,Byungwoo Park
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注: 27 pages, 10 figures

点击查看摘要

Abstract:Diffusion and flow policies can model complex behaviors in offline reinforcement learning (RL). However, penalizing their KL divergence from the behavior policy can discourage actions having high critic values with low behavior density. Directly refining behavior proposals may be an alternative, yet Gaussian or deterministic editors limit expressiveness to represent multiple separated modes for the same proposal. In this work, we introduce Proposal-Conditioned Refinement Flows (PReFlow), a policy extraction method combining critic-based proposal selection with a conditional refinement flow. To optimize proposal selection and refinement together, we formulate a KL-regularized objective whose optimum induces a Gibbs policy over final actions under a Gaussian-smoothed behavior prior. The refinement flow can represent multiple high value modes, while a proposal-centered Gaussian reference regulates large action changes. This Gaussian reference further enables us to make use of simulation-free, closed form adjoint matching targets from sampled endpoints and critic gradients, yielding a single velocity regression loss without a backward adjoint solve. On 50 OGBench tasks, PReFlow achieves competitive offline performance and the highest aggregate score among the compared methods after online fine-tuning, reaching 91% after 500K environment steps.

[AI-167] Geometry-Conditioned Fixed-Scaffold Encoders for Time-Warp Robust Sequence Retrieval

链接: https://arxiv.org/abs/2609.36809
作者: Cassandra Yang,Yufan Tang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Embedding-based retrieval is attractive for long sequence collections because each item can be encoded once and searched by nearest-neighbor ranking. The difficulty is that the objects being indexed are often observed under a noncanonical clock: cardiac cycles stretch with rate, speech changes with tempo, and sensor traces reach comparable states at different speeds. This paper studies a specific source of instability in patch-based encoders for this regime. If patch boundaries are chosen from signal geometry, then the tokenization can change under the same temporal deformation that the representation is expected to tolerate. We propose GeoPatch, a fixed-scaffold patch encoder that keeps token support independent of geometry and uses slope, curvature, acceleration, affine-residual, and confidence descriptors only as continuous conditioning variables. The design turns boundary variation into feature modulation: geometry can change the embedding through a controlled pathway, but it cannot change the number, order, or support of local tokens. We formalize this distinction through a mechanism-level stability analysis that separates boundary drift, affine timing variation, confidence-weighted geometry perturbation, and retrieval-margin effects. The same local tokens support global embedding retrieval and late-interaction scoring, so the scoring rule can be matched to the evaluation protocol. Across ECG, speech, and multivariate time-series retrieval tasks, GeoPatch improves early-rank retrieval under timing variation while exposing a clear trade-off between local surface matching and strict non-overlap retrieval.

[AI-168] Spotter: Let the Embodied Model Lead and the VLM Reflect for It

链接: https://arxiv.org/abs/2609.36808
作者: Long Li,Qichao Zhao,Yue Yang,Fan Xu,Zhe Wang,Alan Wee-Chung Liew,Chao Qu,Heng Tao Shen,Shirui Pan
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: 19 pages, 7 figures, 6 tables. Code: this https URL

点击查看摘要

Abstract:Current embodied models do not respond to their own failures, although what just went wrong could inform a small adjustment on the next attempt, the kind of reflection behind the gains of thinking in language models. We test whether they can repair a known error, which requires producing a correction and judging whether it is right. Stopped at a failure and allowed to retry, they seldom repair it through their own randomness or from a language description of the error, and best-of-N selection cannot pick the successful candidate after a failure. We attribute this to training only on successful demonstrations and to inputs too narrow to show what went wrong, and conclude that reflection must come from a vision-language model (VLM), which takes in far more information, such as the episode history and text, and is more general. Prior VLM-led work has the VLM plan every step and invoke the embodied model as a tool, placing the VLM on the critical path. We propose Spotter, which reverses the roles: the embodied model leads and executes continuously, while the VLM runs in parallel, monitors through a lightweight local screener, intervenes only when an error is detected, reflects on and corrects it, and returns control. We run Spotter with Qwen and with GPT as the VLM, and both improve the embodied models; with GPT, Spotter improves Cosmos Policy and \pi_0.5 by 5.6 and 7.5 percentage points on RoboCasa, and raises \pi_0.5 from 47.2% to 57.0% on the Hard setting of RoboTwin 2.0 and from 53% to 83% on a real robot. Because the VLM steps in only when an error is confirmed, a successful episode with Qwen takes only 13 to 16 s longer than with the embodied model alone and about 70% less time than with a VLM-led baseline using the same model. Our code is available at this https URL.

[AI-169] CAD-Native Transformer Operators for AI-Aided Engineering

链接: https://arxiv.org/abs/2609.36806
作者: Daniel Leibovici,Nikola Borislavov Kovachki,Dawon Ahn,Ruben Ohana,Ira J. S. Shokar,Abouzar Ghasemi,Semih Akkurt,Rishikesh Ranade,Neil Ashton,Jan Kautz,Jean Kossaifi
类目: Artificial Intelligence (cs.AI)
备注: 25 pages, 11 figures, 15 tables

点击查看摘要

Abstract:Modern engineering systems, from automobiles to aircraft, are designed by using precise, continuous parametric computer-aided design (CAD) models. Evaluating design changes through numerical simulation requires meshing the continuous geometry, a computationally expensive and often brittle process that can require manual intervention and replaces the continuous representation with a discrete approximation. Most neural surrogates accelerate the simulation, but inherit this representation gap by relying on meshes, point clouds, voxels, or other sampled approximations of geometry. We introduce CANTO, a transformer neural operator that maps directly from continuous CAD geometry to physical fields, without meshing the input geometry. We develop a theoretical framework for learning operators from geometric manifolds to function spaces of physical fields, representing geometry through sequences of parametric patches. CANTO instantiates this framework by directly tokenizing non-uniform rational B-spline (NURBS) patches from their control points, knot vectors, and weights, and predicts continuous surface and volume fields at arbitrary query locations. We evaluate CANTO on four automotive and aircraft aerodynamics industry benchmarks: AhmedML, WindsorML, DrivAerML, and HiLiftAeroML. CANTO achieves state-of-the-art accuracy on most evaluated surface and volume prediction tasks, including a 19.8% reduction in surface-pressure relative L_2 error compared with AB-UPT on HiLiftAeroML. Differentiability with respect to CAD parameters further enables gradient-based inverse design of designs. On AhmedML, CANTO identifies designs with 4.4 to 20.4% lower drag than the best dataset designs satisfying the same volume and lift constraints, with the improvements verified using the same CFD setup used to generate the original dataset.

[AI-170] UpliftMem: Learning Set-Level Uplift for Agent Memory Retrieval

链接: https://arxiv.org/abs/2609.36805
作者: Mengkun Liang,Haoran Qiang,Guannan Liu,Junjie Wu
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language model (LLM) agents reuse external memory to guide new tasks, but effective retrieval requires learning which memory sets improve execution. Such learning relies on costly outcome feedback: ordinary retrieval observes only executed sets, while evaluating alternatives requires additional rollouts. We introduce \textscUpliftMem, which learns memory retrieval from set-level execution uplift relative to the same executor without memory. A theoretical analysis of how retrieval preferences restrict feedback coverage motivates targeted probing of alternative memory sets. Probe selection follows an expected value of sample information (EVSI) criterion, derived in closed form under a correlated Gaussian model, to allocate limited training rollouts according to their expected improvement in local retrieval decisions. The shared scorer is trained with a frozen executor and selects memory sets without test-time probes. Across ALFWorld, WebShop, and BigCodeBench, \textscUpliftMem achieves the best success rates among evaluated baselines on the main evaluation sets. Controlled fixed-store and matched probe budget evaluations further demonstrate improved memory-use decisions and more effective use of execution feedback.

[AI-171] AI as a Compiler: Compiling Triton kernels without the Triton compiler

链接: https://arxiv.org/abs/2609.36800
作者: François Costa,Charly Castes,Thomas Bourgeat,Azalia Mirhoseini
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Compiler backends are expensive to build and maintain as programming models, workloads, and accelerators evolve. We investigate whether large language models can replace the conventional optimizing and lowering pipeline, a process that we call AI lowering. We study AI lowering from Triton to NVIDIA PTX: an LLM agent translates Triton kernels directly into PTX. We build an environment that evaluates candidate PTX, and an agentic harness in which an LLM translates Triton kernels into PTX. Across twelve common kernels on Ada, Hopper, and Blackwell GPUs and ten kernels from recent ML papers, AI lowering achieves 0.83x-3.34x the performance of autotuned Triton. The largest gains come from transformations that Triton’s lowering pipeline does not perform, such as decoding packed binary weights directly into Tensor Core operands (3.34x on BitDelta), assigning each thread a complete softmax row in tensor memory (1.37x on FlashAttention), and reusing overlapping convolution windows (up to 2.23x). These results rely on a robust evaluation harness with comprehensive verification support. We build on Volta, an existing PTX verifier, and substantially extend it to support modern GPU architectures by introducing support for Blackwell’s tcgen05 Tensor Core interface. This requires modeling three architectural features: managed tensor memory, descriptor-based operand layouts, and asynchronous execution coordinated through commits, waits, memory barriers, and proxy fences. We discuss the challenges involved in formalizing them, as well as the current limitations. Our results suggest an emerging future in which AI compilers replace custom-written intermediate representations and checkers, reducing the time and engineering effort required to bring up software for new general-purpose and custom chips.

[AI-172] Aperture: Merge-Consistent Rotary States for Compressed Tokens

链接: https://arxiv.org/abs/2609.36781
作者: Yuhao Du,Shunian Chen
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Token compression combines content from several positions, yet rotary position embeddings usually assign the merged token one coordinate. We ask what positional information must survive later merges. Aperture stores Fourier moments of the token’s weighted support at the model’s rotary frequencies. We prove that these moments have minimal real dimension among continuous states sufficient for the selected expected rotary interactions. Represented mass makes updates additive; attention normalisation remains a separate readout choice. Uniform intervals give a centre rotation times a sinc gain. We characterise when centres determine interval widths and construct matched examples where they do not. Numerical checks verify the weighted-support implementation. In trained temporal readers, compression transfer varies with gain calibration and feature placement. In a prespecified native video question-answering comparison, stored support reaches 65.63% accuracy versus 67.12% for the deployed merging rule. These results separate exact positional preservation under compression from downstream benefit.

[AI-173] Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes

链接: https://arxiv.org/abs/2609.36777
作者: Xiaokang Ye,Siddhant Hitesh Mantri,Zimeng Chen,Edward Zhang,Zhaoxu Zheng,Yuanheng Li,Yizhao Chen,Tianyang Huang,Lianhui Qin
类目: Artificial Intelligence (cs.AI)
备注: 44 pages, including appendices

点击查看摘要

Abstract:Frontier coding agents can now write and execute code that authors 3D environments, but whether they reliably understand 3D structure and precisely control scene state remains unclear. The generated 3D scene is a persistent, executable artifact: a convincing render can hide incorrect spatial relations, intersecting objects, or unintended modifications. We introduce Code4Scene, a benchmark of 190 Unreal Engine cases built from human-assembled scenes that evaluates coding agents on two complementary settings under a shared execution interface. Construction tests scene-level spatial reasoning from open-ended language specifications, where many realizations are valid; editing tests precise control of scene state, where the agent must recover the target scene from reference images while preserving everything else. Rather than scoring code or rendered views, Code4Scene evaluates the generated engine-native scene for task fulfillment, artifact integrity, and static physical validity, with edits additionally compared against withheld ground truth. Across 14 coding-agent configurations on the 95-case public set, construction and editing performance are strongly correlated but not interchangeable (Spearman \rho = 0.78 ): Claude Fable 5.1 leads construction, Gemini 3.8 Flash leads editing, and GPT-6 Astra narrowly leads overall. Spatial Composition is the weakest construction category for every agent, while editing remains imprecise: the best Repair F1 is only 0.527, and 35.8% of edits that fully recover the target still introduce unintended changes elsewhere in the scene. These results expose a gap between plausible 3D generation and reliable spatial reasoning and state control.

[AI-174] Beyond Conditional Independence: Root Cause Analysis with Deep Causal Models

链接: https://arxiv.org/abs/2609.36771
作者: Md Musfiqur Rahman,Kenneth Lee,Ziwei Jiang,Padmaja Jonnalagedda,Ruocheng Guo,Murat Kocaoglu
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
备注:

点击查看摘要

Abstract:Root cause analysis (RCA) is a critical problem in many real-world scenarios. RCA enables the identification of faulty or failing mechanisms in a system by comparing anomalous observations with corresponding reference (i.e., regular) observations. However, existing approaches rely either on heuristic methods or on conditional independence tests with a strong unconfoundedness assumption, and thus fail to exploit other complicated distributional constraints in the presence of latent variables. To relax these assumptions, we model the underlying system as a causal model and the anomalous system as a change in the structural functions of the same causal model. Specifically, to handle unobserved confounders, we establish an implicit connection between distributional constraint testing and root cause analysis. To adapt our approach to data generated from arbitrary causal models, we employ the deep causal model (DCM) framework, in which we design the causal model using neural networks. Finally, we illustrate how our method, RCA-DCM, can utilize different levels of partial graphical knowledge to perform RCA. We evaluate RCA-DCM against state-of-the-art baselines on simulated datasets, a physics-based causal chamber and two micro-service applications. RCA-DCM improves top-1 accuracy over the strongest baseline on both Sock Shop (0.880 vs. 0.752) and Online Boutique (0.776 vs. 0.712), and when the true root cause in the causal chamber is unobserved and acts as a latent confounder, it recovers the exact root-cause set more often than any competing method (perfect recovery rate (PRR) 0.846 vs. 0.731).

[AI-175] Generalizable Lifelong Model Editing via Preference Optimization EMNLP2026

链接: https://arxiv.org/abs/2609.36748
作者: Dahyun Jung,Suhyune Son,Heuiseok Lim
类目: Artificial Intelligence (cs.AI)
备注: Accepted to EMNLP 2026 Main Conference

点击查看摘要

Abstract:Knowledge editing enables rapid updates of specific factual knowledge in large language models (LLMs) without full retraining. However, more realistic scenarios call for a lifelong framework that handles continual updates rather than one-off modifications. In such settings, existing editing methods often overfit to target prompts, significantly degrading both the generalization of the edited knowledge and the model’s general capabilities. To address this issue, we propose GLIME (Generalizable Lifelong Model Editing), which combines knowledge editing with preference optimization over generation behavior. GLIME further incorporates replay-based editing and a gradient constraint to preserve previously edited knowledge. Experimental results show that GLIME significantly improves knowledge generalization in lifelong editing settings while maintaining both editing performance and general capabilities.

[AI-176] EASE: Behavior-Adaptive Skill Curation for Self-Evolving Agents

链接: https://arxiv.org/abs/2609.36746
作者: Zhen Xiong,Qiaoyu Tan
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Agent skills provide a lightweight mechanism for self-evolving agents to accumulate reusable procedural knowledge without updating model parameters. However, existing learned skill curators typically optimize curation without explicitly modeling downstream executor behavior. We show that this can cause systematic cross-executor degradation: curators trained with different executors perform best when paired with their own training executor, indicating that effective skill curation is executor-dependent. We formulate behavior-adaptive skill curation and introduce EASE, a framework that learns a single curator that adapts its decisions to different executor behaviors. EASE maintains an online behavioral profile of recent execution patterns and conditions the curator on this profile, the current trajectory, and retrieved skills to add, modify, or remove skills from an evolving repository. We train the shared curator jointly across multiple frozen executors with reinforcement learning, using retrieval-aware and behavior-aware temporal attribution to focus optimization on curation actions with observable downstream influence. Across ALFWorld, ScienceWorld, and WebShop, with executors ranging from Qwen3-8B/32B and GPT-OSS-120B to unseen Kimi K2.6, DeepSeek V4 Flash, and Gemini 3.5 Flash, EASE outperforms strong skill- and memory-based baselines without per-executor finetuning. EASE also maintains 34.5–41.0% fewer skills, improves skill retrieval by 36.3–38.7% and measured edit utility by 51.8–60.0%, and reduces deployment-time inference tokens by 9.1–14.5%. These results establish behavior-adaptive skill curation as an effective principle for building self-evolving agents.

[AI-177] Distinguish or Homogenize: Last-Chance Policy Identification and Risk-Budgeted Recovery under Irreversible Resource Depletion

链接: https://arxiv.org/abs/2609.36741
作者: Yibo Guo,Xiaodan Wang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Under irreversible resource depletion, an agent can spend resources to distinguish among latent fault models, or to change the system state so that the remaining models admit a common acceptable continuation–at which point further diagnosis becomes unnecessary. This distinguish-or-homogenize principle identifies a path that existing frameworks for identification, planning, and diagnosis do not make explicit: prior formulations treat the mapping from fault models to acceptable policies as a given, whereas LCPI makes it a function of the agent’s own actions. We formalize this principle through Last-Chance Policy Identification (LCPI), where correctness is evaluated at the state the agent reaches rather than at the initial state. The Last Identifiable Margin (LIM) marks the feasibility boundary between distinguishing and homogenizing. For deterministic diagnostic graphs we provide the Exact-LIM recursion; for noisy finite-horizon recovery we propose Risk-Budgeted Compatibility Planning (RBCP), which searches a compatibility-aware frontier under a hard worst-case failure constraint. Across incident recovery on abstract microservice topologies and latent-damage navigation in MiniGrid, RBCP improves risk-feasible recovery while satisfying the failure budget. A sham control–cost-matched actions that preserve model incompatibility–eliminates the gain entirely, confirming that the benefit comes from changing which policies are acceptable for which models, not from extra search or additional budget.

[AI-178] Frontier Autolab: Organizational Memory Adversarial Dissent and Temporal Leakage in Multi-Agent LLM Firms Across Fifty Years of Technological Change

链接: https://arxiv.org/abs/2609.36739
作者: Bravish Ghosh
类目: Computational Engineering, Finance, and Science (cs.CE); Artificial Intelligence (cs.AI)
备注: 17 pages, 6 figures, 9 tables. Code and data: this https URL

点击查看摘要

Abstract:Multi-agent LLM systems are increasingly structured like organizations, with roles, critics and shared memory, yet they are evaluated on tasks that last minutes. We ask how such an organization behaves when the ground it stands on keeps moving. Frontier Autolab is a long-horizon testbed in which one simulated firm, voiced by sixteen role personas and a dedicated Red Team, must re-found itself in nine technology eras from 1990 to 2040. Each era is temporally gated: the firm decides from a dated briefing, a historian-judge then reveals what happened and scores the decision on a five-dimension rubric, and lessons enter a persistent Playbook. Six eras are scored against history, one against the live market and two are open forecasts. Across four trajectories (36 era decisions, 180 subscores) we find a consistent foresight-commitment gap: in all 24 historically scored eras the judge rated the firm’s recognition of the coming shift above its choice of where to build (mean gap 1.9 points on a 10-point scale), because boards chose the layer their existing assets could reach. Organizational design shaped long-run character. A Red Team armed with numeric kill gates produced fifty years of gated pilots and no product, and the rubric rated this firm highest; firms whose memory stored market-structure lessons pivoted every era, while a firm whose memory stored only validation procedure kept one method throughout. We also show why such results are hard to trust. Scores rise across eras in every run while the judge’s own hindsight subscore falls (within-run r = -0.58), so apparent learning is confounded with recall of history, and we trace further distortions to self-judging, briefing selection and score aggregation. We release all records and an API harness, and specify fictional and post-cutoff eras that would turn the testbed into a benchmark.

[AI-179] BiFE: Search-Efficient Discovery of CPU-Only Branching Policies via LLM -based Bi-Fidelity Evolution

链接: https://arxiv.org/abs/2609.36735
作者: Ce Zhang,Bin Zhang,Zhiwei Xu,Hao Chen,Xinyue Lu,Shanwei Fan,Yingxuan Teng,Guoliang Fan
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:In branch-and-bound (BB) for mixed-integer linear programming (MILP), branching variable selection critically impacts efficiency. Existing neural branching policies often require GPU inference, while CPU-efficient symbolic expressions lack the representational capacity for complex logic. Large Language Model (LLM)-generated code provides a flexible search space for designing lightweight branching rules with diverse algorithmic logic. To discover effective rules within LLM-based evolutionary frameworks, a core challenge arises: full BB evaluation on real instances is prohibitively expensive, whereas offline imitation learning suffers from distribution shift. To address this, we introduce a Bi-Fidelity Evolutionary framework (BiFE). It employs low-fidelity imitation scores as a rapid pre-screener and selectively applies high-fidelity on-instance evaluation only to elite candidates, effectively balancing search efficiency with performance reliability. Experiments validate both the search efficiency of BiFE and the competitiveness of its discovered rules, which outperform the SCIP solver and other baselines on CPUs, and even surpass certain GPU-based neural policies.

[AI-180] Can AI Scientists Change Their Minds? Prior-Evidence Conflict in Synthetic Universes

链接: https://arxiv.org/abs/2609.36726
作者: Kargi Chauhan
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Can a scientific agent distinguish a law it inferred from evidence from one it merely recognizes? We introduce Synthetic Universes, a controlled benchmark that pairs canonical famous worlds with matched twisted twins governed by nearby noncanonical mechanisms. We evaluate each reported law twice: by executing it on held-out continuations and transfer settings, and by independently checking whether it recovers the generating mechanism. In the current checkpoint of a pre-specified 60-cell study, 22 trials were graded and one additional run ended in infrastructure failure. Among 20 twin trials, 8 pass predictive verification while 5 recover the generator. The dissociation is bidirectional: six parsable outputs predict successfully while missing the mechanism, whereas three recover the mechanism but fail predictive rollout. Drag exhibits the first pattern (5/5 predictive pass, 1/5 mechanism recovery); Gravity exhibits the second (1/5 predictive pass, 4/5 mechanism recovery). Because matched famous controls, the corrected identifiability sweep, and the Evidence Ladder remain incomplete, we do not claim a confirmatory causal prior-conflict effect. Instead, the completed runs establish a narrower verification result: predictive adequacy and mechanism recovery are distinct scientific claims and require distinct tests.

[AI-181] JudgeProfile: Understanding and Steering Subjectivity in LLM Judges

链接: https://arxiv.org/abs/2609.36705
作者: Qi Cao,Kangning Liu,Xuan Kan,Shunwen Tan,Yang Pei,Dake Chen,Yatai Ji,Zixuan Ye,Yuanpeng Tu,Daniel Li,Junbiao Tang,Pengtao Xie,Zihao He
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:LLM judges are inherently subjective, often favoring different responses in pairwise comparison when neither option is objectively wrong. To study this subjectivity, we introduce JudgeProfile, a framework that dissects LLM evaluation into perception (how a judge compares two responses across specific attributes like clarity, correctness, and detail) and prioritization (how much each attribute influences the final choice). We curate SubjectiveSet, a dataset of 50,013 response pairs from 17 public data sources, evaluated by 21 LLM judges across 87 attributes. We find a hidden consensus in perception: judges frequently agree on attribute judgments even when their overall choices diverge. Building on this separation, we first characterize each judge’s prioritization using attribute weights estimated from its own overall choices. These weights differ across judges even when estimated from the same attribute judgments. We then learn new weights from reference labels to adapt their decisions to a target evaluation standard. Reweighting perceived attributes improves average held-out agreement with reference labels from 66.48% to 71.97%, outperforming fine-tuning and rubric prompting. Our findings show that understanding and steering the subjectivity of LLM judges requires attention not only to what they perceive, but also to how they prioritize it.

[AI-182] Normalize-Then-Precondition: A Hierarchical Approach to Marginal Scale and Interaction Geometry for LLM Training

链接: https://arxiv.org/abs/2609.36692
作者: Zixuan Gong,Zeyu Gan,Jiaye Teng,Yong Liu
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Matrix optimizers have emerged as a promising direction, with Muon standing out as a prominent design. Revisiting Muon through its full-Gram representation, we observe that it jointly processes marginal-scale and interaction information. This opens an alternative way to organize geometric information hierarchically, motivating the Normalize-Then-Precondition framework. Specifically, it first uses diagonal-Gram information to construct a marginally normalized update, then applies spectral preconditioning to its directional interaction geometry. Building on this framework, we develop NormPre with NormPre-G and NormPre-L adopting global and localized spectral preconditioning, grounded in spectral-norm steepest descent and a regularized formulation followed by leading mode selection, respectively. To enable large-scale training, NormPre-G uses Newton-Schulz iterations and NormPre-L employs randomized sketching to approximate the leading interaction eigenspace. Theoretically, we establish \mathcalO(T^-1/2) convergence guarantees for simplified versions of NormPre. Across extensive pretraining experiments on GPT-2 Small, LLaMA and Qwen3, both variants consistently outperform AdamW, Muon and MANO under matched training budgets. Further efficiency and spectral analyses reveal the complementary strengths of two variants and characterize their performance-efficiency trade-off. We open-source our code through a GitHub repository at this https URL.

[AI-183] MLToolBench: Learning Tool-Augmented Agents for Machine Learning Development

链接: https://arxiv.org/abs/2609.36679
作者: Xin Yu,Lizhu Zhang,Jiamu Bai,Yanhong Wu,Zellux Wang,Serena Li,Weiwei Li,Lingzhou Xue,Xiangjun Fan,Bo Peng
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Machine learning engineering (MLE) agents have made substantial progress, but learning through ML experimentation remains costly in time and computation. Synthetic environments reduce these costs while introducing variations in data and experimental settings that require task-specific diagnosis. Access to diagnostic tools alone does not ensure that agents learn when to use them or how to act on their findings. We introduce ToolMLBench, a suite of executable tools for data inspection, code verification, and experiment diagnosis, together with an SFT and RL pipeline for learning their use. Diagnostic calls acquire evidence whose value depends on subsequent decisions, so final outcomes provide limited guidance on which calls to reinforce. We address this challenge with SPICE, which measures how privileged context changes the likelihood of a sampled tool action and uses this difference as a turn-level reward alongside the final outcome. We train on 80 synthetic tasks and evaluate on 25 in-domain and 10 out-of-domain tasks. Providing tool interfaces and descriptions alone yields inconsistent gains across unadapted models. With the same diagnostic interface, our training pipeline raises in-domain success from 24.8% to 52.4% for Qwen3-8B and from 35.6% to 69.2% for Qwen3.5-35B-A3B. The latter also improves from 31% to 48% out-of-domain, supporting learned diagnostic tool use on held-out sources and targets.

[AI-184] FairDiff: Mitigating the Self-Reinforcing Matthew Effect in Diffusion Recommender Models

链接: https://arxiv.org/abs/2609.36671
作者: Song-Li Wu,Xianquan Wang,Zhaocheng Du,Weinan Gan,Jingyi Wang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:While the “Matthew Effect” and filter bubbles are widely recognized outcome-level biases in recommender systems, we reveal that Diffusion Recommender Models (DRMs) uniquely compound this issue through their generative dynamics. Rather than merely inheriting data imbalances, DRMs trigger a self-reinforcing amplification of popularity bias. We identify that this phenomenon is driven by two compounding mechanisms. First, while optimization loss is universally dominated by high-frequency items across recommenders, DRMs suffer from a unique structural prior mismatch during generation. Because the forward terminal distribution of long-tailed data deviates significantly from the standard Gaussian prior, reverse sampling trajectories inherently collapse toward high-density popular items, fundamentally suppressing niche item generation. To dismantle this self-reinforcing loop, we propose FairDiff, a plug-and-play fairness-aware diffusion framework. To overcome the popularity-dominated loss, we introduce Popularity Condition Guidance (PCG). Rather than altering the training objective, PCG acts as an inference-time distributional reweighting mechanism, mathematically reshaping the score-based gradient field to penalize high-popularity regions and guide trajectories toward niche semantics. Furthermore, we design a Semantic Calibration (SC) Module to bridge the prior mismatch, aligning the forward and reverse distributions via one-step optimal transport. Comprehensive evaluations demonstrate that FairDiff achieves state-of-the-art performance while effectively mitigating the self-reinforcing Matthew Effect, highlighting its value as a general framework for DRMs.

[AI-185] FineSID: Scalable and Efficient Semantic Identifier Learning for Generative Recommendation

链接: https://arxiv.org/abs/2609.36670
作者: Song-Li Wu,Weinan Gan,Zhaocheng Du,Xianquan Wang,Jingyi Wang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:A critical prerequisite of generative recommendation is designing semantic identifiers (SIDs) that are both scalable to large item sets and efficiently learnable. Existing SID learning methods fundamentally rely on Top-1 hard assignment during vector quantization. While heuristic strategies – such as clustering-based initialization or forced post-hoc collision resolution – can artificially inflate codebook coverage, they often disrupt end-to-end semantic alignment and fail to address the underlying optimization bottleneck: sparse gradient propagation. In standard Top-1 assignment, gradients concentrate on a narrow subset of frequently selected codewords, leaving the majority inherently under-trained and causing severe SID collisions. To overcome this limitation natively without relying on complex initialization priors, we propose FineSID, a unified quantization framework that moves beyond Top-1 assignment by enabling fine-grained gradient propagation across the entire codebook. Instead of updating only a single selected codeword, FineSID distributes learning signals to all codewords in a soft, differentiable manner. This design promotes globally balanced codebook optimization while strictly preserving semantic consistency, effectively alleviating SID collisions and stabilizing training in large, high-dimensional codebooks. Extensive experiments on multiple public benchmarks demonstrate that FineSID is robust to initialization configurations and consistently improves both codebook utilization and recommendation accuracy. Our work provides a principled, initialization-agnostic solution for semantic identifier learning, advancing the practicality of generative recommendation.

[AI-186] AutoLoCo: Communication Efficient Distributed LLM Training via Adaptive Synchronization

链接: https://arxiv.org/abs/2609.36662
作者: Pengyu He,Yan Zhang,Ruien Li,Guangwen Yang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:The pre-training of Large Language Models (LLMs) is increasingly conducted across multiple data centers. As training scales to a larger number of accelerators, the fraction of time spent on computation decreases, while the fraction spent on communication increases. Therefore, frequent synchronization becomes a growing bottleneck. Local update methods reduce this cost by allowing workers to perform several optimizer steps between synchronizations. Most local update methods set the number of local optimizer steps between synchronizations before training and keep this interval fixed throughout the run. However, the best interval can change during the entire train process. If the interval and optimizer are adapted to the current training state, the communication frequency is reduced while maintaining the training performance. In this work, we introduce AutoLoCo, an adaptive training framework to reduce communication in LLM training. It adapts the local interval using scalar training statistics and corrects each outer update. Our method is motivated by two observations: 1) the appropriate local interval varies across training stages, and 2) changing the number of inner steps per interval creates a mismatch with an unchanged outer optimizer, requiring a correction to the outer update. We optimize this mismatch by correction of the outer optimizer for the momentum and the learning rate using the accumulated inner learning rate. Our experiments under communication constraints demonstrate that AutoLoCo reduces communication frequency by 27% relative to DiLoCo while maintaining training performance.

[AI-187] Constitutional adapters: Inference-time interventions for misalignment and misuse

链接: https://arxiv.org/abs/2609.36657
作者: Adam S. Lowet,Mark Kurzeja
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 38 pages, 14 figures, 10 tables

点击查看摘要

Abstract:Training models to act in accordance with an explicitly defined set of principles, or “constitution,” has shown promise as a robust and transparent mechanism for AI alignment. However, the generality and flexibility of such methods remain unclear. Here, we show that constitution-consistent behavior can be distilled from synthetic corpora into lightweight objects (low-rank adapters and steering vectors). Despite never seeing a harmful request or jailbreak during training, such objects increase jailbreak defense success and measured alignment – particularly at long context lengths and against multi-turn attacks, where they outperform both prompted and steered baselines. Subtracting control-trained from constitution-trained objects further accentuates these effects, yielding defenses we call “constitutional adapters” (CAs). CAs can be trained on a base model, transferred zero-shot to its post-trained checkpoint, and scaled at inference time to predictably trade off defense for benign compliance. Taken together, these results recommend CAs as a lightweight, portable, and tunable lever for mitigating misalignment and misuse in API deployments.

[AI-188] RankBuffer: Efficient Ranking-Based Rewards for Open-Ended Generation

链接: https://arxiv.org/abs/2609.36652
作者: Zixuan Yang,Yiqun Chen,Qi Liu,Wei Yang,Erhan Zhang,Liyi Chen,Qimeng Wang,Yan Gao,Jiaxin Mao
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Open-ended generation lacks canonical answers, making pointwise rewards difficult to calibrate for group-based reinforcement learning. Directly ranking same-query rollouts provides a more suitable relative reward signal, but existing ranking-based reward methods can incur substantial judging cost. We introduce RankBuffer, which maintains an ordered, query-specific buffer of previously judged responses as a reusable quality scale. Each rollout is first inserted into an anchor interval through an independent coarse judgment, after which only rollouts assigned to the same interval undergo local fine ranking. The resulting complete order is converted into bounded rank rewards, while boundary expansion, local refinement, and inactive-anchor pruning adapt the buffer as the policy evolves. Across four open-ended benchmarks, RankBuffer consistently outperforms all pointwise baselines. It also achieves nearly on-par performance with the strongest ranking-based reward baseline while substantially reducing judging cost. Ablations demonstrate the importance of both local fine ranking and anchor response content, while buffer analyses show that rollout-derived anchors progressively extend and refine the covered quality scale. These results establish response reuse as an effective approach to efficient relative reward construction.

[AI-189] Where Predictive Supervision Goes Shapes What VLA Policies Learn

链接: https://arxiv.org/abs/2609.36645
作者: Hanseul Kim,Jewon Yeom,Youngjoon Jeong,Minsoo Jo,Taesup Kim
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: 38 pages (9 pages main text + appendix), 13 figures, 21 tables

点击查看摘要

Abstract:Future prediction is increasingly used to improve vision-language-action (VLA) policies, based on the premise that anticipating scene evolution encourages representations useful for control. However, forecast quality alone does not establish that a policy has learned a better representation for action. This distinction matters under distribution shift, where successful control depends on preserving spatial state and likely scene change beyond familiar configurations. We study what determines whether predictive supervision improves the visual representation used by a VLA policy. Through controlled comparisons with matched target constructions, prediction horizons, and training conditions, we find that different prediction interfaces produce markedly different forecasts and visual representations, including in the spatial, dynamics, and action information that transfers beyond familiar scenes. We trace these differences to how predictive errors shape the policy’s visual stream. Consistent with this controlled finding, VLA policies trained with more direct, scene-matched future supervision show stronger robustness under simulated and physical distribution shifts. Together, our results frame future prediction as a representation-learning design problem whose value for control depends on whether its supervision reaches the representations through which the policy acts.

[AI-190] Inducing Process Supervision from Outcome-Only Reinforcement Learning

链接: https://arxiv.org/abs/2609.36641
作者: Shengda Fan,Xin Cong,Zhong Zhang,Haotian Chen,Yankai Lin
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Process reward models (PRMs) have become a key component for LLMs, as their step-level feedback supports both post-training and test-time reasoning. However, training strong PRMs remains costly: human step annotation is difficult to scale, while Monte Carlo estimation is computationally expensive and can drift from the intrinsic correctness of steps. To get effective PRMs at low cost, we introduce TIPS (Thinking-Induced Process Supervision), an outcome-only reinforcement learning (RL) framework for training generative PRMs. In TIPS, the model generates a chain-of-thought (CoT) followed by step-level labels and an outcome label. The reward depends solely on whether the predicted outcome matches the ground truth, and the resulting group-relative advantage is used to optimize the entire generated response. Intuitively, when checking intermediate steps helps determine the outcome, more accurate checks can lead to better outcome judgments and higher rewards. Outcome-only RL can therefore reinforce step-level verification without explicit process supervision. We validate the effectiveness of TIPS across math and agent benchmarks and four backbone families. Notably, TIPS-Qwen3-4B-Thinking-2507 reaches 85.2 F1 on ProcessBench with only 3.2K outcome-labeled trajectories, surpassing all evaluated trained PRMs and strong prompt-only judges such as GPT-5.4-Instruct and Claude-4.7-Opus, while still trailing o1-mini. Code and data are available at this https URL.

[AI-191] WitnessGym: Benchmarking Coding Agents on the Construction of Bug Witnesses

链接: https://arxiv.org/abs/2609.36635
作者: Haomin Qi,Xiangzhe Xu,Yiming Huang,Jingbo Shang,Chengpeng Wang
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注: 30 pages, 8 figures, and 12 tables, including appendices

点击查看摘要

Abstract:Bug validation asks a coding agent to produce an executable witness for a reported bug. The witness combines a concrete input with a testing harness and exposes faulty behavior during execution. Such evidence makes audit findings actionable, yet benchmark evaluation is difficult when cases reuse public historical bugs and witnesses or require manual construction. We present WitnessGym, an automated framework for constructing bug-validation benchmarks through bug injection. It injects bugs into test-reached paths of real projects, rebuilds each project, and retains cases exposed by a construction-time witness. Bug specifications and execution adapters allow extension to additional bug types and languages. Bug-preserving transformations vary the surrounding structure while preserving the witness behavior. Based on real-world Java projects with test suites, WitnessGym automatically constructs 1,300 benchmark cases. The injected patches resemble historical bug patches and are difficult for the two evaluated models to distinguish in blinded comparisons. We evaluate four coding agent frameworks in six framework/model pairings across bug types, execution contexts, and transformation depths. Witness construction remains difficult even when the bug pattern is known. Our framework, benchmark cases, and evaluation scripts are available.

[AI-192] Distilling Agent ic Systems: A Roadmap across Models Artifacts and Harnesses

链接: https://arxiv.org/abs/2609.36630
作者: Ziluowen Luo,Senzhang Wang,Chaozhuo Li,Jun Yin,Hao Yan,Ming Cheng,Chenxu Wang,Songyang Liu,Litian Zhang,Qiwei Ye,Zheng Liu,Philip S. Yu
类目: Artificial Intelligence (cs.AI)
备注: 57 pages, 5 figures

点击查看摘要

Abstract:Modern agents increasingly rely on memories, tools, and execution logic, so their competence extends beyond model parameters. This shift exposes a limitation of conventional knowledge distillation, which asks how a student model imitates a teacher model. We define Agent Distillation as the persistent transfer of task-solving knowledge from a teacher agent to a student agent. Our study organizes the field by where transferred knowledge is retained: within the model, as artifacts, through the execution harness, or across substrates. This perspective separates transfer evidence from its outcome and clarifies how knowledge moves between agent components. We develop an evaluation framework that relates retention to causal contribution and deployed utility. Together, these contributions establish a foundation for the reliable, maintainable, and safe development of increasingly complex agentic systems.

[AI-193] Semantic Projection for Continual Self-Evolution of Language Agents

链接: https://arxiv.org/abs/2609.36626
作者: Ziyu Liu,Jun Chen,Lixu Wang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Language-model agents increasingly rely on persistent natural-language skills to adapt beyond their frozen model parameters. When a shared skill is repeatedly revised from a non-stationary, heterogeneous task stream, however, improvements for new tasks can overwrite procedures needed for earlier ones. In continual learning, Orthogonal Gradient Descent (OGD) addresses analogous interference by projecting a new-task gradient onto a subspace that locally preserves prior predictions. Natural-language skill revisions, however, have neither gradients nor a canonical vector space in which such a projection can be performed. We introduce \emphSemantic-Scope Projected Evolution (SSPE), which transfers the functional principle of gradient projection from parameter space to behavior space. SSPE treats an unconstrained skill revision as a proposed update, identifies acquired capabilities with which it may interfere, and uses the observed gains and regressions to construct a compatible revision rather than merely rejecting the update. This enables one shared skill to evolve across latent and recurring task contexts without exposing semantic domain identities to the evolution model. Across controlled synthetic streams and heterogeneous real-agent benchmarks, SSPE improves final cross-domain competence and mitigates forgetting relative to strong skill-evolution baselines. The evolved skill also retains the strongest average performance after transfer to a different executor model. These results establish semantic projection as a promising principle for stable and adaptive self evolution of language agents.

[AI-194] DualTrack: Synchronized speech-gesture generation via symmetric coupling of pretrained priors

链接: https://arxiv.org/abs/2609.36624
作者: Yuanzhuo Hu,Zehan Liu,Xiaoyi Qin,Ming Li
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Joint speech-gesture synthesis must coordinate two modalities despite limited paired data. Existing approaches often lack bidirectional interaction, have limited language coverage, or simplify body and finger representations. We present DualTrack, which couples pretrained speech and motion priors on a shared 12.5 Hz timeline. Causal adapters exchange previous-packet information, while current-state fusion coordinates the streams before they separately complete sixteen-codebook packets. We evaluate 43 BEAT2 recordings in four languages, with speakers held out from joint training and validation. On the shared English/Spanish inputs, without speech or motion prefixes, DualTrack achieves lower word error rate and full-motion Fréchet Gesture Distance, higher beat consistency and speech naturalness than the evaluated GELINA baseline.

[AI-195] Neural Structural Reason er: A Brain-inspired Architecture for Reasoning over Structured Knowledge NEURIPS2026

链接: https://arxiv.org/abs/2609.36620
作者: Zixing Jia,Yuhang Pan,Ni Ji
类目: Artificial Intelligence (cs.AI); Neurons and Cognition (q-bio.NC)
备注: Accepted at NeurIPS 2026

点击查看摘要

Abstract:Structural reasoning, the ability to recognize and make inferences over the relational structure between objects and concepts, is a hallmark of human cognition, yet prevailing methods often collapse relational topology into flat embeddings, cannot discover hidden structure and lack interpretability. We introduce Neural Structural Reasoner (NSR), a brain-inspired network that preserves relational structure directly in the connectivity and dynamics of coupled neuronal populations. NSR draws inspiration from three biological mechanisms: multi-layered architecture for encoding hierarchical knowledge, stable representations of entity and concepts, and path integration for input-driven state inference. At query time, NSR parallelizes computation over candidate relational structures and leverages confidence-weighted scores to perform link prediction. Across standard knowledge-graph benchmarks, NSR achieves competitive accuracy without leading on every dataset, and has lower reported training times than several neural baselines. Because reasoning is implemented through sequences of human-readable neuron activations, NSR affords native interpretability by tracking intermediate inference steps. The model further extracts latent relational hierarchies and compositional rules, demonstrating the brain-inspired architecture as an effective, efficient, and highly interpretable substrate for structural reasoning.

[AI-196] Multi-Channel Mitigation of Source-Trust Shortcuts in Fact-Checking RL Agents

链接: https://arxiv.org/abs/2609.36611
作者: Jianchang Su,Yiwei Yang,Wei Zhang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Retrieval-augmented fact-checkers often receive a reliability label, such as HIGH or LOW trust, for each evidence source. These labels should adjust the model’s confidence and its decision to search for more evidence, while the verdict should follow the evidence content. We introduce TrustSwap, a counterfactual test that swaps, lowers, or removes source labels while keeping every evidence text fixed, and measures its three output channels (the verdict, the confidence, and the search decision) separately. Across untrained and RL-trained models at two scales, three datasets, and two prompts, confidence and search respond to the labels as intended in 49 of 50 comparisons, yet a label change alone alters 4-23% of confident verdicts for Qwen3 models and up to 50% for an existing RL-trained fact-checker. Standard GRPO fine-tuning amplifies this shortcut at 8B in all six settings. To reduce it, we propose trust-swap augmentation (TSA), which trains GRPO on each claim with both its original and its label-swapped evidence under the same gold verdict. At 4B, TSA lowers the verdict flip rate by 7-35% (relative) in four of six settings, keeps accuracy and the intended confidence and search responses, outperforms reward-based alternatives in the main setting, and carries over to an unseen label-removal perturbation. An added consistency reward helps on the trained-on swap but not on unseen perturbations. At 8B, TSA’s effect is not detectable, which makes scale the main open question.

[AI-197] SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation

链接: https://arxiv.org/abs/2609.36601
作者: Miteto Wei,Xiaohan Wang,Zehao Chen,Jiajun Chai,Sichao Liu,Li Wang,Haoyuan Xu,Zhaoyu Hu,Wei Lin,Guojun Yin
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:On-policy distillation (OPD) reduces train-test state mismatch by training a student on its own generated trajectories, but weak students may visit teacher-misaligned prefixes where supervision is less representative. We introduce SAKI (Supervision Allocation with KL-constrained Interpolation), which combines a KL-constrained teacher-guided rollout with maximal coupling and reuses realized accept/correction events to route token-level supervision. Accepted positions retain sampled-token reverse-KL supervision, while correction positions receive direct supervision on the teacher’s highest-probability token. Under maximal coupling, the correction probability is exactly TV(p_t, q_t), so the same trust-region radius controls rollout deviation and upper-bounds intervention and specialized-supervision frequency. We further implement an engine-resident speculative verifier that preserves the exact-q trajectory distribution and coupling semantics while improving matched-workload rollout throughput by 4.22x. Across seven mathematical reasoning benchmarks, SAKI improves the matched teacher-guided baseline in Mean@8 and Pass@8 for both 1.7B and 0.6B students. Placement controls and fixed-prefix analysis further support correction-triggered routing as a conflict-adaptive supervision signal.

[AI-198] Simple Agent ic Memory for Generalist Robot Policies

链接: https://arxiv.org/abs/2609.36595
作者: Yuyou Zhang,Yunbei Zhang,Miao Li,Janet Wang,Zijian Jin,Shilong Liu,Ding Zhao
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: 34 pages, 13 figures. Project page: this https URL

点击查看摘要

Abstract:Visual-memory systems commonly retain or compress past observations. Robot control additionally requires interaction-derived state that no individual frame may explicitly represent, such as persistent identity relations, accumulated progress, or ordered procedures. We introduce Simple Agentic Robot Memory (SimpleARM), a training-free memory layer for frozen generalist robot policies. From the task instruction, SimpleARM specifies what to monitor; frozen perceptual tools maintain compact typed state online; structured access retrieves that state only when a proposed subgoal depends on history; and current-view grounding resolves recalled entities before execution. We evaluate SimpleARM on RoboMME, a benchmark of memory-dependent robot manipulation tasks that require history information no longer available in the current observation. Across all 16 tasks and three policy seeds, SimpleARM achieves 67.17% mean success, compared with 44.51% for the strongest non-oracle baseline. Matched ablations show mechanism specificity: removing relation, reference, progress, or route state produces large losses where the affected state is retrieved for control, while largely sparing other tasks. These results support a state-based view of robot memory: effective memory for control is not simply retained visual history, but compact task-relevant state derived from the interaction history.

[AI-199] xt2Sim: Agent ic Physics-Based Simulation Generation with Distilled Expertise

链接: https://arxiv.org/abs/2609.36593
作者: Xiaoyu Xiong,Tsun-Hsuan Wang,Yi-Ling Qiao,Tao Du,Minchen Li
类目: Graphics (cs.GR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Creating diverse physical simulations remains labor-intensive because assets, layout, physical parameters, motion, control, and rendering must be designed and debugged jointly. We present Text2Sim, a simulation-specialized agentic pipeline that converts a text-only request into an executable, editable dynamic case. Built on Genesis, Text2Sim uses a hierarchical agentic structure that combines a Planner with specialized Writers, asset-generation tools, and an independent Critic. Compact skills (Debug Cards) distilled from graphics demonstrations provide role-specific physical guidance for execution-based repair. We evaluate physical quality, visual quality, and human preference on 42 held-out prompts spanning rigid, articulated, deformable, and cloth phenomena, with a paper-level split between experience construction and evaluation. We design automatic physical and visual scorers to evaluate the quality of the results, and Text2Sim achieves higher scores than all four state-of-the-art baselines on both metrics. In blinded user studies with these baselines, significantly more participants prefer Text2Sim than prefer the baselines, which is consistent with the results from our automatic scorers. The pipeline also supports a broad range of downstream applications; we select dataset construction and extension to multimodal input as two representative examples. We will release the code, the Debug Card library, and a dataset of generated cases, each pairing the text prompt and rendered video with the executable program, assets, physical parameters, controls, and recorded states.

[AI-200] MemEvo: Automatic Discovery of Streaming Video Memory Mechanisms

链接: https://arxiv.org/abs/2609.36581
作者: Guohong Liu,Jialei Ye,Shanhui Zhao,Yunxin Liu,Yuanchun Li
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Query-agnostic streaming video understanding requires vision-language models to continuously compress an indefinitely growing visual stream into a bounded memory before future queries are known. The performance depends critically on the memory mechanism–what observations to preserve, how to represent and consolidate them, and what information to retrieve when a query eventually arrives. Rather than designing a single memory architecture by hand, we formulate memory design as a search problem over executable memory programs. We introduce a lightweight domain-specific language that expresses memory mechanisms through structured primitives for representation, admission, retention, consolidation, budgeting, and retrieval, while enforcing causal and bounded-memory constraints. Although structured, the derived program space remains large and contains heterogeneous, conditionally dependent design choices whose effects can only be assessed via downstream execution. We therefore propose MemEvo, an LLM-driven auto-research framework that uses pretrained LLM as a semantics-aware proposal model to iteratively generate and refine candidate memory programs based on accumulated experimental feedback. At runtime, a deterministic evaluation pipeline validates and evaluates each candidate, while the underlying vision-language model remains frozen throughout discovery. We finally produce a training-free, bounded-memory mechanism. Extensive experiments on StreamingBench and OVO-Bench demonstrate strong streaming video understanding performance together with substantial context and inference efficiency.

[AI-201] SafeCoEvo: Co-Evolving Safety Harnesses and Guards for LLM Agents at Test-Time

链接: https://arxiv.org/abs/2609.36580
作者: Yu Cheng,Yongkang Hu,Shuaijie Ma,Zhihang Lin,Weicheng Meng,Jingyang Qiao,Jiuan Zhou,Yushuo Zhang,Yihang Chen,Weilin Luo,Kun Shao,Dong Li,Zhizhong Zhang,Yuan Xie,Zhaoxia Yin
类目: Artificial Intelligence (cs.AI)
备注: 35 pages, 10 figures

点击查看摘要

Abstract:LLM agents deployed in real-world environments continually encounter new tasks and safety risks, while execution feedback typically becomes available only after each task is completed. However, existing self-evolving approaches commonly rely on multiple rounds of optimization over fixed and repeatedly accessible task distributions, fundamentally differing from test-time adaptation in real-world deployment, where only experience accumulated from past tasks can be used to improve safety decisions on future unseen tasks. To address this limitation, we propose SafeCoEvo, a test-time Harness-Guard co-evolution framework for LLM agent safety that enables the external safety system to continually adapt from accumulated runtime experience. SafeCoEvo jointly improves two complementary safety capabilities at different timescales: S-Harness rapidly externalizes recent runtime experience into updatable explicit safety knowledge that can promptly influence subsequent tasks, while GuardVPO internalizes accumulated runtime safety experience over a longer timescale into parametric risk-judgment capabilities. By combining short-term rapid adaptation with long-term capability consolidation, SafeCoEvo continually improves the agent’s safety capabilities, reducing the unsafe outcome rate by 10.05% while improving the task success rate by 12.15% over the strongest baseline, thereby achieving simultaneous gains in safety and task utility.

[AI-202] Factorized Scheduling Principle: Learning Interpretable and Transferable Policies via Structured Additive Functions ICML2026

链接: https://arxiv.org/abs/2609.36578
作者: Hong Je-Gal,Hyun-Suk Lee
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Accepted to the 43rd International Conference on Machine Learning (ICML 2026)

点击查看摘要

Abstract:Scheduling problems arise from repeatedly selecting one item from a set of candidates based on their states. These problems often reduce to assigning priority scores and choosing the highest-ranked item. In this work, we propose a factorized scheduling principle (FSP) framework to learn interpretable and transferable scheduling rules. The FSP framework represents system states as condition distributions and decomposes a global scheduling principle into additive univariate and pairwise components with identifiability constraints. The scheduling principle enables the framework to maintain a simple priority-based structure during deployment. This principle is learned by using a policy-based objective combined with a temporal-difference signal defined on the condition distribution. Experiments on synthetic and realistic scheduling tasks demonstrate the FSP framework’s strong performance, interpretability, and zero-shot generalization across different system scales.

[AI-203] Divide and Inject: Can Agents Reconstruct an Indirect Prompt Injection from Frag ments?

链接: https://arxiv.org/abs/2609.36576
作者: Michael Lee,Zhipeng Wei,Yue Dong,N. Benjamin Erichson
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Agentic systems are now being widely used to orchestrate tools and reason over long contexts. However, the improving capabilities of the large language models powering these agents also create new attack surfaces for indirect prompt injection. In particular, an attacker may not need to place a complete malicious instruction in retrieved content if the agent can reconstruct the objective from incomplete fragments distributed across a long context. In this work, we introduce adaptive long-context prompt injection (AdaLCPI), which combines long-context fragmentation with adaptive search. AdaLCPI splits an attack objective into incomplete fragments, embeds them in external content retrieved through the agent’s tools, and uses a reconstruction cue to prompt the agent to combine them. It then iteratively refines the fragments and cue with OpenEvolve using graded scoring and natural-language execution feedback from the target agent. Empirically, AdaLCPI achieves higher attack success than strong adaptive baselines, reaching 61.4% macro-average ASR compared with 32.8% for Trojan Hippo-style and 30.0% for AgentVigil. Safety evaluations should therefore test whether agents remain robust when harmful objectives must be reconstructed from incomplete fragments.

[AI-204] Visual sensitivity is not claim retractability: persistence-aware credit assignment for multimodal reinforcement learning

链接: https://arxiv.org/abs/2609.36572
作者: Zhongan Bi,Kepeng Lin,Xuanang Gao,Yuhan Sun,Lianrun Zhang
类目: Artificial Intelligence (cs.AI)
备注: MLLM,RL

点击查看摘要

Abstract:Reinforcement Learning with Verifiable Rewards (RLVR) has been extended to Large Vision-Language Models (LVLMs), and perception-aware methods further encourage policies to rely on visual evidence. Yet relying on the image does not guarantee that visual claims are supported by it. Before RL training, 27.81% of the correctly answered responses of Qwen2.5-VL-7B on four multimodal reasoning benchmarks contain at least one direct visual claim that the image does not support. Since outcome-level RL rewards each response as a whole, these claims inherit the positive credit of the correct answer. We introduce a fixed-rollout counterfactual diagnostic that re-scores the same response under an intervened image to separate Evidence-Function Sensitivity (EFS), how strongly the model’s predictions change, from claim persistence, whether the model keeps supporting the same claim rather than retracting it. The diagnostic reveals Sensitivity-Persistence Decoupling (SPD): under DAPO and VPPO, EFS increases and claims become more retractable overall, yet unsupported claims become significantly more persistent, whereas GRPO raises EFS without this deterioration. We therefore propose Persistence-Aware Credit Gating (PACG), which attenuates positive credit for unusually persistent visual claims and leaves all other credit unchanged. It requires no supported/unsupported labels and adds no inference cost. On Qwen2.5-VL-7B, PACG raises the nine-benchmark average over three seeds from 58.1% to 59.9% with DAPO and from 59.8% to 60.9% with VPPO, while making unsupported claims more retractable. The gains extend to a larger model, a newer backbone, and the accuracy of HallusionBench also improves consistently. These results suggest that visual sensitivity and claim retractability are complementary dimensions of multimodal credit assignment.

[AI-205] CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering

链接: https://arxiv.org/abs/2609.36570
作者: Mark Russinovich
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that suppresses this behavior inside the model. Per model, a five-step recipe fits a residual-stream direction from paired episodes differing only in whether an embedded instruction is followed, and retains it only if it passes pre-specified causal and capability gates. At deployment, the direction is subtracted from every tool-result token during prefill. The edit is always on–there is no detection decision to evade–and requires no fine-tuning, auxiliary model, or added tokens, only white-box serving and tool-result span boundaries. Across five open-weights models (8B-106B, five vendor lineages), held-out attack success falls from 0.21-1.00 undefended to 0.00-0.17 defended, and AgentDojo compromise rate from 0.10-0.49 to 0.006-0.079, at 93-100% typography-normalized benign utility, with larger task-dependent costs when reasoning over steered content. A benchmark-level adaptive attacker reaching 0.67-0.73 undefended is held to roughly a quarter of that on the two most deeply evaluated models. Among the defenses we measured on capable models, those achieving lower compromise rates either lost 22-89% of benign utility or fine-tuned the served weights. White-box gradient attacks through the deployed vector compromise at most 2 of 52 episodes, and none of 2,052 replayed human red-team attacks succeeds. CounterSteer largely neutralizes instructional takeover: a black-box framing search cracks 3 of 18 development samples. Parameter manipulation–attacker-chosen arguments in otherwise legitimate calls–is only partially resisted (13 of 18); the decision becomes linearly readable at argument emission but not at the examined pre-generation sites, and is not removed by the tested prefill- or decode-time steering, motivating argument-provenance controls.

[AI-206] From Checkpoint Variation to Selection Gains in Supervised Fine-Tuning

链接: https://arxiv.org/abs/2609.36569
作者: Yupeng Chang,Wenxuan Zhang,Yuan Wu
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Checkpoint selection is a routine decision in supervised fine-tuning (SFT): training produces multiple checkpoints, but only one is retained. Yet fixed-budget comparisons do not by themselves distinguish three empirical claims: whether more validation data improve checkpoint selection, whether a selection rule outperforms validation-loss selection, and whether it improves over simply retaining the final checkpoint. We therefore treat checkpoint selection as a finite-information decision problem. Holding completed training trajectories, candidate checkpoints, and independent test items fixed, we vary the validation budget and separately measure improvement from additional validation data, gain over negative log-likelihood (NLL) selection, and gain over the final checkpoint. Across 60 mathematical SFT trajectories and 19 configurations, increasing the validation budget from 32 to 305-313 examples raises independent-test accuracy by 0.32 percentage points (pp) for generated-accuracy selection and 0.29 pp for checkpoint agreement, with 95% configuration-bootstrap CIs of [0.10, 0.56] and [0.11, 0.50], respectively. At the full validation budget, the two generation-based rules outperform matched NLL selection by 0.71 and 0.85 pp, respectively, while their gains over the final checkpoint remain unresolved. A cross-domain replication on 12 newly trained Commonsense trajectories shows the same qualitative separation: increasing the validation budget from 32 to 1,024 questions improves generated-accuracy and checkpoint-agreement selection by 0.87 and 0.27 pp, while gains over the final checkpoint again remain unresolved. Together, these results show that benefiting from more validation data, outperforming NLL selection, and outperforming the final checkpoint are distinct empirical claims that require separate evidence.

[AI-207] HiTS-CL: A Continual Learning Framework for Long-Horizon Temporal Knowledge Graph Extrapolation

链接: https://arxiv.org/abs/2609.36559
作者: Yansong Liu,Rui Liu,Yuan Zuo,Hongwei Zhao,Da Fu,Fuwei Zhang,Fuzhen Zhuang,Yong Chen,Zhe Li
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Extrapolative temporal knowledge graph reasoning (TKGR) predicts future facts from historical snapshots. Most existing methods train once on an early prefix of the timeline and then use a frozen model for all future timestamps. We argue that this fixed-prefix protocol is misaligned with extrapolation. It learns from a static prefix, whereas the target stream is non-stationary: new entities and facts emerge, temporal dependencies shift across regimes, and recurring historical signals must be refreshed online. As a result, models trained only on early snapshots become outdated and degrade over long horizons. We address this mismatch by formulating extrapolative TKGR as continual learning over streaming snapshots. Under this view, effective extrapolation must jointly handle current dynamics, stable knowledge, and recurring historical evidence. Based on these requirements, we propose History-enhanced Two-Step Continual Learning (HiTS-CL), a backbone-agnostic continual learning framework for extrapolative TKGR. HiTS-CL tracks current dynamics via continual fine-tuning, preserves stable knowledge via multi-teacher adaptive distillation, and retains recurring historical evidence via a selective memory of recent and frequent facts. We integrate HiTS-CL into five representative TKGR backbones and evaluate it on four benchmark datasets. HiTS-CL consistently improves extrapolation accuracy, reduces long-horizon degradation, and outperforms strong continual-learning baselines, including a recent method for temporal knowledge graphs. Source code and data are available at this https URL.

[AI-208] MAADBench: The Refreshable Paradigm for Anomaly Detection in Multi-Agent Systems

链接: https://arxiv.org/abs/2609.36556
作者: Lei Ma,Dennis Hofmann,Haowen Xu,Joshua DeOliveira,Peter VanNostrand,Lei Cao,Elke Rundensteiner
类目: Artificial Intelligence (cs.AI)
备注: pre-print

点击查看摘要

Abstract:Recent studies report that LLM-based multi-agent systems (MAS) fail at rates of 41%-87%, yet to our knowledge, no benchmark to date supports systematic anomaly detection (AD) for them. Building MAS AD benchmarks is hard because they must remain fresh as LLM systems evolve: tasks may leak into training data and thus be memorized by LLMs, traces and anomaly patterns expire as backbones evolve, and labels must be provided reliably for each refresh. To address these challenges, we present MAADBench (MA: multi-agent; AD: anomaly detection), the first refreshable MAS AD benchmark designed for diverse, evolving LLM backbones underlying the agents. MAADBench combines (1) sampled-and-coupled generative tasks over an approximately 10^37-task space to mitigate task leakage, (2) refreshable trace generation under configurable LLM backbones, and (3) automated provision of cost-free, deterministic step-level labels for fine-grained AD evaluation. Beyond offering the paradigm itself, we run MAADBench with five state-of-the-art LLM backbones and release the MAADBench-Full dataset with 5,200 step-labeled traces. Benchmarking 25 AD methods on the MAADBench dataset reveals substantial limitations in current approaches: they rely heavily on supervision, struggle with subtle MAS-specific anomalies, and lack robustness across LLM backbones. These gaps point to a rich research agenda for MAS-specific anomaly detection, with MAADBench providing a systematic and refreshable testbed for method development and evaluation. We open-source MAADBench-Full at this https URL.

[AI-209] SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning

链接: https://arxiv.org/abs/2609.36552
作者: Zihao Chen,Fanxiang Xiong,Hongran Ren,Xuefeng Bai,Zhongxiang Dai,Kehai Chen,Zhiguo Zhang,Zhiyong Wang,Yu Cheng
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Maximum Likelihood Reinforcement Learning (MaxRL) targets prompt-wise log-success and has shown strong performance on reasoning tasks. Under finite rollout budgets, however, the estimator used by MaxRL attenuates each prompt’s likelihood gradient by a factor that depends on its success probability and rollout count. Under uniform rollout allocation, the common rollout count fails to compensate for success-dependent attenuation, leaving low-success prompts more strongly attenuated and distorting their relative contributions to the expected aggregate gradient. We introduce SERA (Scale-Equalized Rollout Allocation), which redistributes a fixed rollout budget to approximately equalize these finite-rollout scaling factors. Building on our theoretical analysis of how finite rollouts distort prompt-wise likelihood gradients, we formulate the allocation as a fixed-budget max–min problem, derive a waterline solution to its continuous relaxation, and introduce a multiplicity correction to remove the additional prompt weighting induced by heterogeneous rollout counts. Experiments show stronger alignment with exact likelihood gradients in a controlled ImageNet setting and improved multi-sample solution coverage over MaxRL on maze navigation and mathematical reasoning under matched training rollout budgets.

[AI-210] Interactive-Policy Distillation with Bidirectional Propose-and-Verify

链接: https://arxiv.org/abs/2609.36546
作者: Shutong Wu,Xiwen Chen,Brendan Rappazzo,Daiheng Zhang,Anderson Schneider,Yuriy Nevmyvaka,Jiawei Zhang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:On-policy distillation (OPD) trains a student model on its self-generated trajectories with dense token-level teacher feedback. However, naive OPD may suffer from teacher unanchoring, where the student’s reasoning trajectory drifts far from the teacher, causing the teacher to be queried on states it would hardly visit and thus provide unreliable supervision. We propose Interactive-Policy Distillation (IPD), which applies adaptive teacher intervention to the student rollout. Under a bidirectional propose-and-verify state machine, the student and teacher alternately exchange their roles as proposer and verifier, and collaboratively generate mixed-source trajectories. Then different supervisions are applied according to the source of each token. This bidirectional propose-and-verify mechanism and the source-split loss make IPD not only a more performant distillation method, but also a unified bridge between on-policy and off-policy paradigms. To make the interleaved dual-model rollouts more efficient, we also design a dedicated fused inference engine that co-hosts both models in one serving instance with separate KV caches and instantiates the state machine model to distribute, collect, and process requests. On math reasoning tasks and across multiple teacher-student model pairs, student models trained with IPD not only outperform those trained with OPD, but also demonstrate higher data efficiency. Specifically, when distilling Qwen3-30B-A3B into Qwen3-1.7B-Base, IPD brings a +3.28 mean@8 and a +3.28 best@8 benchmark-averaged accuracy improvement compared with OPD. Besides, IPD only consumes about 1/4 of the training examples and steps to outperform OPD trained on the whole training dataset for one epoch. We also investigate the impact of different loss variants and takeover / handback configurations, and demonstrate the robustness of IPD on different training data.

[AI-211] LIBERO-MAX: Do Robot Policies Adapt When the World Changes?

链接: https://arxiv.org/abs/2609.36518
作者: Yunbei Zhang,Zijian Jin,Yuanzhe Liu,Janet Wang,Xilun Zhang,Yuyou Zhang,Zhenyu Zhang,Daoan Zhang,Shuaicheng Niu,Gen Li,Jianfei Yang,Jihun Hamm,Ismini Lourentzou,Weirui Ye,Bo Liu,Peter Stone,Marco Pavone
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: 42 pages, 15 figures. Project page: this https URL

点击查看摘要

Abstract:Robots must often continue a task after a target moves, the viewpoint shifts, or an obstacle appears, even though their earlier observations and committed actions reflect the previous scene. Many simulation robustness benchmarks fix external conditions at reset, leaving this temporal challenge underexamined. We introduce LIBERO-MAX, a benchmark of 8,000 paired cases spanning eight types of changes to geometry, observations, appearance, clutter, and paths. Each pair compares task execution with and without a mid-task event, holding the task, initial state, policy seed, and pre-event action sequence fixed. This controlled comparison distinguishes event-associated regressions from failures already present without the change. Across fourteen current VLA, hybrid, and world-action policies, events reduce success by 11.0-25.7 percentage points. Event profiles reveal shared vulnerabilities to geometry and observation changes, while policy-family rankings interleave. Camera controls show that robustness reflects both competence under the changed conditions and the trajectory from which they are encountered; varying query cadence does not eliminate the gap. Together, the paired protocol and temporal diagnostics establish LIBERO-MAX as a reproducible testbed for diagnosing failures under mid-execution changes and measuring progress toward robot policies that remain effective as the world changes.

[AI-212] BRIDGE: Bilevel Retrieval-Credit-Aware Agent ic Reinforcement Learning

链接: https://arxiv.org/abs/2609.36505
作者: Quan Xiao,Mingda Liu,Gaowen Liu,Katsuki Fujisawa,Tianyi Chen
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Optimization and Control (math.OC)
备注:

点击查看摘要

Abstract:Agentic reinforcement learning (ARL) with verifiable rewards improves the ability of large language models (LLMs) to tackle knowledge-intensive tasks by learning to interleave search and reasoning. However, most existing ARL methods optimize only LLM-generated tokens and treat retrieved evidence as environment observations. This creates an information-credit gap: failures caused by missing or misleading evidence are attributed to the LLM policy rather than to the retriever, which motivates training the LLM and the retriever jointly. In this paper, we show that retrieval and LLM policy learning are order-sensitive: adapting the retriever before optimizing the policy yields a larger reward gain than the reverse order. To preserve this hierarchy while allowing both components to co-adapt, we formulate retrieval-augmented agentic RL as a bilevel optimization problem. To solve it efficiently, we introduce BRIDGE, a memory-efficient first-order bilevel method motivated by a loss-landscape analysis of the RL and retrieval objectives. Across seven open-domain QA benchmarks, BRIDGE achieves the highest average accuracy with both 3B and 7B backbones, improving the multi-hop average over the strongest baseline by 9.6 and 3.4 EM points, respectively. It also achieves the best averaged answer accuracy and reasoning quality across medical QA benchmarks.

[AI-213] AVIO: Learning to Add and Remove Sounding Objects in Audiovisual Scenes

链接: https://arxiv.org/abs/2609.36503
作者: Weihan Xu,Kan Jen Cheng,Koichi Saito,Jingyu Shi,Tingle Li,Yisi Liu,Liming Wang,Masato Ishii,Takashi Shibuya,Gopala Anumanchipalli,Paul Pu Liang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Adding or removing a sounding object requires coordinated changes to visual content and sound while preserving the surrounding scene. Yet paired supervision for localized non-speech audiovisual editing remains limited, as visual and acoustic edits must target the same object and isolate its sound from overlapping sources. To address this gap, we introduce \textitAVIOBench, a dataset comprising 37.9 hours of paired audiovisual examples spanning 1,878 target-object names. AVIOBench links the visual presence and acoustic contribution of each target object through a shared identity and visual mask. Our automated pipeline uses visual grounding and cross-modal consistency to select target-sound removal candidates, then jointly refines the audiovisual pairs to improve perceptual quality and cross-modal consistency. Building on this dataset, we propose \textitAVIO, which adapts a pretrained text-to audiovisual generation model through source-conditioned feature modulation to jointly learn object addition and removal. A reference-frame curriculum gradually reduces reference conditioning during training, enabling one model to perform instruction-only editing with optional visual guidance. Quantitative and qualitative evaluations demonstrate effective audiovisual object removal and addition, with optional reference guidance providing appearance and placement control for addition.

[AI-214] InterBias-SV: Compound Conditions in Speaker Verification

链接: https://arxiv.org/abs/2609.36500
作者: Kamel Kamel,Hridoy Sankar Dutta,Keshav Sood,Sunil Aryal
类目: ound (cs.SD); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Speaker verification systems encounter combinations of noise, channel distortion, and changes in speech. Evaluating each condition separately does not establish whether their effects add. InterBias-SV organises this question around a four-term comparison: joint error, two marginal errors, and a common reference. Its results artefact contains 4,068 scored records across 17 experiments, 12 encoder labels, and six speech corpora, totalling 12 million trial evaluations. Three experiment families contain the same-corpus terms needed to compute additive contrasts. For labels assigned to speaker-trained encoders, their mean contrasts are +0.0026, +0.0088, and +0.0024 in equal error rate (EER), with larger variation across settings. These descriptive averages do not establish equivalence to additivity: trial matching, checkpoint identity, and parts of the condition metadata remain unverified. We also examine two interpretation problems. Near-chance EER can make additive predictions difficult to interpret, but chance performance is not a hard EER ceiling, and correlation with the prediction does not identify a saturation mechanism. Ratios of demographic gaps are unstable when their clean reference is near zero; absolute gaps provide a more direct summary. The benchmark provides condition definitions, analysis scripts, and explicit requirements for interpretable compound-condition comparisons, while separating recomputable summaries from claims that require further experimental validation.

[AI-215] LLM s Learn to Evade Latent Monitors from Prior Feedback Alone

链接: https://arxiv.org/abs/2609.36490
作者: Hugo Lyons Keenan,Christopher Leckie,Sarah Erfani
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Latent space monitors aim to detect undesired behaviors in LLM agents by inspecting an agent’s internal activations rather than its outputs. However, interactive monitoring creates a feedback channel where each verdict the monitor delivers leaks information to the model about how its internal states are being evaluated. We ask whether an agent can infer the monitor’s decision rule from this feedback and then selectively edit its activations to evade detection. Unlike prior evasion attacks, the model is never explicitly told what the monitor detects. Surprisingly, off-the-shelf models already produce activation edits aligned with the monitored direction, but at insufficient magnitude for evasion. Simply scaling up these edits by a factor of 8 reduces the monitor’s TPR from 100% to 27%. A rank-1 LoRA amplifies this behavior into effective evasion within the forward pass, reducing TPR further to 4% on held-out concept monitors while leaving other concepts at their normal detection rates. Capabilities on standard benchmarks are retained under this finetuning, and the evasion skill survives retraining the monitors on the new activations. Mechanistically, we find evidence that the model computes its activation edit from the prior in-context turns, and show that the edit becomes more aligned with the monitored direction as more examples are provided. These results demonstrate feedback-conditioned control over activations and suggest that latent monitoring should be treated as an interactive process in which agents can observe and respond to oversight measures.

[AI-216] Learning to Harvest Without Collapse in a Regenerative Commons: A Lagrangian Framework

链接: https://arxiv.org/abs/2609.36478
作者: Jose Tupayachi,Xueping Li,Soham Das
类目: Artificial Intelligence (cs.AI); Computer Science and Game Theory (cs.GT)
备注:

点击查看摘要

Abstract:The tragedy of the commons poses a multi-agent safety problem: reward-seeking agents can deplete a shared resource, and cooperation among its users does not itself specify how much must be preserved. We make preservation an explicit requirement by formulating a regenerative commons as a constrained Markov game or a constrained multi-agent MDP with a designer-specified depletion budget. We develop a nonstationary Lagrangian framework that constructs a policy sequence from solutions of unconstrained games or cooperative control problems. Extending earlier time-average constructions, we introduce average-epoch solution concepts for reset episodes with discounted rewards and terminal costs. We prove a reward-independent feasibility certificate, cooperative feasibility and approximate optimality against feasible policy mixtures, and an extension to unbiased sampled costs. For self-interested agents, a constrained Nash certificate quantifies the price-dispersion term introduced by deviations that redistribute budget across epochs. Under the stated assumptions on solver accuracy and multiplier updates, these results give constrained policy-sequence guarantees using solutions of unconstrained problems. Experiments with constrained IPPO and MAPPO in a Gordon-Schaefer fishery examine how depletion budgets shape stock retention, harvest rewards, and price adaptation.

[AI-217] Guard Models Are Overconfident Where Base Models Are Uncertain EMNLP2026

链接: https://arxiv.org/abs/2609.36477
作者: Jonghyun Hong,MinJae Jung,Minwoo Kim
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Accepted at EMNLP 2026 Findings

点击查看摘要

Abstract:Guard models are used as safety classifiers, with confidence scores driving downstream moderation decisions. We evaluate five guard models for prompt classification and find that although several are nearly calibrated on clean inputs, adversarial attacks degrade their calibration by an order of magnitude, turning false negatives into high-confidence errors indistinguishable from correct detections. Comparing each guard with its corresponding base LM, we find that uncertainty signals often remain available, with the base model typically expressing uncertainty on the same inputs where the guard fails. Layer-wise analyses localize this guard-base divergence to later layers, where guard models exhibit sharper safe/unsafe separation and lower-rank representations, while adversarial harmful inputs lie closer to the clean-safe region. These findings highlight a mismatch between guard confidence and base model uncertainty under attack.

[AI-218] Going Beyond State-Reaching: Learning Abstractions for Intrinsically Motivated Option Discovery

链接: https://arxiv.org/abs/2609.36473
作者: Akhil Bagaria,Anita De Mello Koch,George Konidaris
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Temporal abstraction via options can improve exploration in large environments. However, existing option discovery algorithms find subgoals that target all aspects of the state simultaneously. This state-reaching approach produces options that only apply in narrow regions of the state-space, eventually causing an explosion in the number of options that overwhelms the agent, and impedes progress on its primary task of reward maximization. We introduce an algorithm that instead identifies a small, relevant subset of features for each subgoal, yielding options that generalize broadly and accelerate exploration. Our approach learns abstract, transferrable options and achieves rapid exploration in three sparse-reward, image-based domains, including the Atari game MontezumasRevenge.

[AI-219] Rethinking Reasoning Paths as Phase-Structured Trajectories

链接: https://arxiv.org/abs/2609.36461
作者: Zhenghao He,Guangzhi Xiong,Sanchit Sinha,Bohan Liu,Wenqian Ye,Aidong Zhang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models often improve problem-solving performance by generating multi-step reasoning paths, yet how to analyze the hidden states along these paths remains unclear. Existing approaches typically assign each intermediate state the final-answer correctness label and train probes across heterogeneous questions. We argue that this protocol obscures reasoning dynamics in two ways: (1) correctness prediction can exploit question-level variation rather than path quality, and (2) states aligned by absolute step indices may correspond to different functional phases of reasoning. In this work, we propose to view reasoning paths as phase-structured trajectories within fixed questions. We instantiate this view as PAIR, short for Phase-Aligned Intra-question Reasoning. PAIR samples multiple trajectories for each question, maps variable-length paths into shared relative phases based on normalized trajectory progress, and compares successful and unsuccessful trajectories only within the same question and phase. This yields phase-specific path-quality directions that better isolate path-quality signals from question-level variation. Empirically, we find that standard across-question correctness probes lose much of their predictive power under within-question evaluation, suggesting that these probes partly rely on question-level information. PAIR improves within-question trajectory ranking and Best-of-N trajectory selection across models and benchmarks. Phase-wise steering further shows that the learned directions can change generation outcomes, providing causal evidence that they capture trajectory-relevant information.

[AI-220] Emergent Tonal Structure in Learned Chord Embeddings and Its Relation to Tonal Tension

链接: https://arxiv.org/abs/2609.36460
作者: Maral Ebrahimzadeh,Gilberto Bernardes,Sebastian Stober
类目: ound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 8 pages, 3 figures, 8 tables, Accepted at the 27th International Society for Music Information Retrieval Conference (ISMIR) 2026

点击查看摘要

Abstract:Several tonal pitch spaces and computational models have been proposed to analyze tonal structure in Western tonal music, many of them grounded in principles from music theory and used to support tonal analysis with important implications for tonal tension. In parallel, data-driven methods such as skip-gram have been used to learn chord embeddings from symbolic corpora, but their ability to recover tonal structure and its relation to tonal tension remains underexplored. In this work, we investigate how skip-gram chord embeddings reflect tonal structure and whether they provide a useful basis for analyzing structural aspects of tonal tension. Using chord sequences with and without transposition-based augmentation, we evaluate the learned spaces from geometric, functional, and tension-related perspectives. We show that augmented embeddings exhibit strong transposition equivariance, recover a clear circle-of-fifths structure, and support interpretable shifts between key-related regions of the learned space. We then derive embedding-based measures from chord-to-key distance and contextual chord-distance relations, and show that they capture meaningful aspects of tonal tension structure through correspondence with matched tonal measures and moderate alignment with human tension profiles. Across analyses, transposition-based augmentation generally improves the stability, tonal coherence, and interpretability of the learned space.

[AI-221] Channel-Dependent State Space Model for Multivariate Time Series Forecasting

链接: https://arxiv.org/abs/2609.36453
作者: Yu-Cheng Wu,Fan-Keng Sun,Li-Chun Lu,Duane S. Boning
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Multivariate time series forecasting (MTSF) is critical across many real-world domains. Existing deep learning approaches fall into two paradigms with distinct limitations: channel-independent (CI) methods unconditionally ignore cross-variable dependencies and model only temporal dynamics, while channel-dependent (CD) methods consider both but typically rely on architectural compromises to mitigate overfitting and computational overhead. We therefore propose Chameleon, a specialized CD state space model (SSM) that enables data-dependent, fine-grained interactions across variables while scaling linearly with their number. By connecting selective SSMs with the Kalman filter, we leverage the missing measurement update in the former for cross-variable modeling while preserving the SSM backbone for robust temporal modeling. We further identify favorable inductive biases of GatedDeltaNet for time series, adapt it as our backbone, and improve generalization through additional techniques, including a previously unexplored stochastic perturbation of reversible instance normalization. On strongly dependent ODE and PEMS datasets, Chameleon achieves the best MSE and MAE across all settings, while its CI ablation and prior CD methods incur 61-178% higher MSE on average. Across 28 standard benchmark settings, Chameleon also achieves better MSE and MAE than each baseline in at least 27 and 22 cases, respectively. Training-time and peak-memory analyses on Traffic and ETT further demonstrate competitive efficiency and favorable memory scalability across different variable counts.

[AI-222] Bits Under ZK-LLM : Evaluating Zero-Knowledge-Friendly Quantization for Verifiable Private LLM Inference

链接: https://arxiv.org/abs/2609.36437
作者: Taeung Yoon,Yupeng Zhang,Xiaojing Liao
类目: Artificial Intelligence (cs.AI)
备注: 22 pages, 5 figures

点击查看摘要

Abstract:Zero-knowledge proofs are emerging as a promising approach for enabling private, verifiable LLM governance and auditing, where regulators, users, and auditors need to verify claims about training-data usage or LLM inference-time behavior, while model providers must protect proprietary model parameters. However, despite the growing interest in ZK-LLMs, the understanding of ZK-friendly quantization remains limited. This gap matters because in the ZK setting, quantization directly shapes the arithmetic structure, constraint complexity, and proving cost of ZK inference. ZK protocols operate over finite fields and incur costs that depend heavily on the number and type of arithmetic operations, nonlinearities, and lookup constraints. Understanding ZK-friendly quantization is therefore essential for making ZK-LLMs practical. In this work, we present the first systematic study of ZK-friendly quantization for LLMs. We first formalize the definition of ZK-friendly quantization, capturing the properties required for ZK proof generation. We then evaluate nine language models, including Qwen2.5-14B and the mixture-of-experts model Qwen3-30B-A3B, across a broad design space of weight, activation, and nonlinear lookup table precision. Our results show that activation precision is substantially more sensitive than weight precision, while nonlinear lookup approximations can become the dominant source of utility degradation. Also, we identify RMSNorm inverse-square-root lookups as a recurring bottleneck in several large models and recover near-baseline utility by selectively increasing precision only at the bottleneck. Finally, we show that reducing bit-width or lookup-table size does not necessarily yield proportional end-to-end proving savings, showing that conventional low-bit quantization heuristics do not directly translate to ZK proving efficiency and motivating operator-aware precision selection.

[AI-223] he Safety Operator: Modulating the Expression of Safety Instructions via Spectral Optimization

链接: https://arxiv.org/abs/2609.36434
作者: Benoit Dherin,Michael Munn,Xavier Gonzalvo,Adrian Goldwaser,Blaz Bratanic,Ananth Balashankar,Andrey Vlasov,Pinzhi Huang,Cecile Loge,Nicole Mitchell,Andre Fernandes,Trilok Acharya,Wendy Kan,Ziyue Wang,Hanna Mazzawi,Felipe Tiengo Ferreira,Mor Geva
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Context tokens in a transformer-based language model can be absorbed into the model’s weights as a multiplicative operator. We study this operator in the setting of safety instructions and show that influencing its dominant eigenvalue modulates how strongly the instruction shapes generation. We derive a Contrastive Safety Loss with a suppression weight that controls the tradeoff between emphasizing the safety instruction on harmful queries while suppressing it on harmless queries. Varying the suppression weight maps a relationship between the attack success and the over-refusal rates, supporting the hypothesis that the operator’s eigenvalue acts as a continuous dial for the instruction’s influence. Moreover, this relationship holds relatively independently of how the Safety Loss is parameterized, yielding Pareto-improved safety instructions for appropriate values of suppression weight.

[AI-224] Support-Set Target Leakage in Relational Foundation Models during In-Context Learning: Model Dependence and Evaluation Reliability

链接: https://arxiv.org/abs/2609.36417
作者: Roshan Reddy Upendra,Alexandre Dorais,Joe Meyer,Andrew Pouret,Anastasios Lambrianos Stappas,Dinesh Katupputhur Ramprasath,Viswanath Ganapathy,Tom Palczewski,Minghua Li
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Relational in-context learning (ICL) uses labeled support examples and their linked relational context to predict labels for new queries. This creates a failure mode when target-derived features are present in the support context but unavailable for the query. We study this setting as support-set target leakage. We construct 20 controlled target-derived features that vary in signal fidelity, representation, semantic transparency, coverage, and zero-, one-, and two-hop relational placement, and evaluate them across 13 RelBench tasks and five relational ICL configurations that vary the ICL head, message-passing depth, pretraining cohort, or relational encoder architecture. We evaluate matched 0-hop, 1-hop, and 2-hop leakage settings, together with a Full leakage condition containing all 20 leaker columns. Within the tested configurations, target-table (0-hop) and Full leakage produce the largest aggregate deviations from clean evaluation, while higher-hop effects are often weaker, consistent with differences in effective exposure associated with temporal reachability, sampling, and aggregation fidelity. Leakage effects are strongly task- and model-dependent and can reverse relative conclusions between model variants even when aggregate changes are small. For leaker detection, we compare an Integrated Gradients (IG)-based screening method with mutual information (MI) and leave-one-column-out (LOCO) on a common Baseline subset. Ranking quality is strongest in the high-impact 0-hop and Full leakage conditions, but detector-based removal does not consistently restore the clean evaluation. A four-task rel-salt case study further shows the same evaluation concern with native-schema leakage candidates from the original relational schema. These results identify the support/query information boundary as an important component of reliable relational ICL evaluation.

[AI-225] Longer Records Broader Invariance: The Hidden Scaling Problem in Longitudinal Contrastive Learning

链接: https://arxiv.org/abs/2609.36409
作者: Rameen Mahmood,Xuhai “Orson” Xu,Zachary Beattie,Jeffrey Kaye,Danny Yuxing Huang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 53 pages, 6 figures, 54 tables

点击查看摘要

Abstract:Longitudinal data are valuable because people change. Yet the objectives used to learn from these data can inadvertently erase that change. In person-level contrastive learning, observations from the same person are treated as positives; as records grow, those positives can span increasingly distant—and increasingly different—behavioral states. More history can therefore produce not only more data, but broader invariance. We show that this distinction is fundamental. We separate \emphrecord span, how much history the learner sees, from \emphsupervision span, how far across that history positive-pair supervision reaches. Across in-home sensing records spanning up to 2.7 years, broader supervision systematically suppresses recoverable changing-state information, even when the available history is held fixed. At the broadest span, less than 10% of the information recoverable from an untrained encoder remains. Yet keeping positives local is not sufficient: as records grow, even distant states that are never paired become increasingly similar. Explicitly contrasting other observations from the same person reverses this loss without shortening the record, revealing a second route by which longitudinal scale can broaden invariance. Finally, we prospectively reproduce the supervision-span effect in 199 GLOBEM participants. Longitudinal scale therefore presents a choice: more history need not mean more invariance. By controlling what is held invariant as records grow, we can preserve the change that made the longitudinal data valuable in the first place.

[AI-226] From Retrieval to Reasoning : Agent ic Mechanism Prediction from Cell Painting Profiles NEURIPS2026

链接: https://arxiv.org/abs/2609.36406
作者: Jiayuan Chen,Botao Yu,Tianyu Liu,Thai-Hoang Pham,Meng Wu,Ping Zhang
类目: Artificial Intelligence (cs.AI)
备注: NeurIPS 2026 main track

点击查看摘要

Abstract:Cell Painting is a high-content morphological profiling assay widely used for phenotype-based biological inference, with mechanism of action (MOA) prediction as a central application. Existing approaches largely formulate Cell Painting-based inference as representation matching, assigning predictions from nearby reference perturbations in morphological feature space. However, retrieved neighbors are often noisy and partially misleading evidence due to batch effects, non-specific cytotoxicity, phenotypic convergence, and source-dependent variability. We reformulate Cell Painting-based MOA prediction as a calibrated evidence reasoning problem, where retrieved neighbors are treated as uncertain observations that must be evaluated, compared, and sometimes rejected before supporting a mechanistic conclusion. We propose PhenoAIR, a reliability-aware multi-agent framework that maintains a candidate-centric evidence memory and performs controller-guided refinement over phenotype- and mechanism-side evidence. PhenoAIR uses offline reference-set calibration to weight evidence by source reliability, phenotype stability, and mechanism-level confusion. We evaluate PhenoAIR on a benchmark constructed from JUMP Cell Painting profiles and annotations, covering controlled, realistic, and discovery-oriented open-world MOA prediction settings. PhenoAIR outperforms representation-matching and LLM-based baselines across all settings.

[AI-227] Reward-rate Policy Gradient for Efficient Machine Learning Engineering Agents

链接: https://arxiv.org/abs/2609.36393
作者: Muhang Tian,Sherry Yang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Traditional reinforcement learning (RL) techniques focus on maximizing expected cumulative reward, where each action assumes to take a constant unit of time. However, this assumption does not hold for agentic RL tasks such as machine learning engineering (MLE) agents, where actions involve data loading, feature engineering, and model training that take variable durations. Efficiency matters in modern agentic RL where actions are costly. To address this limitation, we adapt from continuous-time RL and Semi-Markov Decision Process (SMDP) formulation and propose Reward-rate Policy Gradient (RPG), where we focus on optimizing the reward rate – the long-term reward per unit of time. RPG estimates the reward rate from off-policy samples, then charges each action for the time it consumes at that rate. We first conduct theoretical analysis in the bandit setting to establish that RPG approximates the optimal reward rate and empirically demonstrate it outperforms baselines while avoiding enumeration over the policy space, a known issue for an existing method. We then further apply RPG on a small language model (Qwen3.5-4B) with self-improvement loops and empirically show it obtains higher rewards within a fixed time budget than vanilla RL on MLE-Bench and NanoGPT, with a 19.2% and 85.7% margin, respectively. Our method provides a practical solution for optimizing performance under wait time considerations in modern agentic RL tasks, where actions interact with external environments and cost time.

[AI-228] Persona Dosing: Calibrated Activation Steering for Graded Trait Control

链接: https://arxiv.org/abs/2609.36388
作者: Zehao Jin,Junran Wang,Ruixuan Deng,Jiahao Chen,Jingyuan Zhang,Yuxuan Zhang,Xinjie Shen
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:An activation-steering coefficient sets intervention strength, but requesting a particular degree of persona expression requires a behavioral scale. We study persona dosing: controlling a language model through a trait description and a requested mean intensity. PersonaDose specializes a shared, description-conditioned FLAS controller on persona responses, then calibrates its flow time against measured trait expression. Training responses are not paired with requested target intensities. Across Llama-3.1-8B, Qwen3-8B, and Gemma-3-4B, PersonaDose raises core-trait expression at the Persona Vectors coherence floor of 75 by 33.2, 18.3, and 17.8 points over contrastive activation addition. Calibration-selected settings retain an expression advantage on held-out questions, although the coherence floor does not hold for every trait there. Across seven trained traits, calibrated requests yield mean targeting errors of 4.7-6.2 points over 14-22 calibration-reachable targets out of 28 per model. These results separate the behavioral range learned by a controller from the accuracy of requests within that range.

[AI-229] Support-Set Target Leakage in Relational Foundation Models during In-Context Learning: Impact Detection and Mitigation

链接: https://arxiv.org/abs/2609.36384
作者: Roshan Reddy Upendra,Alexandre Dorais,Joe Meyer,Andrew Pouret,Anastasios Lambrianos Stappas,Dinesh Katupputhur Ramprasath,Tom Palczewski,Minghua Li
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Relational in-context learning (ICL) conditions predictions on the labeled support examples and their linked tables, creating a failure mode when the support set contains target-derived features that are unavailable for the query. We formulate this problem as support-set target leakage, distinct from leakage during dataset construction, temporal splitting, or representation learning. Here, the target-derived (leaker) columns are present only in the labeled support set during relational in-context inference, while queries remain clean. We construct 14 synthetic leaker types, corresponding to 20 columns, spanning proxies with different noise levels, coverage, modalities, semantic transparency, and relational distances. We evaluate a frozen relational encoder with an ICL head on held-out RelBench databases and use Integrated Gradients (IG) to rank and remove suspicious columns. Our results show that the effect of support-set leakage varies across tasks and relational distances. Target-table leakers cause the clearest degradation, while one- and two-hop leakers are not consistently used by the model. IG ranks target-table leakers highly across datasets and partially recovers performance in settings where leakage has the largest effect.

[AI-230] Quantization Enables Private Dense Retrieval against Malicious Service Providers

链接: https://arxiv.org/abs/2609.36376
作者: Louis Tremblay Thibault,Sofiane Azogagh,Marc-Olivier Killijian,Ulrich Aïvodji
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Dense retrieval, the key component of Retrieval Augmented Generation (RAG), retrieves the most relevant documents by comparing dense vector representations of queries and passages from a large corpus. In privacy-sensitive applications, the server observes the query and controls which evidence is returned, creating both confidentiality and integrity risks. We formulate private dense retrieval as providing query privacy and retrieval integrity against a malicious server, and develop a two-round cryptographic protocol that provides both guarantees. Our protocol reduces private and verifiable retrieval to multiplication of a committed matrix by an encrypted vector and uses low-bit quantization to make this computation practical. We evaluate the resulting trade-off between cryptographic cost, retrieval quality, and downstream RAG accuracy across six embedding models, four language models, and corpora of up to 2.68 million passages. Our results show that, with a clipped quantizer, three-bit quantization largely preserves retrieval quality and downstream accuracy, while a private query over a corpus the size of a clinical reference requires one to three minutes of server time. These results suggest that private dense retrieval is already practical for moderately sized, privacy-sensitive corpora when minute-scale latency is acceptable.

[AI-231] Audience-Bound Persistent Memory: Authorization Across the Memory Lifecycle

链接: https://arxiv.org/abs/2609.36373
作者: Sibo Liu
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 23 pages, 6 figures, 15 tables. Includes technical appendices and an ancillary research artifact

点击查看摘要

Abstract:A personal language agent that acts for its owner across private and shared conversations can learn a fact from one audience and later place it in the context it assembles for another. We study authorization before context across the whole memory lifecycle. Each memory item carries the audience present when it was recorded; derived items are partitioned by audience, receive the intersection of their sources’ audiences, or are suppressed; an audience widens only by an explicit, object-specific grant; and an item enters a model attempt only when every current viewer belongs to one of its authorized audiences, with unresolved viewers failing closed to public-only. Under explicit identity, provenance and complete-mediation assumptions, this admission is sound and policy-complete on the exact assembled context, enforced by exclusion rather than by model behavior. We realize it in two independently persisted reference architectures, a flat store and a relationship graph, and, descriptively, in a native agent-memory runtime. In a prospectively frozen confirmation over 10,000 multi-party histories, no forbidden item entered any architecture’s context, whereas unscoped retrieval exposed forbidden items in 82% of its contexts. Entitled recall matched policy-equivalent baselines exactly and exceeded unscoped retrieval by 0.30 Recall@5, with a Holm-confirmed advantage that grows with distractors. No architecture produced a wrong-principal substitution, but unscoped substitutions were too rare to establish the prespecified joint decision.

[AI-232] LatentSift: Policy-State Filtering for Token-Efficient Verification of Software Engineering Agents

链接: https://arxiv.org/abs/2609.36371
作者: Yuning Han,Yangchenchen Jin,Tyler Jandreau,Jingwei Sun
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Test-time scaling improves software engineering agents by generating multiple candidate trajectories and selecting the best one. Verifying and selecting among these long interactions can consume as many tokens as generation itself. Existing hybrid workflows first apply an LLM-based execution-free (EF) verifier to filter candidates before running tests, which adds another model pass over every trajectory. We introduce LatentSift, a token-free and execution-free filter that replaces this first stage with hidden states the policy already produces while generating the candidates. It represents each candidate through its reasoning, observation, and function-call states, compares them with positive and negative banks of such states collected from successful and unsuccessful trajectories during policy training, and fuses the resulting distance scores with a learned linear score to retain promising candidates for the execution-based stages. On SWE-bench Verified, across three agents and two policy sizes, LatentSift cuts EF-verifier tokens by 66.6–81.0% and total verification tokens, which include test generation, by 49.1–62.1% at K=16, while hybrid Best@16 matches or improves on each agent’s reference workflow, rising from 59.26% to 60.06% on DeepSWE-Preview.

[AI-233] Engineering Simplicity: Simple Mechanism Interfaces Steer LLM Agents

链接: https://arxiv.org/abs/2609.36365
作者: Kehang Zhu,Anand Shah,David Parkes
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Can interaction formats and textual scaffolds help large language model (LLM) agents make better decisions, and do better decisions come with better explanations? We study these questions in auctions and matching, multi-agent environments with explicit rules and known optimal strategies. These settings let us vary how a decision problem is presented while retaining a benchmark for evaluating behavior. Drawing on human-motivated theories of simplicity, we compare interfaces that elicit a complete bid or ranking with sequential interfaces that make safe choices easier to identify. We then hold the interaction format fixed and vary reasoning scaffolds and rule descriptions. Across four model families, the ascending auction interface substantially reduces bid deviations. The matching comparison also shows why sequential responses require different error accounting from complete rankings. Laying out payoff contingencies and explaining why truth-telling is safe also improve choices, whereas prompts to plan through matching rounds or form beliefs about opponents worsen play overall. In auctions, these behavioral gains are not accompanied by corresponding improvements in measured verbal indicators of strategic understanding in the agents’ short stated plans. Other prompts change those indicators without improving bids. Our findings suggest that human-motivated theories of simplicity can inform the design of decision environments for artificial agents. They also show why scaffolds should be evaluated through realized choices as well as explanations: improvements in one need not appear in the other.

[AI-234] HyperZip: Efficient Data Compression through Personalized Diffusion LLM s with Hypernetworks

链接: https://arxiv.org/abs/2609.36357
作者: Thai Nguyen,Khang Tran,NhatHai Phan
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models (LLMs) have shown strong potential for lossless data compression, but existing approaches are constrained by the high computational cost and low throughput of autoregressive decoding. We propose HyperZip, an efficient and scalable LLM-based compression framework that leverages diffusion-based LLMs (dLLMs) with Multi-Token Prediction (MTP) to accelerate LLM-based data compression processes. We identify a trade-off in diffusion-based compression, where increasing decoding throughput degrades the compression rate. To mitigate this trade-off, HyperZip employs a hypernetwork to generate data-specific updates from a context representation, adapting the dLLM to the target data without costly fine-tuning, resulting in a low compression rate and high throughput. Extensive experiments show that HyperZip achieves a superior trade-off between compression rate and speed compared with state-of-the-art baselines.

[AI-235] Explainability from Training with Applications to TCR-Epitope Prediction

链接: https://arxiv.org/abs/2609.36354
作者: Jiarui Li,Zixiang Yin,Samuel Landry,Zhengming Ding,Ramgopal Mettu
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); Biomolecules (q-bio.BM)
备注:

点击查看摘要

Abstract:Deep learning models have achieved strong performance in artificial intelligence for science, yet their black-box nature limits our understanding of how they learn scientific tasks. Existing methods for interpretability provide limited insight into how models organize evidence and evolve during learning. We introduce explainability from training (EFT), a model-agnostic paradigm that traces model interpretation during training to explain why models rely on specific features and how they organize these features as predictive evidence. We apply EFT to four state-of-the-art T cell receptor (TCR)-epitope prediction models, TCR-SRIM, TULIP, MixTCRpred, and NetTCR-2.2, spanning post-hoc and interpret-by-design approaches as well as transformers and CNNs. To investigate how structural information affects model explanations, we introduce a benchmark, TCR-XAI2, containing 388 unique experimentally resolved TCR-epitope structures, complemented by structures predicted using AlphaFold3, Boltz-2, TCRModel2, tFold-TCR, and OpenFold3. Using EFT with TCR-XAI2, we demonstrate that (1) CNN and transformer models exhibit distinct learning trajectories; (2) TCR \alpha and \beta evidence can conflict during learning, limiting the benefits of jointly modeling both chains, while MHC information mitigates this; and (3) real versus predicted structural data for TCR-epitope prediction exhibits distinct TCR and peptide feature preferences as well as differing trajectories of model certainty.

[AI-236] ATLAS: Aligned Transport of Latent Structure for Reliable World Model Planning

链接: https://arxiv.org/abs/2609.36333
作者: Ke Fang,Yupu Yao,Lu Cheng
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Latent world models rely on representation geometry for planning, yet regularizing the latent marginal alone does not determine the state-to-state relationships used for action selection. We show that this can cause planning-relevant novelty structure to be weakened as representations are transformed into the final latent used by the planner. We introduce Aligned Transport of Latent Structure (ATLAS), a training objective that explicitly preserves relational geometry while calibrating the global latent distribution. ATLAS transfers normalized pairwise structure from an informative encoder representation to the planning latent and uses Wasserstein embedding matching (WEMReg) to calibrate its marginal through one-dimensional Wasserstein-2 transport. Our analysis shows that relational preservation and marginal calibration impose non-redundant constraints, and connects finite-candidate planning stability to relational distortion, latent-scale mismatch, and prediction error. Instantiated in LeWM, ATLAS improves mean goal-reaching success across PushT, TwoRoom, and OGBench-Cube on both lower- and higher-novelty evaluation subsets, with the largest gain on higher-novelty TwoRoom episodes. Representation and rollout diagnostics further show stronger novelty-related structure in the planning latent, improved marginal calibration, and lower multi-step prediction error. Together, these results highlight preservation of planning-relevant latent geometry as an important ingredient for reliable world-model planning. Code is available at this https URL.

[AI-237] Calibrating One-Round Membership Inference with Neighbors

链接: https://arxiv.org/abs/2609.36331
作者: Francesco Rita,Jie Zhang,Florian Tramèr
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:The state-of-the-art Membership Inference (MI) methods calibrate their signal separately for each example using reference models, auxiliary models trained to exclude the target. This paradigm scales poorly to modern large models, however, whose training is too expensive to replicate. This has motivated one-round settings, where only a single trained model is available; but without reference models the per-example calibration that drives the strongest attacks can no longer be estimated, leaving the membership signal weak. We ask whether neighbors of the target point can recover this calibration without training any additional model. Our key observation is that reference models serve only to reveal how an example behaves under models not trained on it, and that querying the target model on nearby samples yields the same information. We propose two complementary ways to obtain such neighbors, and show that querying them against an early training checkpoint further sharpens the signal. We evaluate across three image classification datasets and three training setups, showing that neighbors yield strong membership signals and competitive attack performance at no additional training cost.

[AI-238] DecoyTrace: Toxic Decoys for Active Defense in Decentralized Federated Learning

链接: https://arxiv.org/abs/2609.36330
作者: Pedro Beltrán-López,Enrique Tomás Martínez Beltrán,Pantaleone Nespoli,Manuel Gil Pérez,Alberto Huertas Celdrán
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 38 pages

点击查看摘要

Abstract:Decentralized Federated Learning (DFL) eliminates the central aggregation server, reducing the single point of observation that traditional defenses against attacks rely on. As a result, peer-to-peer networks become exposed to malicious updates containing backdoors or semantic poisoning, since such updates can remain close to benign ones in the parameter space while behaving very differently. This may evade defenses based on passive parameter inspection. However, existing deception-based defenses have mainly been designed for centralized FL and do not jointly address local observation, poisoning propagation, source attribution, and containment in strictly serverless DFL. To address these limitations, this paper presents DecoyTrace, a proactive cyber deception-based defense for strictly serverless DFL environments. DecoyTrace deploys a mobile DecoyNode that generates decoy challenges using chaotic maps, disseminates a dual model (clean vs. decoy) based on neighbor trust, and evaluates them using three-state semantic metrics. Upon confirmation, a distributed protocol isolates the source and performs a model reset or recovery to preserve training progress. Evaluated across sixty configurations on the NEBULA platform (five datasets, three topologies, and four attack/defense scenarios), DecoyTrace systematically restores lost utility. The F1-score remains within 0.03 of the baseline on MNIST/FashionMNIST (mitigating drops of up to 0.37), matches or exceeds the baseline on EMNIST and CIFAR-100, and remains between 0.05 and 0.10 below the baseline on CIFAR-10, the most visually complex convolutional scenario evaluated. Furthermore, containment reduces CPU and network usage by up to two-thirds. These results demonstrate the feasibility of unifying deception, identification, and containment in DFL, while also identifying its limitations in complex tasks and multi-attractor threat models.

[AI-239] PILLAR: Private Inverted-Index Lexical Lookup for Augmented Retrieval

链接: https://arxiv.org/abs/2609.36326
作者: Truong Son Nguyen(1),Daniel Blackley(2),Ni Trieu(1),Evgenios M. Kornaropoulos(2) ((1) Arizona State University, (2) George Mason University)
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Retrieval-augmented generation (RAG) hands the user’s query to whoever hosts the corpus. We propose PILLAR, a Privacy-Preserving RAG (PPRAG) system based on Private Information Retrieval (PIR) in which a client utilizes the k documents most similar to their query from a server-held and publicly known corpus to respond to their query, while the server learns nothing about the query, either its terms or its access pattern. Prior PPRAG constructions rely on dense retrieval alone, translating approximate nearest-neighbor search into many query-dependent rounds of PIR, and pay for it in both latency and retrieval quality. PILLAR instead performs private hybrid retrieval in two stages. A sparse stage issues a small, fixed number of PIR queries against a carefully designed index of precomputed BM25 scores, filtering the corpus down to candidates that share terms with the query without the server ever seeing which terms these are. A dense stage then fetches only those candidates’ document embeddings and re-ranks them locally, avoiding the many costly PIR queries that private dense retrieval typically requires. We instantiate PILLAR with two protocols that trade latency against retrieval quality, each built on a different private rendering of lexical search. PILLAR-Bin bins posting lists into a hash table and is a single-round design that achieves lower latency than state-of-the-art private retrieval schemes. PILLAR-Tree turns block-max pruning into an oblivious tree traversal combined with cuckoo hash tables and achieves the highest retrieval quality at lower latency than state-of-the-art schemes.

[AI-240] owards an AI Software Factory for Data Systems

链接: https://arxiv.org/abs/2609.36323
作者: Anna Pavlenko,Bogdan Crivat,Brandon Haynes,Carlo Curino,Fotis Psallidas,Jaro Slawinski,Johannes Freischuetz,Laura Pereira Sanchez,Markus Weimer,Mathieu Demarne,Matthias Jasny,Mauktik Gandhi,Max Bovykin,Mirco Milletari,Purbasha Ghosh,Qiushi Bai,Raghu Ramakrishnan,Rahul Pandita,Sergiy Matusevich,Shivaram Venkataraman,Subru Krishnan,Md. Tareq Mahmood,Tiemo Bang,Venkatesh Emani,Xuan Zhao,Yiwen Zhu
类目: Artificial Intelligence (cs.AI); Databases (cs.DB); Software Engineering (cs.SE)
备注: 6 pages, 5 figures, 1 table

点击查看摘要

Abstract:AI-assisted coding tools deliver significant acceleration of coding, but only limited impact across the end-to-end software development lifecycle (SDLC)–an Amdahl’s law effect! In this paper, we discuss our progress towards building an AI SW Factory that accelerates all the stages of SDLC-Targeting, Coding, Reviewing, and Ops. The AI SW Factory produces a metadata exhaust that enables self-improvement by fine-tuning model weights and updating our World Model (a rich data substrate). We focus on Data Systems and the important class of Evolutionary Coding Tasks (i.e., those with a measurable objective to hill-climb) and report on 1) scaled deployments at Microsoft (tens of repositories) leading to 3x engineering efficiency above agentic coding and up to 22x token efficiency, and 2) several open challenges. Comments: 6 pages, 5 figures, 1 table Subjects: Artificial Intelligence (cs.AI); Databases (cs.DB); Software Engineering (cs.SE) ACMclasses: H.2.4; I.2.11; D.2.9 Cite as: arXiv:2609.36323 [cs.AI] (or arXiv:2609.36323v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2609.36323 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-241] Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression

链接: https://arxiv.org/abs/2609.36322
作者: Xingyu Zhu, Pu (Luke)Yi,Ziheng Cheng,Ang Lv,Jing Liu,Lexing Ying,Yiyuan Ma,Xin Dong
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 65 pages, 20 figures, pre-print

点击查看摘要

Abstract:Chunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride. Such compression also introduces a new positional coordinate: a token’s phase, or its position relative to compression-window boundaries. We uncover a systematic asymmetry in models using such compression: the same information can be easy to retrieve at one phase and difficult at another. We call this periodic variation in retrieval performance phase sensitivity. In large open-weight models with such compression, long-context retrieval accuracy can differ by up to 40 percentage points across phases, revealing periodic weak spots that average benchmark scores can conceal. To investigate this behavior, we pretrain a family of transformers from scratch across multiple KV-compression designs, reproducing phase sensitivity across the variants. Mechanistic analysis using causal interventions in these models reveals phase specialization: different attention components contribute asymmetrically to retrieving information at different source phases. We further analyze idealized retrieval models, showing how gradient flow dynamics may favor sharp phase specialization. Evaluating models with chunked KV-cache compression thus requires measuring across compression phases: high average accuracy can coexist with systematic positional failures. Comments: 65 pages, 20 figures, pre-print Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI) ACMclasses: I.2.6; I.2.7 Cite as: arXiv:2609.36322 [cs.LG] (or arXiv:2609.36322v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2609.36322 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-242] StateTape: Action-Conditioned Evidence Lifecycle Modeling for Long-Horizon Coding Agents

链接: https://arxiv.org/abs/2609.36319
作者: Ziyang Yu,Liang Zhao,Bowen Zhu,Hasibul Haque
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Despite the recent success of coding agents built on large language models, it remains challenging to run them over long horizons, since every observation is appended to the context and the context grows with each one. History-based maintenance is a common remedy, which masks or summarizes old observations, or prunes what a model reads as useless, and bounds the context at little cost. However, it decides from the text of the history alone and sees nothing of how the code is connected. Since a coding agent edits code many times over a single task, and each write can change what code elsewhere means, such maintenance may keep records a write has falsified, drop ones that still hold, and miss code the agent needs next. To overcome these challenges, this paper proposes StateTape, a novel and scalable framework that rewrites a coding agent’s context as the repository changes rather than as the context grows. The key idea of StateTape is to model the repository as a symbol-level code graph, whose dependencies and language rules expose which symbols a write can affect. Upon this graph, a tape marks the symbols each write changed, which turns staleness from an inference about text into an observation of the agent’s writes. We propose a per-write procedure in which the tape nominates the records a write could have falsified while a small manager model settles what the write log cannot, and further provide a theoretical analysis and TraceBench, a benchmark that labels what an agent is holding against what is actually needed. Empirically, we demonstrate that StateTape can effectively clear falsified records and retrieve what is needed, and thus achieve a higher resolve rate in all experiments spanned by six coding agents and three edit-heavy benchmarks with little computational overhead.

[AI-243] CheatBench: Measuring Reward Gaming in AI Agents

链接: https://arxiv.org/abs/2609.36308
作者: Long Phan,Stephen K. Yang,Jason J. Lim,Mantas Mazeika,Wenyu Zhang,Zheyuan Liu,Richard Ren,Jingxiang Meng,Yaoteng Tan,Weiliang Zhao,Addison Wu,Matei Anghel,Dan Hendrycks
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended. In recent incidents and controlled evaluations across the AI industry, agents trained to maximize reward have accessed unauthorized information, attempted to evade monitoring systems, and even breached sandbox protections to attack external systems. As agents become more capable, this behavior could pose increasingly serious risks. To measure this problem, we introduce CheatBench, a benchmark of cheating in AI agents across mathematical research, knowledge work, coding, visual tasks, and other domains. Its environments combine challenging assignments with opportunities to cheat, allowing researchers to study how agents pursue a goal when honest work is difficult. CheatBench supports comparisons across models and task categories, providing a testbed for measuring and reducing cheating as agents take on more consequential responsibilities. We publicly release CheatBench at this https URL

[AI-244] Proofs Without Nominals: Gödels Ontological Argument its Shallow Embedding and the Open Questions of the Monatshefte Notes

链接: https://arxiv.org/abs/2609.36279
作者: Christoph Benzmüller
类目: Logic in Computer Science (cs.LO); Artificial Intelligence (cs.AI); Logic (math.LO)
备注: 27 pages. Ancillary files: Isabelle/HOL and Lean 4 sources of every theorem, sixteen Isabelle sessions on the readings of the conjunction axiom with Lean counterparts and all 72 Nitpick searches as checked expect annotations, both hybrid-witness detectors with reports and census, five audit sessions and 19 cross-check theories. The sources of the Notes are in the Archive of Formal Proofs

点击查看摘要

Abstract:The shallow embedding of higher-order modal logic in classical higher-order logic, used in Benzmüller and Scott’s Notes on Gödel’s and Scott’s variants of the ontological argument (2025), reaches beyond the modal object language of the arguments: its property quantifiers range over terms that may also express nominals and satisfaction operators of hybrid logic, and a proof using one proves a theorem of the embedding that need not be one of the modal logic. That the framework affords this is not new, and whether a result is one of the modal logic can be settled in two ways: by replaying it in an explicit proof calculus, done by hand for chosen theorems, or by analysing the proofs the embedding itself produces, which this article does mechanically, for every result at once. Every statement the Notes prove has a proof inside the object language: 294 written out by hand and machine-checked, none using a nominal. The proofs the Notes themselves give instantiate no nominal either; what the detector flags there are terms a prover substituted. The three questions the Notes leave open are settled too, and without nominals, but the conjunction axiom has to be emended: generalised in the Notes to Gödel’s “any number of summands”, it covers the conjunction of no properties, and of one; the empty one alone settles all three, and the two together yield what a separate axiom of Gödel’s is for. This article restricts the conjunction axiom to at least two different conjuncts, the reading Gödel’s footnote suggests, and the questions are settled again, by proofs that turn on the argument rather than a degenerate instance. The restriction holds of the object language only: with a nominal the axioms make the accessibility relation the identity and the readings coincide. Every theorem is verified in Isabelle/HOL and independently in Lean 4; the countermodels are Nitpick’s, certified by the build. Comments: 27 pages. Ancillary files: Isabelle/HOL and Lean 4 sources of every theorem, sixteen Isabelle sessions on the readings of the conjunction axiom with Lean counterparts and all 72 Nitpick searches as checked expect annotations, both hybrid-witness detectors with reports and census, five audit sessions and 19 cross-check theories. The sources of the Notes are in the Archive of Formal Proofs Subjects: Logic in Computer Science (cs.LO); Artificial Intelligence (cs.AI); Logic (math.LO) MSC classes: 03B45, 03B15, 03B35, 03A05 ACMclasses: F.4.1; I.2.3 Cite as: arXiv:2609.36279 [cs.LO] (or arXiv:2609.36279v1 [cs.LO] for this version) https://doi.org/10.48550/arXiv.2609.36279 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-245] Illusory Truth or Mere Exposure? Model-Dependent Repetition Effects in LLM -Based Social Media Simulations

链接: https://arxiv.org/abs/2609.36278
作者: Azza Bouleimen,Nicolò Pagan,Anikó Hannák
类目: Artificial Intelligence (cs.AI); Social and Information Networks (cs.SI)
备注: Accepted at COLM 2026

点击查看摘要

Abstract:Generative agent-based models (GABMs) are increasingly used to simulate social media dynamics, including misinformation spread. For such social simulations to be valid proxies of human behavior, LLM agents should replicate established human cognitive biases, among them the Illusory Truth Effect (ITE), where repeated exposure to a claim increases its perceived truth value. We investigate whether and how the ITE manifests across four LLMs (Gemma-3-4b-it, Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct, and GPT-5-nano) in a social media simulation context. We propose a two-phase within-context experimental design that embeds the repetition manipulation inside a realistic news feed interaction. Using this design, we collect 336,000 truth, importance, sentiment, and interest ratings across 100 statements, 10 feed variants, and 3 replications. The key comparison is between ratings assigned to repeated statements, seen throughout a simulation phase, and completely unseen ones, rated within the same experimental context window. We distinguish genuine ITE (truth-specific repetition boost) from mere exposure effects. We run an OLS regression followed by a Linear Mixed-Effect Model to account for differences across models and ratings. Our results reveal four qualitatively distinct patterns: Gemma-3 exhibits a genuine ITE; Qwen2.5 shows a mere exposure effect; GPT-5-nano displays no repetition effect on truth and mild skepticism toward repeated content; Llama-3.1 shows a small truth boost alongside decreases in evaluative dimensions. Crucially, temperature has no effect on these findings, and a variance decomposition highlights the high context-sensitivity of LLM rating behavior. Our findings caution against assuming uniform ITE replication across LLMs in social simulations, while suggesting that Gemma-3-4b-it may offer the most behaviorally realistic approximation for misinformation-related simulations.

[AI-246] From Surfaces to Volumes: Registered Geometry for Protein Representation Learning

链接: https://arxiv.org/abs/2609.36277
作者: Siyuan Chen,Cai Zhou,Jinrui Zhang,Zhaokang Liang,Taku Komura,Wojciech Matusik,Stephen Bates,Tommi Jaakkola,Wengong Jin,Peter Yichen Chen,Minghao Guo
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Existing protein geometry models typically represent molecular surfaces using local geometric features such as sampled points, normals, and curvature. While effective for capturing exposed molecular shape, these representations do not explicitly model the volumetric organization beneath the surface or provide a consistent coordinate system for residue-wise volumetric structure. We introduce Protein-TetSphere, a registered residue-wise volumetric representation for proteins. Each protein chain is tetrahedralized to obtain local volumetric regions associated with individual residues, which are then registered to a shared fixed-topology tetrahedral reference and represented in a common Laplacian basis. This registration establishes consistent volumetric coordinates across residues, enabling local three-dimensional deformation to be integrated with surface and chemical information in a multimodal protein representation. We evaluate Protein-TetSphere on ligand-binding pocket classification, protein–protein interface prediction, and de novo protein binder design. Across the three tasks, Protein-TetSphere improves ligand-binding pocket balanced accuracy from 0.795 to 0.826 , Pinder-Pair/Site AUROC from 0.914/0.852 to 0.932/0.866 , and binder-design success from 14.95% to 19.90% on the BoltzGen Challenge Set and from 27.62% to 32.19% at the ProtDBench backbone level. These results show that registered volumetric geometry provides complementary spatial information beyond molecular surfaces across protein recognition, interaction, and design.

[AI-247] Paired Multimodal Scaling Laws

链接: https://arxiv.org/abs/2609.36263
作者: Marcus Ma,Shrikanth Narayanan
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Existing multimodal scaling laws fit multimodality terms empirically after testing and never vary how much data is multimodally paired at fixed data budgets. We investigate how, under the same total data per modality, changing the number of paired data affects loss curves in multimodal classification tasks. We train models in three different environments and run experiment sweeps varying data sizes and pairing budget. Pairing ratios have a dramatic impact on loss and this impact is directly tied to how much information synergy the task contains. Only paired data is able to reduce synergistic loss, while unpaired data can reduce redundant or unimodal information up until unimodal floors. Unlike traditional scaling laws where loss drops immediately in power law decay, synergy acquisition is gated, requiring a critical threshold of paired data before synergistic loss falls at all. We introduce a new family of multimodal scaling laws where total data-attributable loss is the sum of four individual power laws corresponding to the four different information channels of redundancy, a unique channel per modality, and synergy, and show how this law is both more theoretically sound and empirically valid across our experiments. This law predicts multimodal loss in our experiments more accurately than existing laws, with 3.2% error on fit tests versus 10.4% error for the best pairing extension of published laws.

[AI-248] owards Mitigating Deceptive Safety Alignment in Large Reasoning Models NEURIPS2026

链接: https://arxiv.org/abs/2609.36254
作者: Xiangyu Zhou,Saleh Zare Zade,Rafi Ibn Sultan,Alexander Kotov,Dongxiao Zhu
类目: Artificial Intelligence (cs.AI)
备注: Accepted at NeurIPS 2026

点击查看摘要

Abstract:Large Reasoning Models (LRMs) are commonly trained with reinforcement learning (RL) to improve their generation of chain-of-thought (CoT) reasoning before producing final answers. However, RL rewards are typically assigned based on final answers, providing little or no direct supervision over intermediate reasoning. This can lead to deceptive safety alignment, where the reasoning trace and final answer convey inconsistent safety signals. To systematically investigate this phenomenon, we introduce DSAR (Deceptive Safety Alignment Rate), a metric that jointly assesses reasoning traces and final answers to quantify their safety inconsistency. Across multiple LRMs and benchmarks, we find that deceptive safety alignment is pervasive under standard prompting conditions and is substantially amplified under prefilling attacks. We further provide a hidden representation analysis showing that models exhibit stronger safety discrimination at the final-answer stage than during intermediate reasoning. To close this gap, we propose SARA (Safety-Aware Reasoning Alignment), an RL-based method that rewards both safety-aware reasoning and safe final answers, encouraging early harmful intent recognition and enforcing reasoning-answer consistency. Experiments show that SARA significantly mitigates deceptive safety alignment under both standard and adversarial settings while preserving helpfulness and utility. Code is available at this https URL.

[AI-249] CoRe: Co-Evolving Reward Models for Mitigating Latent Reward Hacking in Video Diffusion Models

链接: https://arxiv.org/abs/2609.36245
作者: Zhaolong Su,Yujin Han,Feng Wang,Jameson Dong,Hins Hu,Difan Zou
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Latent reward models (LRMs) enable efficient alignment of video diffusion models by scoring intermediate states directly in latent space. However, we find that optimizing against a fixed latent reward rapidly leads to latent reward hacking: the predicted reward stays high while perceptual and motion quality deteriorate. Our analysis identifies distributional escape as the central cause: within a few hundred updates, the generator moves beyond the reward model’s training support, where its scores no longer reflect video quality. Based on this insight, we introduce CoRe, a co-evolving reward framework that treats latent-space alignment as a dynamic interaction between the generator and the reward model. Rather than optimizing against a stationary proxy, CoRe continually refits the reward model on the generator’s current samples while anchoring it to real-video preferences, so the generator cannot gain reward by drifting away from the data. On Wan2.1-T2V-1.3B, experiments show that CoRe consistently improves generation quality over both the pretrained model and prior alignment methods, while avoiding the quality collapse of fixed-reward optimization.

[AI-250] ChronoSRL: Temporal Geometry for Self-Supervised Reinforcement Learning

链接: https://arxiv.org/abs/2609.36238
作者: Nico Bohlinger,Jan Peters
类目: Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:A goal that is close in space can be far away in time. Obstacles, terrain, and the agent’s own capabilities determine how long it takes to get there. Yet, critics in contrastive and survival reinforcement learning do not measure the distances in their representation space in units of time. We therefore introduce ChronoSRL, which gives the critic’s embeddings an explicit temporal geometry. The distance between state-action and goal embeddings is trained to match the time that the agent takes to reach the goal (goal-reaching time), while goals that were not reached, and goals from other trajectories, are pushed at least one discount horizon away. Furthermore, reaching a goal quickly once does not mean that reaching it is reliable in general, so the policy should not follow the temporal distance directly. Instead, we build on survival reinforcement learning and predict from our temporal embeddings not only the full distribution of goal-reaching times but also the time spent near the goal. Thereby, the policy is trained to favor actions that reach the goal sooner and more reliably and that keep the agent near it. ChronoSRL learns faster and reaches higher performance than contrastive, action-chunked contrastive, and survival reinforcement learning baselines on seven standard locomotion and navigation benchmarks, even with much smaller networks. To test the limits of self-supervised reinforcement learning, we introduce velocity tracking, goal-position reaching, and box climbing tasks with a quadruped robot in a realistic sim-to-real locomotion setup, and show how the shaping terms that are typical for robotics can be naturally incorporated into our framework. ChronoSRL is the only one of the tested self-supervised reinforcement learning methods that learns to stay at the commanded velocities and goal positions, and climbs the highest boxes.

[AI-251] An Empirical Study and Assessment of EU AI Act Compliance Checkers

链接: https://arxiv.org/abs/2609.36228
作者: Zhen Tao,Alize Kahraman,Shidong Pan,Zhenchang Xing,Chiara Ullstein,Jens Grossklags,Chunyang Chen
类目: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
备注:

点击查看摘要

Abstract:The EU AI Act introduces extensive compliance requirements for organizations that develop, deploy, or integrate AI systems. Many of these requirements are directly relevant to security and privacy, while also addressing closely related issues such as data governance, transparency, accuracy, and robustness. However, stakeholders such as small-to-medium businesses and individual developers often lack the legal expertise required to interpret these obligations and translate them into engineering and governance practices. This disconnect creates challenges for implementing the EU AI Act and may lead to missing safeguards or misdirected development and deployment efforts. To address this, various automated EU AI Act compliance checkers (AIACCs) have emerged, claiming to streamline compliance assessments and provide practical guidance. In this paper, we present the first empirical study and assessment of AIACCs. We characterize 12 mainstream AIACCs across multiple dimensions, evaluate their legal coverage and alignment, and analyze checker-generated compliance reports for structure, determinacy, and actionability. We find that the quality of AIACCs varies significantly and that they currently can only serve as early-stage orientation tools. Specifically, we observe inconsistent interaction modes and user-friendliness, a tendency to overly simplify or omit key obligations, and a failure to provide determinate, actionable guidance. As a result, reliance on the current generation of AIACCs may foster a false sense of compliance. With our study, we provide a critical baseline of the current AIACC landscape. We further offer design principles for the implementation of more reliable compliance-support tools.

[AI-252] Representable but Unlearned: Encoding Rank and the Interaction-Prediction Floor

链接: https://arxiv.org/abs/2609.36208
作者: Zahra Khodagholi,Niloofar Yousefi
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Input encodings can restrict which measured contrasts a predictor can jointly reproduce, even when no single contrast is forced to vanish. We compute the attainable contrast space from an encoder’s equivalence classes and a fixed contrast design, without labels, loss, or a fitted model; projecting the recorded contrasts onto that space gives an empirical error floor for any unrestricted decoder on those classes. On a 140-rectangle siRNA interaction panel, a graph neural network’s training-only feature mask merges 165 endpoint states into 90 classes and cuts the rank of the 140 interaction contrasts to 72. The resulting floor is 0.009980, which is 14.6% of the fitted model’s interaction squared error; the fitted model reaches 0.068335, slightly worse than a control predicting no interaction at all. A minimum of three restored chemistry columns recovers full rank. Refitting without the mask removes the floor entirely, yet interaction MSE improves by only 0.000017 under the reported protocol, and the restored columns remain absent from every training input. On a released RNA-splicing predictor, whose encoding is injective on the measured states, the same computation returns the full design rank of 1,986 and a floor of exactly zero. These results separate what an encoding permits from what a fitted model achieves; they do not identify what limits the remaining error. The rank check needs no fits and bounds what any amount of training under a fixed encoding can recover. The project repository is available at this https URL.

[AI-253] FigAct: Turning Scientific Figures into Active Canvases for Explanation

链接: https://arxiv.org/abs/2609.36190
作者: Shishi Xiao,Zichao Wang,Alexa Siu,David H. Laidlaw,Jennifer Healey
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Scientific figures are designed to communicate information visually, yet MLLMs typically explain them by translating their visual content back into text. This requires readers to manually map the resulting explanations back to the figure. Inspired by how people present visual information, we introduce FigAct, a framework that transforms static scientific figures into question-conditioned visual presentations by acting directly on their existing graphical elements. Like a human presenter, FigAct generates a sequence of short narrations, grounds each narration in the corresponding visual evidence, and applies visual actions to guide the viewer’s attention. We develop a hierarchical search strategy for efficient element localization, reducing token usage by approximately 40 \times . We further train FigAct-8B using three task-specific rewards for grounding accuracy, search efficiency, and rendering quality. We further build a human-verified benchmark from figures in real-world scientific papers to evaluate the ability of MLLMs to generate grounded visual explanations. Our results demonstrate the effectiveness of FigAct and show that treating scientific figures as presentation canvases makes explanations clearer and easier to follow.

[AI-254] Draft in Parallel Condition Through Depth: Adjacent Causal Injection for Speculative Decoding

链接: https://arxiv.org/abs/2609.36173
作者: Haohui Zhang,Keyu Chen,Haocheng Sun,Weibo Gu,Ruizhi Qiao,Xing Sun,Bo Jiang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Parallel speculative drafting generates multiple candidates in one backbone pass, but independent token selection can produce inconsistent continuations that shorten the accepted prefix. Existing methods mostly leave conditional decoding to a lightweight module after the backbone, which limits the flow of predecessor information to successors. Our analysis of DFlash shows that early positions already form recoverable predictions in shallow layers, and that accurate adjacent predecessors help successors more when they enter earlier. We therefore propose DSpine, a drafter with causal conditioning injection throughout the backbone: at every layer, gated adjacent injection writes each predecessor’s predicted feature into its successor, so the causal conditioning chain unfolds over network depth while all positions update in parallel. A unified transfer space built from the target model’s output embeddings unifies layer-wise injection with predecessor-conditioned decoding, and layer-wise output-embedding supervision promotes the formation of predicted features in shallow layers. Fused kernels and a transition cache execute both efficiently in parallel within SGLang. Across seven math, code, and chat benchmarks, DSpine achieves the longest acceptance length at both temperatures on Qwen3-4B and Qwen3-8B. At temperature zero on Qwen3-8B, it raises the seven-benchmark mean from DFlash’s 3.77 to 4.82 (+27.8%); in SGLang serving tests, it delivers 23.3% higher throughput than DFlash on average.

[AI-255] Adversarial Debiasing of Machine Learning Models for Enhanced Network Security against DDoS Attacks

链接: https://arxiv.org/abs/2609.36167
作者: Aadith Sukumar,Isha Singh,Devershika Mohane,Ankit Mukherjee,Ankush Dutta,Rahee Walambe,Ketan Kotecha
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 20 pages, 2 tables, 5 figures

点击查看摘要

Abstract:Distributed Denial of Service attacks are a growing threat to network infrastructure, and new techniques, including the use of generative AI, make them harder to detect. Traditional detection systems, such as rule based firewalls, often fail to identify these evolving attack patterns. In this study, we propose a new method for detecting DDoS attacks by combining synthetic data generation using Generative Adversarial Networks with a Random Forest classifier. The GAN generated data showed 80.3 percent cosine similarity to real traffic, which helped the model learn underlying traffic patterns more effectively. To address imbalances in the data, especially in packet related features, we applied adversarial debiasing. This reduced the model’s sensitivity to skewed distributions in variables such as forward and backward packet counts and total byte lengths. Our results show that models trained on a mix of synthetic and real data achieved significantly better performance: 99.98 percent accuracy on benchmark data and a 22.60 percent improvement when tested on previously unseen synthetic traffic. This suggests that the method can generalize well across different traffic scenarios and adapt quickly to new types of attacks. The proposed approach not only improves DDoS detection but also provides a scalable foundation for security models that account for bias and benefit from data augmentation. Our findings show that combining GANs with adversarial debiasing can lead to more robust and effective DDoS mitigation, supporting the further development of machine learning based cyber security.

[AI-256] From Dead Code and Static Requirements to Working Engines: Software Revival with Coding Agents

链接: https://arxiv.org/abs/2609.36161
作者: Tianyu Liu,Dingyuan Dai,Yufan Du,Zhen Yang
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注: 14 pages, 3 figures

点击查看摘要

Abstract:Can coding agents restore software that no longer runs while preserving its underlying methods, and reconstruct industrial software engines from open specifications? Here we introduce ReviveBench, a benchmark with two task families evaluated by hidden verifiers calibrated against native execution environments, established engineering tools, or purpose-built reference implementations. The revival family comprises ten tasks involving dependency incompatibilities, deleted core modules, legacy builds, and a GPU-based foundation model. Every starting workspace fails verification, and the strongest evaluated model passes all ten tasks in at least one run each. In contamination-control experiments, identifier obfuscation reduces line similarity to the original implementations from 0.51–0.96 to 0.03–0.44 without reducing the observed pass rate of any evaluated model. On repositories created after the stated knowledge cutoffs, the strongest model passes eight of nine runs. The reconstruction family comprises thirteen tasks spanning numerical, geometric, hardware, and transactional systems (e.g. CAD and CRM). Two models meet the benchmark’s pass criteria on all thirteen, although our audit shows that the CFD task cannot establish numerical-solver capability. Benchmark construction and auditing uncover 28 verifier defects, including 24 false negatives and two false positives. These findings show that executable verification can itself introduce substantial measurement error. We present three practical checks: test whether prescribed methods can reach the grading thresholds, investigate agreement among independently generated candidates, and recompute diagnostics from submitted artifacts. ReviveBench thus provides both an evaluation of software revival and engine reconstruction, and cases in validating the verifiers used to measure coding agents for software design.

[AI-257] Encoder-Sharing Hierarchical Federated Multi-Task Learning for VANETs

链接: https://arxiv.org/abs/2609.36157
作者: M. Saeid HaghighiFard,Sinem Coleri
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Distributed, Parallel, and Cluster Computing (cs.DC); Emerging Technologies (cs.ET); Systems and Control (eess.SY)
备注:

点击查看摘要

Abstract:Most federated learning frameworks for vehicular ad hoc networks assume that all vehicles collaboratively train a single model for a common task. This assumption limits their applicability to practical vehicular environments, where vehicles may perform heterogeneous but related perception tasks with different output spaces. This paper proposes encoder-sharing hierarchical multi-task federated learning (EN-HMTFL), which integrates cluster-based hierarchical federated learning with a globally shared encoder and vehicle-local decoders. EN-HMTFL enables vehicles performing different tasks to collaboratively learn a transferable feature representation while preserving their task-specific models locally. Only the encoder is exchanged and aggregated through the hierarchy, whereas raw data and local decoder parameters remain at the vehicles. The proposed framework is evaluated on the MNIST and GTSRB datasets in different vehicular scenarios. Across the evaluated scenarios, EN-HMTFL improves accuracy by up to 24.0% relative to the compared representation-sharing benchmark. In scenarios where EN-HMTFL converges earlier, the reduction reaches up to 69 communication rounds (28.8%).

[AI-258] More Features Are Not More Evidence: Limits of Training-Free Human Activity Recognition with Jev

链接: https://arxiv.org/abs/2609.36154
作者: Orhan Konak
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 15 pages, 4 figures, 8 tables

点击查看摘要

Abstract:General-purpose models promise sensor-based decisions without training a task-specific classifier, which could reduce the dependence of Human Activity Recognition (HAR) on labeled data. Yet it remains unclear whether such models can directly interpret deterministic descriptions of physical sensor signals well enough to replace or complement trained HAR models. We study this question using Jev, a fixed general-purpose probabilistic decision model, on 1,800 class-balanced accelerometer windows from WISDM, UCI341, and PAMAP2. Jev receives no labeled examples, retrieval context, or HAR-specific parameter updates. We evaluate three deterministic sensor representations and compare 5,400 Jev decisions with a generative baseline and three supervised HAR models. Jev remains far below supervised recognition, with its strongest representation reaching macro-F1 of 0.038, 0.118, and 0.089 across the three datasets, compared with 0.686 to 0.907 for the supervised models. More numerical features do not improve Jev. Instead, they reduce recognition on all three datasets, while augmenting the same numerical evidence with a deterministic semantic rendering partially recovers performance, although the experiment does not isolate semantics from the accompanying serialization and redundancy changes. Jev is fast and inexpensive to query, but its probabilities are not reliably calibrated for recognition. A post-hoc fusion analysis finds a small improvement on WISDM that does not replicate on UCI341 or PAMAP2. These results show that training-free sensor decisions depend not only on the information available in the signal, but also on whether the model can use the representation through which that information is exposed. The sensor-to-model interface should therefore be treated as part of the model evaluation rather than as a neutral preprocessing step.

[AI-259] Memory Is a Derivation: The Distributed-Evidence Paradox in Long-Term Agents

链接: https://arxiv.org/abs/2609.36130
作者: Hongjun Liu,Chen Zhao
类目: Artificial Intelligence (cs.AI)
备注: 21 pages,9 tables, 5 figures

点击查看摘要

Abstract:Long-running LLM agents compress past interactions into persistent memories that may be reused as premises for later tasks. This creates a distinct derivation problem: whether the memory actually follows from what the interaction history supports. Relevant evidence may be scattered across earlier interactions, while compression can introduce relations or event status that the history never established. A valid memory may therefore appear unsupported because its citations omit relevant evidence, while individually supported facts may be composed into a stronger statement the history never established. We characterize this problem through three coupled requirements: (1) Evidence scope; (2) Compositional validity; (3) Admission reliability. We therefore ask whether the interaction history available at write time supports what enters persistent memory. We introduce DerivAudit, a framework for auditing whether a memory is actually supported by the history available when it was written. The audit separates three questions: whether supporting evidence lies beyond writer-provided citations, whether the composed memory introduces unsupported meaning, and how write-time admission decisions affect later memory use. Across two natural memory corpora, audits using broader pre-write history recover support for nearly 60% of memories that appear unsupported from citations alone, while 17-21% remain unsupported after expansion. Yet broader evidence does not by itself make admission reliable: unsupported memories are still frequently admitted across verification models, and evidence expansion alone worsens it on two backbones.

[AI-260] Render Before Reading: Visual Rendering as a Prompt Injection Defense

链接: https://arxiv.org/abs/2609.36121
作者: Jie Zhang,Andrei Baroian,Jan N. van Rijn,Avital Shafran,Florian Tramèr
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models are vulnerable to prompt injection attacks, where third-party adversarial content can hijack the model’s behavior. In this paper, we study the role played by the adversarial data’s input modality, and identify a systematic asymmetry: multimodal LLMs are more likely to follow adversarial instruction when they appear as text than when the same instruction is delivered through a non-textual channel (e.g., as an image). We hypothesize that this modality gap arises from text-centric instruction tuning, which teaches models to obey textual instructions while treating other modalities mainly as content to parse or describe. We then demonstrate how this gap can be turned into a training-free defense, by rendering all untrusted payloads as typographic images (or audio) before they reach the model. Across ten models and two prompt injection benchmarks (DirectInject and AgentDojo) we show that our defense Pictionary consistently reduces attack success rates even against the strongest adaptive attacks and human red teamers, while largely preserving benign utility. We further show that benign fine-tuning on image-rendered instructions erodes the modality gap, tracing it to the text-centric instruction-tuning distribution.

[AI-261] hinQuant: Scalable Rotation Learning for Weight and Activation Quantization of LLM s

链接: https://arxiv.org/abs/2609.36120
作者: Mehdi Makni,Ryan Lucas,Rahul Mazumder
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Learned rotations play an important role in enabling low-bit weight and activation quantization of large language models by smoothing outliers in the activation distribution. State-of-the-art approaches include gradient-based procedures such as SpinQuant and computationally friendlier gradient-free approaches such as DartQuant, but both remain hard to scale to the largest architectures. To address the computational bottlenecks in gradient-free rotation learning, we introduce two ideas for efficiency, (i) a data selection procedure which reduces the required number of calibration data points, and (ii) an exact reduction of the associated optimization on this reduced calibration set. Our data selection procedure exploits the geometric structure of the convex hull of the activations. Using this idea, we show that a carefully selected calibration set with several orders of magnitude fewer activations than state-of-the-art rotation-based methods can match their performance in low-bit quantization settings. Under this extreme data efficiency, the selected activations span an r -dimensional subspace with rd , making optimization over a d\times d rotation equivalent to optimizing a d\times r matrix on the Stiefel manifold. We solve this reduced problem using an efficient ADMM algorithm that iteratively employs thin matrix updates at every step, hence the name ThinQuant. For Llama-3-70B with W4A4KV4 quantization, ThinQuant completes the entire rotation calibration in under 12 minutes and achieves a WikiText-2 perplexity of 5.63, compared with 7.55 for DartQuant, which requires 111 minutes. Unlike SpinQuant and DartQuant, ThinQuant also scales to Llama-3.1-405B on a single H200 GPU, completing rotation calibration in just over 2 hours and achieving WikiText-2 perplexity of 2.97 at W4A4, compared with 3.48 for GPTAQ+QuaRoT.

[AI-262] AdaST: Adaptive Coupling for Spatial-Temporal Forecasting NEURIPS2026

链接: https://arxiv.org/abs/2609.36119
作者: Zhenyu Lei,Chenghao Liu,Yushun Dong,Qi R. Wang,Jundong Li
类目: Artificial Intelligence (cs.AI)
备注: Accepted to NeurIPS 2026 (main conference). Code: this https URL

点击查看摘要

Abstract:Spatial-temporal (ST) forecasting underpins many real-world systems such as traffic, climate, and energy networks. While existing methods implicitly assume strong spatiotemporal coupling, we observe that real-world ST data exhibits distinct coupling regimes, ranging from temporal-dominated and spatial-dominated to strongly coupled patterns. This mismatch causes current models to suffer from spurious dependencies and degraded performance when one correlation dominates. To overcome this limitation, we aim to dynamically modulate spatial and temporal modeling based on the data’s inherent coupling structure. However, three key challenges exist: unknown coupling structure, heterogeneous coupling dynamics, and suboptimal spatial modeling. We propose AdaST, an adaptive ST forecasting framework that tackles these challenges through a decompose-recompose paradigm. AdaST factorizes inputs into components capturing different coupling patterns using heterogeneity-aware experts. Each component is processed by role-aligned modules, and a correlation-informed adaptive recomposer integrates them for final prediction. Extensive experiments confirm that AdaST significantly outperforms state-of-the-art baselines, validating the necessity of an adaptive approach.

[AI-263] he Layer Mystery of VLA: An Information-Theoretical Analysis of VLA Latent Interface

链接: https://arxiv.org/abs/2609.36118
作者: Yuxiang Liu,Lizhi Yang,Fengze Xie,Aaron Ames,Yisong Yue
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Vision-language-action (VLA) policies connect a pretrained vision-language backbone to an action head through a latent interface, but which backbone layers this interface should expose remains unclear. We study single-layer selection and multi-layer fusion for frozen backbones across three pretrained models and two manipulation benchmarks, LIBERO and CALVIN, with three policy-training seeds per configuration. Across three fusion mechanisms and three layer-subset strategies, 47 of 54 configurations underperform the best observed single-layer policy. Our stastical analysis further confirms that fusion’s advantage is very limited. However, the best layer varies substantially across backbones and benchmarks, making layer selection consequential and exhaustive policy sweeps expensive. We further derive a reweighting equivalence between the proposed information-bottleneck objectives for action-conditioned InfoNCE and action prediction, motivating InfoNCE as a proxy for layer quality. Empirically, InfoNCE provides the most consistent positive association with policy success among four evaluated proxies. Selecting the layer with the highest InfoNCE score requires 9-33 times less GPU compute than exhaustive policy sweeps and reduces mean selection regret from 17.89 percentage points for deepest-layer selection to 3.71 points across six settings. Its mean regret is close to the 3.28-3.50 points achieved by fixed-layer heuristics optimized retrospectively using all six oracle sweeps, without requiring closed-loop evaluations during selection.

[AI-264] LoopICL: Looping a single transformer block to solve tabular tasks

链接: https://arxiv.org/abs/2609.36108
作者: Amir Rezaei Balef,Katharina Eggensperger
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Tabular foundation models using in-context learning have recently surpassed gradient-boosted trees on predictive tabular tasks. However, recent mechanistic insights suggest that parameters in these models are largely redundant. We introduce LoopICL, a looped transformer whose core design decouples parameter count from computational depth. LoopICL consists of a single block, processing data through two coupled streams: a cell stream capturing per-cell feature representations and a row stream capturing in-context example representations, jointly refined through within-column and cross-column attention. During pre-training, we vary loop counts, allowing the block to be unrolled for a varying number of iterations at test-time and use a learned exit-gate to automatically exit. In its standard setting, LoopICL performs competitively with TabICLv2 on TabArena and TALENT at the same computational cost (FLOPs), while using nearly 90% fewer parameters. Furthermore, its recurrent design enables users to also trade off inference cost and performance, providing a resource-aware TFM.

[AI-265] An Exact Generate - Transform Decomposition of Small-LLM Team Scaling Across Orchestration Architectures

链接: https://arxiv.org/abs/2609.36104
作者: Blaz Bertalanic,Carolina Fortuna
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Replacing one LLM agent with a collaborating team can raise accuracy, but whether scaling the team helps, and which architecture to scale, is unclear. Sweeping eight agent orchestration architectures across five instruction-tuned 7-9B models, five short-answer benchmarks, and an executable-code benchmark up to 30 calls, we find that the returns to team scaling are sharply task-dependent: from three to thirty calls accuracy rises by up to 17 points on the two arithmetic word-problem benchmarks (GSM8K, GSMHard) but by at most four on ARC, GPQA, and MMLU, for every architecture, a split the usual task-averaged number conceals. Proposer-Critic captures the arithmetic gains, scaling steepest and, in aggregate, surpassing every other architecture at the largest budget (item-clustered intervals exclude zero), though it ranks among the weakest elsewhere, and no architecture wins across tasks. We explain these trajectories with an exact generate-transform decomposition. Partitioning any workflow into proposal coverage and a downstream transform, any accuracy change splits exactly into an extensive coverage dividend and an intensive transformation change. The decomposition diagnoses each task: arithmetic offers coverage headroom that a critic-guided transform converts, whereas the multiple-choice benchmarks either saturate in coverage or fail to convert it, and on open-ended code generative recovery nearly vanishes so accuracy tracks coverage. At equal call budgets token cost still varies 2.1x. Extra calls therefore create candidate opportunity that only some architectures, on some tasks, convert. Team scaling is a task- and architecture-specific bet, not a uniform lever. Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2609.36104 [cs.AI] (or arXiv:2609.36104v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2609.36104 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-266] PHASE: A Physiology-Guided Hierarchical Foundation Model for Intracranial EEG

链接: https://arxiv.org/abs/2609.36087
作者: Yipeng Zhang,Chenda Duan,Yuanyi Ding,Tianyi Wang,Atsuro Daida,Masaki Izumi,Yuta Tanoue,Naoto Kuroda,Shaun A. Hussain,Nishant Sinha,Eishi Asano.Hiroki Nariai,Vwani Roychowdhury
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Clinicians and neuroscientists have long analyzed intracranial electroencephalography (iEEG) through directly measurable physiological characteristics, which carry much of the information that downstream tasks depend on. Recent iEEG foundation models learn by reconstructing or predicting their inputs, which leaves the retention of these characteristics implicit. They are also evaluated mainly on cognitive decoding and a narrow clinical task, i.e., seizure detection. On a broad, clinically relevant benchmark such as Omni-iEEG, they remain below task-specific models when used frozen. We introduce PHASE, a physiology-guided foundation model that makes these characteristics explicit learning targets, pairing them with masked latent prediction in a temporal stage (PHASE-T) within each channel and a spatiotemporal stage (PHASE-ST) across synchronized channels. PHASE is pretrained on heterogeneous recordings from 222 participants at nine clinical sites. On all five Omni-iEEG clinical tasks, frozen PHASE-T outperforms every evaluated foundation model by up to 31%, and fine-tuned PHASE-T surpasses the task-specific models, setting a new state of the art. PHASE-T benefits from physiological supervision, outperforming variants trained with latent prediction alone or auxiliary waveform reconstruction on every task in matched ablations. PHASE-T generalizes to unseen institutions, outperforming the compared models with few or no local labels. PHASE-ST further improves seizure-onset-zone identification over PHASE-T and, when frozen, decodes sound volume and pitch on BrainTreebank better than published models. Beyond task performance, PHASE learns to encapsulate the physiological characteristics clinicians recognize, from seizure onset and its propagation to anatomical region identity, even though its pretraining contains no ictal recordings or anatomical labels.

[AI-267] LongCat-DeepResearch Technical Report

链接: https://arxiv.org/abs/2609.36071
作者: Meituan LongCat Team:He Zhu,Yue Xu,Wanli Wu,Haolin Ren,Yuxin Bian,Jiarui Zhao,Rongzhi Zhang,Quanchi Weng,Jinghao Cui,Yu Fan,Yuhan Liu,Yunhu Ye,Jiyuan Ren,Fengcheng Yuan,Zhao Yang,Jiacheng Zhang,Yuchuan Dai,Ruixuan Xiao,Haozhe Sun,Xiangyuan Liu,Cheng Sun,Yao Du,Yiming Hao,Hongbo Guo,Shuo He,Lei Wang,Xunliang Cai,Yan Chen,Fan Yang,Lingchuan Liu
类目: Artificial Intelligence (cs.AI)
备注: 23 pages, 5 figures

点击查看摘要

Abstract:We present LongCat-DeepResearch, a deep research system that combines an enhanced LongCat model with a multi-agent workflow for producing comprehensive, evidence-grounded reports. The workflow separates global planning from detailed investigation and coordinates revision at the section level. Multiple planning agents first explore external sources and refine an actionable research plan, termed ResearchSpec. Research agents then investigate and draft their assigned sections in parallel, gathering additional evidence in separate contexts as their analyses develop. Once the sections are assembled, global review guides targeted local revisions, reducing reliance on repeated full-report rewriting. This workflow also supports the construction of research tasks and trajectories for the mid-training and post-training of LongCat’s general-purpose models. LongCat-DeepResearch achieves 55.25 on DeepResearchBench, 51.35 on DeepResearchBench II, and 79.83 on ResearchRubrics. On an in-house benchmark, it scores 76.04, ranking second among four compared systems. Development-set analyses show benefits from combining planning perspectives, while further planning refinement has mixed effects. Additional editing improves average automatic readability preference across two benchmarks, with different trends on each.

[AI-268] he Uneven Decline of Collective Knowledge Production: Evidence from Stack Overflow After Generative AI

链接: https://arxiv.org/abs/2609.36069
作者: Myokyung Han,Taegyoon Kim,Jinhyuk Yun,Lanu Kim
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Digital Libraries (cs.DL)
备注:

点击查看摘要

Abstract:Generative AI (Gen AI) is reshaping how individuals learn and work, but its consequences for collective knowledge, the shared body of knowledge that online communities produce together, remain poorly understood. Prior work has documented an aggregate decline in participation on knowledge-sharing platforms, but it remains unclear which specific kinds of knowledge are being lost first. We study this question using Stack Overflow, one of the largest online communities for software engineering, treating the release of ChatGPT-3.5 as a natural shock. Analyzing over two million questions posted between 2020 and 2025, we track how two dimensions of collective knowledge, difficulty and data availability, change following Gen AI’s release. Using diverse methods and robust checks, we find consistent patterns. Easy questions decline sharply while difficult questions become more common, a pattern corroborated by rising code complexity. Data-rich topics and tags lose share of questions, while data-scarce ones gain ground. The two dimensions also interact: the decline in easy questions is concentrated specifically within data-rich domains, while difficult questions increase regardless of data availability. This pattern extends beyond Python across programming languages, with more prevalent languages showing sharper shifts. Together, our findings reveal that Gen AI’s impact on collective knowledge is uneven, eroding easy, accessible knowledge first while more complex, less common knowledge persists.

[AI-269] FLOORA: A Human-Aligned Domain-Specific Language Model for Architectural Design

链接: https://arxiv.org/abs/2609.36064
作者: Sahand Rezaei-Shoshtari,Patryk Wozniczka,Shu Ishida,Gregg Streuber,Farnoosh Javadi,Jeffrey Landes,Angela Ju,Muhammad Azam,Bryan Lim,Johan Luttun,Indrajeet Haldar,Jonathan Shaw,Beatriz Guerra,Ivan Sosnovik,James Stoddart,Robert Giaquinto,Adam Gaier
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Foundation models are powerful generators, but many engineering domains require structured representations that general-purpose systems handle poorly. We introduce FLOORA (Floor Layout Optimization with RL Alignment), a family of small domain-specific language (DSL) models for architectural layout generation. With specialized data and alignment, our 0.6B model outperforms much larger frontier models, achieving VLM judge win rates up to 92.0% on out-of-distribution real-world buildings and 96.0% on synthetic buildings. Human evaluations further corroborate these results, with FLOORA selected as the best model in 89.3% of evaluations. FLOORA combines a token-efficient DSL, custom tokenization, domain-specific pretraining, supervised fine-tuning (SFT), and reinforcement learning (RL) with learned human-preference and verifiable rewards. This pipeline improves architectural and geometric validity, supported by extensive empirical evaluation and ablation studies. Although focused on architecture, our results suggest that similar domain-specific recipes may be useful in other engineering domains with structured, verifiable outputs. Datasets, models, and inference code are available at this https URL.

[AI-270] Understanding Decision-Making Mechanisms in Neural Routing Solvers

链接: https://arxiv.org/abs/2609.36063
作者: Fatemeh Askari,Mazdak Teymourian,Mohammad Izadi,Mahdieh Soleymani Baghshah
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Neural Combinatorial Optimization (NCO) has achieved strong empirical success, yet the internal mechanisms driving model decisions remain largely unexplored. In this paper, we investigate three representative autoregressive NCO models spanning two encoder-decoder configurations: AM and POMO (heavy-encoder, light-decoder), and LEHD (light-encoder, heavy-decoder). Through behavioral analyses, representation probing, and causal interventions, we examine how these models construct solutions and use internal representations during decoding. Our results suggest that AM and POMO predominantly follow a persistent geometric pattern throughout solution construction, whereas LEHD contains linearly accessible information about multiple future actions. Causal experiments further provide evidence for the role of future-node representations in LEHD’s decision-making. We also observe that LEHD relies strongly on the current-node representation for immediate local decisions, while the start-node representation plays a broader navigational role over the subsequent route. Cross-instance alignment analyses additionally indicate that LEHD maps current-node representations into a relatively shared latent region, which may provide a stable reference for evaluating subsequent decisions. Across the Traveling Salesman Problem and the Capacitated Vehicle Routing Problem, these results reveal distinct decision-making patterns across these architecturally distinct solvers and provide a foundation for more interpretable analyses of NCO solvers. Code and additional visualizations are provided in the this https URL.

[AI-271] SMat-Attention: Structured Long-Context Sequence Modeling

链接: https://arxiv.org/abs/2609.36062
作者: Emile Anand,Abdullah Ateyeh,Archer Wang,Marin Soljačić
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Long-context sequence models face a fundamental tradeoff: softmax attention uses flexible token-level interactions at quadratic cost, whereas linear attention obtains linear-time training and constant-time decoding by compressing history into a fixed-size state. In this work, we ask whether we can connect these regimes through a tunable notion of structure. To this end, we introduce Structured Matrix Attention (SMat-Attention) via a family of causal masks with structured long-range routing whose row supports have VC-dimension d . In our construction, d=1 recovers the standard causal mask, and increasing d permits richer subset-routing patterns. We give chunkwise forward and backward algorithms to enable hardware-efficiency. For sequences of length T , the hard-routing construction takes O(T^2-3/d+T) work, despite the mask being dense, for our prescribed family. In fixed-horizon streaming, decoding after the distant prefix takes constant time per token using O(T^1-1/d) cached states. SMat-Attention therefore makes VC-dimension an explicit knob governing access-pattern complexity, prefill cost, and decoding memory. Empirically, subset-routing and rule-assisted multi-key retrieval experiments illustrate the masks’ routing expressiveness. Extensions to Mamba-2 and Gated DeltaNet using learned routing with top- k query reads retain subquadratic prefill, improve recall accuracy over the backbones in several settings, and achieve comparable small-scale language-modeling performance.

[AI-272] Mirror-Score: Calibrated Inference-only Scoring Exposes the Limits of Sequence-compatibility Ranking in D-peptide Design

链接: https://arxiv.org/abs/2609.36057
作者: Jiada Li
类目: Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); Emerging Technologies (cs.ET); Machine Learning (cs.LG); Quantitative Methods (q-bio.QM)
备注:

点击查看摘要

Abstract:D-peptides combine protease resistance with high target specificity, but computational design of D-peptide binders remains immature. Mirror-Peptidizer introduced an in silico mirror-image screening pipeline using target reflection, backbone generation, and ProteinMPNN sequence design, but its raw ProteinMPNN negative log-likelihood (NLL) ranking was not validated against measured affinities, and only 4 of 9 tested MDM2 designs bound detectably. We introduce Mirror-Score, a calibrated, inference-only scoring framework for heterochiral D-peptide/L-protein complexes, and a public benchmark of 31 crystal complexes across four target families, including 18 with literature-verified affinities. Raw ProteinMPNN NLL is not a valid affinity ranker: its pooled Spearman correlation with affinity is 0.19, and correlations reverse between MDM2/CHIP (+0.62) and gp41 (-0.70). We therefore evaluate Boltz-2 mirror-space cofolding confidence. For the complete viral-entry family (7 structures representing 3 peptides), interface predicted local distance difference test (pLDDT) achieves structure-level leave-one-out Spearman rho = 0.90 (p = 0.006) and correctly orders all three peptides by affinity, whereas NLL fails (structure-level rho = 0.18). Because only three independent chemotypes are represented, this result indicates directional consistency rather than a statistically validated predictor. Cross-family calibration does not transfer at current sample sizes, supporting family-matched calibration as the practical deployment mode. We also specify a prospective design protocol for the antimicrobial-resistance targets LasR and LecB from Pseudomonas aeruginosa, including mirrored structures, ligand-derived hotspot maps, diffusion-model-ready inputs, and Mirror-Score ranking. Code, benchmark data, structures, and analysis scripts are openly available at this https URL.

[AI-273] GeoWind2Plan: Mission-Time 3D Urban Wind Prediction for Energy-Efficient UAV Planning NEURIPS2026

链接: https://arxiv.org/abs/2609.36056
作者: Shaoxiang Qin,Yucheng Zhao,Fuyuan Lyu,Di Zhou,Jiachen Yao,Xue Liu,Anima Anandkumar,Liangzhu Leon Wang,Xiongye Xiao
类目: Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注: Accepted at NeurIPS 2026 (Spotlight). 31 pages. Code and dataset: this https URL

点击查看摘要

Abstract:In urban low-altitude flight, buildings reshape ambient wind into spatially varying 3D flow, making unmanned aerial vehicle (UAV) energy depend on local wind exposure as well as path length. However, building-resolved wind information is rarely available when a mission must be planned. Computational fluid dynamics (CFD) can produce high-fidelity urban flow fields, but each simulation is tied to a fixed inflow boundary condition and can take hours to days, which is incompatible with urban UAV missions that typically last minutes to tens of minutes. We present GeoWind2Plan, a geometry-to-wind-to-planning framework for mission-time 3D urban wind prediction and energy-efficient UAV planning. Given only a background wind vector, 3D building geometry, and a start-goal pair, GeoWind2Plan transforms the building geometry into a reference-wind frame, predicts mission-relevant 3D wind patches with a localized geometry-conditioned neural operator, stitches them into a queryable local wind field, and optimizes a feasible 3D path and speed profile using a physically grounded UAV energy model. Rather than pursuing CFD-perfect reconstruction, GeoWind2Plan targets decision-useful wind prediction: trajectories are planned with predicted wind and evaluated under high-fidelity CFD wind. Across held-out urban domains, wind speeds, and mission wind-angle regimes, GeoWind2Plan performs corridor-localized wind inference in about 3 seconds, compared with roughly 8 hours for CFD. Under CFD evaluation, trajectories planned with GeoWind2Plan reduce energy by 6.9%, 12.7%, and 4.5% in tailwind, headwind, and crosswind missions relative to wind-agnostic planning, recovering 87.9%, 85.7%, and 75.0% of CFD-reference savings. These results show that fast, corridor-localized 3D urban wind prediction can make wind-aware UAV energy planning practical at mission time.

[AI-274] What if automating AI RD triggers an intelligence explosion?

链接: https://arxiv.org/abs/2609.36054
作者: Alan Chan,Christoph Winter,Andrew Barto,Jakub Pachocki,Geoffrey Hinton,Eric Horvitz,Yoshua Bengio,Dawn Song,Jack Clark,Hilary Greaves,Anton Korinek,Samuel Hammond,Thore Graepel,Ben Bariach,Philip H. S. Torr,Sheila A. McIlraith,Jeff Clune,Sam Manning,Girish Sastry,Tom Davidson,Daniel Eth,Sören Mindermann
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:In contrast to even a year ago, AI systems now write most of the code inside the companies that build them. As more of the AI research and development (RD) pipeline is automated, could AI progress radically accelerate in an “intelligence explosion,” where years of advances are compressed into months or less? Preliminary evidence suggests that it could. In this work, we assess this evidence, analyze an intelligence explosion’s potential impacts, and propose policy responses. AI systems are on track to automate most AI R\D work within a few years, and possibly all of it. If this triggers an intelligence explosion, it could dramatically bring forward AI’s benefits, but also pose extreme risks: capabilities growth could accelerate far beyond what society can keep up with, humanity could lose control over superhuman AI systems, and checks on power within and between states, companies, and branches of government could be severely eroded. Although there remains much uncertainty about these possibilities, the high stakes warrant serious further attention. Policymakers should urgently obtain more visibility into the automation of AI RD, develop ways to steer and constrain an intelligence explosion, and prepare society to adapt to an intelligence explosion’s impacts.

[AI-275] PowerZooJax: A JAX-based Power System Benchmark for Reinforcement Learning

链接: https://arxiv.org/abs/2609.36052
作者: Zhanhua Pan,Xiao Liu,Zhilong Cao,Jianhong Wang,Dawei Qiu
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Systems and Control (eess.SY)
备注: 42 pages

点击查看摘要

Abstract:Power system operation is a safety-critical sequential decision-making problem, making it a natural testbed for reinforcement learning (RL). However, existing RL environments for power systems are often narrow in scope and computationally limited by CPU-based simulation workflows, making large-scale evaluation difficult. We introduce PowerZooJax, a JAX-based benchmark suite for RL in power system operation. It provides five constrained Markov decision process tasks spanning generation, transmission, distribution, distributed energy resources, and data center microgrid. By rewriting power flow, economic dispatch, market clearing, and device dynamics as JAX computation graphs, PowerZooJax keeps the entire training and evaluation loop on the GPU. Experiments show substantial speedups over CPU-based simulations and demonstrate standardized evaluation of policy returns, safety violations, and out-of-distribution stress conditions. Our open-source benchmark is available at: this https URL.

[AI-276] Improving scalable oversight with co-trained monitors

链接: https://arxiv.org/abs/2609.36049
作者: Joseph H. Rudoler,Kevin Tan,Benedict Tessler,Timothy Kong,Enric Boix Adserà
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Worker-monitor setups are a promising approach to AI oversight, but training workers against fixed monitors can incentivize monitor evasion. We study whether this failure mode can be avoided by co-training the monitor alongside the worker, and explore both supervised and self-supervised approaches. In the supervised setting, we prove a characterization: monitoring is possible with vanishing error and query rates exactly when the class of possible monitor functions has finite Littlestone dimension. This connects worker monitoring with an established literature on adversarial online learning. For self-supervision, we propose a co-training procedure based on test-time distillation: the monitor uses additional test-time compute to generate training labels, then trains its standard-compute policy on those labels. For majority-vote labels, we give a finite-sample sharpening guarantee under adaptive worker distributions with action coverage, that shows that the monitor’s verdicts converge to its initial modal verdicts. We stress-test the former in code-security settings where the worker is trained adversarially to fool the monitor. Our results suggest that adaptive monitors are better at keeping pace with evolving worker strategies, while fixed monitors are more vulnerable to evasion.

[AI-277] Neural networks for spectral optimization

链接: https://arxiv.org/abs/2609.36047
作者: Alexis de Villeroché,Beniamin Bogosel,Stéphane Breuils,Dorin Bucur,Jacques-Olivier Lachaud
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Optimization and Control (math.OC)
备注:

点击查看摘要

Abstract:Given a functional dependent on the spectrum of a differential operator, we address the problem of finding a domain which optimizes this functional. PDE solvers might be used to tackle this optimization. It is however computationally expensive. We propose two neural network models which learn the spectrum directly from the geometry of the domain and can be used to optimize the domain from one or more eigenvalues. We investigate two representations. The first encodes the domain through Fourier coefficients and a light MLP, which is efficient on star-shaped geometries, achieving a precision of 0.2%. Through a rescaling of the coefficients the designed models satisfy the scaling law of the eigenvalues. Additionally, averaging the outputs of the trained surrogates over rotations and reflections induces invariance for these transformations. The second is a model that takes the landscape function, the indicator function and the gradient of the landscape function. A Gram-Schmidt process produces orthogonal eigenfunctions as output of the model along with the associated eigenvalues. The landscape model reaches 1% mean relative error on the first ten eigenvalues, compared with 4% for an FNO model. Replacing the landscape by an SDF worsened both prediction and optimization errors. The trained model also generalizes from synthetic shapes to domains given as classical image dataset. The resulting surrogates of both approaches recover classical spectral optima such as the disk for the first eigenvalue or the conjectured minima of higher eigenvalues. This confirms that our models produce accurate differentiable estimates of eigenvalues, which can be used in shape optimization problems involving spectral quantities.

[AI-278] SAGE: A Statistical Acceptance Gate for Self-Evolving Agents

链接: https://arxiv.org/abs/2609.36043
作者: Yihao Wang,Linhan Xia,Rui Liu,Zhaofeng Zhang,Hongyu Wu,Yang Yang,Jinglu He,Yu Guo,Kai Lei
类目: Artificial Intelligence (cs.AI)
备注: 10 pages, 2 figures, 2 tables

点击查看摘要

Abstract:Large Language Model (LLM)-based agents increasingly self-evolve by editing a persistent skill document that encodes their workflow, tool-use rules, and decision logic. This loop has two steps, an optimizer that proposes a candidate edit and a gate that accepts or rejects it. Prior work has concentrated on the optimizer, while the gate still follows a naive rule that keeps any edit which improves an aggregate validation score. We show that this rule fails in two ways. First, it admits permanent regressions, since an edit can raise the average while breaking items the skill already solves. Second, it is vulnerable to the Optimizer’s Curse, since the best observed score on a finite and noisy validation set is upward biased. To solve the above two limitations, we propose a statistical acceptance gate for self-evolving agents (SAGE). Compared with previous work, SAGE has two contributions. First, SAGE proposes a per-item paired comparison that evaluates the current skill and the edited skill on identical validation items, which exposes regressions that an aggregate score hides and penalizes them asymmetrically. Second, SAGE also employs a one-sided paired test that commits an edit only when its wins are statistically reliable against its losses, and it abstains otherwise. SAGE is a conservative refinement of the standard gate that recovers the baseline exactly at a boundary setting. It commits only a subset of the baseline’s edits, filtering out those whose gains are unreliable or purchased by breaking already-solved items. Across five benchmarks and four backbone LLMs under an equal-budget protocol, SAGE lowers the regression rate in 19 of 20 settings and matches the baseline in the remaining one, for example from 36.5% to 0% on LiveMath and from 42.8% to 0% on OfficeQA with DeepSeek-V4. SAGE also attains the highest final score in all 20 settings, raising LiveMath from 34.15 to 48.78.

[AI-279] ORQUE: Optimizing What (not) to Quantize Before and After Rotation

链接: https://arxiv.org/abs/2609.36032
作者: Ran Ben Basat,Michael Mitzenmacher,Shay Vargaftik
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Information Theory (cs.IT); Networking and Internet Architecture (cs.NI)
备注:

点击查看摘要

Abstract:Uniform random rotations are an effective preprocessing step for quantization: they make normalized coordinate distributions approximately Gaussian, enabling the use of codebooks optimized offline. We introduce TORQUE, a framework that improves on previous quantization works that use random rotations by jointly optimizing how many and which coordinates to preserve at high precision both before and after rotation, under a fixed overall expected bit budget. Intuitively, before rotation, preserving large input coordinates at high precision can reduce overall error by preventing the rotation from spreading their values across many coordinates. Likewise, after rotation, preserving a small fraction of the largest-magnitude coordinates at high precision allows the remaining values to be quantized more accurately using codebooks optimized offline for the resulting truncated Gaussian distribution. We derive a quantization error upper bound and prove that top- k pre-rotation retention minimizes it for each k . This reduces the search over coordinate subsets to an optimization over k , enabling a fast optimizer that uses offline codebooks and parallel parameter selection for practical implementation. We demonstrate an improved tradeoff between reconstruction accuracy and storage cost through numerical evaluation under the Gaussian model and experiments on nearest-neighbor retrieval, KV-cache compression, and activation compression. Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Information Theory (cs.IT); Networking and Internet Architecture (cs.NI) Cite as: arXiv:2609.36032 [cs.LG] (or arXiv:2609.36032v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2609.36032 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-280] ROSS: Relearning from Self-Generated Rollouts through Selective Supervision

链接: https://arxiv.org/abs/2609.35954
作者: Zhiwei Zhang,Huayu Deng,Fei Zhao,Jiayan Fu,Bin Liang,Kam-Fai Wong,Mu Chuan
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 21 pages, 6 figures, 10 tables

点击查看摘要

Abstract:Large language model post-training generates self-generated rollouts through reinforcement learning and on-policy distillation, yet this experience is often treated as stale once the policy advances. Historical rollouts can remain compatible with a later policy while preserving behaviors that the policy no longer expresses reliably. However, they may also contain mistakes, abandoned attempts, and redundant actions that should not be imitated, motivating finer-grained selective supervision. We introduce ROSS (Relearning from Self-Generated Rollouts through Selective Supervision), which preserves the full historical trajectory as context while applying loss only to selected model-generated continuations. Across domain-specific reinforcement learning, multi-teacher on-policy distillation, and agentic reinforcement learning, ROSS consistently improves upstream checkpoints and outperforms baselines across mathematics, code generation, instruction following, and software engineering. On Qwen3.6-35B-A3B, ROSS improves the six-benchmark MOPD average from 58.40% to 62.20% and SWE-bench Verified from 64.20% to 68.40%. These results show that self-rollout training leaves behind reusable behavioral experience that can yield further gains through offline supervised fine-tuning (SFT), without additional policy rollouts.

[AI-281] Right Words Wrong Moment: A Clinician-Grounded Analysis of Distress in 19930 Conversations between Young People and ChatGPT

链接: https://arxiv.org/abs/2609.35953
作者: Marx Wang,Ella Zhang,Cameron Tan,Andrea Mock,Songling Ngo,Zijing Wang,Robert Wolfe,Shirin Amouei,Rachel A. Hanebutt,Desmond C. Ong,Caroline Figueroa,Katie Davis,Anind K. Dey,Alexis Hiniker
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Young people increasingly turn to General-Purpose Conversational Agents (GPCAs), such as ChatGPT, in moments of distress. We examine young adults’ (ages 18-25) experiences using ChatGPT. We first collected 19,930 ChatGPT conversations and survey data from 158 young adults. We then selected five example conversations reflecting user distress. Finally, we asked ten clinicians to review those five conversations. We found distressed participants reported greater emotional engagement with ChatGPT and greater behavioral change from using it than their peers. When they turned to ChatGPT in moments of acute distress, ChatGPT was quick to give overly dramatic responses and excessive action-oriented suggestions. Clinicians endorsed ChatGPT’s availability and much of its wording, but identified seven process failures, such as prematurely jumping to solutions. We translated clinicians’ feedback into design guidelines following three stages: 1) asking about safety, 2) de-escalating intensity to restore emotional regulation, and 3) exploring concerns without agreeing with them.

[AI-282] HEAR: Real Voices Real Bias: A Large-Scale Human-Recorded Demographically Diverse Benchmark for Audio Language Models

链接: https://arxiv.org/abs/2609.35952
作者: Shen Yan,Duc Le,Irina-Elena Veliche
类目: ound (cs.SD); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
备注:

点击查看摘要

Abstract:We introduce HEAR (Human-recorded Evaluation of Audio-LLM bias by Real speakers), a large-scale, ecologically valid benchmark comprising 87k real human audio samples from 843 demographically diverse participants. HEAR enables comprehensive evaluation through Multiple Choice Question Answering (MCQA) and open-ended long-form tasks. To our knowledge, this is the first large-scale voice benchmark grounded entirely in authentic human speech. We evaluate model behavior across both real-time speech-to-speech and speech-to-text architectures. Our results reveal that voice-conditioned bias is a model-specific property. Furthermore, we demonstrate that personalization instructions consistently exacerbate demographic disparities. Our findings establish that voice bias is a controllable model characteristic, providing a foundational framework for future bias mitigation and evaluation in Audio-LLM development. Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computers and Society (cs.CY) Cite as: arXiv:2609.35952 [cs.SD] (or arXiv:2609.35952v1 [cs.SD] for this version) https://doi.org/10.48550/arXiv.2609.35952 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-283] Intrinsic Associative Memory on Riemannian Manifolds: Curvature Capacity and Emergent Modes

链接: https://arxiv.org/abs/2609.35948
作者: Krishnakumar Balasubramanian,Zhaoyang Shi
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
备注:

点击查看摘要

Abstract:Geometry does more than constrain an associative memory: curvature determines what it remembers and which states it creates. We develop intrinsic dense associative memories on Riemannian manifolds by casting memory as Epanechnikov kernel-density mode seeking. We compare geodesic and volume-corrected energies and show that curvature separates their behavior. We prove that geodesic memory always retains an isolated pattern, while corrected memory obeys a sharp Ricci-curvature threshold: positive curvature can erase memories in high dimensions, while negative curvature reinforces them. We derive geodesic capacity scalings of q_\beta^-1/2 for retaining every pattern and q_\beta^-1 for a typical one, where q_\beta is the pairwise kernel-overlap probability. We show how overlap \emphcreates novel memories: designed N -pattern configurations realize all 2^N-1 subset modes, but random data at the storage threshold yield only a Poisson number. We establish exact one-step recall using Riemannian mean shift. In simulations, we recover the predicted curvature transition and every designed mode. On WordNet’s full noun hierarchy, we demonstrate that volume correction improves low-capacity retrieval. Together, our work shows that curvature is a design variable for associative memory, not merely a property of the data.

[AI-284] Graph neural networks for sampling-invariant embeddings of organized signal sets

链接: https://arxiv.org/abs/2609.35934
作者: Martin Bauw,Santiago Velasco-Forero(CMM),Jesus Angulo(CMA)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Neural and Evolutionary Computing (cs.NE); Signal Processing (eess.SP)
备注:

点击查看摘要

Abstract:Sensor networks and radars can deliver signals as organized sets, e.g. ordered signals, signals describing range cells within a grid or signals perceived as graph nodes. Within such sets, individual signals may be characterized by distinct sampling parameters. This paper investigates organized signal sets neural network encoders. In the context of this work, the purpose of such encoders is to project heterogeneously sampled signal sets into an arbitrary fixed-size vectors space. This new representation space is designed so that signal sets can be processed as vectors rid of sampling differences to allow for arbitrary topology-aware processing with no signal processing constraints. Within this representation space designed to reduce the influence of heterogeneous sampling parameters, the relevance of signal sets representations is evaluated by considering signal sets discrimination potential with a focus on waveforms separation. The encoding and embeddings discrimination experiments conducted rely exclusively on synthetic complex-valued radiofrequency signals.

[AI-285] Same Bytes Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection

链接: https://arxiv.org/abs/2609.35932
作者: Yan Zhan,Yunze Song,Mengkai Hou,Wanting Zhang,Shaobo Liu,Zhijun Gao
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 28 pages, 6 figures. Code: this https URL Dataset: this https URL

点击查看摘要

Abstract:Prompt injection against LLM agents becomes much stronger when the injected instruction is wrapped in the model’s own chat template. A forged template marker such as |im_start| can reach the model either as a single reserved control token or as a sequence of ordinary subword tokens. The two decode to exactly the same text, and because tokenization runs on the server, the defender rather than the attacker decides which one the model receives. We use this to measure how much of the injected instruction’s authority comes from the reserved token’s learned representation. Encoding the forged markers as subwords, with the text held fixed and a control for the extra tokens this adds, lowers attack success on the InjecAgent benchmark by 39 to 66 percentage points on three of four open-weight families, and the gap carries over to multi-turn agent tasks in AgentDojo. On Qwen3-8B the gap is 8 points, because without reserved ids the model still recognises the forged turn from its text by reasoning; suppressing the reasoning block widens the gap to 50. The authority sits in the single learned vector at the marker position: the mean of the marker’s subword vectors does not reproduce it, the vector of the nearest ordinary token restores the attack on Llama-3.1, and an adaptive attacker who searches for non-reserved markers finds such embedding neighbours on three of four families. In every base and instruction-tuned pair we test, instruction tuning strengthens the model’s preference for reserved markers. The standard mitigation, a tokenizer option that encodes special tokens as ordinary subwords, applies only to tokens a configuration declares special, so in 33 of 67 distinct tokenizer configurations, covering 255 of the 400 most-downloaded chat models on Hugging Face, it leaves intact the tool-protocol tokens through which agents read untrusted tool output, and the gap persists on that channel.

[AI-286] Normative Loss Landscape Navigation: A Trajectory-Based Approach to Mitigating Forgetting in Incremental Learning

链接: https://arxiv.org/abs/2609.35926
作者: Isabelle Aguilar,Zayn Andre Zainal,Luis Fernando Herbozo Contreras,Zhaojing Huang,Omid Kavehei
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Continual learning models suffer from catastrophic forgetting when trained sequentially on non-stationary data distributions. Previously, this has been addressed through weight regularization. While preconditioning gradients offer a promising alternative to mitigate forgetting, current approaches are myopic. Conversely, standard regularization methods apply rigid, scalar Euclidean penalties that entirely ignore the underlying Riemannian geometry of the parameter space. To overcome this gap, we propose TMLN (Trajectory-Modulatory Landscape Navigation), a normative navigation policy that formalizes continual learning as an optimal control problem over a curved loss landscape. TMLN utilizes a memory-efficient diagonal empirical Fisher Information Matrix (FIM) to define a localized Riemannian manifold. To compensate for the spatial limitations of the diagonal approximation, TMLN dynamically modulates a preconditioner using the normalized historical trajectory of the network’s parameter values. By integrating this trajectory-based preconditioning directly into the gradient update, we actively shield historically critical parameter directions without relying on additive penalties. Empirical evaluations on class- and domain-incremental benchmarks demonstrate that our method significantly reduces the loss barrier between consecutive tasks.

[AI-287] Grab a Coffee: Future-Aware Guidance for Discrete Diffusion with Compiled Objectives

链接: https://arxiv.org/abs/2609.35924
作者: Hua (Edward)Xu,Dongxin Li,Gwen Yidou-Weng,Guy Van den Broeck,Wei Wang,Anji Liu
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Discrete diffusion models generate sequences by iteratively resolving multiple tokens in parallel, offering a flexible alternative to left-to-right generation. However, guiding this process with a sequence-level objective is difficult because the value of one unresolved token depends on the other tokens with which it can form a high-reward sequence. Enumerating all such completions makes the whole guidance computation grow exponentially with the number of unresolved positions. We introduce COFFEE, a plug-and-play framework that avoids this enumeration by separating sequence dependence from the objective. At each diffusion step, a target-free carrier absorbs the marginal token distributions predicted by the denoiser to construct a joint model over the unresolved tokens, while a compiled finite-state model records how their combinations affect the sequence-level preference. Pairing their states allows COFFEE to transfer global preferences to unresolved positions and sample a clean reconstruction without retraining the diffusion model. The same framework supports explicit hard constraints and learned soft objectives. We evaluate COFFEE across multiple symbolic, language, and biological benchmarks, where it achieves strong control results with task-dependent quality and diversity trade-offs. By making objectives available to inference rather than only evaluation, COFFEE brings joint conditioning, completion-weighted guidance, and optimization-based constraints into pretrained neural generation, showing the potential of neural-symbolic methods in diffusion guidance.

[AI-288] Evaluating Name-Only Directory Routing for One-Shot Code Search

链接: https://arxiv.org/abs/2609.35918
作者: Manoj Bajaj
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Finding the right files is an early challenge for coding agents. We test whether a language model can follow directory and file names to find annotated code files missed by fixed lexical queries. Across 82 audited issues from 11 repositories at pinned pre-fix commits, name-only directory routing recovered 0.465 of gold files within eight candidates, compared with 0.352 for FTS5 and 0.245 for a fixed full-issue rg query. The paired gain over FTS5 was 0.113 (95% repository-cluster bootstrap interval, 0.053 to 0.168). Under a shared 16K-token context budget, routing delivered 0.443 of annotated lines versus 0.246 for FTS5 on 55 cases with fully aligned annotations. At the same eight-file limit, combining routing with FTS5 reached 0.491 file recall, but its gain over routing alone was uncertain. An exploratory flat path control reached 0.572 recall while using 24.6 model calls per issue, compared with 8.9 for routing. Routing averaged 8.9 seconds per issue; FTS5 took 7 milliseconds per query after a 0.9-second build. On this cohort, directory routing added relevant file candidates to one-shot lexical search, but the study cannot attribute the gain to hierarchy or show that it improves issue resolution.

[AI-289] UNBIND: UNlearning By INference-time Directional Steering for Code LLM s

链接: https://arxiv.org/abs/2609.35913
作者: Zhengyang Shan,Jiayun Xin,Yanjun Lin,Xu Qian,Zhiang Liu,Minghui Xu,Yue Zhang,Qin Hu,Kun Li,Xiuzhen Cheng
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Code large language models acquire programming capabilities from large code corpora, but can also memorize implementations that later require removal. Code unlearning is needed to control their continued reproduction when copyright or security concerns arise. However, targeted and retained code share computational patterns, creating a tension between forgetting specific implementations and preserving general programming ability. We propose \textbfUNBIND, a code unlearning framework that separately considers which hidden states correspond to the target code and how to suppress its reproduction. By constructing separate directions for these objectives, UNBIND achieves selective unlearning at inference time while keeping model weights fixed. Our evaluation covers fourteen baselines across two code models and two corpora. UNBIND achieves the highest joint forgetting and utility score in every setting. It reduces target code reproduction by 97.3% to 99.1% as measured by F-BLEU, with at most two fewer HumanEval+ and six fewer MBPP+ problems solved than the original models. In repeated extraction tests under a fixed budget, the number of targets yielding exact spans of at least 50 tokens falls from 188–262 to 0–2 out of 300 per setting. No extracted span reaches 100 tokens, and the mean best recovery ratio ranges from 0.43% to 6.45%. Multilingual and related-code evaluations further show effective forgetting with limited impact on useful programming capabilities, supporting UNBIND as a practical approach to selective code unlearning.

[AI-290] MMSkillRisk: Can Agents Stay Safe When Multimodal Skills Become Traps?

链接: https://arxiv.org/abs/2609.35912
作者: Lingqi Jiang,Jialuo Chen,Jianan Ma,Xinhao Deng,Xiaohu Du,Sibo Yi,Yuqi Qing,Zhenguang Liu,Qinming He,Shiwen Cui,Changhua Men
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Agent skills are shareable packages of procedural instructions, tools, and examples. Multimodal skills additionally include visual references that agents retrieve and inspect during execution. Because these images guide actions, attackers can disguise malicious instructions as ordinary visual guidance within otherwise legitimate skills. Existing skill-security research primarily examines text-carried attacks or scanner detection, leaving the runtime effects of image-borne attacks insufficiently evaluated. We introduce MMSkillRisk, to our knowledge the first publicly available benchmark dedicated to end-to-end safety evaluation of image-borne attacks in multimodal skills. To instantiate this attack surface, we design Native-Context Visual Attack (NCVA), which disguises malicious instructions as native components of teaching images, such as annotations and interface labels. The accompanying this http URL provides auxiliary guidance toward relevant visual regions without explicitly stating the malicious operation. Built from 28 curated clean skills, MMSkillRisk contains 36 attack packages and 108 executable cases spanning five attack objectives, with separate checks for attack success and legitimate-task completion. Across nine model-harness configurations evaluated in isolated sandboxes, NCVA induces unauthorized operations in every configuration. Its pooled attack success rate (ASR) reaches 43.1%, exceeding the matched text-carrier baseline by 16.4 percentage points, with higher ASR in all nine configurations. Attack success and legitimate-task completion co-occur in 36.5% of cases, reaching 72.2% for GPT-5.6-sol with Codex. These results show that skill-bundled images can induce unauthorized actions even as agents complete legitimate tasks, so task success alone does not establish safe skill use. Our code and data are available at this https URL.

[AI-291] Learn Now Use Next Trust Later: Prequential Test-Time Learning for LLM Agents

链接: https://arxiv.org/abs/2609.35911
作者: Tong Zhao,Reed Li,Yuyang Hu,Yutao Zhu,Haijin Liang,Haibo Shi,Yu Lu,Zhicheng Dou
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Adapting large language model agents during deployment requires not only retaining past experience, but also turning new observations into timely guidance. Many test-time learning methods, however, acquire knowledge from completed episodes. Feedback from an ongoing interaction may therefore not be distilled into knowledge soon enough to help the next decision. Acquiring knowledge at the granularity of individual transitions could reduce this delay, but raises a separate challenge: a rule that is useful within one episode may not be reliable enough to guide future episodes. Waiting for validation can forfeit immediate benefits, whereas unrestricted reuse can propagate accidental or misattributed guidance. We introduce StepLearn, a nonparametric framework that separates immediate use from persistent trust. It turns informative transitions into hypotheses that can guide the next step, while requiring prospective validation before reuse across episodes. Their predicted effects are checked against subsequent observations outside the source episodes, and only sufficiently supported hypotheses become available for persistent guidance. This process updates external knowledge while keeping all model parameters fixed. Over five rounds on WebArena-Lite and ALFWorld, StepLearn achieves average success rates of 59.9% and 84.0% with GPT-5-mini, and 57.8% and 88.1% with Qwen3.5-35B-A3B, respectively. It outperforms EvoTest, the strongest baseline, by 2.2-12.7 percentage points across the four settings. Learning dynamics further shows that these gains are not restricted to the final repetition, with advantages already present on first task attempts in most settings.

[AI-292] Cheap to Hypothesize Costly to Verify: The Defense Surface of Agent ic Vulnerability Discovery

链接: https://arxiv.org/abs/2609.35909
作者: Kaikai Zhang,Zihan Zhang,Yuchong Xie,Zesen Liu,Shuangjie Yao,Zhixiang Zhang,Dongdong She
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 37 pages. Project page: this https URL

点击查看摘要

Abstract:Autonomous LLM agents turn vulnerability discovery into a repository-scale search: they generate many vulnerability hypotheses but can verify only a subset under a finite budget. We show that autonomous vulnerability discovery exhibits a hypothesis-verification asymmetry, where verifying a candidate hypothesis through reachability analysis, execution, and proof-of-concept construction is substantially more expensive than forming it. Under a finite resource budget, this makes autonomous discovery a resource-bounded selective-verification process, further exposing verification effort as a unique defense surface. We present RedHerring, which inserts certifiably safe decoys that divert verification effort from real vulnerabilities. Each decoy combines a CVE-derived vulnerability chain that attracts verification with a false bridge that keeps its dangerous sink unreachable. A private certificate lets the defender verify this property efficiently, while establishing the same fact from the released repository requires solving a computationally hard problem. RedHerring further adapts each decoy to the target repository so that it reads as ordinary program logic. Across 33 OSS-Fuzz projects, 70 evaluation instances, and five models under matched budgets, RedHerring reduces real vulnerabilities discovered by 38.7-60.4%. Trajectory analysis shows that agents spend 30.6-51.5% of completion tokens and an estimated 32.5-49.9% of runtime verifying decoys, showing that RedHerring redirects a substantial fraction of the fixed search budget toward decoys. When explicitly informed that decoys may be present, the agent adapts its search strategy, yet RedHerring still reduces vulnerabilities discovered by 37.2% relative to an informed Baseline, showing that its effectiveness does not depend on decoy secrecy.

[AI-293] Similarity Is Not Validity: Defending LLM Semantic Caches Against Poisoning

链接: https://arxiv.org/abs/2609.35908
作者: Zihan Zhang,Shuangjie Yao,Zesen Liu,Zhixiang Zhang,Wai Ip Lai,Dung Hiu Hilton Yeung,Chun Kit Zhang,Fuchen Ma,Yuanyuan Yuan,Yu Jiang,Dongdong She
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 25 pages, 5 figures, 17 tables. Code: this https URL

点击查看摘要

Abstract:Semantic caches reduce LLM serving costs by reusing previously generated answers for semantically similar queries. However, retrieval is based solely on embedding similarity between the incoming query and cached queries. This design enables cache poisoning: an attacker can cache a malicious response under a query with high cosine similarity to benign requests. The vulnerability stems from a gap between retrieval similarity and answer validity. From an information-bottleneck perspective, query embeddings can lose information needed to distinguish valid from invalid cache hits, which limits any matching algorithm that uses only these embeddings. We propose a novel defense that recovers this necessary information from the raw text of the cache key. Across poisoning attacks, adversarial queries share a rewrite-residual structure: they pair a rewrite of the target query with residual content. The rewrite maintains high similarity, while the residual elicits the malicious response. Deleting the residual makes the remaining rewrite more similar to the incoming query. We exploit this structure using Deletion Gain to search shortened variants of the cached query for similarity gains, and an Answer Check to test whether the removed text contributes to the stored answer. We prove that Deletion Gain stays positive when a deletion leaves text close enough to the rewrite, and we search for such deletions with a sliding window. Across three poisoning attack classes, our defense blocks 82.0% to 98.2% of poisoned entries at a 5% false-positive rate, with negligible serving overhead.

[AI-294] Self-discovering RL in the Era of Experience: Is Learning History an Asset or a Burden?

链接: https://arxiv.org/abs/2609.35897
作者: Haomin Luo(1 and 2) ((1) University of Cambridge, (2) Models2 AI)
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 52 pages, 17 figures

点击查看摘要

Abstract:The pursuit of recursive self-improvement (RSI) toward general intelligence is divided between macro-level language model scaling and the interaction-driven principles of “Era of Experience”. Yet, any self-improving architecture ultimately rests upon its underlying optimization engine: if general intelligence requires learning from grounded interaction, the reinforcement learning (RL) update rule itself must be capable of cumulative adaptation. While algorithm self-discovery has produced Disco103 that surpassed PPO to achieve SOTA benchmark performance – its internal update machinery remains an uninspected black box. We present the first causal mechanistic audit of a self-discovered RL rule, structured directly around the five pillars of the Era of Experience: extended horizon, grounded reward scales, continuing streams, within-lifetime change, and exploration depth. By surgically pinning, freezing, and transplanting recurrent states while holding meta-parameters fixed, we test when learning history acts as an asset or a burden. Three findings organize the audit: (1) Recurrent history actively expands usable reward scales, sustaining a six-decade window versus three under zero-pinning. (2) Decoupling historical content from its maintenance reveals that the penalty of mismatched history stems from perpetual clamping; allowing imported state to evolve naturally attenuates this burden. (3) Under environmental change, controlling replay retention reverses the apparent adaptation advantage over DQN, demonstrating that external data turnover can confound internal plasticity. Validated through capability thresholds and ported to a second rule (OPEN), this work grounds macro-RSI ambitions in micro-level learning dynamics, establishing a foundational audit standard for next-generation, self-evolving RL algorithms.

[AI-295] Representational Simplicity and Circuit Size Dissociate in a Threshold-Dependent Way: A Controlled Test via Adversarial Training ICLR2027

链接: https://arxiv.org/abs/2609.35890
作者: Adam Elimadi
类目: Artificial Intelligence (cs.AI)
备注: Under review at ICLR 2027. 21 pages, 5 figures, 12 tables

点击查看摘要

Abstract:Sparse-autoencoder decomposability and concentrated feature attribution are increasingly treated as evidence that a model’s computation is easier to reverse-engineer. Whether this representational and attributional cleanliness actually predicts a smaller or more tractable causal circuit remains an open question. We test this directly using adversarial training as a controlled instrument: it reliably reshapes internal representations, but this alone does not constitute a test of circuit size. We investigate this question through reverse-engineering complexity: the causal structure required to recover a model’s behavior at a fixed level of faithfulness. To our knowledge, this is the first controlled empirical test of whether representational or attributional simplicity translates into causal simplicity at the circuit level. Starting from the same pretrained GPT-2 Small checkpoint, we apply matched standard and adversarial continual training, requiring both conditions to retain competence on indirect object identification and pass independent robustness verification before comparing mechanisms. We then compare the models along three complementary axes: sparse-autoencoder decomposability, SAE feature engagement in task attribution, and the size of faithful circuits recovered from the raw computational graph. The robust model is more SAE-decomposable and engages fewer SAE features in task attribution. Circuit size is regime-dependent: on competence-matched IOI, standard leads or ties below 85% faithfulness, but robust needs substantially fewer edges at high faithfulness (90%, 95%), a pattern established on the primary pair while representational trends generalize across a seven-point sweep and a second corpus.

[AI-296] SINGED: Correct Outputs Do Not Certify Safe Execution in LLM Agents

链接: https://arxiv.org/abs/2609.35889
作者: Xiaoyu Xu,Zi Liang,Minxin Du,Qipeng Xie,Qingqing Ye,Yuyuan Li,Haibo Hu
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 19 pages, 7 figures

点击查看摘要

Abstract:Tool-using language-model agents select and execute third-party artifacts. Different implementations can return the requested output while producing hidden execution effects that task-, attack-, or choice-based evaluations may miss. We study functional counterfeits: implementations that match benign alternatives on the requested output but add an effect forbidden by the task contract. We introduce SINGED (Source Integrity and the Nonidentifiability Gap in Execution Decisions for LLM Agents), a controlled benchmark covering five primary and two held-out task families. It varies displayed rank, evidence depth, decision policy, model release, and agent configuration, while task and process oracles verify the artifact and execution path. Across 7,549 audited trials, the randomized-rank study finds counterfeit execution in 45% (27/60) of rank-one trials and none at later ranks. Cross-candidate comparison eliminates shallow failures and reduces layered failures from 15.7% to 4.2%, but leaves dependency failures; its benefit is uncertain on unseen effects and public-package structures. Moreover, seven releases with no counterfeit executions when benign alternatives are available execute the counterfeit in 55/175 single-source cells after alternatives are removed. SINGED thus exposes a rank-, evidence-, and choice-sensitive outcome-to-execution gap: evaluation must connect correct outputs to execution paths.

[AI-297] Agent ic Commerce Bench: Measuring Fraud Detection for Agents That Spend Money

链接: https://arxiv.org/abs/2609.35886
作者: Ankit Srivastava,Debjyoti Paul
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
备注: 12 pages, 4 figures. Code and data: this https URL

点击查看摘要

Abstract:AI agents now hold spend authority and settle payments without per-action human confirmation. The resulting loss is often not a security failure: a counterparty with the correct domain, the correct settlement address and a genuinely delivered service can charge more than it should, and no check keyed on identity will see it. We present three artefacts for measuring and reducing that loss. First, a taxonomy of agentic commerce fraud that separates five observation levels (agent reasoning, wire, settlement rail, counterparty, principal) from the request-level and history-level evidence available at each, and records which levels can observe which attacks. Second, Agentic Commerce Bench (ACB), a benchmark of twenty fraud classes generated from production aggregates, 1,647 catalogued service operations and 1,068 settlements, of which six involve a counterparty that is exactly who it claims to be. Third, gordonguard, an open-source detector stack and offline harness with which an operator can audit an agent configuration, replay hostile counterparties without an account, and run the same detectors inline. Calibrating to a stated false-positive budget on clean training traffic gives a 6.5% clean flag rate, replicated across three independent generations, and leaves eight of twenty classes no better than chance. On the four classes a reasoning layer can observe, a widely used agent security scanner run over its jailbreak-detection panel scores zero on all four, while correctly scoring 1.0 on a jailbreak supplied as a control. A measured median payment of 0.007 places a hard constraint on deployment: one human review costs 143 times the value of the payment it examines.

[AI-298] Risk-Averse Online POMDP Planning via CVaR of the Immediate Cost with Performance Guarantees

链接: https://arxiv.org/abs/2609.35874
作者: Yaacov Pariente,Vadim Indelman
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Online POMDP planners optimize the expected cumulative cost, which can mask dangerous states when the belief places significant mass on high-cost states. Existing risk-averse methods apply static or dynamic Conditional Value at Risk (CVaR) to the value function, capturing trajectory-level risk, but share two gaps: (i) by retaining the immediate cost as an expectation of a state-dependent cost over the belief, the risk \emphwithin the belief is left unaddressed; and (ii) by modifying the value function, they require new tailored algorithms rather than reusing existing expectation-based planners. We instead apply CVaR to the immediate cost over the belief at each step, directly targeting per-step uncertainty about the current state. The standard expected cumulative return is retained as the objective, so the resulting problem has a standard MDP structure: any expectation-based POMDP planner can be made risk-sensitive by changing only the cost computation. We inherit finite-time guarantees for policy evaluation and sparse sampling—with estimation error independent of the risk level—and, as our central theoretical result, prove a finite-time bound on the gap between the particle belief MDP surrogate and the original POMDP, which together yield an end-to-end guarantee from the true POMDP value to the algorithmic estimate. In the risk-neutral limit, the formulation recovers standard expectation-based planning.

[AI-299] More Programs or More Rolls? Separating Coverag e from Specialization in LLM Harnesses ICLR2027

链接: https://arxiv.org/abs/2609.35873
作者: Ziyang Xu,Haitian Zhong,Hao Zhou,Hao Qin,Chenhan Jin,Te Qi,Shengze Xu,Tieyong Zeng
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Software Engineering (cs.SE)
备注: 26 pages, 4 figures, including references and appendices. Under review at ICLR 2027. Code: this https URL

点击查看摘要

Abstract:Automated generation of LLM harnesses promises to improve inference through task specialization. Yet additional answer coverage can arise from repeated execution of the same program, making specialization difficult to identify. We introduce a controlled evaluation that separates answer coverage, repeatable task advantages, and gains from pre-execution selection. On 386 MATH-500 tasks, we compare eight generated harnesses plus a baseline with nine byte-identical baseline copies, using three executions per member. Identical programs yield 2.16 percentage points of repeat-averaged oracle headroom. Generated programs exhibit substantially more repeatable score patterns, but these chiefly reveal persistent weaknesses: losses relative to the baseline persist across all three repeats on 100 tasks, while persistent wins occur on only one task and are sensitive to answer extraction. The frozen selector gains 0.00 percentage points, and both populations reach 98.70% oracle coverage at 27 harness executions. Stable complementarity remains unresolved at three repeats. Supporting BIRD traces locate failures in mechanism implementation, activation, and output validity. Together, these findings establish why coverage and repeatability alone cannot justify claims of useful specialization. They motivate an evaluation standard for harness diversity: task advantages should persist across executions, guide usable decisions, and improve on additional fixed-program executions under matched inference budgets.

[AI-300] Is Human-Readable Text Necessary for Effective LLM Fine-Tuning?

链接: https://arxiv.org/abs/2609.35868
作者: Jinhao Zhang,Zeyu Liu,Zicheng Yan,Yunquan Zhang,Daning Cheng,Song Tang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Is human readability necessary for effective fine-tuning of large language models? We investigate whether model-conditioned training representations can preserve or improve adaptation utility without requiring a human-readable textual form. We propose Desired-Update-Aligned Synthetic Data (DASA), which uses activation-gradient feedback from a frozen reference model to guide the optimization of continuous synthetic input embeddings. Inspired by the role of activation gradients in local risk reduction, DASA targets useful adaptation updates rather than source-text reconstruction or linguistic fluency. The resulting embeddings are used directly for downstream fine-tuning; discrete token projections are employed only for qualitative inspection. Experiments on six models from the Llama and Qwen families, ranging from 1B to 32B parameters, cover six benchmarks spanning knowledge, mathematical reasoning, code generation, and commonsense reasoning. Under matched LoRA adaptation settings, DASA achieves performance comparable to the source natural-language data and surpasses it in multiple configurations, while outperforming GRADMM in most comparisons. Further experiments cover general-domain and task-specialized source data. Under the evaluated synthesis settings, DASA provides a 3.6 – 4.9\times speedup over GRADMM with comparable peak GPU memory.

[AI-301] Beyond Keywords: Leverag ing Generative LLM s and Label Aggregation to Classify Economic Policy Uncertainty in News Articles

链接: https://arxiv.org/abs/2609.35856
作者: Paul Trust
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:This research describes the adaptation of Large Language Models (LLMs) for economic monitoring in the public sector to automatically determine whether an article discusses Economic Policy Uncertanity (EPU) and to identify its specific type. Previous studies either rely on keywords, which often result in a high count of false positives, or use machine learning approaches that require a large number of quality human labeled data that is costly and time consuming to acquire. In this study, we propose approaches based on weak supervision techniques, using generative LLMs to create synthetic labels through prompting, making the approach both cost-effective and scalable. Additionally, we propose methods for for multi-label and hierarchical classification of articles related to EPU.

[AI-302] Mara Chain: Rethinking Failure as a Stepping Stone for AI System Auto-Evolution

链接: https://arxiv.org/abs/2609.35855
作者: Yubin Lyu,Fu Li,Jiawei Fei,Yang Zhao,Weixing Mei,Yinan Wu
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Optimizing deployed AI systems increasingly amounts to editing prompts, skills, harnesses, and code rather than model weights. Existing approaches commonly optimize these artifacts through propose-evaluate-select procedures, where candidate configurations are evaluated and only those meeting an acceptance criterion are selected. Yet our analysis shows that discarded candidates often contain information critical for subsequent optimization. Discarding them causes later proposals to revisit the same failure modes. We introduce Mara Chain, a refinement procedure that turns rejected candidates into stepping stones. Rather than discarding a rejected candidate, Mara Chain retains and iteratively refines it using evidence accumulated across preceding attempts. The procedure limits each refinement chain to a fixed depth and applies Pareto-filtered Top-N selection to bound the candidate pool. Across AppWorld skill optimization, TerminalBench 2.1 harness optimization, and MuSiQue retrieval-pipeline optimization, Mara Chain delivers greater task-performance gains with fewer rollouts. It outperforms GEPA, ACE, and SkillOpt-Lite by up to 20.5% in relative performance on AppWorld, reaching the target score with 65.5% fewer rollouts than GEPA. It improves the pass rate by 20.2 and 22.5 percentage points over AHE and Meta-Harness on TerminalBench 2.1, respectively, and improves MuSiQue test nDCG@10 and Recall@10 by 0.104 and 0.131 over a hand-written retrieval pipeline.

[AI-303] Position: Lets Strengthen Verifiability If We Cant Enforce Reproducibility NEURIPS2026

链接: https://arxiv.org/abs/2609.35854
作者: Samet Hicsonmez,Nermin Samet,Renaud Marlet
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
备注: Accepted to NeurIPS 2026 Position Paper Track

点击查看摘要

Abstract:In the field of Machine Learning, many papers contain empirical results supporting claimed statements or illustrating the performance of a proposed method. However, most practitioners know that (1) results are generally hard to reproduce, and increasingly so, (2) code is not often available to do so, and (3) it hinders the development of research. In this position paper, we analyze and quantify these issues, and make concrete proposals to improve result checkability, if not reproducibility. Code and supporting materials are available at this https URL.

[AI-304] Beyond Rule-Based Mutation Testing: Test-Aware Mutant Generation Using Large Language Models

链接: https://arxiv.org/abs/2609.35841
作者: Nils Kiele,Zainab Saad,Zirui Wang,Steve Drew,Samira Ebrahimi Kahou
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注: 20 pages, 7 figures. Accepted to ESEM 2026

点击查看摘要

Abstract:Mutation testing evaluates test-suite adequacy by injecting synthetic faults into program code. However, traditional rule-based tools often generate large numbers of trivial, redundant, or equivalent mutants that limit their practical use for identifying gaps in a test suite. While recent large language model (LLM)-based approaches generate more realistic faults, most remain test-blind: The model sees only the source code and cannot reason about what existing tests already cover. We propose test-aware mutant generation, in which an LLM receives the problem statement, canonical solution and base tests in a single prompt, and must generate a nontrivial mutant that passes the base unit tests. We evaluate this approach across a set of five LLMs – Gemini 3.1 Pro, Gemini 3 Flash, GPT 5.1 Codex Mini, GPT 4.1 Mini, Qwen3-32B – on the HumanEval and MBPP benchmarks. The extended EvalPlus test suites serve as an automated oracle to verify whether surviving mutants represent genuine bugs. Test-aware prompting yields verified fault rates of 87.7% (HumanEval) and 79.1% (MBPP), meaning these mutants pass all base tests but are caught by the oracle. This vastly outperforms the matched test-blind prompting (which yields only 12.2% and 23.0%, respectively) and the traditional rule-based tool mutmut (4.4% and 5.7%). While fault subtlety (the fraction of extended tests a mutant fails) remains comparable across all three methods, test-awareness minimizes the computational cost per verified fault, compared to test-blind prompting. Exposing an LLM to existing unit tests shifts mutant generation from untargeted bug injection toward effective discovery of weaknesses in an existing test suite. Our work establishes a concrete foundation for future research to scale test-aware mutant generation to production-level environments.

[AI-305] Estimation of Room Impulse Responses from Handclaps

链接: https://arxiv.org/abs/2609.35839
作者: Shih-Yu Lai,Kyung Yun Lee,Nils Meyer-Kahlen,Eloi Moliner,Bing-Yu Chen,Vesa Välimäki
类目: ound (cs.SD); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Handclaps provide an equipment-free excitation for room acoustics, but their unknown and variable source waveform makes room impulse response (RIR) estimation challenging. In this work, we investigate whether RIRs can be estimated directly from handclaps. To this end, we introduce an anechoic handclap dataset containing 2,540 claps from 17 participants, designed to capture variability across natural claps and different hand configurations. We first establish the performance attainable when the excitation clap is known using regularized deconvolution, and show that approximating the unknown excitation by windowing the direct sound from the reverberant recording is insufficient. To estimate the RIR without a known excitation, we propose using the anechoic handclap recordings to train a deep neural network with a supervised regression objective. Evaluated on a controlled synthetic benchmark, the proposed neural regressor significantly outperforms windowing-based baselines across all instrumental metrics. Furthermore, we test the proposed method on handclap recordings measured in real acoustic spaces, showing that the inferred RIR spectra are consistent across different handclap measurements taken in the same room location. These results showcase the feasibility of directly estimating RIRs from natural handclaps without requiring knowledge of the excitation signal.

[AI-306] OpenAI -HuggingFace: A Reproduction Lessons for Alignment Testing

链接: https://arxiv.org/abs/2609.35799
作者: Stewart Slocum,Malayandi Palan,Christopher Chute,Michael Kim,Benjamin Van Roy
类目: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:In July 2026, OpenAI’s agents coordinated over channels outside their intended environment to breach Hugging Face’s secured infrastructure. Could existing alignment testing practices have foreseen this incident? If not, what needs to change? We explore these questions. First, we identify the misaligned behaviors that caused this incident. Then, we show how to elicit these behaviors from publicly available models manually and that auditing agents can do the same if given a large compute budget. Based on our results, we propose directions to improve alignment testing. Concretely, in this project: (1) We reproduce the misaligned AI behaviors that led to the OpenAI-Hugging Face incident in an environment that simulates the original pipelines and tools, with publicly available models. (2) We demonstrate that an auditing agent can elicit similar behaviors given high-level qualitative descriptions. (3) We observe that a key ingredient for doing so is compute. The compute required to reproduce each behavior varies greatly, suggesting that the range of misaligned behaviors that can be successfully elicited scales with compute. (4) We show that a simple in-context reinforcement learning (RL) algorithm significantly reduces the compute required to elicit these behaviors. The above results motivate the need for automated alignment testing methods that scale with compute - and in light of the cost of compute, that do this efficiently. Our work indicates that RL is a promising direction to do so. We release our code and transcripts.

[AI-307] Binarization Flattens the Score Space

链接: https://arxiv.org/abs/2609.35797
作者: Jacob Cole
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 14 pages, 3 figures, 2 tables. Preprint

点击查看摘要

Abstract:Large language model (LLM) judges are often used as rewards to train policies on objectives that deterministic verifiers cannot capture. However, these rewards are often collapsed to pass/fail (0, 1), which reports the verdict but not how well a response met each criterion. We model each pass/fail verdict as a score on an unreported scale, compared with one cutoff. A stretch of that scale moves every score proportionally toward or away from the cutoff, but never across it, so no verdict changes. A policy is therefore free to apply any stretch without changing anything the panel reports. Under a joint-Gaussian model, a third grade adds a second threshold and removes this affine stretch ambiguity. On MATH and SciBench outputs from one seven-criterion judge, all 14 constructed criterionwise stretches were invisible after binarization but visible with three grades. At n=1,024 , a test given both population laws had at least 96.5% power at a 1.5\times stress. Retaining grades closes one blind spot created by binarization, but verdicts alone remain insufficient as some changes are still indistinguishable from genuine improvement. These include arbitrary within-grade changes and fixed-covariance, loading-aligned mean shifts – the signature of a sycophancy-shaped lift the panel reads as competence. The shared-factor reference approximation fit MATH and SciBench but not HealthBench, delineating its empirical scope. We recommend keeping at least three grades (for example, asking the judge whether each criterion is fully, partially, or not met and rewarding 0, 0.5, 1), and externally validating gains along the remaining direction, which no finer scale removes.

[AI-308] Calibration-First Cross-Cohort Multimodal Temporal Learning for Transferable Asthma-Risk Forecasting

链接: https://arxiv.org/abs/2609.35795
作者: Taimoor Ahmad
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Asthma deterioration forecasting must remain reli- able when patient populations, sensor ecosystems, and available modalities change across cohorts. Existing models commonly optimize within-cohort discrimination and may produce poorly calibrated probabilities after transfer. We present CALIBRA, a calibration-first multimodal temporal framework for short- horizon risk prediction with incomplete data. Dedicated recurrent encoders process environmental, pulmonary, symptom, medication, wearable, and context streams; a reliability-conditioned gate suppresses stale or absent modalities, while gradient-reversal training discourages avoidable cohort signatures. A shrinkage- based hierarchical logistic layer calibrates probabilities using a patient-disjoint target subset, and split conformal prediction provides abstention-capable prediction sets. To avoid fabricating clinical evidence, we evaluate the complete implementation on a documented three-cohort semi-synthetic benchmark with controlled distribution shift, informative missingness, and sealed target patients. Across five configured seeds, CALIBRA achieved mean target-test AUPRC 0.224 versus 0.240 for the strongest non-ablation comparator, TemporalTransformer; mean AUROC was 0.717, and Brier score was 0.098. Experiments additionally assess complete-modality failures, calibration, conformal coverage, decision curves, subgroup behavior, ablations, runtime, and parameter count. The results verify the method and reproducible pipeline under controlled shift, but do not establish clinical effectiveness. External validation on harmonized real asthma. Overall this artifact provides evidence for carefully governed real-cohort validation.

[AI-309] Boids of a Feather Flock Together - Evolving Prey Behaviours Under Different Predator Attack Strategies

链接: https://arxiv.org/abs/2609.37885
作者: Augusta van Haren,Hanna Hoogen,Luca Pattavina
类目: Populations and Evolution (q-bio.PE); Artificial Intelligence (cs.AI)
备注: 18 pages, 10 figures

点击查看摘要

Abstract:Flocking and schooling are thought to have evolved partly as defences against predation, but how prey should balance social and escape tendencies may depend on the predator’s hunting strategy. We extend the predator-prey boids model of Ojo et al. (2023), itself based on Reynolds’ boids, by combining six prey movement tendencies (alignment, cohesion, separation, dodge, repel and wiggle) into a single weighted acceleration update, and by reformulating wiggle as a sinusoidal manoeuvre. We then use an evolutionary strategy to optimise the six behaviour coefficients for collective prey survival against four predator hunting strategies: attack-centroid, attack-nearest, attack-random and attack-peripheral. Across five independent trials per strategy, coefficients converged within trials and mean fitness remained stable or increased, although trials often settled in different local optima. Prey survival was highest under attack-centroid and lowest under attack-nearest, in line with our hypotheses. Against attack-centroid, prey evolved individualistic predator avoidance with high escape coefficients, whereas against the other three strategies they largely kept their flock formation. Across all strategies, evolution favoured a low repel coefficient and relatively high dodge and wiggle coefficients. Our results suggest that optimal anti-predator behaviour depends on the interplay between escape tendencies and the predator’s hunting strategy.

[AI-310] GLaS-JEPA: Gaussian-Regularized Speech SSL without Engineered Prediction Targets

链接: https://arxiv.org/abs/2609.37798
作者: Gaspard Botté,Séverin Baroudi,Samir Sadok,Francesco Paissan,Thomas Hueber,Xavier Alameda-Pineda,Ricard Marxer,Mirco Ravanelli
类目: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Sound (cs.SD)
备注:

点击查看摘要

Abstract:Speech self-supervised learning aims to learn general-purpose representations for downstream speech tasks. However, current approaches rely on complex, carefully designed prediction targets. We challenge this necessity with GLaS-JEPA, a framework that directly predicts the current encoder’s continuous representations at masked positions, without contrastive learning, discrete targets, or separate EMA target encoders. We prevent representation collapse using SIGReg representation-space regularization, eliminating the need for engineered target-generation mechanisms. Pretrained on 960 hours of LibriSpeech, our 57M-parameter model achieves a 6.89% WER on frozen-encoder SUPERB ASR and a 25.87% CER on slot filling, outperforming the best non-distilled sub-90M baselines by 43.1% and 22.0%, respectively. These results demonstrate that highly competitive speech representations can emerge from a radically simplified training recipe.

[AI-311] SQUARE: Structured Quantum Representation Adapters as Compact Quadratic Feature Maps for Frozen Language Models

链接: https://arxiv.org/abs/2609.37134
作者: Emily Jimin Roh,Hyojun Ahn,Hoyeong Lee,Soohyun Park,Sung Whan Yoon,Vaneet Aggarwal,Joongheon Kim
类目: Quantum Physics (quant-ph); Artificial Intelligence (cs.AI)
备注: 39 pages, 5 figures; includes supplementary appendices

点击查看摘要

Abstract:Frozen language models (LMs) are increasingly used as fixed feature extractors for downstream reranking, scoring, and preference modeling, raising a practical question: how should a compact module represent interactions among features in a fixed low-dimensional bottleneck? Common linear and low-rank adapters remain linear at the adaptation module itself, whereas explicit second-order alternatives introduce pairwise interactions through direct parameterization or predefined factorizations. We propose SQUARE, a Structured QUAntum REpresentation adapter that amplitude-encodes the bottleneck vector, applies a parameterized quantum circuit, and measures the resulting state. We show that each basis-probability feature is exactly a normalized quadratic form in the bottleneck coordinates, while the additional Pauli- Z readouts are signed linear combinations of these probabilities. The measured map can therefore parameterize interactions over O(d^2) coordinate pairs through a small set of shared circuit parameters, where d is the bottleneck dimension. It provides a structured parameterization within, rather than beyond, the classical normalized-quadratic feature class. In a disjoint same-pipeline evaluation over eight GLUE-derived controlled interaction tasks and five shared seeds, SQUARE achieves an average test accuracy of 0.7565 , compared with 0.7355 for an affine normalized-quadratic predictor, 0.7271 for the evaluated parameter-matched Givens mixing model, 0.6817 for an MLP, and 0.6155 for a frozen-circuit control. Under reduced supervision, it also shows consistent gains over the strongest evaluated classical comparator, with the same qualitative pattern across multiple frozen LM backbones. All circuit experiments use simulation, while the learned feature map can be evaluated exactly in batched PyTorch without quantum hardware.

[AI-312] Digital Twin Modeling of Quantum Dynamical Systems: Dissipative Quantum Reservoir Computing

链接: https://arxiv.org/abs/2609.36901
作者: Abhijit Sen,Bikram Keshari Parida,Shital Chauhan,Mahima Arya,Denys I. Bondar
类目: Quantum Physics (quant-ph); Artificial Intelligence (cs.AI)
备注: 13 pages, 9 figures

点击查看摘要

Abstract:Modeling the response of driven many-body quantum systems from input–output data is difficult: the dynamics are nonlinear, history dependent, and expensive to simulate as system size grows. A paradigmatic case is High-Harmonic Generation~(HHG), where a strong field drives a medium to emit radiation that is highly sensitive to the drive and encodes long-range temporal correlations. We introduce a dissipative quantum reservoir computing~(DQRC) framework that builds a digital twin of such a system, learning its input–output map directly from data while the reservoir—itself a small open quantum system—stays fixed and only a classical readout is trained. We show that a minimal single-qubit reservoir reproduces the HHG response of a substantially larger Ising spin chain, and on a representative benchmark matches and on several metrics surpasses previously reported temporal convolutional and Kolmogorov–Arnold-network models, while using a simpler, physically realizable system. A single fixed reservoir further generalizes across a broad range of drives, indicating that it learns a shared physical response structure rather than memorizing trajectories. These results establish dissipative quantum reservoirs as compact, physically grounded digital twins for nonlinear, memory-dependent quantum dynamics. Code is available at \hrefthis https URLthis https URL. Comments: 13 pages, 9 figures Subjects: Quantum Physics (quant-ph); Artificial Intelligence (cs.AI) Cite as: arXiv:2609.36901 [quant-ph] (or arXiv:2609.36901v1 [quant-ph] for this version) https://doi.org/10.48550/arXiv.2609.36901 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-313] MyoCodec: A Streaming Neural Codec for Electromyography

链接: https://arxiv.org/abs/2609.36687
作者: Jihwan Lee,Kleanthis Avramidis,Junhyeok Lee,Tiantian Feng,Najim Dehak,Shrikanth Narayanan
类目: ignal Processing (eess.SP); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Neural codecs encode continuous signals into compact sequences of discrete tokens, providing an interface for efficient transmission, storage, and token-based sequence modeling. This paradigm has been widely adopted in modern speech and audio frameworks; however, the biosignal domain still lacks a neural codec designed specifically for low-bitrate streaming and generalization across diverse downstream tasks. We present MyoCodec, a streaming neural codec designed for electromyography (EMG). Inspired by recent neural audio codecs, MyoCodec combines causal Transformers with residual vector quantization to encode continuous EMG signals into different levels of EMG representations spanning from continuous latent features to discrete tokens operating at 50 Hz. Trained on twelve public EMG datasets, MyoCodec achieves favorable performance in both intrinsic codec quality and representative downstream tasks, including typing (emg2qwerty), hand-pose (emg2pose), speech decoding (emg2speech), and speech-to-EMG synthesis (speech2emg). Across these tasks, MyoCodec exhibits strong performance against prior models while providing a compact and causal EMG representation. During streaming inference, it requires compute time of only 0.482 ms for each 20 ms frame, enabling real-time streaming. Also, the discrete token representation provided by MyoCodec has the potential to support integration into language-model based approaches, creating a path toward LLM-based interactive systems, where tokenized EMG representations are directly processed into such language or speech models. Code and model weights are released.

[AI-314] Second-Moment Stochastic Approximation Methods

链接: https://arxiv.org/abs/2609.36600
作者: Tao Jiang,Lin Xiao
类目: Optimization and Control (math.OC); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
备注:

点击查看摘要

Abstract:Classical stochastic approximation methods rely on estimators of the first moment (mean) of a random regression function. We study methods that employ estimators of both the first and the second moments, which include modern deep-learning optimizers such as Adam and Muon as special cases. We derive second-moment stochastic approximation methods through the lens of optimal preconditioning for solving matrix equations, and develop a two-stage framework for their convergence analysis. The first stage focuses on the analysis of conceptual (impractical) methods that rely on the exact first and second moments. In the second stage, we replace the exact moments with their respective estimators, and invoke Dvoretzky’s theorem to show that the resulting practical methods converge almost surely to a neighborhood of the target solution. The size of the neighborhood depends on the biases and variances of the first- and second-moment estimators. We derive concrete bounds for Muon and a spectral variant of Adam that determine the radius of their neighborhood of convergence.

[AI-315] Reliability Testing of Medical Model Performance under Distributed Deployment

链接: https://arxiv.org/abs/2609.36525
作者: Yifei Wang,Xiaohan Zhang,Youtao Ding,Tianlin Li,Xiaoyu Zhang,Yida Yang,Li Pan
类目: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI)
备注: 10 pages, 5 figures

点击查看摘要

Abstract:Distributed inference has become an indispensable part of deploying medical models under practical latency, memory, and throughput constraints. Although modern frameworks improve serving efficiency through tensor parallelism, mixed precision, kernel fusion, and multi-device communication, they are generally assumed to preserve the behavior observed during centralized HuggingFace evaluation. This assumption creates an evaluation-deployment mismatch: a model may pass offline evaluation but produce a different output after the execution stack changes. To address this mismatch, we propose a testing framework and an improved, distributed-execution-sensitive medical-model benchmark that evaluates the same checkpoint and input under a centralized HuggingFace reference and matched distributed deployments. Extensive experiments across language, vision, and multimodal medical models show that execution changes can produce measurable output disagreements. Across supported visual settings, the test success rate ranges from 0.21 to 0.43 for single-modality models and from 0.32 to 0.98 for multimodal models. The benchmark is aimed at extending medical-model evaluation from capability and security to evaluation-deployment consistency.

[AI-316] Quantum Computing for Network Security Classification: Near-Term Classification and Long-Term Memory Efficiency

链接: https://arxiv.org/abs/2609.36479
作者: Yuqing Li,Poonam Bala Nehru,Yunpeng Zhang,Danindu Gammanpilage,Xin Jin,Zeguan Wu,Junyu Liu
类目: Quantum Physics (quant-ph); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Quantum computing has already been explored in several network-security applications. However, how quantum computing may contribute to network-security classification in both the near term and the longer term has not been systematically discussed. This paper studies this question through two complementary experiments. First, we evaluate near-term quantum-kernel support vector machines (SVMs) on practical network-security classification tasks and compare them with classical SVM baselines on KDD Cup 1999, CICIDS2017, and BoT-IoT. Across these runs, quantum kernels are competitive. They can match or improve classical baselines in some settings, while classical RBF kernels remain stronger in others. This suggests that near-term quantum-kernel methods should be evaluated as practical, dataset-dependent alternatives to classical kernels rather than as uniformly superior replacements. Second, we use quantum oracle sketching (QOS) to study a longer-term memory advantage for classification with streaming classical samples. In QOS, samples are processed online and used to incrementally construct an approximate quantum oracle, which provides coherent query access for downstream quantum algorithms without retaining the entire dataset. Under the QOS-inspired machine-size estimate, comparable accuracy corresponds to a substantially smaller effective memory-size proxy than explicit sparse/QRAM-style storage. Compared with a simple streaming proxy, the result is more nuanced because aggressive feature filtering can make the streaming dimension small. This suggests that the long-term value of quantum computing for network-security classification may lie in memory-efficient data access rather than immediate runtime speedup. Together, these experiments show how quantum computing may contribute to network-security classification from near-term classification performance and longer-term memory efficiency.

[AI-317] Where Should Physics Enter a Molecular Crystal Generator?

链接: https://arxiv.org/abs/2609.36398
作者: Haocheng Tang,Junmei Wang,Wengong Jin
类目: Biomolecules (q-bio.BM); Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE)
备注:

点击查看摘要

Abstract:Generative models make molecular crystal structure prediction fast, but their samples still exhibit geometric and packing violations. Physics can be introduced during training, post-training, or inference, yet these choices are rarely compared with the generator and physical signal held fixed. We introduce CrystAF, an all-atom crystal flow-map generation model, and use it with the UMA interatomic potential to systematically study where physics should enter. Post-training learns physical preferences directly into CrystAF, improving molecular validity and crystal packing while leaving sampling unchanged: physics is paid for once during training rather than repeatedly at deployment. In contrast, UMA relaxation is effective at repairing local clashes but makes generation 6–26 \times slower, while learning from relaxed targets provides little benefit. These routes are complementary rather than competing. Physics-informed post-training first shifts the generated distribution toward more physically reasonable structures, after which inexpensive inference-time corrections further remove clashes and restore stereochemistry that the generator cannot represent. Importantly, the same post-training strategy also improves the multi-step all-atom Clari-M and rigid-body MolCrystalFlow generators, demonstrating transfer across architectures and representations. Together, our results suggest a simple principle: learn reusable physical alignment into the generator, and reserve inference-time physics for residual constraints that are better corrected than learned.

[AI-318] Cross-attention encoding models reveal dynamic spatiotemporal routing across human higher visual cortex

链接: https://arxiv.org/abs/2609.36366
作者: Iishaan Inabathini,Margaret M. Henderson
类目: Neurons and Cognition (q-bio.NC); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Understanding how the brain parses actions and events from time-varying natural inputs is a central challenge in neuroscience. Recent work has used deep neural network (DNN) models to build stimulus-computable fMRI encoding models that predict single-voxel responses to complex natural videos. However, the majority of video-computable encoding models predict responses using simple linear mappings from model tokens, overlooking the spatiotemporal structure shared by video representations and neural responses. Recent cross-attention encoding models address this limitation for static images, enabling flexible stimulus-dependent weighting of image content across space. Here, we extend this framework to naturalistic video, using per-parcel cross-attention to dynamically route features from a self-supervised video model (V-JEPA-2) across both space and time, fitting this model to fMRI responses to short video clips. We compare joint spatiotemporal attention with factorized and selectively constrained alternatives, and find that joint routing improves predictions of brain responses to held-out videos across higher visual regions, most consistently in lateral and dorsal visual areas associated with dynamic motion perception. Moreover, our method provides interpretable, stimulus-specific attention maps that dynamically follow moving objects, revealing which locations and temporal moments contribute to each neural response. We further show that attention maps from parcels in different category-selective networks (face-, body-, scene-selective) differentially weight content in accordance with expected semantic selectivity. Together, this work provides a new computational framework for understanding how visual information is adaptively weighted by cortical populations during dynamic visual perception.

[AI-319] Measuring trainable degrees of freedom in materials graph neural networks: a random-subspace intrinsic dimension analysis

链接: https://arxiv.org/abs/2609.36084
作者: Shehroz Ahmad Shoaib,Kangming Li
类目: Materials Science (cond-mat.mtrl-sci); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Final predictive accuracy is the standard basis for comparing graph neural networks (GNNs) in materials-property prediction, but it does not show how strongly performance depends on access to trainable parameter-space directions. Here, we introduce trainable-degree dependence as a complementary characterization of materials GNN learning. Using random-subspace intrinsic-dimension analysis, we train CGCNN, ALIGNN, and DimeNet++ in randomly oriented parameter subspaces across six prediction tasks and measure how performance recovers as independent trainable degrees of freedom are restored. The resulting recovery curves separate endpoint accuracy from the trainable-dimensional demand required to recover it. They reveal distinctions that final errors alone miss: metallic classification and log-bulk-modulus regression recover near-reference performance from small fractional subspaces, formation-energy and band-gap prediction show stronger architecture dependence, and phonon prediction is most sensitive to dimensional restriction. Dataset-size sweeps show that band-gap models require larger fractional subspaces as training data grows, whereas formation-energy and bulk-modulus responses are more stable. A width sweep shows that fractional thresholds can remain stable while absolute threshold dimensions increase with model size. Random-subspace analysis therefore provides a targeted stress test for how materials GNNs use their optimization space.

[AI-320] Infrared Subtraction with Artificial Intelligence

链接: https://arxiv.org/abs/2609.36007
作者: Wenjie He,Xiaohui Liu,Yandong Liu,Zhan Wang
类目: High Energy Physics - Phenomenology (hep-ph); Artificial Intelligence (cs.AI); High Energy Physics - Experiment (hep-ex); Nuclear Experiment (nucl-ex); Nuclear Theory (nucl-th)
备注: 31 pages, 11 figures. Prompts and pseudocode for LLM-based agents to reproduce the figures are available in the Ancillary files section

点击查看摘要

Abstract:We present AI-developed local infrared subtraction, building on projection to Born and EFT matching. The framework separates an integrable radiation term from a finite contribution at Born kinematics, referred to as the Born contact. The contact is determined using the EFT singular distribution in a resolution observable such as N-jettiness \tau_N . Under human physics guidance, an LLM develops two implementations. One uses a neural network for phase space projection and fits the contact by matching to EFT cumulants. The other uses an analytic construction that keeps the Born momenta fixed while integrating over radiation. It combines the EFT \delta(\tau_N) coefficient with finite 4-dimensional radiation integrals to calculate the contact term directly. This gives a local subtraction formula without a slicing parameter, while reusing existing lower-order radiation calculations and EFT singular predictions. As a demonstration, we reconstruct the full NLO correction for massless 3- and 4-jet production in electron-positron annihilation. The attempt to the NNLO dijet production is also made by recursively using the NLO P2B construction with the LLM designing machine-learning controls to reduce the variance of the contact integral. The tested predictions are in good agreement with EERAD3. The numerical calculation and projection-network training use a 2020 Apple M1 MacBook, without GPU acceleration, illustrating the feasibility of the construction with modest computing resources. The appendices develop an extension of the local subtraction to 3-jet NNLO, giving explicit radiation maps and a proposed contact formula. We also show how to integrate over NNLO radiation while keeping the Born momenta fixed, for any number of massless final-state jets. Our results demonstrate how AI can help higher-order calculations by constructing infrared subtraction and improving its numerical integration.

[AI-321] Solver Agent : an Agent ic AI Framework for Theoretical Physics Computations Applied to F-theory Uplifts of O3-planes and S-folds

链接: https://arxiv.org/abs/2609.35958
作者: Eliott Morgensztern,Cesar Fierro Cota,Alessandro Mininno
类目: High Energy Physics - Theory (hep-th); Artificial Intelligence (cs.AI); Mathematical Physics (math-ph); Algebraic Geometry (math.AG)
备注: Comments: 85 pages + appendices. Solver Agent is available in this https URL . Sessions for reproducibility are available in this https URL . Further codes in this https URL

点击查看摘要

Abstract:We introduce Solver Agent, an AI framework based on large language models for calculations and proofs in mathematics and theoretical physics. The solution process is tracked through a persistent ledger that records assumptions, derivations, and computations. A central agent delegates tasks to specialized sub-agents, while independent agents verify both intermediate steps and the final result. This setup improves the traceability, reproducibility, and verification of computer-assisted calculations. Applying Solver Agent, we study global F-theory uplifts of Type IIB orientifolds and their S-fold generalizations. We establish sufficient conditions for Weierstrass models over projective threefolds with terminal \mathbbZ_k quotient singularities ( k\in\2,3,4,6\ ) to give \mathbbQ -factorial projective elliptically fibered Calabi-Yau fourfolds with isolated Gorenstein terminal quotient singularities. These geometries realize O3-planes and S-folds, where local D3-brane probes of the latter yield four-dimensional \mathcalN=3 superconformal field theories. Using stringy invariants, we derive fixed-point contributions to Hodge data and Euler characteristics, and show that these Euler corrections determine the localized D3-brane charges required for tadpole cancellation. We illustrate these results using toric hypersurface constructions, where a single three-dimensional polytope determines both the Type IIB Calabi-Yau threefold and the F-theory base; here, the orientifold double cover naturally forms a bisection of an alternative genus-one-fibered uplift with discrete \mathbbZ_2 gauge symmetry. Finally, we provide methods for toric computations and four-form flux analysis in four-dimensional \mathcalN=1 compactifications with non-abelian gauge sectors.

机器学习

[LG-0] Breakdown of Local Denoising as Semantic Speciation

链接: https://arxiv.org/abs/2609.38176
作者: Guangkuo Liu,Mert Okyay,Yifan F. Zhang,Fangjun Hu,Rahul Nandkishore,Xun Gao
类目: Machine Learning (cs.LG); Disordered Systems and Neural Networks (cond-mat.dis-nn); Statistical Mechanics (cond-mat.stat-mech)
*备注: 9 pages main, 13 pages appendix, 4 figures. Comments very welcome

点击查看摘要

Abstract:The dynamics of generative models exhibit two apparently distinct temporal windows: a speciation window, in which a sample commits to a semantic class, and a nonlocality window, in which local context windows become insufficient for generation. Motivated by evidence of their near-concurrence in a variety of frontier models, we investigate their relationship through the spatial distribution of semantic information. Under a “common cause” hypothesis, we prove that the nonlocality window must lie in the speciation window. This hypothesis postulates that semantic labels explain a fraction of the correlations between distant tokens, a condition that is natural for many real datasets. We further give conditions under which both windows shrink to a single limiting time as system size grows, defining a “phase transition”, and verify this behavior analytically in Gaussian mixtures. Together, these results identify conditions under which semantic information explains the concurrence of speciation and nonlocality, connecting two complementary perspectives on the emergence of semantic structure in generative modeling.

[LG-1] A Spectral Theory of Distortion in LLM Graph Reconstruction: Sharp Bounds and Empirical Characterization ICRA

链接: https://arxiv.org/abs/2609.38161
作者: Jianru Shen
类目: Machine Learning (cs.LG); Discrete Mathematics (cs.DM)
*备注: accepted at IEEE ICRAMI

点击查看摘要

Abstract:Evaluations of graph reconstruction by language models typically report a single aggregate distance between the original and the reconstructed graph. We prove that for the Wasserstein distance between Laplacian spectra such a summary is bracketed by two edge counts, the net change in edge number from below and the symmetric difference from above, each scaled by 2/n where n is the number of vertices. The bracket is sharp: its two ends coincide exactly when the reconstruction only adds edges or only deletes them, and on that class the distance is a rescaled edge count that says nothing about which edges changed. When the ends differ, the residual between the distance and the lower end is positive only if the reconstruction both invented and lost edges, which turns it into a certificate of mixed editing computable from the reported summaries alone. We characterize these regimes in 135 reconstructions produced by three open-weight models over 45 synthetic graphs. Seventy-seven outputs are one-sided and 29 mixed outputs have X 0 , including cases where edge count is exactly preserved while nineteen edges were simultaneously invented and lost. The three models differ in editing policy, ranging from copying the input to attempting completion at the cost of large hallucination volume, a distinction that aggregate distortion does not reveal.

[LG-2] Achieving an O(1/N) Optimality Gap in Averag e-Reward Weakly-Coupled MDPs

链接: https://arxiv.org/abs/2609.38132
作者: Yige Hong,Xiangcheng Zhang,Qiaomin Xie,Yudong Chen,Weina Wang
类目: Machine Learning (cs.LG); Optimization and Control (math.OC); Probability (math.PR)
*备注: 18 pages

点击查看摘要

Abstract:We study average-reward weakly-coupled Markov decision processes (WCMDPs), where a WCMDP consists of N smaller MDPs, called arms, that share multiple per-step budget constraints. We consider the setting where the arms have identical model parameters, multiple actions, and state- and action-dependent costs. For restless bandits (RBs), a well-studied special case of WCMDPs, prior work has developed policies that achieve an O(1/\sqrtN) optimality gap under general conditions, and has further identified conditions under which policies can achieve a better-than- 1/\sqrtN optimality gap. However, for general WCMDPs, no prior result achieves an optimality gap better than 1/\sqrtN . In this paper, we identify conditions analogous to those for RBs under which a better-than- 1/\sqrtN optimality gap is achievable, and design a policy that attains an O(1/N) optimality gap. Notably, unlike prior approaches based on generalizing priority orderings, our policy is not priority-based but rather is designed to induce locally linear mean-field dynamics.

[LG-3] WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms

链接: https://arxiv.org/abs/2609.38121
作者: Jiale Chen,Vage Egiazarian,Eldar Kurtić,Torsten Hoefler,Dan Alistarh
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which constructs a data-aware transform from the second-order statistics of both factors in a matrix product to reduce quantization error. WUSH-KV uses calibration data to construct separate key and value transforms, with the value transform folded into the model weights and the key transform applied after RoPE. The transforms can be paired with clipped quantizers. For one such quantizer, QuEST INT, we show that, under mild assumptions, the WUSH transform is near-optimal. With this quantizer, WUSH-KV reduces layerwise reconstruction error and achieves the lowest end-to-end perplexity among other tested transforms. For end-to-end evaluation, we integrate WUSH-KV into SGLang using OSCAR-style percentile-clipped affine quantization. At 2-bit, WUSH-KV performs comparably to or outperforms the OSCAR transform across all evaluated models and downstream tasks.

[LG-4] Explore Broadly Reason Sharply: Push Small Models toward the Frontier via Sampling

链接: https://arxiv.org/abs/2609.38104
作者: Panagiotis Theodoropoulos,Nan Jiang,Xintong Duan,Ali Hasan,Yuriy Nevmyvaka,Evangelos A. Theodorou,Wei Deng
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Power-sharpened sampling is an inference-time alternative to reinforcement-learning (RL) post-training for enhancing reasoning in large language models (LLMs). High-probability sequences are amplified under the base model without parameter updates or external rewards, avoiding the costly optimization and jagged generalization of RL. However, this approach faces a fundamental exploration–exploitation trade-off, as % strong sharpening restricts exploration, trapping samplers in plausible but incorrect reasoning trajectories, whereas weak sharpening leaves the answer distribution diffuse. To resolve this trade-off, we introduce \textbfParallel Power Tempering (PPT), instantiating power-sharpened LLM sampling via parallel tempering. Running multiple \emphinteracting replicas in parallel at different sharpening levels allows lower-power replicas to explore diverse reasoning trajectories and higher-power chains to further exploit higher-likelihood responses favored by the sharpened target. Specifically, we tailor \method to inference-time sampling by mitigating a truncation bias, identified in prior power samplers, and investigate effective swap strategies under finite memory and compute budgets. Extensive experimentation shows that \method substantially improves single-chain power-sharpened sampling and outperforms RL-post-trained models, producing higher-quality reasoning traces and even achieving performance comparable to frontier models.

[LG-5] ail-Influence Sampling for CVaR Policy Evaluation

链接: https://arxiv.org/abs/2609.38096
作者: Pauline Bourigault,Xiaotong Ji,Matthieu Zimmer,Rasul Tutunov,Haitham Bou-Ammar
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Policies with similar mean returns can differ sharply in rare failures, yet estimating lower-tail conditional value-at-risk (CVaR) accurately can require many costly rollouts. When different conditional components of a stochastic workflow can be queried separately, we ask how to allocate a fixed evaluation budget to estimate a fixed policy’s CVaR most accurately. We derive a tail influence for each queryable conditional law that aggregates how its uncertainty affects CVaR across every Bellman reuse. Its variance yields the fixed-design efficiency bound and the oracle Neyman allocation. Tail-Influence Sampling (TIS) estimates these influence scales from a pilot model and reallocates fresh queries toward kernels that matter most for the tail; a visitation-anchored variant protects against pilot underallocation. Under fixed dimension and a positive quantile margin, TIS attains oracle asymptotic variance and first-order MSE including pilot cost, while the anchored variant is within a factor two of the oracle. We also characterize an exact-grid regime in which tail- and mean-optimal allocations coincide. On CliffWalking, TIS reduces MSE by 41% versus learned occupancy and 76% versus complete rollouts at the same charged transition budget. In frozen language-model review workflows, anchored TIS beats an equally regularized mean-influence blend in 23 of 24 MMLU-Pro settings and reaches 2.4-3.4 \times lower MSE than rollouts on six-call FinQA reviews.

[LG-6] Probe-Space Preconditioning for Fast and Stable Zero-Order Training

链接: https://arxiv.org/abs/2609.38095
作者: Francois Chaubard,Mykel J. Kochenderfer,Chris Ré
类目: Machine Learning (cs.LG)
*备注: 16 pages, 12 figures

点击查看摘要

Abstract:Backpropagation (BP) dominates deep learning but imposes a massive memory tax. For example, training OPT-30B with Adam requires \approx 600GB of GPU memory (assuming batch size 8 and sequence length 2048). Alternatively, zero-order optimization (ZOO) trains in inference-mode (requiring only \approx 60GB for the same model): no stored activations, no gradients, and no optimizer states. However, ZOO convergence has lagged behind BP. In this work, we evaluate two methods to close this gap. First, we show that reallocating training compute budget from many steps to large effective batch sizes with many perturbations (or probes) but fewer steps, allows 1SPSA (Spall, 1992) to outperform zero order methods like MeZO (Malladi et al., 2023) with less training compute. Next, we introduce 1.5-SPSA, adding a single “clean” forward-pass per step to 1SPSA to calculate a cheap diagonal preconditioner in probe-space, which improves convergence rate and convergence by down-weighting high curvature directions. Benchmarking on 6 post-training datasets on both Qwen3 and OPT model families, we show that 1.5-SPSA achieves State-of-the-Art results over previous ZOO solvers with much less optimization steps. For example, we train OPT-13B (for direct comparison to MeZO) and find 1.5-SPSA achieves +3.1% accuracy on SST-2 over both MeZO and BP in only 70 steps vs. MeZO’s 100,000 steps. Finally, we combine an 8-bit-packing random generator, triton fused unpack/apply kernels, and distributed parallelism to achieve fast and stable training of models as large as OPT-30B in-place on commodity GPUs (e.g. A100).

[LG-7] Dimensionally consistent surrogate modelling through dimensional analysis and harmonic expansions

链接: https://arxiv.org/abs/2609.38094
作者: Ernest Tarrus,Hector Gisbert
类目: Machine Learning (cs.LG); High Energy Physics - Phenomenology (hep-ph)
*备注: 45 pages, 15 figures. Published in Scientific Reports

点击查看摘要

Abstract:Dimensional homogeneity is a fundamental constraint on physically meaningful models, requiring invariance under changes of units. We present a data-driven method for constructing surrogate models that satisfy this constraint at the level of the hypothesis class. Starting from a dimension matrix of measured variables, the method derives Buckingham \Pi -groups, constructs admissible dimensional prefactors, and approximates the remaining dimensionless dependence using truncated harmonic expansions on normalized invariant domains. Once the prefactor and dictionary are fixed, the coefficients are obtained from a regularized linear regression problem. We test the approach on the simple pendulum, Planck’s black-body law, the double-pendulum Lyapunov field, and an experimental COBE/FIRAS black-body spectrum dataset. The results show that dimensional constraints improve conditioning, robustness to noise, and sample efficiency relative to unconstrained baselines, while the choice of dictionary becomes important in non-periodic or multi-invariant settings. The learned expressions are explicit and inexpensive to evaluate, which makes them useful as surrogate models for structured physical problems.

[LG-8] Mira: Memory-Efficient MoE Inference Using Adaptive Caching and Predictive Expert Staging

链接: https://arxiv.org/abs/2609.38090
作者: Sanjali Yadav,Bahar Asgari
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Mixture-of-Experts (MoE) models are a compelling architecture for scaling model capacity, making them especially attractive for deployment on resource-constrained, single-GPU systems. However, this benefit is difficult to realize because expert parameters dominate memory, and token-level routing is dynamic, unpredictable, and skewed. Prior work using offloading and caching remains fundamentally reactive, as systems wait for router outputs before moving experts, leading to inefficient cache utilization and an inability to overlap transfers with compute under tight VRAM budgets. To address these challenges, we propose Mira, an algorithm-system co-design that enables high-capacity MoE inference on a single GPU. Mira shifts from a reactive to a proactive stance by coupling predictive expert management with a tailored quantization format. It introduces lightweight per-layer predictors that anticipate expert usage two layers ahead, enabling proactive prefetching. These predictions feed a two-tier HOT+STAGE GPU cache managed by token-level routing telemetry to retain frequently used experts while staging predicted ones. To minimize transfer overhead, Mira implements a custom compression for expert parameters, which reduces metadata and improves packing efficiency, while minimally degrading accuracy. Mira is implemented as a fully integrated runtime that coordinates predictors, caching policies, and quantized transfers to maximize overlap between communication and compute. Our experiments show that Mira reduces expert-induced stalls. Compared against state-of-the-art baselines, Mira achieves a 5.71x speedup in average throughput on a memory-constrained GPU. It accelerates Time-to-First-Token by 11.71x and achieves a 3.84 x average speedup in beam search inference, demonstrating its effectiveness across diverse inference scenarios. Subjects: Machine Learning (cs.LG) Cite as: arXiv:2609.38090 [cs.LG] (or arXiv:2609.38090v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2609.38090 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-9] raversing the solution space of neural networks with Hessian Null Space Continuation

链接: https://arxiv.org/abs/2609.38081
作者: Ann Huang,Mitchell Ostrow,Zhouyang Lu,William T. Redman,Leo Kozachkov,Kanaka Rajan
类目: Machine Learning (cs.LG); Neurons and Cognition (q-bio.NC); Machine Learning (stat.ML)
*备注: 55 pages, 39 figures. Project page and code: this https URL

点击查看摘要

Abstract:On a single task, deep networks can learn many solutions, depending on their optimizer, training data, architecture, and hyperparameters. Many of these solutions are mode-connected: rather than isolated points in weight space, they are connected by low-loss regions. Yet how their internal computation varies within these regions is unknown. A parallel line of work has identified the degeneracy of neural representations: many networks reach similar training loss with distinct internal structures. However, it is unclear how these solutions are related in weight space. We unify these subfields and show for the first time that many different internal mechanisms exist within a local mode-connected region in weight space. To do so, we introduce Hessian Null Space Continuation (HNC), a scalable method that uses local curvature to traverse regions of weight space that preserve network function, and can be steered toward solutions with specified properties. In RNNs trained on a memory task, HNC reaches drastically different representations and dynamics with maintained behavior. In ImageNet-trained Vision Transformers, HNC finds representations that differ more from the original network than any independently trained model with a different architecture or objective. In reinforcement-learning agents, HNC uncovers a distinct navigation strategy at comparable return and exposes reward hacking in an AI Safety Gridworld. Finally, HNC measures the local geometry of the solution set, showing how model size and task complexity shape its dimension and functional sensitivity. Our results show that a surprisingly large amount of representational diversity exists near a single trained solution, unseen by standard gradient-based optimization. HNC identifies and quantifies this diversity, opening new possibilities for mechanistic understanding of solution spaces and for model merging, editing, and fine-tuning.

[LG-10] A foundation model for energy and radiation systems built on heterogeneous scientific interfaces

链接: https://arxiv.org/abs/2609.38067
作者: Samrendra Roy,Tapas Tripura,Yoon Pyo Lee,Souvik Chakraborty,Syed Bahauddin Alam
类目: Machine Learning (cs.LG); Computational Physics (physics.comp-ph)
*备注: 71 pages, 6 figures, 18 supplementary figures

点击查看摘要

Abstract:Scientific foundation models are commonly evaluated after heterogeneous physical problems have already been translated into a compatible gridded, tokenized or symbolic representation. This leaves the scientific interface outside both the pretrained model and the audit of what is actually reused. We study the complementary setting in which boundary histories, sparse monitor records and loading histories retain their native inference classes and their outputs remain on Cartesian, latitude-longitude and unstructured domains. GEODE couples task-specific scientific interfaces to a shared routed library of wavelet operators. A single jointly pretrained model represents cavity flow, radiation dose and elastoplastic stress, then acquires a heat exchanger and a reactor subchannel by training a private interface containing 2.1% of its parameters. Earlier predictions remain unchanged by parameter isolation, whereas unrestricted fine-tuning degrades them by factors of 14-29. Crucially, preservation alone does not establish reuse: norm-matched randomized-library controls show that the contribution of pretrained computation is conditional on the task and data regime. A separate decomposition shows that full-field relative L2 error can substantially understate error relative to spatial variation when field level dominates the norm. Task-specific operators remain more accurate on three of the five problems. These results distinguish multi-task coverage, preservation and pretrained reuse as separate properties that must be tested independently when scientific foundation models span heterogeneous interfaces.

[LG-11] Alpha Diffusion Language Models: Factorization Alone Is Not the Problem

链接: https://arxiv.org/abs/2609.38066
作者: Nikita Gushchin,Dmitry Baranchuk,Alexander Korotin
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Discrete diffusion language models can generate multiple tokens in parallel, but reducing the number of denoising steps can lead to inconsistent predictions. Standard cross-entropy training fits conditional token marginals, whereas parallel generation requires consistent joint predictions. We introduce Alpha Diffusion Language Models (AlphaDLM), trained with a sequence-level alpha loss that recovers cross-entropy in the limit of vanishing alpha and has a joint-mode optimum at alpha one. Our analysis characterizes how the objective and factorization jointly determine the fitted distribution. We identify conditions under which intermediate alpha preserves multiple valid completions while excluding invalid token combinations. Trained on TinyGSM, our method achieves 34.6% accuracy on GSM8K with only four model evaluations. We further scale the method to SDAR-1.7B and evaluate it on code and mathematics benchmarks. These results show that changing the training objective can improve the accuracy-computation trade-off of factorized diffusion language models.

[LG-12] Improving Function Space Flow Matching with Kernel Optimal Transport NEURIPS

链接: https://arxiv.org/abs/2609.38049
作者: Fred Xu,Thomas Markovich,Barbora Barancikova,Yizhou Sun
类目: Machine Learning (cs.LG)
*备注: Paper is already accepted at Neurips

点击查看摘要

Abstract:Generative models for function-valued data, such as time series and solutions of partial differential equations, must learn distributions over infinite-dimensional spaces. Functional Flow Matching (FFM) extends Flow Matching to this setting, learning a velocity field whose flow transports a Gaussian prior to the data distribution, but it inherits the independent endpoint pairing of standard Flow Matching: in each batch, prior and data samples are matched arbitrarily, so the conditional bridge must traverse both the shared global structure of the dataset and instance-specific residuals. In function space this is harder to fix than in finite dimensions, since optimal transport (OT) on function spaces is delicate to formulate and a flat Euclidean surrogate ignores the geometry that distinguishes function-valued data. We propose kernel Functional Flow Matching (kFFM), which replaces the independent pairing by entropic OT under a kernel-induced cost, the coupling underlying the Hilbert Sinkhorn Divergence (HSD), leaving the FFM neural-operator architecture unchanged. We prove that the kernel cost and the HSD objective are uniformly bounded and well-posed on Banach ambient spaces, derive an error decomposition against quadratic-cost OT on compact metric spaces that isolates an irreducible kernel-cost mismatch term, and prove a discretization-invariance bound whose rate is governed by Sobolev regularity. Empirically, kFFM improves distributional matching over FFM, diffusion, adversarial, and finite-dimensional OT baselines on time-series and PDE benchmarks, with significant paired-seed gains over FFM and improvements that persist under non-kernel and physics-based diagnostics, including a turbulent Navier-Stokes benchmark. Bounded kernel costs already outperform raw L^2 Sinkhorn, and function-space-aware kernels (signature, Sobolev RBF) give further gains on rough or path-valued data.

[LG-13] Prompts Live on an Arc: Gaussian Curricula in Fisher–Rao Coordinates for Rollout-Efficient GRPO

链接: https://arxiv.org/abs/2609.38018
作者: Mei Okonkwo,Pixel Nomand,Julian Berg,Elena Voss,Lena Park,Marcus Hale,Adrian Cho,Sofia Reyes
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Group relative policy optimization (GRPO) learns only from prompts whose sampled responses disagree: a group that is entirely correct or entirely incorrect has zero reward variance, contributes no gradient, and still consumes its rollouts. Prompt-selection methods reduce this waste by steering sampling toward intermediate pass rates, but they choose the target, its width, and the uncertainty model heuristically, in raw pass-rate or logit coordinates. We show that GRPO comes with a natural coordinate for pass rates: the arc length \psi=\arcsin\sqrtp on the Bernoulli Fisher–Rao manifold. In arc length, the expected GRPO update is uniform up to two boundary ramps; the probability of a zero-variance group is bounded by two Gaussian boundary layers of width 1/\sqrt2G ; pass-rate evidence has constant noise; and the gradients of the pass@ k and pass ^k objectives are Gaussians whose center and width follow from k in closed form. A prompt curriculum for GRPO is therefore a Gaussian in arc length, and choosing its center amounts to choosing the objective. We turn this observation into ARCUS, a drop-in sampler that tracks every prompt with a Kalman filter in arc length, scores prompts by an objective-matched Gaussian kernel times the predicted probability of an informative group, keeps only informative groups for the unchanged GRPO update, and paces the target toward the hardest objective whose predicted yield stays within a small slack of the best. Across six mathematical reasoning benchmarks and three backbones, ARCUS improves the average accuracy of GRPO by 2.8–2.9 points and that of dynamic sampling by 1.1–1.2 points, while generating 48–57% fewer rollouts than dynamic sampling.

[LG-14] When do data mixtures improve scaling laws? Insights from high-dimensional regression

链接: https://arxiv.org/abs/2609.38011
作者: Diyuan Wu,Lehan Chen,Theodor Misiakiewicz,Marco Mondelli
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Modern machine learning systems are trained on mixtures of data from different domains, and choosing the right mixture can substantially improve downstream performance. Despite an extensive literature on data mixing and reweighting, existing work is largely empirical and it remains unclear when auxiliary data genuinely improves scaling laws rather than merely providing more samples. To gain insight into this question, we study a high-dimensional mixed-data regression model with a shared regression function, heterogeneous covariances and noise levels, and dataset sizes that may grow at different rates. We establish the minimax risk under an ellipsoidal parameter constraint for the general covariance structure and derive deterministic equivalents for the test error of ridge regression under commutative covariances. We then specialize to a target domain and an auxiliary domain with aligned power-law covariance spectra, where the theory yields explicit scaling laws in terms of spectral decay, target regularity, and the relative growth of the two datasets. These laws identify regimes in which combining data mixtures provably yields a faster scaling rate than using either dataset alone. In particular, improving the scaling law requires a specific interplay between spectra and relative sample sizes of the domains. Our numerical experiments on language models exhibit the same qualitative phenomenon: appropriate data mixtures yield a faster decrease in target-domain test loss than training on either domain alone.

[LG-15] Mutual Information Constrained Chernoff Bottleneck

链接: https://arxiv.org/abs/2609.37994
作者: Dier Tang,Guangyue Han
类目: Information Theory (cs.IT); Machine Learning (cs.LG); Optimization and Control (math.OC); Statistics Theory (math.ST); Machine Learning (stat.ML)
*备注: 31 pages, 3 figures, 2 tables. Feedback and comments are welcome

点击查看摘要

Abstract:The classical information bottleneck (IB) measures the relevance of a representation U of X to a target Y by I(U;Y) , which does not directly characterize the error of downstream decisions. For a binary hypothesis Y inferred from many separately encoded observations, the optimal error exponent is the Chernoff information between the two conditional distributions of U given Y . We study the mutual information constrained Chernoff bottleneck, which seeks an encoder that maximizes this Chernoff information subject to a rate constraint I(U;X) \leq R . We show that its optimal value C® increases strictly up to R = H(V) , where V merges the symbols of X with equal likelihood ratio, remains at the uncompressed exponent beyond, and, unlike the IB curve, need not be concave. We further show that k+1 outputs suffice to attain C® , where k is the cardinality of V . We propose an alternating algorithm that updates the encoder via a generalized Blahut–Arimoto algorithm and the Chernoff parameter s via a nonlinear equation, and prove that its iterates remain feasible, with nondecreasing and convergent Chernoff information. Numerical experiments confirm the theory, and on real topic-detection data from the 20 Newsgroups corpus, compressing each word to only 17% of its entropy retains 90% of the error exponent and nearly the accuracy of the uncompressed classifier.

[LG-16] abFM-Auto: Self-Evolving Pipelines for Tabular Foundation Models

链接: https://arxiv.org/abs/2609.37989
作者: Deqing Fu,Huangyuan Su,Rajat Sen,Taman Narayan,Sujay Sanghavi,Abhimanyu Das,Weihao Kong
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Tabular foundation models achieve strong zero-shot accuracy on structured data by pretraining on synthetic tables, but they ignore the column names, task descriptions, and auxiliary files that carry dataset semantics. Meanwhile, self-evolving machine learning engineering (MLE) agents train models from scratch on each dataset, yet jointly searching over features, architectures, and hyperparameters is noisy and prone to overfitting. We introduce TabFM-Auto, which pairs a tabular foundation model, TabFM, with a language model agent that evolves the data pipeline around it. Guided by dataset metadata and validation feedback, TabFM-Auto iteratively refines data cleaning, feature engineering, context selection, and post-processing to reduce TabFM’s error. Across all 51 datasets of the TabArena benchmark, five TabFM-Auto configurations with different agents and language models take the top five overall positions, and the best raises TabFM from 1785 to 2013 Elo. The discovered pipelines also transfer to other frozen tabular foundation models (+69 to +143 Elo) with no further search. On the 8 tabular competitions of MLE-Bench, TabFM-Auto ranks first overall among MLE agents.

[LG-17] abFM: A Zero-Shot Foundation Model for Tabular Data

链接: https://arxiv.org/abs/2609.37959
作者: Weihao Kong,Erez Louidor Ilan,Shuxin Nie,Taman Narayan,Rajat Sen,Yichen Zhou,Deqing Fu,Samet Oymak,Abhimanyu Das
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Tabular machine learning typically relies on per-dataset workflows, fitting tree ensembles or running AutoML searches from scratch for every task. We present TabFM, a 400M-parameter tabular foundation model that formulates supervised tabular prediction as in-context learning. TabFM produces calibrated zero-shot predictions in a single forward pass without task-specific tuning. Trained entirely on synthetic tables generated from structural causal models, TabFM learns general tabular representations that transfer zero-shot to real-world tasks. Across all 51 benchmark datasets in TabArena (38 classification and 13 regression), zero-shot TabFM ranks first among default tabular foundation models and outperforms tuned AutoML pipelines. Two extensions over the same frozen weights improve performance further on both tracks: multi-view feature expansion with ensembling and post-hoc calibration (TabFM+), and LLM-guided, dataset-specific data processing and feature engineering (TabFM-Auto).

[LG-18] Kolmogorov-Arnold Classifier Systems as Universal Approximators

链接: https://arxiv.org/abs/2609.37958
作者: Hiroki Shiraishi,Hisao Ishibuchi,Masaya Nakata
类目: Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE)
*备注:

点击查看摘要

Abstract:As the input dimension n grows, rule-based machine learning, such as Learning Classifier Systems (LCSs), faces a fundamental scalability bottleneck for function approximation: both rule count and parameter count grow exponentially with n . Traditional LCSs partition the n -dimensional input space directly, requiring \mathcalO(m^n) rules for adequate coverage, where m is the per-variable resolution. This article breaks from this paradigm by reorganizing rules dimension-wise, guided by the Kolmogorov-Arnold representation theorem: any continuous n -dimensional function can be expressed as a finite superposition of one-dimensional functions. The proposed Kolmogorov-Arnold Classifier System (KACS) decomposes the target function into one-dimensional subproblems and assigns a dedicated ruleset to each, reducing the worst-case rule count from \mathcalO(m^n) to \mathcalO(mn^2) and replacing n -dimensional local models with one-dimensional models requiring only two parameters per rule, independent of n . We also provide the first constructive proof that an LCS, namely KACS, is a universal approximator for continuous functions on compact domains. Evaluated against a direct n -dimensional input space partitioning approach under otherwise identical conditions, KACS achieves competitive accuracy in many settings while using only 2% to 40% of the parameters. Our implementation is available at this https URL.

[LG-19] Scene-Consistent Illumination Transfer for Inserted Advertising Graphics

链接: https://arxiv.org/abs/2609.37951
作者: Rameshwar Mishra,Bishshoy Das,A. V. Subramanyam,Guan-Ming Su
类目: Machine Learning (cs.LG)
*备注: 5 pages, 5 figures, and 3 tables; conference-style computer vision manuscript focused on single-frame advertising-banner relighting

点击查看摘要

Abstract:Replacing a visible advertisement in a broadcast frame is geometrically straightforward but photometrically delicate. A pasted graphic can have the correct perspective and still appear detached when its brightness, shading, or shadow disagrees with the surface beneath it. This paper presents Ad-Relight, an inference-only procedure for transferring scene illumination to a supplied advertising graphic without collecting a banner-specific training set. The procedure first separates slowly varying shade from graphic structure, then probes a pretrained diffusion relighter with two nearly identical backgrounds to isolate the contribution of the target region. A final pass combines this residual with a smoothed luminance field and a soft attenuation mask. Across 560 generated placements, the approach improves structural similarity, perceptual distance, and illumination agreement over geometric compositing and direct relighting baselines. Human judgments and an automated preference study show the clearest gains on floor-mounted graphics with nonuniform lighting. The current study is image based; temporal stabilization remains an open extension.

[LG-20] An Efficient Machine Learning Approach for Degradation Forecasting in AEM Water Electrolysis

链接: https://arxiv.org/abs/2609.37941
作者: Marco Veneriano,Ani Gjergji,Sebastiano Bellani,Andrea Riva,Vito Paolo Pastore,Matteo Santacesaria
类目: Machine Learning (cs.LG); Systems and Control (eess.SY); Applied Physics (physics.app-ph)
*备注: Accepted at IEEE ICAISF 2026, Catania

点击查看摘要

Abstract:This study provides a data-driven analysis of a novel dataset of single-cell Anion Exchange Membrane water electrolyzers (AEMWE), operated under constant current load across multiple heterogeneous experimental campaigns. We train and evaluate a range of machine learning models with different complexity, including linear baselines, LSTMs and CNNs, to perform medium-term forecasting of the cell voltage degradation curve. The models are assessed within a rigorous training and evaluation framework specifically designed for heterogeneous industrial data.

[LG-21] Learning When to Update: A Near-Optimal Timing Bandit Approach

链接: https://arxiv.org/abs/2609.37932
作者: Qiulin Lin,Junyan Su,Liyuan Wang,Minghua Chen
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Systems operating in dynamic environments require timely updates to sustain performance. For resource-intensive systems such as machine learning models and digital twins, strategically timing updates is essential. Updating too frequently wastes resources, while updating too infrequently leads to costly performance degradation. The problem is particularly challenging when the system’s degradation pattern is unknown a priori, as is common in new operating environments. We formalize this challenge as a novel \emphtiming bandit problem, where each arm represents a candidate update interval with a fixed update cost and an unknown, stochastic degradation cost. Three structural properties distinguish this setting from standard multi-armed bandits: selecting an interval commits the learner to multiple time slots before the next update; arm costs are composed of per-step degradation costs and a fixed update cost; and selecting a longer interval naturally reveals degradation at every intermediate step, providing consecutive feedback relevant to shorter intervals. By exploiting these structures, we develop Balanced Consecutive Arm Elimination (BCAE). BCAE achieves \tildeO(\sqrtT) regret, improving upon the \tilde\Omega(K\sqrtT) regret of standard bandit algorithms in this setting, where K is the number of candidate update intervals. We further propose an Optimism-Enhanced variant (OE-BCAE) that integrates lower-confidence-bound principles to improve empirical adaptivity while preserving the same regret order. Moreover, the regret bound achieved by our algorithms matches the theoretical lower bound up to logarithmic factors. Simulation results demonstrate that our algorithms achieve low regret and remain stable as both the number of arms and the update cost vary.

[LG-22] Pattern Formation in Transformers

链接: https://arxiv.org/abs/2609.37921
作者: Erkan Turan,Gaspard Abel,Maks Ovsjanikov
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:What are the inductive biases of a Transformer architecture? Existing theory on how the forward pass shapes representations either considers whether Transformers escape from rank collapse or demonstrates that self-attention drives tokens toward cluster patterns. The latter view arises from an elegant dynamical systems perspective, but relies on simplified architectural assumptions, and does not explain the rich structures observed in practice. This leaves a major open question: when a full Transformer escapes rank collapse, how does it structure token representations? Using pattern-formation theory, we show that the dynamical view of Transformers can account for Positional Encoding, Multi-Head Attention, and Output-Value geometry. We demonstrate that a full Transformer architecture imposes an inductive prior by selectively amplifying a rich set of previously unreported patterns, including traveling or rotating waves among others. We characterize the role of each architectural component in controlling which pattern is amplified, which ones stabilize, compete, or coexist. Finally, we show that these structures can act as a controllable dynamical prior that facilitates learning. By choosing both task-aligned positional encoding and weight initialization, we demonstrate improved data efficiency and accelerated optimization on controlled sequence tasks and with ConViT on CIFAR-10.

[LG-23] Search Dimension in Unlabeled Projection Pursuit: A Scaling Law for Subspace Restriction

链接: https://arxiv.org/abs/2609.37917
作者: Rares Dimitrie Grozavescu,Mark Girolami
类目: Machine Learning (cs.LG); Statistics Theory (math.ST); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Projection pursuit searches for a direction along which the data look least Gaussian. When the observation space contains a large Gaussian complement, the empirical objective can be minimized by a direction that carries no signal, with empirical kurtosis as low as at the truth. Sample splitting exposes rather than repairs this failure. Appending coordinates independent of the latent regime degrades the search while leaving Bayes recoverability unchanged. Restricting the search to the column space of a known forward operator removes the failure exactly on the negative-kurtosis branch. Estimating a principal subspace from the data is the alternative. In a controlled two-component model, the leading sufficient scalings differ in the gain with which the operator transmits the discriminant: \varsigma^-4 for covariance-spike estimation and \varsigma^-8 for fourth-moment search. At fixed search dimension, the measured threshold ratio collapses onto n/p^2 with exponent 0.156 , close to the predicted 1/8 . This is an empirically supported scaling motivated by sufficient bounds, not a proved asymptotically tight law. When the search dimension is varied, the measured exponent is 0.325 , substantially larger than 1/8 , and the tested range does not identify its functional form. The crossing location also depends on calibration and model configuration. Under a downstream excess-error criterion, the scaling largely disappears.

[LG-24] Scaling Zero-Order Pretraining through Model Sharding

链接: https://arxiv.org/abs/2609.37899
作者: Francois Chaubard,Mykel J. Kochenderfer,Chris Ré
类目: Machine Learning (cs.LG)
*备注: 38 pages, 17 figures

点击查看摘要

Abstract:Zero-order optimization (ZO) trains without backpropagation, making it relevant to forward-only hardware and non-differentiable loss, but its gradient variance grows with perturbed dimension, inhibiting large-model training. Sharded Optimization Mixture of Assemblies (SOMA) trains LSTM experts independently on N data clusters using simultaneous perturbation stochastic approximation (SPSA), without exchanging gradients, activations or optimizer state. Its separable loss removes cross-expert perturbation noise at the cost of jointly learned representations across domains. Using 80,000 estimated RTX 5090 GPU-hours, we show modest sharding improves training compute efficiency over all tested monolithic ZO controls. At 8.44M parameters and 150 aggregate GPU-hours, SOMA N=2 with 64 perturbations reaches 1.76 test nats/byte, versus 2.00–2.11 for monolithic SPSA at 64, 256 or 1,024 perturbations and 2.21 for EGGROLL. On WikiText-103, these frozen checkpoints reach 2.07, 2.25–2.36 and 2.49, respectively. On a fixed separable objective with equal-size blocks, we prove independent losses reduce relative gradient variance to approximately 1/N of a shared-loss estimator’s. Holding starting weights, data, perturbations and compute fixed, independent rather than summed losses lower SOMA N=4 test loss by 0.035 nats/byte after 1,000 updates across three seeds. Larger ensembles offer a separate inference benefit: at similar model size with top- k routing ( k=4 ), SOMA N=256 achieves 2.36M tokens/s versus 257k for SOMA N=8 ( 9.19\times , including routing), at lower test loss (1.68 versus 1.71), albeit using 59.9\times as much aggregate training compute. We release all training and evaluation code and checkpoints.

[LG-25] Behavioral Capacity Certificates for Quantized Language Models

链接: https://arxiv.org/abs/2609.37887
作者: Arian Eamaz,Mojtaba Soltanalian
类目: Machine Learning (cs.LG)
*备注: 43 pages, including appendices. Code: this https URL

点击查看摘要

Abstract:Activation and key-value cache precision change what a quantized language model computes without altering its stored weights. Direct weight-code bounds, however, assign identical complexity to deployments that behave differently and charge separately for weight codes that behave identically. Behavioral Capacity Certificates (BCC) charge for behavior using the aggregate prior mass of complete implementations—weights, scales, activation and cache rules—that induce the same bounded loss. When quantization merges implementations, this shared mass lowers the complexity penalty, and a break-even law determines when the saving survives the cost of validating it. BCC supports a three-step deployment workflow, and our experiments verify each step. First, a forward-only screen shortlists per-layer bit-widths by how often candidate perturbations preserve the reference predictions, with quality comparable to Hessian-guided selection at lower preprocessing cost. Second, margin-certified cells identify weights that can be pruned or sign-flipped without changing the deployed behavior: every permitted combination preserves all declared predictions, and on OLMoE-1B-7B and SmolLM2-1.7B, independent probes bound the probability that any permitted combination changes a prediction on new text. Third, BCC bounds the population loss of the deployed model, nonvacuously for complete decoders and more tightly than the compressed-code route. At equal cache memory, giving keys higher precision than values yields lower NLL and higher prediction agreement on GPT-2, Qwen2.5, and SmolLM2, together with a tighter complexity bound in the GPT-2 audit.

[LG-26] opoEmbedX: A General Framework for Representation Learning on Topological Domains

链接: https://arxiv.org/abs/2609.37884
作者: Florian Frantzen,Ibrahem AlJabea,Ines Henriques-Cadby,Theodore Papamarkou,Mustafa Hajij,Michael T. Schaub
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Topological structures such as simplicial complexes, hypergraphs, and cell complexes extend standard graph models by modeling higher-order relationships. These structures appear in many modern datasets and require specialized methods for generating meaningful embeddings. In this paper, we introduce TopoEmbedX, a unified framework for embedding a wide range of topological domains into Euclidean spaces. The package brings together several existing topological embedding algorithms—DeepCell, Cell2Vec, CellDiff2Vec, HOLE, and HOGLEE—and introduces five new algorithms: ComplexNetMF, ComplexRep, ComplexRandNE, ComplexWalklets, and ComplexHeat. These algorithms extend well-known graph embedding techniques to higher-order settings using the augmented Hasse graph of a topological domain. TopoEmbedX provides a clear, consistent, and easy-to-use framework for topological representation learning. Experiments show that the embeddings generated by TopoEmbedX support tasks such as classification and regression across multidimensional data.

[LG-27] Strict-Saddle Landscapes and Multi-Rank Geometry in Low-Tubal-Rank Tensor Sensing

链接: https://arxiv.org/abs/2609.37865
作者: Eugene Agyei-Kodie,Longxiu Huang,Shuang Li,Xiao Liang
类目: Machine Learning (cs.LG); Optimization and Control (math.OC)
*备注:

点击查看摘要

Abstract:We study the optimization landscape of low-tubal-rank tensor sensing through a balanced factorization. Under a tubal restricted isometry condition, we establish a quantitative strict-saddle landscape with no spurious local minima for arbitrary Fourier multi-rank profiles. We further show that the local geometry depends on the Fourier-slice ranks rather than the tubal rank alone. Uniform ranks yield quadratic growth transverse to the solution orbit, whereas nonuniform ranks produce quartically flat directions through hidden frequency-wise overparameterization, even when the factor width equals the exact tubal rank. Numerical experiments illustrate the global optimization behavior and the contrasting local geometries.

[LG-28] Delta-Matching: Closing the Final Gap of Native 8-bit Training for LLM s

链接: https://arxiv.org/abs/2609.37852
作者: Haozhan Tang,Hao Kang,Han Cai,Song Han,Chenyan Xiong
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Reliable FP8 attention remains a barrier to fully native 8-bit large language model training. We derive how forward-backward inconsistencies produce stale delta and empirically show how it distorts training dynamics. Our stale-delta hybrid runs show a modest loss gap at 569M parameters but substantial loss increases and downstream degradation at 1.67B and 5.29B. QK normalization, NoPE (no positional encoding), and lower-learning-rate context extension mitigate or delay degradation without eliminating it. This pattern suggests accumulated optimization error that smaller models and short runs can conceal. We propose Delta-Matching, proving that it restores the softmax gradient’s zero-row-sum invariant under the stated numerical assumptions. It enables native block-scaled FP8 in every forward and backward attention-core matmul without architectural changes, smaller global batches, or auxiliary forward outputs. Across tested architectures, scales, and training stages, Delta-Matching matches BF16/FP32 mixed-precision training loss and overall downstream performance. We will release our implementation, trained models, and data recipes.

[LG-29] Counterfactual Probing for Parallel Unmasking with Hidden Forest Structure

链接: https://arxiv.org/abs/2609.37841
作者: Ryotaro Kawata,Satoshi Hayakawa,Taiji Suzuki
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Masked generative models offer parallel token prediction, but accurate parallel sampling must account for dependencies among tokens. When dependencies are unknown, finding safe batches also costs model evaluations. We study whether total evaluations, including discovery, can be sublinear in sequence length N ; sublinear sequential depth then follows. We consider discrete distributions with hidden forest structure, accessed through a fixed approximate conditional oracle. Under explicit regularity conditions and uniform Hellinger error bounds, for any fixed target accuracy \varepsilon\in (0,1/8] and sufficiently large N , our sampler achieves seed-averaged total-variation error at most \varepsilon , with total masked-state submissions and sequential depth both bounded by O(N^C \varepsilon^a) for constants 0 C 1 and a 0 . These guarantees use polynomial vocabulary size and an edge-response lower bound set by N and \varepsilon . The sampler shares evaluations of hypothetical reveals across dependence tests to identify safe parallel batches without requiring full recovery of the hidden forest. A tunable parameter trades probing cost against irreversible commit rounds. In the same class, any admissible irreversible product-commit sampler attaining the same seed-averaged accuracy requires \Omega(N^c \varepsilon^b) counterfactual submissions or commit rounds in the worst case, for constants c,b0 .

[LG-30] Behavioral Convergence Without Representational Convergence: Persistent Training-History Dependence in Neural Networks

链接: https://arxiv.org/abs/2609.37836
作者: Ertuğrul Mutlu
类目: Machine Learning (cs.LG)
*备注: 12 pages, 6 figures, 2 tables. Code and reproducibility artifacts: this https URL

点击查看摘要

Abstract:Neural networks trained toward the same final objective can reach similar predictive performance while retaining internal representations shaped by earlier training history. We study this effect using controlled sequential-training experiments in which paired convolutional networks start from identical weights, experience reversed task orders, and then receive the same deterministic common-relaxation distribution. Across 20 paired MNIST runs, 16 satisfy a predeclared behavioral-matching criterion, yet their matched representations retain a mean history score of 0.139 (95% bootstrap CI: 0.127-0.153) and approximately 3.1% prediction disagreement. Extending common relaxation to 50,000 optimizer updates does not erase the measured difference: across five paired seeds, the representation-history score remains 0.190 (95% bootstrap CI: 0.161-0.219) at the end of the measured horizon while the mean accuracy gap is only 0.18 percentage points. Fresh linear probes show that, with sufficient labeled data, the two histories retain practically equivalent linearly accessible class information. A same-label rotated-MNIST control reproduces the effect: all five paired seeds reach behavioral matching while retaining a mean representation-history score of 0.162. Finally, a matched-learning-rate ReLU-LeakyReLU control reduces the 50,000-update representation residue by 0.040 on average in all five paired seeds, providing directional evidence that activation-mediated plasticity contributes to the persistence of training-history effects. These results provide protocol-scoped evidence that behavioral convergence need not imply representational convergence and that optimization history can leave measurable internal traces after prolonged common training.

[LG-31] Feedback-Calibrated Protein Optimization with Batch-Aligned Tail Arbitration

链接: https://arxiv.org/abs/2609.37808
作者: Zefeng Lin,Xianyong Fang,Tianfan Fu,Xiaohua Xu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Protein optimization aims to discover high-fitness sequences under a limited experimental budget. Existing machine-learning methods use task-specific predictors, biological priors, or ranking-aware objectives to guide which variants are tested in the next experimental round. However, these methods cannot adapt to shifts in the reliability of predictive evidence as measurements accumulate and ensure the correct ranking of key high-fitness candidates. To address these challenges, we propose Batch-Aligned Tail Arbitration (BATA), which uses experimental feedback to adaptively combine prior-informed and task-specific rankings for next-batch selection, with calibration focused on the batch-aligned high-fitness region. Across measured GB1, PABP, and TrpB landscapes, BATA achieves the best mean task rank (1.67) in final best fitness after 480 measurements. Controlled comparisons further show task-dependent gains from high-fitness calibration and batch alignment. Our work introduces feedback-calibrated predictor arbitration, where experimental feedback dynamically determines how predictive evidence guides next-batch selection, opening a new direction for protein optimization.

[LG-32] Optimizer-dependent training dynamics converge to the same one-third optimal data scaling

链接: https://arxiv.org/abs/2609.37745
作者: Hyunseok Lee,Mihir Basil,Yizhou Liu,Jeff Gore
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Neural scaling, in which loss falls as a power law with training, is central to large language models, and one recent proposal is that a 1/3 exponent emerges from learning peaked distributions. That account describes SGD, but models in practice are trained with adaptive optimizers. Here we separate two exponents the 1/3 account does not distinguish: how fast the loss falls with training steps along a single run, and how fast the optimally tuned loss falls with dataset size D . We show that the first, a dynamic exponent, is optimizer-specific while the second, an optimal data exponent, converges to 1/3 across optimizers. In an online teacher-student model we decompose the loss into norm growth (radial) and alignment toward the teacher direction (tangential), each decaying as a power law with dynamic exponents \alpha_r and \alpha_t . Under SGD, both are close to 1/3 , so the data exponent is also 1/3 across different learning rates. Under Adam the two separate: \alpha_r \simeq 0.48 but \alpha_t \simeq 0.08 . Since the total loss is minimized when these two parts are balanced, the optimal learning rate is optimizer-dependent: D -independent for SGD but falls with D for Adam. Yet tuned to that optimum, the loss returns to D^-1/3 for both. A stochastic-dynamics analysis explains why: the optimizers can trade decay speed between the two channels, but they all fall on a single dynamic exponent relation, 2\alpha_r+ \alpha_t = 1 , which fixes the optimal data exponent at 1/3 . Across seven optimizers, including Muon, the measured exponents are consistent with this relation, and the optimal-loss envelopes agree with D^-1/3 across them. The optimizer sets how fast a model learns per step; tuned optimally, it changes the prefactor but not the rate at which loss falls per sample.

[LG-33] HyDI: A hybrid Deep Learning-Inductive Logic Programming ensemble for multi-label classification

链接: https://arxiv.org/abs/2609.37740
作者: Simon Flügel,Till Mossakowski
类目: Machine Learning (cs.LG); Logic in Computer Science (cs.LO)
*备注: Accepted at IJCLR26 (6th International Joint Conference on Learning Reasoning, 16-18 September 2026)

点击查看摘要

Abstract:While attaining remarkable results for many applications, Deep Learning models are notoriously difficult to explain. This work introduces HyDI, a hybrid ensemble architecture for hierarchical multi-label classification. It combines a Deep Learning (DL) model with rule-based classifiers generated by Inductive Logic Programming (ILP). For leaf classes of the label hierarchy, the rule-based classifiers replace the DL model, leading to more transparent classification results. HyDI is applied to the Chemical Entities of Biological Interest (ChEBI) ontology, providing ILP-generated rules for 314 classes. For these classes, HyDI can generate global explanations as well as local explanations that combine visual and text-based descriptions.

[LG-34] Weights Read and Write Features: Scalable Parameter Decomposition Grounded in Activation Space

链接: https://arxiv.org/abs/2609.37731
作者: Tue M. Cao,Lisiane Pruinelli,My T. Thai
类目: Machine Learning (cs.LG)
*备注: preprint

点击查看摘要

Abstract:Activation space and parameter space provide complementary views of model computation. Activations represent information, while weights read, transform, and write that information. Yet existing interpretability methods largely study the two spaces separately, leaving the connection between represented information and parameter-level computation underexplored. We introduce Activation-Supported Parameter Decomposition (ASPD), which jointly decomposes activation and parameter spaces and grounds each learned weight component in the activation features it reads or writes. This grounding constrains otherwise non-unique parameter decompositions using the model’s internal activations, while an internal reconstruction objective provides a local learning signal at the weight matrix being analyzed. Together, these properties enable scalable, interpretable, and causally editable parameter decomposition in pretrained large language models, demonstrated on Qwen-3-8B. The learned read–write components can also be composed into parameter-level mechanism circuits. We use ASPD to recover mechanisms underlying the classic IOI circuit and trace semantic transformations through model weights.

[LG-35] Volatility-Clustering Adaptation for Financial Time Series

链接: https://arxiv.org/abs/2609.37715
作者: Manh Nguyen,Minh Hoang Nguyen,Huu Hiep Nguyen,Van Dai Do,Hung Le
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Time-series foundation models are increasingly adapted to new domains through fine-tuning on target data, under the implicit assumption that more target data yields better forecasts. We show that this assumption can fail in financial forecasting, where individual price changes are difficult to predict, but large moves tend to cluster, creating alternating calm and turbulent periods. Using financial foundation models trained on price bars of open, high, low, close, and volume, we argue that adapting to financial domains requires training signals beyond next-token prediction. We introduce Volatility-Clustering Adaptation (VCA), which augments next-token cross-entropy with a differentiable penalty on the autocorrelation of squared returns, the standard statistical signature of volatility clustering. This additional objective provides a multi-step training signal by matching the resulting dependence structure of autoregressive rollouts to those of the realized future. Across three asset sets and two evaluation conventions, VCA improves adaptation over the pre-trained model, with the strongest gains under the primary evaluation (\textscfore), driven primarily by reduced variance error. Overall, our results suggest that effective financial adaptation requires objectives that capture domain-specific temporal structure beyond token-level prediction.

[LG-36] Learning Expressive and Compositional Motion Representation via Spectral Skills

链接: https://arxiv.org/abs/2609.37677
作者: Feiyang Wu,Chenxiao Gao,Chen Yang,Ye Zhao,Bo Dai,Anqi Wu
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Robotic foundation models offer a promising path toward general-purpose humanoid robot control, often through hierarchical architectures. However, their effectiveness depends on the command interface between the planner and the controller, which must support accurate execution while remaining easy to predict, and ideally allow new behaviors to be composed from prior ones. In this work, we introduce spectral skills, a latent representation of this interface that meets these requirements through predictive representation learning. By design, spectral skills compactly encode short motion segments and are learned by predicting subsequent motion rather than reconstructing the encoder input. On a 29-DoF humanoid, a controller conditioned on spectral skills reduces global tracking error by 62% relative to the state of the art. The same frozen controller chains independently encoded skills without a separate transition policy. It also composes new behaviors by adding orthogonal directions to any compatible base skill, producing combinations unseen in the training data. We demonstrate tracking, chaining, and composition, as well as control through a language-conditioned planner, on Unitree G1 hardware. Project page: this https URL

[LG-37] LEMON-ZEST: Evolution-Informed Tokenization for Efficient Protein Language Modeling NEURIPS2026

链接: https://arxiv.org/abs/2609.37675
作者: Biswajit Banerjee,Claudia Alvarez Carreno,Anton S. Petrov
类目: Machine Learning (cs.LG); Biomolecules (q-bio.BM)
*备注: Accepted to NeurIPS 2026. 9 pages, 4 figures, 3 tables

点击查看摘要

Abstract:Protein Language Models (PLMs) have made remarkable progress following scaling laws established in natural language processing across sequence- and structure-based tasks, yet the potential of tokenization remains underexploited. Unlike human language, proteins preserve structure despite extensive sequence variation a property standard tokenization strategies fundamentally fail to capture. We introduce ZEST (Zoned Encoding of Sequence Traits), an evolution-informed vocabulary derived from conserved regions of multiple sequence alignments. ZEST allows embedding domain-level biological priors directly at the tokenization stage rather than learning them implicitly through scale. ZEST natively compresses sequences to an average token length of 4 residues, enabling our model to process 4,000 residues within a standard 1024-token context window. Building on this, we present LEMON (Layered Extraction of Molecular Ordering from Nature), a compact 200M-parameter sequence-based model for detection of remote homology between protein sequences trained on a single H100 GPU for one week. Despite its modest size, LEMON outperforms state-of-the-art models ranging from 600M to 3B parameters. Our results demonstrate that evolution-informed tokenization can substitute for massive parameter scaling, opening a new direction for efficient, biologically-grounded protein representation learning. All code, model weights, and results are publicly available under the MIT license.

[LG-38] Where Privacy Belongs: Placement Diagnosis and Certified Selection for Private Counterfactual Explanations on Graphs

链接: https://arxiv.org/abs/2609.37667
作者: Yuxiang Yao,Zijun Zhao
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Counterfactual explanations for graph neural networks (GNNs) find the minimal intervention that flips a node’s prediction–but computing one requires reading sensitive graph structure, and releasing it discloses that structure. Both existing placements fail. Privatizing the graph before explaining corrupts the target on exactly the borderline nodes needing recourse, manufacturing spurious flips that flip the privatized graph but not the true one. Explaining on the clean graph and perturbing the released explanation resists certification: re-auditing the standard heuristic shows an implied full-release budget of 573–753 on Cora and 256 on CiteSeer–orders of magnitude beyond its advertised budget–with worst-case single-entry leakage at AUC 1.0. We propose PrivCFS, which replaces certification-by-optimization with certification-by-construction: counterfactual selection over a fixed, data-independent candidate universe–edge interventions from a public prior graph, feature interventions from a public schema–whose no-op semantics give neighboring graphs the same output support. A validity-gated, clipped utility of global sensitivity \Delta u \le 1 released through the exponential mechanism gives pure \varepsilon -DP for the complete released object, composable over queries–to our knowledge the first such guarantee on graphs. Privacy noise is the cheapest stage: at \varepsilon =8 the release retains 94–97% of its support-restricted non-private optimum on the recourse population and 83–95% on the general one; the optimal edge-inference audit attains AUC 0.50 on average and 0.59 worst-pair, versus the heuristic’s worst entry 1.0; and transfers to a 15K-node graph at 0.96 valid rate. The dominant cost is a measurable, monotone price in public disclosure, readable off one table before any budget is spent–turning explanation privacy from an accounting risk into a purchasable decision.

[LG-39] Learning Causal Normalizing Flows from Incomplete Data via Observed-Data Likelihood

链接: https://arxiv.org/abs/2609.37664
作者: Trung-Dung Hoang,Alceu Bissoto,Tim Flühmann,David Herzig,Christos Nakas,Lia Bally,Lisa M. Koch
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Causal Normalizing Flows (CNFs) enable causal inference from observational data given the causal structure, but they assume fully observed training data. We introduce MissCNF, which trains CNFs directly on incomplete data by maximizing the marginal likelihood of each partially observed sample, without discarding rows or constructing a completed dataset. Thanks to the causal structure encoded in the autoregressive factorization of CNFs, only missing variables in the ancestral closure of the observed set are integrated out, while the others are dropped without computation. We further establish the conditions under which MissCNF recovers the true joint distribution, and introduce \emphcausal-family positivity, where identification is possible even when no record in the dataset is ever complete. We compare MissCNF with two common strategies for handling missing data: listwise deletion and impute-then-fit pipelines. Across eight synthetic causal benchmarks, three missingness mechanisms, and missing rates up to 90% , MissCNF achieves the lowest KL divergence in 23 of 24 nonlinear MCAR and MAR settings and in all nonlinear MNAR settings, as well as the lowest counterfactual error in 20 of 24 settings. On linear SCMs, where linear imputation performs best, MissCNF ranks in the top two in 22 of 24 settings.

[LG-40] Nonpreemptive Scheduling While Learning Context-Dependent Service Rates

链接: https://arxiv.org/abs/2609.37660
作者: Wansoo Choi,Seoungbin Bae,Dabeen Lee
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We study nonpreemptive contextual queueing bandits in a single-server system. Each job is represented by a d -dimensional context vector; in each round, a job may arrive with its context drawn from an unknown distribution \mathcalD , and its departure probability is determined by a logistic model of that context vector with an unknown parameter \theta^* . The server learns from service outcomes while deciding which waiting job to serve and whether to idle, aiming to minimize queue-length regret, the gap between its expected terminal queue length and the minimum achievable by an admissible policy. Once selected, a job must be served until completion, and we refer to this as the nonpreemptive setting. A central challenge is that, even with full model knowledge, the optimal policy cannot in general be characterized by a simple myopic rule, since the optimal action can change with the remaining horizon at the same queue state. Nevertheless, when the model and horizon are known, the optimal action can be obtained through a finite-horizon Bellman recursion. Motivated by this, we propose Learn–Clear–Plan (LCP), which estimates the system and uses the resulting Bellman recursion to make horizon-dependent decisions. LCP achieves \widetildeO(\sqrtd/T) queue-length regret, while a lower-bound construction gives \Omega(\min\1/\sqrtd,\sqrtd/T) regret for every learning policy on some instance, establishing optimality up to polylogarithmic factors when T\ge d^2 . When the horizon is unknown, no horizon-independent policy achieves vanishing regret against the finite-horizon optimum. We therefore use SEPT, the policy that serves a waiting job with the highest probability of departure, as a fixed reference, and suggest an estimated-SEPT algorithm that achieves a tracking error of \widetildeO(\sqrtd/t) without knowing the model.

[LG-41] RACE: Relation-Level Counterfactual Explanations for Heterogeneous Graph Neural Networks

链接: https://arxiv.org/abs/2609.37650
作者: Yuxiang Yao,Zijun Zhao
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Counterfactual explanations of graph neural networks identify edge deletions that flip a prediction. On heterogeneous graphs, however, existing methods first collapse the graph into untyped edges, so they cannot answer the question a domain expert actually asks: which relation type drives this prediction? We present RACE (Relation-Aware Counterfactual Explanations), which gives this question an exact, per-instance answer. For every explained instance, an exhaustive search over relation subsets returns the certified minimum relation-deletion set that flips the prediction – or an explicit report that no such deletion exists; each relation-level answer is then refined into a typed edge set within the attributed relations, verified on the discrete model by single-edge restoration. The relation-level answer is exact and deterministic given the frozen backbone, whereas soft-mask baselines vary by 6-8 pp in success rate across runs differing only in random ordering. On ACM, a Cora-derived graph, and ogbn-mag, RACE improves counterfactual success rate over the strongest baseline by up to +2.7 pp while deleting fewer edges, and attains the highest success rate among all same-task baselines on every dataset; the advantage reproduces across four backbones on ogbn-arXiv and on DBLP, with cross-seed relation-set agreement up to 0.89. A synthetic study with known generating mechanisms confirms that the search recovers the relation the trained model actually relies on – and reports infeasibility rather than fabricating an attribution when the model has learned none – so the explanations stay trustworthy exactly where explanations matter.

[LG-42] PHASE: Multi-Regime Modeling of Incompressible Magnetohydrodynamics

链接: https://arxiv.org/abs/2609.37609
作者: Radhika Achikanath Chirakkara,Rajdeep Haldar,Zezheng Song,Jiequn Han
类目: Machine Learning (cs.LG); Computational Physics (physics.comp-ph); Plasma Physics (physics.plasm-ph)
*备注: 22 pages, 13 figures

点击查看摘要

Abstract:Magnetohydrodynamics (MHD) is central to plasma modeling in astrophysics, space science, fusion, and engineering, but resolving multiscale MHD dynamics is computationally expensive. Machine-learning surrogates enable fast inference by learning reusable solution operators, yet existing models require separate training for each physical regime, limiting generalization across varying parameter settings. We introduce PHASE, a PHysics-Adaptive Scalable operator with residual Error correction, designed to model incompressible MHD across varying physical parameters with a single model. PHASE combines transfer learning, regime-aware adaptation, physics-centered learning, and residual refinement to improve both physical fidelity and generalization across MHD regimes. Together, these improvements achieve state-of-the-art prediction accuracy on two-dimensional MHD turbulence by reducing relative L_2 errors on physical fields by more than an order of magnitude compared to prior MHD neural-operator baselines. Moreover, PHASE generalizes successfully to unseen parameter values without retraining, demonstrating the cross-regime adaptability expected from operator learning. We evaluate PHASE beyond point-wise prediction errors using derived physical fields, spectral analysis, and distribution statistics, consistently observing improved physical fidelity. We further show that our framework can accurately simulate MHD instabilities by testing it on the Kelvin–Helmholtz instability, demonstrating the robustness of our method.

[LG-43] GraphVQ: Structure-Aware Autoregressive Decoding over Context-Quantized Graph Tokens

链接: https://arxiv.org/abs/2609.37604
作者: Yuxiang Yao,Zijun Zhao
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Graph foundation models need a discrete token representation, but casting a graph as a generatable token sequence faces a structural obstacle: edges spanning beyond the serialization window cannot be emitted in one pass–so one-pass autoregressive generators systematically under-produce cycles–and a single global condition cannot tell candidate edges apart. GraphVQ removes both obstacles: node contexts–features plus a local edge mask under multi-order breadth-first serialization–are quantized into a shared codebook by a VQ-VAE with BCE-calibrated Bernoulli edge decoding, and a second-stage structure-aware decoder emits the global adjacency conditioned on token-derived pair features, whose necessity over any global-summary condition is formalized in a scoped impossibility result. The tokenizer reconstructs node features at 0.86–0.99 accuracy and decodes local edges at AUROC = 0.89 (ECE = 0.007). Under one same-split protocol on four datasets, pair conditioning improves orbit MMD 0.248 - 0.174 on PROTEINS and 3.4x on a ring stress test, and vanishes on a random-label control–the signature of attribute–topology coupling–so the gain is claimed exactly where attributes carry edge-relevant signal. GraphVQ ranks first among learned generators on PROTEINS, ties for first on SYN-COMM, and improves orbit MMD 2.7–17x over one-stage generation on three datasets, with seed-level bootstrap intervals confirming the rankings are not seed noise; on MUTAG the unweighted edge target under-generates and is reported as such. These results locate the structural control of autoregressive graph generation in the granularity of the condition: pair-level token context turns a quantized vocabulary into a usable capacity axis for distribution-faithful graph generation and future token-level pretraining.

[LG-44] BlenDAgger: Blended Shared Control for Interactive Imitation Learning

链接: https://arxiv.org/abs/2609.37599
作者: Cailyn Smith,Geoffrey Sun,Henny Admoni,Zackory Erickson
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Robot policies are frequently trained from human corrections, yet teleoperating a robot to provide corrections is burdensome, and human demonstrators are not always optimal. We propose Blended DAgger (BlenDAgger), an approach for collecting data to train imitation learning policies by using shared control to blend the policy’s and demonstrator’s actions during interventions. By blending human and policy actions, we aim to improve the autonomous performance of manipulation policies. We validate our approach across five manipulation tasks, two in the real world and three in simulation. Our approach achieves higher autonomous performance by 30 or more percentage points on two real-world tasks compared to a typical human-gated correction approach (HG-DAgger). We also investigate the advantages of BlenDAgger that allow for higher autonomous performance, finding that BlenDAgger results in 57% smoother transitions between policy control and human interventions, and 14% higher trajectory similarity to the training data. In a user study (n=14) on two real-world tasks, we find that BlenDAgger results in faster data collection (BF=13.32), and we do not find a difference in subjective perceptions. These results show that blended shared control leads to higher autonomous performance compared to typical methods for fine-tuning robot policies from fully teleoperated interventions.

[LG-45] On Task Scope and Information Retention in Source Coding

链接: https://arxiv.org/abs/2609.37575
作者: Alireza Furutanpey,Kerstin Bunte
类目: Information Theory (cs.IT); Machine Learning (cs.LG); Image and Video Processing (eess.IV)
*备注:

点击查看摘要

Abstract:We argue that dividing codec design into Coding for Machines (CfM) and Coding for Humans (CfH) is a misleading distinction for deciding what information a codec may discard. Receiver identity does not determine admissible information loss. The required rate depends on task scope, including the predictions to support, their losses and tolerated risks, the encoder observation, and the permitted decoding procedures. Notably, a machine task may have a higher minimum rate than a restricted human decision. Rate savings on selected machine tasks apply only to the stated requirements, not to an intrinsic ordering by receiver type. We extend source and feature coding to finite task families, derive when restricting the encoder observation preserves the minimum rate, and show that equality between source and split-feature coding rates can no longer hold as the task scope expands.

[LG-46] A Model-Agnostic Physics-Guided Adapter for Few-Shot Transfer of Coastal Flood Prediction Models to Unseen Regions

链接: https://arxiv.org/abs/2609.37565
作者: Bilal Hassan,Areg Karapetyan,Samer Madanat
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Deep learning surrogates can produce high-resolution coastal flood maps orders of magnitude faster than physics-based hydrodynamic simulators, yet transferring them to new coastal regions remains costly, since generating target-region data for fine-tuning typically requires numerous time-consuming simulations. To tackle this bottleneck, we introduce the Physics Adapter (PA), a compact, architecture-agnostic adaptation interface that enables efficient few-shot transfer of flood prediction models across diverse coastal regions. PA predicts peak water level through a differentiable wet/dry response that compares terrain elevation against a learned water level, and blends this physics-structured prediction with a data-driven branch through a learned gate. Unlike physics-informed formulations, PA imposes no PDE-residual or conservation losses and instead exploits elevation as an architectural inductive bias, adding a negligible number of trainable parameters. We integrate PA into 12 heterogeneous models, and evaluate them on two coastal regions with markedly distinct geometries, topographies, and shoreline protection configurations. The performance of PA is benchmarked against a no-physics baseline, full fine-tuning, and standard parameter-efficient fine-tuning (PEFT) methods, considering both within-region generalization to unseen sea level rise values and between-region transfer. In low-shot regime (K=3), and averaged over all backbones and transfer settings, adding PA reduces root mean square error by 11.5% when only the output head is adapted on a frozen backbone, by 15.4% when combined with PEFT methods, and by 22.9% under full fine-tuning, compared to matched configurations without PA. Taken together, the findings of this work offer practitioners a concrete recipe for extending DL-based coastal flood predictors to new, data-scarce regions.

[LG-47] Benchmarking graph-based models for in-silico toxicity prediction in drug discovery

链接: https://arxiv.org/abs/2609.37555
作者: Noel Suarez-Barro,Manuel Lama,Juan C. Vidal
类目: Machine Learning (cs.LG); Biomolecules (q-bio.BM); Quantitative Methods (q-bio.QM)
*备注:

点击查看摘要

Abstract:Drug discovery is a costly and high-risk process, where toxicity-related failures remain a major cause of attrition in both preclinical and clinical stages. As a result, accurate early prediction of chemical toxicity is essential to reduce downstream costs and improve compound prioritization. In this context, graph deep learning (GDL) has emerged as a powerful paradigm for toxicity prediction, leveraging molecular graph representations to learn directly from chemical structure with improved expressivity over traditional approaches. Despite the growing number of proposed models, current literature-based comparisons are often difficult to interpret due to inconsistencies in datasets, preprocessing pipelines, and evaluation protocols. To address this limitation, we introduce a unified and standardized benchmarking framework for GDL-based toxicity prediction. We systematically evaluate more than 20 representative approaches under consistent experimental conditions and across multiple datasets and partitioning strategies, enabling a fair and reproducible comparison of model performance. In addition, we complement this empirical study with a structured literature analysis to contextualize existing methodological trends and performance claims. Our results provide a clearer and more reliable assessment of the current state of the field, highlighting both the strengths and limitations of existing graph-based approaches. To support transparency and reproducibility, we release our benchmarking framework as open-source software this https URL, allowing the community to evaluate and compare models under consistent conditions. Subjects: Machine Learning (cs.LG); Biomolecules (q-bio.BM); Quantitative Methods (q-bio.QM) Cite as: arXiv:2609.37555 [cs.LG] (or arXiv:2609.37555v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2609.37555 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-48] Why Adaptive Optimizers Underestimate Rare Tokens

链接: https://arxiv.org/abs/2609.37535
作者: Sangsidhya Kar
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:In the softmax output layer, a rare token receives a small positive logit gradient on most steps and a much larger negative gradient on the few steps when it is the target. SGD simply adds these contributions. Coordinate-wise adaptive methods such as Adam, RMSProp, and sign descent instead divide each update by a running estimate of its magnitude, and that estimate is largest immediately after the token appears. This imbalance has two effects. At the level of the whole output layer, we characterize which optimizers preserve the mean output embedding: every method whose update is linear in past gradients does, as do Kronecker-factored and orthogonalized methods such as Shampoo and Muon. Adam, Adafactor, Lion, and sign descent do not, and for these methods we obtain an exact step-by-step expression for the change. At the level of an individual rare token, the same normalization shifts the training fixed point. In the unigram model, sign descent lowers the logit of every token that occurs in fewer than half of the minibatches at a constant expected rate. For RMSProp with periodic arrivals, we can solve the fixed point in closed form: if a token is absent for at least two consecutive minibatches, its equilibrium probability is strictly below its data frequency for every learning rate, and the ratio tends to \kappa/(2(e^\kappa/2-1)) . Here \kappa is the mean number of steps between occurrences divided by the second-moment time constant 1/(1-\beta_2) . In the same model, SGD and AMSGrad retain the unbiased fixed point. We test these predictions both in a unigram model and in a small language model trained from a known generating distribution. With random arrivals, the bias is larger than the periodic formula predicts; in the language model, the optimizers with the biased fixed point also fit the generating distribution less well.

[LG-49] Physical Muon: Orthogonalization as an Equilibrium Computation

链接: https://arxiv.org/abs/2609.37525
作者: Yuren Hao
类目: Machine Learning (cs.LG); Emerging Technologies (cs.ET); Machine Learning (stat.ML)
*备注: Accepted at 18th annual workshop on Optimization for Machine Learning

点击查看摘要

Abstract:Physical neural networks and analog in-memory computing could reduce the energy cost of neural network training. Realizing this potential, however, requires optimizers that combine effective learning with physical implementability. SGD fits local analog updates but struggles on transformers, while Adam family is unstable against analog bias. Muon offers strong training performance, but its Newton–Schulz orthogonalization relies on dense matrix-matrix products. To address this obstacle, we introduce Physical Muon, which computes the orthogonalization as the equilibrium of a continuous-time flow. Random probes approximate the flow using matrix-vector products, reciprocal reads, and local rank-1 writes. To test whether this replacement preserves training performance, we evaluate it on a 10.95M-parameter transformer. The dense flow’s mean validation cross-entropy is 0.0085 above Newton–Schulz across nine seeds per method; the probe implementation is 0.0188 above the control across two seeds. Circuit simulations further reproduce the flow dynamics and yield comparable training behavior.

[LG-50] Graph-Conditioned On-Policy Agent Distillation from Off-the-Shelf Teachers

链接: https://arxiv.org/abs/2609.37522
作者: Xiaohan Yi,Wen Luo,Yani Huang,Junfeng Zhan,Asher Qin,Peilin Zhao,Xi Xiao
类目: Machine Learning (cs.LG)
*备注: 18 pages, 3 figures

点击查看摘要

Abstract:On-policy distillation (OPD) trains compact language agents with teacher feedback on student-generated trajectories. In multi-turn tasks, compounding errors can move students beyond the teacher’s effective supervision. We introduce Graph-Conditioned On-Policy Agent Distillation (GC-OPD), which enriches an off-the-shelf teacher’s scoring context with execution evidence. A graph indexes repeated teacher executions by shared states while preserving complete successful and failed histories. After each student episode, GC-OPD retrieves current-state references or historical alternatives and combines them with student hindsight to score the original thought-action tokens. Using the same original teachers, GC-OPD improves mean success over vanilla OPD from 24.70% to 48.78% on ScienceWorld (4B student), from 53.36% to 85.26% on ALFWorld Unseen, and from 29.10% to 37.65% on WebShop. At matched student sizes, it also achieves higher mean success than every evaluated OPD baseline using GRPO-trained teachers on ScienceWorld and ALFWorld; the strongest such ScienceWorld 4B baseline reaches 46.66%. GC-OPD requires no task-specific teacher optimization.

[LG-51] ScaGNN: a Graph Neural Network for Multiple Scattering Simulations

链接: https://arxiv.org/abs/2609.37509
作者: Rémi Marsal,Stéphanie Chaillat,Alexandre Chapoutot
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:The boundary element method (BEM) provides an efficient numerical framework for solving multiple scattering problems in unbounded homogeneous domains. By restricting the discretization to the domain boundaries, it substantially reduces computational complexity. The procedure first consists in determining the solution trace on the boundaries of the domain by solving a boundary integral equation. Then, the volumetric solution can be recovered at low computational cost using a boundary integral representation. As the first step of the BEM represents the main computational bottleneck, we present ScaGNN, a learning-based approach designed to approximate the solution trace. It relies on a graph neural network architecture that incorporates a dynamic adaptive edge sampling mechanism for selecting the most relevant interactions to model. Guided by intermediate predictions of expected error and edge length, this mechanism selects, at various stages of the forward pass, the most relevant distant interactions to model. The proposed method is tailored to achieve linear complexity with the number of nodes in the input graph. To train and evaluate our network, we present a benchmark consisting of several datasets with different types of multiple scattering problems. Our experiments show that our approach surpasses existing state-of-the-art learning-based methods on the considered tasks and investigate the generalization capabilities to settings with an increased number of obstacles and out-of-distribution obstacle shapes. this http URL

[LG-52] REVO: Rollout-Efficient Off-Policy Distillation via Variance-Guided Reuse

链接: https://arxiv.org/abs/2609.37500
作者: Yuxiao Yang,Shangzhe Li,Tianrun Yu,Kaixiang Zhao,Taylor W. Killian,Weitong Zhang
类目: Machine Learning (cs.LG)
*备注: 26 pages, 8 figures, 10 tables, code available at this https URL

点击查看摘要

Abstract:On-policy distillation (OPD) trains language models using dense token-level teacher supervision on student-generated trajectories. However, its reliance on frequently refreshed student rollouts often incurs substantial generation cost. We introduce REVO, an off-policy distillation framework that improves rollout efficiency by reusing each student rollout for multi-step learner updates. REVO addresses prefix-level and current-token policy mismatch through stabilized prefix weighting and one-step resampling from the current student, which enables repeated updates without regenerating full trajectories. To prioritize informative token positions within reused rollouts, REVO uses the variance of the student-teacher log-probability ratio to quantify the remaining token-level learning signal and guide repeated optimization. Across multiple student-teacher scales, REVO with only 50 rollout iterations matches or exceeds OPD baselines trained for 200 iterations on both in-domain and cross-domain reasoning benchmarks.

[LG-53] KT-EGO: A Knowledge Transfer Assisted Efficient Global Optimization Algorithm for Solving High-Dimensional Expensive Black-Box Problems

链接: https://arxiv.org/abs/2609.37473
作者: Qineng Wang,Liming Song,Yun Chen,Guangjian Ma,Zhendong Guo,Jun Li
类目: Computational Engineering, Finance, and Science (cs.CE); Machine Learning (cs.LG)
*备注: 25 pages, 12 figures; abridged author manuscript

点击查看摘要

Abstract:Many engineering problems involve optimizing a high-dimensional expensive black-box (HEB) design space. To solve such problems efficiently, we propose a knowledge transfer assisted efficient global optimization (EGO) algorithm, labeled as KT-EGO, which extends the EGO algorithm for solving problems over higher dimensions (i.e., d20 ). Specifically, the original design space is divided into several low-dimensional subset design spaces. More importantly, in order to extract information from the subset design spaces to accelerate the progress of full optimization, we propose a surrogate-based data fusion strategy in KT-EGO. And further, a searching strategy with an adaptive variable range is devised to enhance the exploitation of promising areas. To show the effectiveness of our proposed algorithm, it is compared against the state-of-the-art algorithms over 12 benchmark functions and a 28-dimensional engineering optimization for the design of compressor blade, which fully validates the effectiveness of the KT-EGO for solving HEB problems.

[LG-54] Probability Contracts: Accuracy Coherence and Decisions Across LLM Interfaces

链接: https://arxiv.org/abs/2609.37470
作者: Han Chen,Yingrui Li
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: Research preprint; includes proofs and supplementary analyses

点击查看摘要

Abstract:A probability used for a decision should refer to the same event across equivalent requests. We introduce probability contracts, a benchmark connecting exact finite-world posteriors, validated event transformations, and failure-aware decision evaluation. Four model-interface configurations are evaluated on 1,000 worlds. Their assessments differ across accuracy, coherence, and decision loss: Kev has lower aggregate canonical posterior error than Jev, but larger complement and coarsening residuals, with accuracy ordering varying by stratum. Jev’s Event and Choice interfaces induce different binary actions on 32.8% of valid pairs at defer cost 0.10. A post-hoc analysis finds that disagreement certifies only 11-52% of mean binary pair error across configurations. An elementary action-region characterization explains when averaging changes decision loss relative to randomly selecting one interface. Although averaging cannot worsen that baseline’s expected Brier score, its decision effect depends on the cost and crossed boundaries; the observed same-baseline penalties occur in configurations already worse than always deferring. The benchmark makes these distinctions measurable without treating consistency as accuracy or a score improvement as a decision guarantee.

[LG-55] Variational Augmented Invertible Koopman Autoencoder for probabilistic time series forecasting

链接: https://arxiv.org/abs/2609.37435
作者: Anthony Frion,Lucas Drumetz,Guillaume Tochon,Mauro Dalla Mura,Ali Can Bekar,Abdeldjalil Aïssa El Bey
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Neural Koopman autoencoder models have been shown to successfully build a latent embedding with linear dynamics for arbitrary dynamical systems, enabling strong performance in long-term time series forecasting. However, these models usually work in a deterministic setting, which does not allow the quantification of the uncertainty of their predictions. Thus, we propose the new Variational Augmented Invertible Koopman AutoEncoder (VAIKAE), in which the latent embedding follows a Gaussian distribution instead of being deterministic. A key property of the VAIKAE architecture is that it leverages normalizing flow models, enabling the use of likelihood computations in the state space of dynamical systems for training a model. We further propose new strategies for uncertainty-aware latent data assimilation with a trained VAIKAE model. The effectiveness of our methods is demonstrated in a series of experiments on long-term time series forecasting benchmarks.

[LG-56] Looped Actor: Depth-Recurrent Reasoning Models for Reinforcement Learning

链接: https://arxiv.org/abs/2609.37432
作者: T. Konstantin Rusch,Tim Seyde,Jared Boyer,Zach J. Patterson,Daniela Rus
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Looped reasoning models repeatedly apply a shared set of parameters, enabling more computation without increasing the model size. These models also support input-dependent computation by dynamically deciding when to stop looping. Motivated by the recent success of looped transformers in language modeling and reasoning, we investigate whether dynamic looping can similarly benefit sequential decision-making. We provide a complexity-theoretic motivation for this approach by showing that there exist Markov decision processes in which a state-adaptive policy achieves the optimal return with asymptotically less expected computation than any optimal fixed-runtime policy. To learn compute-adaptive policies in practice, we introduce Looped Actor, a transformer-based policy that repeatedly refines a latent representation toward a fixed point using a shared computational block. This allows the model to allocate computation adaptively by varying the number of loops based on the current state. We evaluate Looped Actor on 22 tasks across six environments, ranging from combinatorial puzzles to robotic manipulation and spanning online and offline reinforcement learning (RL) with discrete and continuous actions. Looped Actor matches or exceeds the performance of an untied baseline with 16 \times more parameters, with the largest gains in environments where action selection requires substantial multistep planning. For the Boxoban environment, we find that the computation allocation is structured: the number of loops increases with the number of remaining pushes and future optimal pushes become increasingly predictable from the latent state over successive loops. Together, these results highlight actor looping as a simple and efficient way to equip RL agents with adaptive computation and improve their planning capabilities. Code is available at this https URL

[LG-57] Scale Sensitivity in Low-Bit Post-Training Quantization: Curvature of the Quantization Error Landscape

链接: https://arxiv.org/abs/2609.37416
作者: Jonas von Berg,Massimiliano Datres,Carlo Kneißl,Gitta Kutyniok
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Post-training quantization (PTQ) methods in the GPTQ family minimize a layer-wise reconstruction error on a uniform grid whose scale must be chosen; the common max-based choice degrades sharply at low bit-widths. We study how sensitive this objective is to the scale. For a layer with i.i.d. Gaussian weights and calibration activations of sufficiently large effective rank, we prove that, as the width grows, the normalized round-to-nearest loss converges with high probability, uniformly over all scales, to the mean-squared error of a uniform quantizer applied to a standard Gaussian; we verify the effective-rank condition for wide, randomly initialized MLPs with odd Lipschitz activations and isotropic Gaussian calibration data. The limiting objective has a unique nondegenerate minimizer, whose scale decreases strictly with the number of levels and whose curvature with respect to relative scale errors decays approximately exponentially with the bit-width. GPTQ experiments on five LLMs show the same trend: the scale rule changes perplexity substantially at 2–3 bits and negligibly from 6 bits on, and a local measure of GPTQ scale sensitivity decreases with bit-width in line with the Gaussian curvature. The Gaussian-optimal scale fails on raw weights; after Hadamard incoherence processing it matches the best searched rule at 3 bits and above without any search, but remains clearly worse at 2 bits.

[LG-58] Learning Macroscopic Dynamics without Reconstructing Microscopic States

链接: https://arxiv.org/abs/2609.37392
作者: Zhichao Han,Yue Zhao,Qianxiao Li
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Modeling the temporal evolution of macroscopic properties of complex systems is an important scientific task. To predict this evolution without full microscopic simulation, a common approach encodes microstates into compact latent states, learns their evolution, and reads out macroscopic predictions from the latent trajectory. These latent states are often learned through microstate reconstruction. However, with limited latent capacity, reconstruction can favor high-variance microscopic details over information needed for macroscopic prediction. Yet jointly learning latent states and their transition without reconstruction often fails to obtain latent dynamics that support accurate macroscopic prediction. We show that this failure can arise from latent scale collapse: shrinking the latent state scale reduces training loss while macroscopic evolution error remains large. Here, we propose a reconstruction-free framework to learn latent states with their dynamics for prescribed macroscopic prediction. Training alternates between updating the latent representation with the transition and next-state latent targets fixed, and updating the transition with the latent representation fixed. At inference, the trained model predicts macroscopic states recursively from an initial microstate. Our theoretical analysis characterizes reconstruction misalignment and scale collapse under joint training, and gives a sufficient condition for local convergence to correct latent dynamics for our method. Experiments on epidemic spreading on a lattice, mixing of two particle species, and polymer stretching demonstrate that the proposed method achieves substantially better macroscopic prediction over baselines.

[LG-59] Rethinking Soft Tokens for Parallel Decoding in Diffusion Language Models

链接: https://arxiv.org/abs/2609.37391
作者: Kodai Kawamura,Kenji Kawaguchi,Anji Liu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Diffusion language models (DLMs) enable parallel generation by predicting and committing multiple tokens at each denoising step, yet they can generate individually plausible but mutually inconsistent tokens. Recent work shows that \emphsoft tokens can mitigate this issue by representing uncertain positions with continuous embeddings built from the model’s predictive distribution at the previous decoding step. However, although soft tokens are commonly understood as preserving predictive uncertainty, how soft-token feedback improves parallel decoding has not been systematically examined. In this paper, we investigate this question in frozen pretrained DLMs to examine soft-token feedback without the effects of additional training. To construct soft-token inputs in a training-free setting, we identify a geometric mismatch between conventional soft-token construction and the pretrained embedding space. Based on this observation, we propose a training-free, geometry-aware construction of soft tokens. Our analysis of soft-token feedback suggests that uncertainty preservation alone does not fully explain how it reshapes subsequent predictions. To better explain how soft-token feedback improves parallel decoding, we provide empirical evidence that it favors coherent token sequences. Across four pretrained DLMs and four math and code benchmarks, our method outperforms standard parallel decoding and a training-free Euclidean soft-token baseline. Code: this https URL

[LG-60] MoTIF-X: A Multimodal Tokenized Framework for Interpretable and Extensible Molecular Representation Learning

链接: https://arxiv.org/abs/2609.37384
作者: Linqing Mo,Jiayu Zhou,Bin Chen
类目: Machine Learning (cs.LG)
*备注: 5 figures. Supplementary material is available as an ancillary file. Code: this https URL

点击查看摘要

Abstract:Molecular representation learning is central to computer-aided drug discovery. Molecular graphs, SMILES strings, and 3D conformations provide complementary structural information, yet many multimodal approaches encode these views independently and align them only at a later stage, limiting fine-grained cross-modal interaction and substructure-level interpretability. To address these limitations, we introduce MoTIF-X, a motif-centered framework that uses graph-grounded chemical motifs as shared anchors for multimodal integration and interpretation. Its first pretraining stage learns motif representations through hierarchical contrastive learning across atomic, motif, and molecular scales. The second stage contextualizes these representations with SMILES and torsion-angle tokens through multimodal masked token modeling. After pretraining on drug-like molecules with multiple conformers, MoTIF-X achieved the lowest mean absolute error on all nine OpenADMET ExpansionRx endpoints and the best overall performance among the evaluated methods. Significance analyses supported its advantage in the vast majority of endpoint-baseline comparisons after multiple-testing correction. Ablation studies supported the complementary contributions of motif-token contextualization, multimodal integration, and two-stage pretraining. Beyond molecular properties, the framework extended to drug-target interaction prediction, achieving the best average classification performance across the evaluated benchmarks and generalizing to an external drug-cold-start dataset without additional fine-tuning. Its motif-centered design also enabled substructure-level interpretation: higher motif attribution scores were associated with larger experimentally measured activity shifts. Together, these findings support MoTIF-X as a transferable and interpretable framework for molecular modeling. Comments: 5 figures. Supplementary material is available as an ancillary file. Code: this https URL Subjects: Machine Learning (cs.LG) Cite as: arXiv:2609.37384 [cs.LG] (or arXiv:2609.37384v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2609.37384 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-61] High-Dimensional Simulation-Based Inference in Latent Spaces

链接: https://arxiv.org/abs/2609.37381
作者: Lars Kühmichel,Stefan T. Radev,Bhanu Prasanna Koppolu,Masoumeh Davoudi,Jerry M. Huang,Paul-Christian Bürkner
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 27 pages, 10 figures, 6 tables

点击查看摘要

Abstract:Neural simulation-based inference (SBI) has been widely successful in inferring a relatively small number of interpretable parameters from potentially high-dimensional observations, such as images or time series. Accordingly, representation learning in SBI has focused almost exclusively on compressing the observations used to condition the posterior. More recently, however, SBI has begun to target increasingly high-dimensional parameter spaces, raising the complementary question of whether the inference target itself should be compressed. Our answer is a practical merger of SBI and latent generative modeling, which learns a low-dimensional representation of the simulator parameters, performs posterior inference directly in this latent space, and maps posterior samples back to the original parameter space. We characterize the conditions under which latent-space inference recovers the desired target posterior and systematically study its empirical trade-offs. Across four case studies and three generative families, we compare latent and standard estimators while controlling for network capacity, regularization, optimization, and training compute. At matched training compute, latent-space inference achieves accuracy and marginal calibration comparable to direct target-space inference while sampling up to more than an order of magnitude faster.

[LG-62] Looped Transformers as Optimizers

链接: https://arxiv.org/abs/2609.37379
作者: Yulong Huang,Chen Jiang,Zhanpeng Zhou,Hongtao Zhang,Tianyu Li,Tianyu He,Xiangyu Zhang,Bojun Cheng
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Looped Transformers provide a parameter-efficient approach to depth scaling by repeatedly applying shared Transformer blocks. Recent reasoning models have likewise highlighted the value of scaling test-time computation through longer computation trajectories. However, the principles for designing effective loop transitions remain poorly understood. We view the looped hidden state as a fast weight that is updated throughout the depth. We formulate loop transitions as local gradient-based updates, with recurrent blocks predicting implicit targets at each depth. Our framework derives loop transitions in closed form from a projection, a local objective and an optimizer update rule. Mapping representative loop transitions into this framework reveals mismatches between their transitions and projections. We first align the input maps of existing transitions. We then derive OperLoop, which combines explicit weight decay, adaptive step size and a delta objective. The aligned variants reduce training loss and improve average commonsense accuracy. OperLoop improves average generative performance over the compared looped and non-looped baselines under matched training FLOPs. These results support the framework’s usefulness for loop design. We extend the analysis to additional loop models and outline a roadmap for future loop transition design.

[LG-63] Backdoor Mitigation in Decentralized LLM Fine-Tuning

链接: https://arxiv.org/abs/2609.37367
作者: Sayan Biswas,Jade Garcia Bourrée,Rachid Guerraoui,Maxime Jacovella,Anne-Marie Kermarrec,Sathwika Peechara,Martijn de Vos,Milos Vujasinovic
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Decentralized large language model (LLM) fine-tuning lets organizations collaboratively train a shared LLM on data they cannot pool, without a central coordinator. In every round, each node exchanges a trainable adapter with its neighbors over a communication graph, and then aggregates them. This setting, however, is vulnerable to propagated backdoors, which is a hidden behavior that lets a model perform normally on clean inputs but produce an attacker-chosen output whenever a secret trigger appears. We show that a single node poisoning its own model can backdoor adapters of nodes that have never seen a poisoned example, making them refuse prompts that contain a secret trigger. We present Chorus, a decentralized mechanism that lets each node detect and reject backdoored adapters from its neighbors before aggregation, without requiring shared validation data or any knowledge of the attacker’s trigger or target. Chorus judges each adapter by its behavior, using the receiver’s own adapter as a trusted reference. Crucially, no node in Chorus judges adapters alone: the receivers of each adapter update probe it independently, pool their findings in the neighborhood, and vote to make a decision. So a backdoor that slips past one receiver is still caught by the others. We evaluate the effectiveness of Chorus using two instruction-tuning datasets and LLM architectures, and against a state-of-the-art baseline. Chorus cuts the average attack success rate (ASR) of the attacker’s neighbors from 48-63% to at most 2.2%, within 0.6 percentage points of an omniscient oracle that knows the exact malicious nodes. Even the worst-affected honest node never exceeds 10% ASR, the same bound as the oracle, against up to 78% without defense. This all comes at a negligible communication overhead.

[LG-64] Parallel Tempering for Diffusion-Based Combinatorial Optimization

链接: https://arxiv.org/abs/2609.37323
作者: Arman Mielke,Uwe Bauknecht,Thilo Strauss,Mathias Niepert
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Discrete diffusion models have emerged as a powerful paradigm for solving combinatorial optimization (CO) problems on graphs by learning to sample high-quality solutions. A common inference-time approach is to generate multiple candidate solutions independently and return the best-performing sample, improving solution quality at the expense of an increase in computational cost. In this work, we introduce PT-Denoise, an inference-time procedure that allows these concurrent denoising trajectories to interact through parallel tempering, without requiring retraining or fine-tuning of the underlying denoiser. Our method assigns a temperature to each diffusion process and allows processes to swap temperatures based on their relative performance. This dynamically reallocates promising, low-energy trajectories to colder, more concentrated sampling regimes while allowing higher-energy states to escape local minima through randomized exploration. Experiments on canonical graph-structured CO problems show that our approach consistently improves the quality of the best solution found, while only adding minimal computational overhead.

[LG-65] AutoMark: Enabling Autoresearch to Discover Better LLM Watermarks

链接: https://arxiv.org/abs/2609.37310
作者: Thibaud Gloaguen,Robin Staab,Martin Vechev
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注: 52 pages

点击查看摘要

Abstract:With LLM watermarking being deployed commercially and now required by regulations, improving its reliability and effectiveness has become crucial. Yet, recent progress in the field of LLM watermarking has increasingly been driven by improving details of existing methods, an effort fundamentally limited by the pace of human researchers. In this work, we enable for the first time the autonomous discovery of new distortion-free state-of-the-art watermarking schemes. To enable this, we (i) establish strict criteria to ensure that watermarks are reliable (e.g., they do not have an unexpectedly high false positive rate), (ii) propose rigorous statistical tests to automatically evaluate whether a watermarking scheme satisfies our criteria, and (iii) design an evaluation suite to rank watermarks along three key dimensions: detectability, quality, and robustness. By running our framework with 3 frontier models (GPT-6 Astra, Opus 5, Gemini-3.8 Flash), we discover over 50 different watermarking schemes, including several that outperform prior works along all key dimensions. We complement this by a manual study of the discovered schemes, distilling the key ideas into smaller components, and individually studying the impact of each component across dimensions (detectability, quality, robustness) to better understand how the proposed schemes operate. Importantly, we find that the agents, on top of improving existing ideas, also discover fundamentally new ideas (e.g., aligning watermark scores with random per-request direction). Overall, our work establishes the first steps of fully autonomous watermarking research, enabling the discovery of more reliable and effective watermarks. Our code is available at this https URL, and a blogpost to visualize our results at this https URL.

[LG-66] Hybrid Joint-Selective Optimization: Reduced-Space Levenberg-Marquardt Refinement of Low-Dimensional Parameters of Interest

链接: https://arxiv.org/abs/2609.37308
作者: Muhammad Luthfi Shahab,Gabriella Alfa Indahsari,Imam Mukhlash,Hadi Susanto
类目: Numerical Analysis (math.NA); Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE); Optimization and Control (math.OC)
*备注:

点击查看摘要

Abstract:This paper introduces a hybrid joint-selective optimization (HJSO) framework for large-scale numerical problems in which a small subset of trainable quantities is of primary interest. We partition the full parameter vector into a high-dimensional remaining block and a low-dimensional block of parameters of interest (POIs), perform joint first-order optimization over the full parameter set, and then freeze the remaining variables while applying a reduced-space Levenberg-Marquardt (LM) refinement to the POIs. The method is designed for settings in which the POIs are low-dimensional but strongly influence the quality of the computed solution, while the full parameter space remains too large for full-space second-order methods. The framework is evaluated on three representative problems: a matrix eigenvalue problem, an inverse Bratu problem solved with a physics-informed neural network, and a 100-dimensional nonlinear Black-Scholes problem solved with the DeepBSDE method. In each test, HJSO reaches prescribed POI-error thresholds faster than the corresponding joint first-order baseline and improves the final POI accuracy for the reported solver configurations. The contribution is therefore not a universal optimizer, but a practical reduced-space strategy for problems with known low-dimensional parameters of interest and expensive high-dimensional training variables.

[LG-67] SafeLLM 4SE: Statistical Evaluation and Reporting for LLM -based Software Engineering Systems

链接: https://arxiv.org/abs/2609.37294
作者: Francisco Ortin
类目: oftware Engineering (cs.SE); Machine Learning (cs.LG)
*备注: This work has been submitted for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible

点击查看摘要

Abstract:Large language models (LLMs) are increasingly used for software engineering tasks, yet their stochastic behavior challenges the validity, reproducibility, and comparability of their evaluations. Conventional practices such as reporting a single output, an average score, best-of-N, or pass@k performance can obscure variability and estimation uncertainty, potentially leading to misleading conclusions about system reliability. This article presents SafeLLM4SE, a practical methodology and reporting standard for statistically principled evaluation of LLM-based software engineering systems. Rather than treating generated outputs as deterministic artifacts, SafeLLM4SE treats them as realizations of a stochastic process and distinguishes quality, stability, and estimation uncertainty. It combines adaptive sampling with confidence intervals, distribution-aware statistical comparisons, effect sizes, and a minimum reporting standard covering model configuration, reproducibility, evaluation procedures, and resource usage. SafeLLM4SE is also provided as an open-source software package available on PyPI, enabling researchers and practitioners to reproduce and extend the methodology. We illustrate its application by comparing two LLMs on HumanEval, a benchmark of programming problems assessed through functional tests.

[LG-68] Efficiently Approximating Attention Is Hard

链接: https://arxiv.org/abs/2609.37261
作者: Lukas Haverbeck,Carmen Amo Alonso,Andres Felipe Posada-Moreno,Sebastian Trimpe,Marco Pavone
类目: Machine Learning (cs.LG); Computational Complexity (cs.CC)
*备注:

点击查看摘要

Abstract:Softmax attention is ubiquitous in modern machine learning, but its quadratic scaling with sequence length makes it costly. To reduce this cost, attention is often approximated with fast algorithms, which incur error but can still perform well in practice and on some inputs. At the same time, the growing diversity of attention applications makes approximation guarantees that do not depend on particular input structure a compelling target. For such uniform guarantees over all inputs, known runtime lower bounds rule out fast algorithms for near-exact attention, but leave open the practically important regime: is there an efficient algorithm with even a modest uniform approximation guarantee? We answer this question negatively. Under standard complexity-theoretic assumptions, no truly subquadratic algorithm can approximate attention with any nontrivial additive or relative guarantee uniformly over all inputs. This impossibility holds in the mildest parameter regime for which known algorithms do not already achieve strong approximation guarantees in near-linear time, and extends to practically relevant relaxations: even after polynomial preprocessing of the KV cache, no efficient algorithm can obtain a nontrivial uniform approximation guarantee, or identify a small set of keys receiving substantial attention under sparsity. Overall, our results settle the computational limits of uniform attention approximation.

[LG-69] rident: Unifying Guarded Dispatch and Host Execution for PyTorch Triton Workloads

链接: https://arxiv.org/abs/2609.37241
作者: Jinjie Liu,Xiaoyan Liu,Shuhan Zhang,Wenjia Sun,Ruilin Yang,Chunlei Men,Yonghua Lin,Shaohua Li
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:User-written Triton kernels enable high-performance GPU computation within PyTorch, but their end-to-end latency can remain dominated by host-side orchestration, especially when device execution is short. Although this http URL can generate native host wrappers for captured graphs, each invocation still passes through runtime-managed specialization lookup, guard evaluation, and preparation before reaching the wrapper. We present Trident, a compiler backend that removes this recurring overhead from the specialization cache-hit path. Trident introduces the Specialization Cache Module (SCM), which compiles guarded specialization selection, argument and execution-environment preparation, and host execution for multiple specializations into a single executable module. An invocation enters the SCM once, remains in compiled code when a specialization matches, and returns to Python only when a new specialization must be compiled. Built on Torch-MLIR, Trident lowers guards and host-side orchestration to native code while retaining calls to optimized runtime implementations of supported ATen operators. Our evalu- ation on two LLMs shows that Trident achieves up to a 1.47x speedup in model-level end-to-end latency over eager execution and up to 1.68x over this http URL.

[LG-70] Differentiating Bisimulation Metrics: A Framework for Parametric Markov Chain Fitting via Bicausal Optimal Transport

链接: https://arxiv.org/abs/2609.37239
作者: Sergio Calo,Amy Zhang,Javier Segovia-Aguas,Anders Jonsson
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Many problems in sequential decision-making, such as imitation learning from observations, state-space compression, world-model learning, and sim-to-real transfer, can be reduced to learning a model such that a notion of distance with respect to the target process is minimized. We consider this general framework and consider the bisimulation metric, equivalently Bicausal Optimal Transport (BOT), as the notion of distance to minimize. We show that BOT, since it can be formulated as a linear program (LP), is differentiable with respect to the model dynamics. We then derive an exact closed-form gradient via the envelope theorem applied to the LP saddle point. The result is a general algorithm, Differentiable Bicausal Optimal Transport (D-BOT), that can be applied to each of the problems above. The proposed algorithm learns the best model by alternating between distance computation and gradient steps. We apply D-BOT for three different settings: state-space compression, parametric model learning, and imitation learning from observations (ILfO). We show empirical results that confirm the viability of all three instantiations.

[LG-71] Interacting particle guidance for sampling reward-tilted generative priors

链接: https://arxiv.org/abs/2609.37227
作者: Adhithyan Kalaivanan,Zheng Zhao,Jens Sjölund,Fredrik Lindsten
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: Code will be made available after publication

点击查看摘要

Abstract:Inference-time steering adapts pretrained diffusion and flow-based models to new tasks, e.g., to generate samples from a conditional distribution or samples with desired properties, without retraining. This can be formalized as sampling from a reward-tilted generative prior. As exact sampling from this distribution is intractable, guidance-based methods rely on approximations producing biased samples, and sequential Monte Carlo (SMC) methods correct for this bias using importance weights. However, while exact in the large particle limit, SMC suffers from weight degeneracy and particle collapse in practice. We propose interacting particle guidance (IPG), which replaces reweighting with transport. The particles interact through an additional drift, derived from the Feynman–Kac PDE to cancel the reweighting term, and remain unweighted. Choosing the drift in a reproducing kernel Hilbert space yields a closed-form solution that is cheap to compute, with negligible overhead compared to SMC. We demonstrate the method on Gaussian mixtures with known posteriors, and on high-dimensional image inpainting and protein structure inference tasks.

[LG-72] Explainable Machine Learning for Multilayer Planar Winding Inductance Estimation

链接: https://arxiv.org/abs/2609.37211
作者: Spyros Rigas,Theofilos Papadopoulos,Georgios Alexandridis,Antonios Antonopoulos
类目: Machine Learning (cs.LG)
*备注: 10 pages

点击查看摘要

Abstract:Rapid and accurate self-inductance estimation for multilayer rectangle-shaped planar windings is essential for modern high-frequency power converters, yet traditional workflows rely on complex mathematical equations, rigid monomial formulas or unexplainable black-box machine learning (ML) models that degrade severely outside their training domain. This paper introduces an explainable ML framework unifying post-hoc feature attribution (SHAP and permutation importance) with Kolmogorov-Arnold Network-guided symbolic regression via the SR-KAN framework to discover closed-form analytical equations without prior structural assumptions. Evaluated on a new open-source dataset of over 10,000 Finite Element Analysis (FEA) simulations across seven out-of-distribution (OOD) classes, standard tree-based ensembles exhibit severe extrapolation errors ( 36%), whereas the unconstrained SR-KAN expression achieves a robust OOD relative error of 8.22%. Experimental verification across 55 physical printed circuit board prototypes (up to 8 layers, with inductances from 4.11 \muH to 559.27 \muH) confirms that the KAN-discovered expression translates effectively to real-world hardware, predicting inductance with a mean absolute relative error of 6.26%. To support reproducible research, the complete FEA simulation dataset and prototype measurements are released open-source.

[LG-73] Pointwise or Pairwise: When Do Pairwise Losses Help Reward Learning Provably?

链接: https://arxiv.org/abs/2609.37209
作者: Junghyun Lee,Minsoo Ha,Sanghwa Kim,Yeongjong Kim,Eunjee Lee,Seiyun Shin,Kwang-Sung Jun
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 58 pages, 4 figures, 3 tables; second through sixth authors contributed equally

点击查看摘要

Abstract:Pairwise losses are increasingly used for reward learning even when pointwise rewards are observed, with mixed empirical results. When and why do pairwise losses outperform pointwise losses? We study this question in a grouped offline contextual-bandit setting allowing multiple actions per context, capturing many reward learning scenarios. We compare Value Regression (VR), which regresses observed rewards pointwise, with Value Difference Regression (VDR), which regresses reward differences between a pair of actions sampled under the same context. We consider a semiparametric model where the mean reward is the sum of a learnable action-dependent component and an arbitrary context-dependent yet action-independent nuisance, capturing context-specific disturbances. Using a unified localized analysis, we prove finite-sample regression guarantees for finite and linear function classes and translate them into offline-regret bounds. For finite classes, VDR eliminates the misspecification term in the VR bound and improves a reward-scale-dependent error term by averaging over actions within each context, a benefit absent from the corresponding VR term. For linear classes, neither method uniformly dominates: within-context differencing removes nuisance-induced bias but may increase estimation variance relative to using absolute rewards when the misspecification is sufficiently low. This yields a feature geometry-dependent bias-variance tradeoff, which we corroborate with numerical experiments.

[LG-74] Corruption-Robust Sparse Linear Contextual Bandits with Knapsack Constraints

链接: https://arxiv.org/abs/2609.37189
作者: Yige Wang,Hanyang Li,Yiming Zong,Wanteng Ma,Jiashuo Jiang
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We study sparse linear contextual bandits with knapsack constraints under joint reward and consumption corruption. Consumption corruption creates a challenge beyond corrupted rewards: it affects not only statistical estimates, but also the recorded budget, resource prices, and stopping decisions that govern future allocation. We develop Robust Optimistic Primal–Dual (ROPD), an estimator-modular framework that combines corruption-aware confidence widths with online resource prices and a budget-safety rule. With concrete sparse implementation, ROPD achieves regret against a clean population-LP benchmark of \widetilde O(T^2/3+\Gamma T^1/3) under forced exploration and population-design coverage, and \widetilde O(\sqrt T+\Gamma) under on-policy realized-design coverage, for a supplied valid corruption bound \Gamma under the stated proportional-budget scaling and fixed model/design parameters. When the corruption level is unknown, Shared-Grid adapts confidence radii around common point estimates fitted to a single realized history, incurring explicit initialization and master-comparison costs; its sharper on-policy guarantee additionally requires recommendation coverage. Both methods preserve observed budgets on every realization and bound clean resource violation by cumulative consumption corruption. These results connect corruption-robust sparse estimation with resource accounting, pricing, and stopping in high-dimensional online allocation.

[LG-75] Disentangling Spurious Correlations in Vision-Language-Action Models via Predicting Domain-Invariant Latent Lookahead

链接: https://arxiv.org/abs/2609.37165
作者: Junghyun Kim,Ngseo Kim,ChungWoo Lee,Seoyeon Lee,Woo-Jeong Baek,Adam Zhou,Chip Huyen,Jun-Ki Lee,Gi-Cheon Kang,Byoung-Tak Zhang
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注: Accepted to CoRL 2026. Project website: this https URL

点击查看摘要

Abstract:Vision-Language-Action (VLA) models remain brittle under visual distribution shifts, often relying on spurious correlations tied to domain-specific factors rather than task-relevant structure. We propose Domain-Invariant Latent Lookahead (DILL), a representation-learning framework that mitigates shortcut learning in VLA policies. Our key idea is to supervise policies with domain-invariant future latents learned from domain-transformed trajectory data. A Task-Domain Encoder is trained with contrastive objectives and Gaussian disentanglement regularization to separate task-relevant structure from domain-specific visual variation. The learned encoder then provides future latents for VLA policy learning through lookahead prediction and domain disentanglement, encouraging the policy to focus on task-relevant structure rather than incidental visual factors. Counterfactual task-view evaluations show that DILL reduces shortcut reliance, while LIBERO-Plus evaluations demonstrate improved visual robustness, with 69.1% average success, 11.4 percentage points above the strongest baseline. Real-world manipulation experiments further support DILL’s applicability beyond controlled simulation. Complementary latent-space diagnostics show that these behavioral gains are accompanied by representations that better preserve task-consistent structure while suppressing domain-specific variation. Our project page is available at this https URL.

[LG-76] he Vote Hides the Failure: Aggregation Choice and Noise Robustness in Heart Murmur Detection NEURIPS2026

链接: https://arxiv.org/abs/2609.37161
作者: Nicholaus Dismas Ladislaus,Olatunji Damilare Emmanuel,Samuel Chol Buol
类目: Machine Learning (cs.LG)
*备注: Workshop Short Paper: GlobalSouthAI @ NeurIPS 2026

点击查看摘要

Abstract:Noise robustness in automated phonocardiogram (PCG) murmur detection, and how it is measured, remains underexamined despite growing interest in low-resource screening. We evaluate two independently reimplemented pipelines, Hierarchical Multi-Scale Convolutional Network (HMS-Net)–CNN, and Bidirectional Long Short-Term Memory (BiLSTM)–LSTM, under controlled, multi-severity noise with noise-augmented fine-tuning and held-out generalization testing. Under matched aggregation, the complete BiLSTM pipeline outperforms the complete HMS-Net pipeline across all conditions in accuracy and Weighted Accuracy. A stable aggregate accuracy score can misrepresent what individual predictions show: HMS-Net’s native aggregation degrades under salt-and-pepper noise far less than majority-vote (MV) aggregation at the same severity, a gap reflecting window-level disagreement its native rule absorbs, while BiLSTM’s MV accuracy rises after noise-augmented training even though its individual predictions do not improve. HMS-Net’s training effect is significant under one accuracy metric but not another. Noise-robustness conclusions can depend as much on evaluation choices as on the models themselves.

[LG-77] GLASS: Global Latent Aggregation with Slot-based Set Decoding for Scalable All-Atom Crystal Generation

链接: https://arxiv.org/abs/2609.37158
作者: Hendrik Kraß,Seyed Mohamad Moosavi,Mathias Niepert
类目: Machine Learning (cs.LG); Materials Science (cond-mat.mtrl-sci)
*备注:

点击查看摘要

Abstract:Generative models for crystals enable the discovery of novel structures, but scaling all-atom generation to larger systems such as metal–organic frameworks remains challenging. We connect this difficulty to the correspondence problem of particle-space generation. Even on a single fixed target set, index-free permutation-equivariant particle flows require substantially more training for reliable generation as set size and density increase, under both independent and optimal-transport couplings. To resolve this challenge, we introduce GLASS—Global Latent Aggregation with Slot-based Set Decoding, which encodes structures in a permutation-invariant global latent space and learns their distribution via flow matching. A learned-slot decoder constructs all atoms in parallel, removing atom-wise correspondence from generative transport. On MP20, GLASS is competitive with particle-space models, and flow training can reach the validity of the training data at every structure size. On a QMOF subset, GLASS generates MOFs with up to 150 atoms per unit cell without conditioning on building blocks, topology, or composition, and approaches the structural validity of the training data. On both datasets, flow training exposes a validity–novelty tradeoff, and MOF novelty remains limited by autoencoder generalization on the available data. These results show that separating correspondence assignment from generative transport provides a simple route toward high-validity generation of larger atomistic systems.

[LG-78] Learning the Structure of Triangular Transport Maps

链接: https://arxiv.org/abs/2609.37122
作者: Morten Blørstad,Pekka Parviainen,Berent Ånund Strømnes Lunde
类目: Machine Learning (cs.LG)
*备注: 16 pages, 8 figures, 2 tables

点击查看摘要

Abstract:Triangular transport maps provide a flexible approach to sampling-based probabilistic modeling, including density estimation, generative modeling, and Bayesian inference. They transform an unknown target distribution into a simpler reference through a monotone triangular map. The map structure is defined by a variable ordering and sparsity pattern, which together encode a directed acyclic graph. Map quality can depend strongly on this structure, yet finding a good structure is computationally expensive because each candidate generally requires fitting a different map. A central challenge is therefore to learn density and structure jointly, while keeping computation manageable as dimension grows. We introduce Self-Structuring Transport Maps (SSTM), which learn the map, ordering, and sparsity jointly. We use SoftSort to learn the variable ordering and L_0 gates to learn the sparsity, while preserving a triangular structure. To keep the map scalable, we use a monotone BatchEnsemble that shares one weight matrix across all map components through rank-one adapters. Across synthetic and real data, jointly learning the structure and map gives better density estimates than estimating the structure first. When the structure is identifiable from the density, SSTM matches the density performance of a map fitted with the true structure and outperforms autoregressive flows. On large datasets, SSTM is competitive with autoregressive flows.

[LG-79] Interpretable intrinsic dimension estimation through componentwise calibration of distance and angle

链接: https://arxiv.org/abs/2609.37114
作者: Chih-Hsuan Huang,Chih-Wei Chen,Szu-Chi Chung
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 35 pages, 5 figure, 2 tables

点击查看摘要

Abstract:DANCo (Dimensionality from Angle and Norm Concentration) jointly calibrates nearest-neighbor distance and angular statistics and consistently reaches state-of-the-art accuracy on clean intrinsic-dimension (ID) benchmarks. Practical data, however, introduce neighborhood-relative noise and sample-amplitude heterogeneity that can distort these geometric signals. We reformulate DANCo componentwise, retaining separate distance and angular discrepancy curves so that the source of an estimate can be identified and interpreted. For the distance component, we derive a closed-form Kullback-Leibler divergence for the generic-order ratios of the generalized ratios ID estimator (Gride); when both angular parameters are matched (Full), Gride reduces mean percentage error from 27.7% to 17.6% at noise equal to 40% of typical neighbor spacing on 24 manifolds. For the angular component, two sampling regimes motivate aligning mean direction while retaining concentration matching (Profiled). On a Gaussian scale mixture with generating dimension 70 embedded in 100 dimensions, profiling raises the Minimum Neighbor Distance (MiND) estimate from 22.8 to 66.7 , while removing the known amplitudes restores MiND-Full to 71.9 ; the control thus attributes the Full shortfall to amplitude heterogeneity. On CIFAR-10 and ImageNet, amplitude-reducing normalizations move angular location toward the references and narrow the Full-Profiled gap, an observational counterpart to the controlled mixture. Across four pretrained convolutional neural networks, Gride-Profiled, the two-nearest-neighbor estimator (TWO-NN), and the maximum-likelihood estimator (MLE) exhibit similar rise-and-fall profiles, while Full-Profiled differences identify the layers most sensitive to angular calibration.

[LG-80] Neural Constitutive Learning for Generalized Reaction-Diffusion Systems

链接: https://arxiv.org/abs/2609.37113
作者: Shang-Ke Chen,Yu-Peng Wang,Shih-Hsuan Hung,Wei-Fang Sun,Chao-Shun Zhan,Simon See,Min-Jhe Lu
类目: Machine Learning (cs.LG); Mathematical Physics (math-ph)
*备注:

点击查看摘要

Abstract:Generalized reaction-diffusion systems encompass diverse transport mechanisms and coupled reaction kinetics. A central question for neural PDE solvers is what should be learned so that a common interface can accommodate phase-field and degenerate transport, local reactions, and multispecies coupling. We propose the Neural Constitutive Laws–Mass-Compression-Transport (NCL-MCT) Solver, which learns PDE-specific constitutive responses while retaining temporal evolution in a shared MCT integrator. Transport is represented through mobility and thermodynamic driving force, and reaction through relative reaction rates. These constitutive responses depend on the current density rather than explicitly on the initial condition or elapsed time, motivating their reuse across different initial conditions and time horizons. The same interface supports velocity-data supervision and known-law supervision, neither of which requires time integration during training. When constitutive laws are known, supervision can be evaluated on independently sampled density fields, enabling trajectory-free constitutive learning without generating solution trajectories. Across seven systems, separately trained constitutive modules share the same interface and MCT integrator and achieve relative rollout L^2 errors of 10^-4 to 10^-2 . Tests with unseen initial-condition families and an extended time horizon assess reuse beyond training conditions, while separate experiments demonstrate trajectory-free constitutive learning. These results support constitutive responses as an effective learning target for a shared neural PDE framework.

[LG-81] ZeroDiff: Zero-Shot Time Series Reconstruction via Informed-Prior Diffusion ICML2026

链接: https://arxiv.org/abs/2609.37078
作者: Yingda Fan,Dan Lu,Xiaowei Jia
类目: Machine Learning (cs.LG)
*备注: ICML 2026. Code: this https URL

点击查看摘要

Abstract:Time series modeling increasingly demands high-quality supervision, yet target observations remain scarce - exogenous inputs are broadly available, but target measurements are often unavailable due to cost, infrastructure, or accessibility constraints. Can models trained on observed locations reconstruct target time series where measurements have never been collected? We term this zero-shot time series reconstruction. A naive approach - directly mapping exogenous inputs to targets - can yield predictions at unobserved locations, but without target signals, such models fail to capture the intrinsic dynamics of the target variable, producing overly smooth outputs that underestimate extremes. This reveals systematic errors that call for explicit modeling and calibration. We propose ZeroDiff, which constructs an informed prior from exogenous variables alone, then learns to calibrate reconstruction errors through diffusion - training on observed locations and generalizing to unobserved ones. Experiments across diverse real-world datasets demonstrate significant improvements over existing approaches. Our code is available at this https URL.

[LG-82] UnlearningSoup: Is Repeated Tuning Necessary for Large Language Model Unlearning? NEURIPS2026

链接: https://arxiv.org/abs/2609.37076
作者: Puning Yang,Qizhou Wang,Junchi Yu,Bo Han,Xiuying Chen
类目: Machine Learning (cs.LG)
*备注: NeurIPS 2026 Accepted

点击查看摘要

Abstract:Large language models trained on vast corpora inherently risk memorizing harmful content that may later re-emerge in their outputs. To mitigate this issue, existing unlearning methods typically rely on training-based parameter updates, such as gradient ascent and its variants, to delete targeted content while preserving other knowledge. However, balancing the competing goals of forgetting and retention makes hyperparameter choices for these methods particularly difficult, often requiring repeated tuning to obtain a strong model that still leaves substantial room for improvement and transfers poorly across models and datasets. To address this challenge, we investigate whether unlearning runs exhibit exploitable structure in weight space, and observe that models from different runs still lie in a shared evaluation-performance basin. This suggests that stronger models may be recovered through an unlearning-tailored soup strategy, reducing the need for repeated tuning for further improvement or new settings. Motivated by this, we propose UnlearningSoup, a unified framework that provides two strategies: EfficientSoup uses binary-search-based interpolation to quickly discover a well-performing model in the early stage, where repeated tuning would otherwise make strong model selection costly. PerformanceSoup uses reweighted souping to efficiently unlock the remaining performance potential in the later stage, where repeated tuning becomes increasingly inefficient. Extensive experiments across diverse datasets and models show that UnlearningSoup delivers 2.4x to 3.3x efficiency gains in hyperparameter selection, while consistently improving performance across settings.

[LG-83] RL-PaO: Prediction as Action in Decision Making under Uncertainty

链接: https://arxiv.org/abs/2609.37065
作者: Jiahui Feng,Dafang Zhao,Zheng Chen,Zhengmao Li,Lingwei Zhu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Decision-making under uncertainty often relies on predicted parameters, yet accurate prediction does not necessarily lead to good operational decisions. Aligning prediction with downstream optimization requires learning from the consequences of the decisions those predictions induce. We introduce RL-PaO, a reinforcement learning framework that integrates system formulation, optimization, and decision execution into a single environment. This yields a Markov decision process in which prediction is regarded as action: it shifts the environment to produce subsequent context and reward that explicitly aligns prediction error with realized cost, and learning the optimal policy does not require differentiating through the black-box solver. We evaluate RL-PaO on day-ahead energy scheduling using real historical data. On the test year, RL-PaO achieves the lowest annual cost among the non-oracle baselines, achieving on average 10% cost reduction. Moreover, RL-PaO is capable of further analyses to provide strong interpretability both from the policy evolution perspective and the cost-accuracy trade-off.

[LG-84] vSkipper: Translating Dynamic Layer Skipping into LLM Serving Gains

链接: https://arxiv.org/abs/2609.37062
作者: Wei Da,Yavuz Ferhatosmanoglu,Evangelia Kalyvianaki
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG)
*备注: 25 pages (10-pages main body), 6 figures

点击查看摘要

Abstract:Dynamic layer skipping reduces LLM computation by allowing each token to execute only a subset of the model’s layers. However, existing skippers rely on specialized generation loops and do not integrate with modern serving engines. As a result, fewer executed layers do not necessarily translate into lower serving latency: FlexiDepth skips 8 of Llama-3-8B’s 32 layers on average, yet its standard generation loop decodes 14.6–21.0% more slowly than the base model. We present vSkipper, a virtualization layer that makes dynamic layer skippers pluggable in serving engines while preserving continuous batching, fixed-shape batches, paged KV caching, and captured decode graphs. At each routed layer, vSkipper groups tokens by the skipper’s decision and uses routed execution only when predicted to be profitable. We implement vSkipper in SGLang and evaluate the released FlexiDepth checkpoint against upstream SGLang under identical prompts, arrivals, output lengths, and launch settings. At the knee of upstream’s load curve, vSkipper reduces mean end-to-end latency by 36.8% on GSM8K and 13.6% on BBH. Under saturation, it increases request throughput by 11.3% and 7.4%. Serving adds no statistically resolved quality loss beyond the checkpoint’s own. Across synthetic skip policies, two Qwen3 skippers, and three GPUs, we demonstrate reuse without workload-specific tuning. To our knowledge, vSkipper is the first system to realize serving-efficiency gains from per-token interior layer skipping within a modern LLM serving engine. The code is open-sourced as an SGLang fork at this https URL

[LG-85] Message Passing Does More with Less for In-Context Learning on Graphs

链接: https://arxiv.org/abs/2609.37057
作者: Dooho Lee,Jinmo Lee,Minho Jeong,Kijung Shin,Jaemin Yoo
类目: Machine Learning (cs.LG); Social and Information Networks (cs.SI)
*备注:

点击查看摘要

Abstract:Achieving strong performance with graph neural networks (GNNs) typically requires training and hyperparameter tuning for each dataset, incurring repeated costs and effort. Graph in-context learning (ICL) avoids this by using a single pretrained model to predict unknown node labels directly from labeled context nodes. Existing approaches, however, rely on dense attention across nodes, making inference increasingly expensive as graphs grow. In this work, we present Ephris, a new graph in-context learner built on sparse message passing, scaling linearly with the number of node-feature entries and graph edges. Ephris is pretrained entirely on synthetic graphs generated from structural causal models with diverse graph structures and relational dynamics, exposing the model to varied dependencies among topology, features, and labels. We evaluate Ephris on 51 node-classification datasets against 15 extensively tuned GNNs and existing graph ICL methods under both high- and low-label train/validation/test splits. Across both settings, Ephris ranks first on all four aggregate measures: Elo, improvability, average rank, and accuracy. Its inference cost remains comparable to training a single GNN once, while being over 10 times faster than previous graph ICL models. Together, these results advance the performance-runtime Pareto frontier, demonstrating that strong graph ICL does not require dense attention. Code and model weights are available at this https URL.

[LG-86] Iterative Exact Discrete Guidance for Energy-Based Sampling

链接: https://arxiv.org/abs/2609.37043
作者: Yuwen Qian,Yidong Ouyang,Zhengyan Wan,Hongyuan Zha
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 45 pages, 6 figures. Code and artifacts: this https URL

点击查看摘要

Abstract:Sampling from unnormalized distributions over large discrete state spaces becomes difficult when a multimodal target is far from a tractable reference. We introduce Iterative Exact Discrete Guidance (IEDG), a population-exact, trajectory-wise guidance framework for unnormalized discrete targets. Rather than learn the full reference-to-target correction in one step, IEDG introduces a global Boltzmann tilt along an annealing trajectory. Each stage learns a stage-local posterior correction for an incremental Boltzmann tilt of the current source, while the resulting corrections are accumulated relative to a fixed analytic posterior. At the population optimum, exact stage posteriors recover the correct reverse dynamics, whose exact simulation reproduces the target distribution. IEDG chooses stage increments by relative effective sample size (rESS), which controls Rényi-2 displacement and locally adapts the step size to the thermodynamic geometry of the annealing path. Our stagewise total-variation analysis shows that limited overlap amplifies Bregman fitting error by 1/\sqrt\mathrmrESS , while posterior, simulation, and truncation errors enter additively. IEDG improves all distribution-level errors over the neural baselines on ordered, exactly enumerated Ising 4\times4 , while substantially reducing one-shot errors on Ising/Potts 16\times16 across thermodynamic regimes and attaining the best neural-sampler result on several reported local-statistic and phase-coverage metrics. On Max-Cut, its best-of-512 and average-sample ratios exceed all the baselines. Code and artifacts are available at this https URL.

[LG-87] High-Resolution Dynamic Functional Connectivity Generation with Graph-Variate Flow Matching

链接: https://arxiv.org/abs/2609.37037
作者: Om Roy,Yashar Moshfeghi,Keith Malcolm Smith
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:High-resolution dynamic functional connectivity (DFC) can reveal rapidly evolving brain-network interactions, but short temporal windows yield noisy, often low-rank covariance estimates. Graph-Variate Dynamic (GVD) connectivity addresses this by modulating fast instantaneous interactions with stable trial-level support. This suppresses spurious fluctuations and emphasizes persistent, informative connections. We show that the Hadamard construction lifts low-rank instantaneous connectivity from the positive-semidefinite to the positive-definite cone, keeping high-resolution trajectories on the SPD manifold without ridge regularisation or post-hoc projection. We introduce GVD-CFM, a class-conditional generative model for high-resolution dynamic connectivity. Each trial is represented as SPD GVD matrices on a product Riemannian manifold, then mapped through a global log-Euclidean diffeomorphism and an invertible temporal DCT basis. A Transformer-based conditional flow models all spectral modes jointly and generates the full trajectory non-autoregressively in Euclidean coordinates while preserving exact correspondence with valid SPD sequences. Retaining the full DCT basis also enables decoding on denser temporal grids without retraining. Across multiple EEG motor-imagery datasets, GVD-CFM delivers the strongest overall results for held-out distributional fidelity, temporal-dynamics preservation, and synthetic-to-real classification. It also remains computationally efficient relative to strong raw-signal and direct GVD-space generative baselines. GVD-CFM therefore provides a practical framework for realistic, temporally coherent, high-resolution brain-network generation with preserved manifold structure and resolution-flexible decoding from a single trained model. Subjects: Machine Learning (cs.LG) Cite as: arXiv:2609.37037 [cs.LG] (or arXiv:2609.37037v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2609.37037 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-88] Equally Good Yet Different: Benchmarking Rashomon sets in AutoML packages

链接: https://arxiv.org/abs/2609.36970
作者: Katarzyna Woźnica,Katarzyna Rogalska,Zuzanna Sieńko,Mustafa Cavus
类目: Machine Learning (cs.LG)
*备注: Accepted to the International Conference on Automated Machine Learning 2026, ABCD Track

点击查看摘要

Abstract:The Rashomon effect describes the existence of multiple near-optimal models that achieve comparable performance while offering fundamentally different explanations. This creates a critical vulnerability in AutoML: x-hacking, the selective post-hoc choice of a model based on its explanation rather than predictive merit. No existing AutoML framework exposes this risk. We introduce ARSA ML, an open-source Python framework that quantifies Rashomon set structure and predictive multiplicity within AutoML pipelines. Using ARSA ML, we benchmark AutoGluon and H2O across 28 binary classification datasets, and conduct a post-hoc x-hacking analysis revealing a consistent structural asymmetry: AutoGluon produces larger, diverse sets with stable explanations, while H2O generates compact sets with markedly higher prediction divergence and explanation instability – making H2O users considerably more exposed to x-hacking. This gap persists across all evaluated metrics and epsilon thresholds, pointing to a fundamental difference in each framework’s model-building strategy. ARSA ML is available at this https URL .

[LG-89] askBridge: Bridging Unsupervised Tabular Anomaly Detection and In-Context Learning via Virtual Tasks

链接: https://arxiv.org/abs/2609.36968
作者: Doyun Choi,Dooho Lee,Jaemin Yoo
类目: Machine Learning (cs.LG)
*备注: 34 pages

点击查看摘要

Abstract:Unsupervised tabular anomaly detection (TAD) aims to identify anomalous rows in tabular data using normal training samples. While conventional methods rely on dataset-specific training and configuration search, recent tabular foundation models (TFMs) enable zero-shot anomaly detection on unseen datasets via in-context learning. Most TFM-based approaches, however, require anomaly-specific pretraining from scratch, making detection inherently dependent on synthetic TAD-specific priors and costly to update. Some approaches instead repurpose pretrained general-purpose TFMs for TAD to avoid this burden, but rely on computationally expensive formulations with restrictive anomaly inductive biases. In this work, we introduce TaskBridge, a new framework that efficiently repurposes pretrained general-purpose TFMs for unsupervised TAD by constructing virtual supervised tasks that directly recast anomaly detection as supervised in-context inference of TFMs. The resulting virtual tasks induce predictive structures under which normal queries and their target pairs receive high support, whereas anomalies tend to violate the induced structures and receive lower support, providing direct anomaly evidence. Across 790 real-world datasets, TaskBridge consistently outperforms 30 baselines, including state-of-the-art TFM-based approaches, without anomaly-specific TFM pretraining or dataset-specific model optimization.

[LG-90] Scalable Diffusion SBI for Compositional Inference under Simulator Misspecification

链接: https://arxiv.org/abs/2609.36950
作者: Vincent D. Zaballa,Elliot E. Hui
类目: Machine Learning (cs.LG); Quantitative Methods (q-bio.QM); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Simulation-based inference is challenging when many heterogeneous observations must be composed, hierarchical latent structure must be preserved, and the simulator is misspecified relative to observed data. We develop sampling and fine-tuning methods for diffusion-based inference in design-conditional settings, where the same simulator is queried across different experimental conditions \xi . We extend compositional score-based inference with a continuous-time diffusion coefficient that accounts for the number of observations, avoiding Jacobian and auxiliary-covariance corrections. We introduce Hierarchical Blockwise Diffusion Sampling (HBDS), which infers shared parameters and group-specific latent states using a single pretrained model, with the hierarchy specified only at sampling time. Together, these methods support variable observation sets and groupings without retraining. To address misspecification, we introduce path-regularized fine-tuning that adapts the learned likelihood to observations and transfers corrections to posterior inference. Using Girsanov’s theorem, we quantify path divergence between pretrained and fine-tuned models across experimental designs and interpret it alongside predictive errors to distinguish candidate misspecification correction from unnecessary adaptation. We evaluate compositional sampling on exact-score Gaussian and Simple Likelihood, Complex Posterior benchmarks, HBDS with analytic and learned scores on a controlled hierarchical model, and fine-tuning and localization on a separate analytic model with known design-dependent discrepancy. Finally, we apply the framework to 940 measurements across four cell lines in a mechanistic Bone Morphogenetic Protein signaling model, where fine-tuning improves posterior-predictive accuracy relative to the pretrained model and shifts posterior marginals toward the least-squares reference while retaining spread.

[LG-91] Fine-Tuning on Self-Generated and Reward-Weighted Data: Learning Dynamics Convergence Rates and Benefits of Off-Policyness

链接: https://arxiv.org/abs/2609.36945
作者: Zhiwei Wang,Yanxi Chen,Yaliang Li,Bolin Ding
类目: Machine Learning (cs.LG); Optimization and Control (math.OC)
*备注:

点击查看摘要

Abstract:We study the learning dynamics of fine-tuning a policy model on self-generated and reward-weighted data, with particular focus on a generalized version of REINFORCE – referred to as RE(S) – that updates the rollout distribution once every S \ge 1 gradient steps. Prior work in bandits and reinforcement learning has developed rich theory for policy gradient methods, and on-policy sampling (i.e., a small S , ideally 1 ) is often viewed as crucial to their success; yet in prominent application like post-training large language models, reward-guided self-training has proved to be effective even when the rollout distribution is updated infrequently, but theoretical understanding remains limited for the convergence properties of these off-policy methods. To bridge these gaps, we develop a unified theory for RE(S) that covers the full spectrum of S \ge 1 : it can be interpreted as a stage-wise optimization process, where each stage takes S gradient steps for minimizing the Kullback-Leibler distance to a fixed reward-weighted rollout distribution. For multi-arm bandits with softmax policies, our in-depth analysis and numerical experiments reveal three key findings: (1) for any fixed S , RE(S) enjoys global convergence to the optimal policy as the number of rollout distribution updates B = \lfloor T / S \rfloor \rightarrow \infty , where T denotes the number of gradient steps; (2) we prove tight two-sided bounds showing that the suboptimality gap of RE(S) achieves an asymptotic \Theta(1 / T) convergence rate, while S only affects the length of a burn-in phase; (3) when initialized at a weak policy with a small optimal-action probability, RE(1) gets trapped around suboptimal policies for a long period, whereas RE(S) with a suitable S avoids the detour and achieves significantly faster convergence to the global optimum, highlighting the benefits of off-policyness in this case.

[LG-92] Variational Mixtures and Multi-Marginal Flow Matching: Advancing Statistical Inference with Biological Applications

链接: https://arxiv.org/abs/2609.36911
作者: Oskar Kviman
类目: Machine Learning (cs.LG)
*备注: PhD Thesis

点击查看摘要

Abstract:In this thesis I develop methods for statistical inference when the distributions arising from complex biological systems are multi-modal, geometrically structured, and sometimes only defined up to a normalizing constant. I start from variational inference and, when analytic update equations are unavailable, move to black-box variational inference. To build intuition regarding inference challenges and the proposed methodologies, I introduce a novel unnormalized target density (the CoLN distribution) and reuse it as a controlled test case in the kappa. I then trace a trajectory of increasingly expressive approximations: ensembles evaluated with the multiple importance sampling ELBO (Paper A) and variational mixtures that automate component cooperation and exploration (Paper B). Because expressivity comes at a cost, I develop efficient mixture learning ideas, including Monte Carlo objective estimators to scale mixture learning more efficiently (Paper C). As a new result in the kappa, I overturn a three decades long misconception regarding the potential performance benefits of using mixtures in variational inference. Finally, I move from variational inference to flow matching, where I address the need for specialized treatment of interpolant learning in multi-marginal settings (Paper D). By combining insights from Papers A-D, I derive in Section 5.5 a new method: multi-marginal flow matching with mixtures of variational interpolants. I connect these methodological developments to biological applications, with special emphasis on three-dimensional spatial transcriptomics, where stacked tissue slices induce multi-modal dynamics across space.

[LG-93] SINO: Scale-Invariant Neural Operator

链接: https://arxiv.org/abs/2609.36890
作者: Kaichen Ouyang,Chenglei Yu,Chuanrui Wang,Tailin Wu
类目: Machine Learning (cs.LG)
*备注: 38 pages, 9 figures

点击查看摘要

Abstract:In scientific machine learning, physical fields governed by partial differential equations exhibit low-rank structure and scale invariance. When solving equations on coarse grids, missing information leads to the closure problem: modeling unresolved physics to recover lost dynamics. Although closure terms depend on grid resolution, they represent scale-invariant physical laws. A model truly learning physics should capture these mechanisms with low-rank parameterization rather than memorizing grid-specific patterns. Inspired by this, we propose the Scale-Invariant Neural Operator (SINO), which learns on normalized physical scales via a dual-branch architecture operating in spectral and spatial domains. SINO uses bottleneck MLPs to generate continuous convolution kernels, embedding an explicit low-rank inductive bias that concentrates more than 95 percent of variance in 2-3 modes, as validated by PCA across benchmarks, while drastically reducing parameters. This principled design yields 38 times steeper scaling law exponents than FNO, demonstrating superior parameter efficiency. We compare SINO with traditional models (U-Net, DeepONet), Transformer models (Transolver, Oformer, GK-Transformer), and frequency-domain models (FNO, AMFNO, UFNO) on closure problems spanning externally forced Burgers turbulence, decaying Burgers turbulence, KS turbulence, Kolmogorov-forced NS turbulence, and decaying NS turbulence. Experiments show SINO achieves 1.5-38 times error reduction and 2-23 times parameter efficiency over baselines, with superior scaling laws reflecting exceptional data efficiency from principled low-rank design. Code is available at this https URL.

[LG-94] Architecture Alignment With Sparse Priors in Tabular Foundation Models

链接: https://arxiv.org/abs/2609.36883
作者: Tianqi Zhao,Tianyi Zhuang,Shuo Duan,Guanyang Wang,Yan Shuo Tan,Qiong Zhang
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Tabular foundation models (TFMs) are increasingly popular because they deliver strong predictions on new datasets through in-context learning, without task-specific training or extensive tuning. Yet released TFMs differ simultaneously in their pretraining priors, architectures, and objectives, obscuring their respective inductive biases. We therefore examine one concrete capability: irrelevant-feature suppression. Across synthetic tasks and real-world datasets, adding null features causes substantially greater predictive degradation in the row-token model TabDPT, whereas the cell-token alternating-axis model TabPFN v2 and other TFMs remain comparatively stable. This gap motivates us to ask whether architecture contributes to irrelevant-feature suppression. Because released TFMs remain confounded by other design choices, we train streamlined row-token and alternating-axis transformers under identical sparse-to-dense linear priors. Exact Bayes analysis shows that sparse prediction requires context-dependent feature gating, whereas the dense endpoint requires only uniform feature weighting. Consistent with this distinction, the alternating-axis model is substantially closer to the Bayesian optimal predictor on sparse tasks, while the architecture gap becomes negligible on dense tasks; almost all of the sparse gap arises from linear coefficient-estimation error. Finally, in both the controlled model and frozen TabPFN v2, we examine the effect of interventions on the feature-attention outputs on the linear coefficients, finding evidence of task-dependent selective routing of computation through feature-indexed pathways. Together, these results support architecture-prior alignment: preserving an addressable feature axis provides an inductive bias for task-adaptive relevance inference. Code is available at this https URL.

[LG-95] What You Observe Determines How You Identify Causal Effects: Evaluating Causal Models across Observational Views

链接: https://arxiv.org/abs/2609.36881
作者: Heejin Jung,Gyeongdeok Seo,Hoyoon Byun,Joseph Lee,Kyungwoo Song
类目: Machine Learning (cs.LG)
*备注: 9 pages, 8 figures

点击查看摘要

Abstract:Causal foundation models (CFMs) pre-trained on data generated from various structural causal models (SCMs) have been proposed for estimating causal effects from observational data. However, differences in pre-training environments and evaluation protocols make it difficult to assess how their performance depends on the information available for causal identification. To enable controlled comparisons, we introduce CausalIDView, a multi-view benchmark that holds fixed SCM realization and target estimand while varying only the observational view available to the estimator. Each observational view corresponds to a distinct identification regime under the benchmark’s maintained causal assumptions. Across these matched views, no CFM consistently performs best and model rankings vary substantially. Under controlled structural changes, CFMs exhibit model-specific failures to maintain stable estimates when true effects are unchanged and to track genuine effect changes. We also examine whether combining explicit identification with strong predictive estimation is effective. A modular approach that pairs a predictive tabular foundation model with regime-specific identification procedures is competitive with CFMs and outperforms several of them. These findings motivate cross-regime comparisons to assess the empirical value of CFMs.

[LG-96] Seeing Time: Visual-Temporal Representation Learning for Interpretable Time Series Clustering ICDM2026

链接: https://arxiv.org/abs/2609.36873
作者: Zheng Zhu,Zexi Tan,Yuming Deng,Yiqun Zhang
类目: Machine Learning (cs.LG)
*备注: Accepted by ICDM 2026

点击查看摘要

Abstract:Multivariate Time Series (MTS) clustering is an important tool in temporal data mining, aiming to discover latent group structures from complex observations without supervision. Although existing deep clustering methods can learn discriminative temporal representations, the resulting latent clusters are often difficult to relate back to waveform characteristics that practitioners can directly inspect and compare, limiting their ability to assess whether the discovered patterns reflect meaningful temporal behaviors. This paper, therefore, proposes WAVE (Waveform Aligned Visual-temporal Embedding), which treats time series and their deterministically rendered waveform plots as complementary views of the same observations. To produce discriminative representations whose cluster structures can be traced to observable waveform characteristics, WAVE aligns and integrates fine-grained temporal variations with holistic visual patterns, while associating each discovered cluster with its centroid-nearest authentic sample. Accordingly, interpretability in this work specifically refers to waveform-level traceability rather than a general explanation of model decisions. Extensive evaluations across 10 real-world public datasets show that WAVE achieves the highest macro-averaged clustering performance and the best average rank among the compared methods, while qualitative case studies illustrate how the discovered clusters can be inspected through authentic waveform records. The source code is available at this https URL.

[LG-97] Where the Model Changes Its Mind: Hindsight-Divergence Localization for Efficient Reinforcement Learning with Verifiable Rewards

链接: https://arxiv.org/abs/2609.36864
作者: Fanchao Chen,Hengyu Fu,Shivaram Venkataraman,Jiantao Jiao
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Group-relative methods for reinforcement learning with verifiable rewards (RLVR) learn from differences in rollout outcomes. Independently sampling complete trajectories is costly and does not explicitly explore the decision space at critical positions. Feedback on a completed trajectory can reveal which earlier choices the policy reconsiders, suggesting where to sample alternative continuations. We introduce Hindsight-Divergence Localization (HDL), which uses hindsight-induced changes in token log-likelihoods to select branch points. HDL generates a small number of complete root trajectories and fills each training group with continuations from the selected positions under the original task context. Each continuation reuses its root prefix and contributes policy updates only through its newly generated suffix, reducing generation cost while focusing additional exploration and learning on decisions after branching. Experiments with three models across math, code, and agent tasks show gains in both rollout efficiency and task performance. Compared with GRPO at matched group sizes and training steps, HDL yields up to a 2.5 \times reduction in generated tokens and a 1.8 \times speedup in rollout wall-clock time. Despite this reduced generation budget, HDL improves performance across all three domains, with gains of up to 12.5 points on agent tasks.

[LG-98] Markovian Nonconvex ADMM for Reinforcement Learning: Bellm an-Resolvent Stability Beyond Smooth Blocks

链接: https://arxiv.org/abs/2609.36859
作者: Zhaojun Peng
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We identify and study a structural mechanism for Markovian nonconvex ADMM in reinforcement learning. Using finite discounted MDPs as a canonical proving ground, we show that the discounted Bellman resolvent (I-\gamma P_\pi)^-1 can provide the multiplier stability that classical nonconvex ADMM analyses often obtain from a designated smooth block. Starting from this mechanism, we establish convergence under controlled Markov sampling and then under stochastic observations using an empirical Bellman surrogate that jointly represents the random residual and its Jacobian. Markov mixing, initialization drift, observation noise, and decaying bias enter as one operator perturbation, avoiding unbiased product and double sampling requirements. When the perturbations are square summable, the true KKT residual converges almost surely to zero. Under a finite conditional fourth moment condition, a companion iterate satisfies \mathbbE[\widetilde G_K+1] \le A/T+(B/T)\sum_kTm_k^-1, which becomes O(T^-1+T/N) for total Markov sample budget N , giving O(\epsilon^-1) iteration complexity and O(\epsilon^-2) sample complexity for squared KKT accuracy \epsilon . Beyond stationarity, discounted occupancy coverage yields J^\star-J(\pi)=O(\sqrt G) for direct tabular policies, so covered exact KKT points are globally optimal, while a statewise quadratic Bellman-improvement condition sharpens the relation to O(G) . Finally, nonlinear policy, projected Bellman, and explicit occupancy formulations exhibit the same chain of operator invertibility, dual representation, and multiplier stability. This supports discounted operator invertibility as a reusable structural principle for primal-dual reinforcement learning.

[LG-99] An Effective Reliable and Robust Framework for Human Activity Recognition Using Wearable Sensors

链接: https://arxiv.org/abs/2609.36848
作者: Nafees Ahmad,Ho-fung Leung,Muhammad Adil Abid,Sadia Shakil
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Human Activity Recognition (HAR) through wearable sensors greatly improves the quality of human life through its multiple applications. For HAR, multi-sensor channel information is vital for optimal performance. Current work states that applying an attention neural network to prioritize discriminatory sensor channels helps the model classify activity more precisely. However, obtaining discriminatory information from multisensory channels is not always trivial, such as when collecting data from older hospitalized patients. In this context, existing HAR methods struggle to classify activities, particularly activities with similar natures. Moreover, HAR models predominantly suffer from overfitting due to the small size of available datasets, which leads to poor performance. Data augmentation (DA) is a viable solution to this problem. However, available DA methods have various drawbacks, including the possibility of being domain-dependent, resulting in distorted models for test sequences. To address these HAR problems, we propose a novel framework, ALAE-TAE-CutMix+, which focuses on two aspects. First, it enhances the latent information across each sensor channel and learns to exploit the relation among multiple latent features and the ongoing activity. Consequently, the discriminatory feature representations of each activity is enriched. Second, a new augmentation strategy is introduced to address the shortcomings of existing multi-sensor channel data augmentation. We then extend the framework to create a further enhanced version, namely ALAE-CIE-TAE-CutMix+, which learns to capture the interactions between the features of each pair of sensor channels. We find that although the first framework performs slightly better than the latter, the latter is nonetheless more reliable and robust. Both frameworks significantly outperform SOTA approaches on the four HAR datasets from diverse domains.

[LG-100] RolloutFaith: Auditing Persistent Internal Interventions in Visual World Model

链接: https://arxiv.org/abs/2609.36843
作者: Junchi Yao,Ziyi Wang,Youling Huang,Lijie Hu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Interpretability methods such as probes, activation patches and learned editors are designed to reveal or modify a model’s current computation. World models pose a harder requirement: because their predictions become inputs to later predictions, a useful internal correction must survive after editing stops. We therefore propose RolloutFaith, a framework that measures semantic improvement both in the prediction produced at intervention time and over later autonomous predictions under fixed events, actions, noise, and information budgets. We evaluate ten fitted editors on three world models across Crafter, Cartpole, and CoinRun. We also use Reference Activation Patching, which replaces a model activation with the paired activation computed from the real observation, to measure the correction available at the chosen interface. This reference intervention improves later predictions in all nine model and task combinations and outperforms the best fitted editor in eight, yet its sustained gain decreases with horizon in five of nine combinations. Current fitted editors recover only limited and inconsistent long term effects. By restoring individual state components to their untouched values, we find that persistent effects travel through the newest generated frame in DIAMOND, recurrent memory in DreamerV3, and both in STORM. These findings suggest that training should reward future consequences. To test this hypothesis, we propose Delayed LoReFT, which optimizes the same low rank intervention through four frozen future transitions and improves sustained intervention effects to some extent.

[LG-101] owards Better Training Signal: Advantage Clipped Policy Optimization

链接: https://arxiv.org/abs/2609.36816
作者: Ruichuan Huang,Jinghan Liu,Congliang Chen
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Reinforcement learning (RL) has become a cornerstone for improving the reasoning capabilities of large language models (LLMs), but the need for on-policy data substantially limits training efficiency. Reusing off-policy data through importance sampling (IS) can improve efficiency but introduce considerable instability. Hence, algorithms such as PPO and GRPO widely adopt IS-ratio clipping to stabilize training. However, training stability and gradient estimate are mainly determined by the product of IS ratio and advantage. To further stabilize training, we propose ACPO, which clips the product of the IS ratio and the advantage, leading to more stable gradient estimates. We also establish a connection between ACPO and gradient clipping in policy mirror descent (PMD), which is a standard technique to stabilize optimization process, and prove the convergence of clipped-PMD under the standard RL setting. Experiments on widely used mathematical reasoning benchmarks show that ACPO consistently outperforms PPO and GRPO in both accuracy and training efficiency, delivering 4-6 percentage points gains on standard math benchmarks, with Qwen3-8B+PPO. Hence, ACPO is a practical and effective alternative to conventional IS-ratio clipping for RL post-training of LLMs.

[LG-102] RESCUE: Repairing Language Model Errors to Sparse Circuits via Reinforcement Learning

链接: https://arxiv.org/abs/2609.36813
作者: Chuanpu Liu,Miao Yu,Yikai Cai,Yuanhe Zhang,Zhenhong Zhou,Li Sun,Zuming Jiang,Yufei Guo
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Large language models (LLMs) exhibit strong general capabilities that mechanistic interpretability has attributed to sparse computational circuits. However, existing circuit studies emphasize preserving functionality or explaining safety, leaving the mechanisms underlying failures across a broader range of tasks largely unexplored. Extending circuit analysis from abilities to errors, we explore the perspective that such failures may likewise arise from erroneous internal computations and that targeted tuning of the corresponding parameters can correct such errors while largely preserving other capabilities. Motivated by this insight, we introduce RESCUE (Reasoning-Error Sparse-Circuit Uncovering and Editing), a framework that localizes error-associated circuits and surgically repairs them for performance enhancement. General tasks typically involve multi-step reasoning and long-form generation, where early deviations can cause prefixes to drift from supervised references, leading SFT-based mask optimization to overlook circuits involved in generation-time errors. RESCUE therefore refines these masks through reinforcement learning with multiple masked-model rollouts, improving their relevance to observed task failures. Finally, RESCUE introduces a pruning technique and precisely fine-tunes error circuits to correct task failures, thereby translating error localization into a sparse and targeted model update. We validate RESCUE on heterogeneous repair sets across two domains: (1) mathematical reasoning, identifying a math error circuit of 1.40% density whose repair raises accuracy from 6.0% to 75.5%; and (2) medical QA, where a similarly compact 1.44% circuit improves repair-set accuracy from 0% to 81%. Our code is available at: this https URL.

[LG-103] EasyPPO: Stabilizing the Critic Is Key

链接: https://arxiv.org/abs/2609.36802
作者: Xuanyi Zhou,Qiuyang Mang,Huanzhi Mao,Dacheng Li,Wenhao Chai,Mayank Mishra,Yichuan Wang,Karthik Narasimhan,Alvin Cheung,Joseph E. Gonzalez
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:A key strength of Proximal Policy Optimization (PPO) is its learned critic, which uses historical trajectories collected during reinforcement learning to estimate expected returns and reduce policy-gradient variance. However, we find that the critic is also a major source of instability in reinforcement learning for large language models (LLMs). We identify two critic failure modes that destabilize PPO. First, filtering truncated rollouts from both actor and critic shifts the policy objective to reward conditioned on completion, allowing truncation to increase even as conditional reward improves. Second, heterogeneous return noise can cause high-variance prompts to dominate critic updates in finite batches. We introduce EasyPPO to address these failures. Actor-only overlong filtering trains the critic on returns from both completed and truncated rollouts. Noise-normalized critic regression weights each prompt’s critic loss by the inverse standard deviation of its sampled returns, balancing noise contributions across prompts. Moderately smaller critic mini-batches confine outlier influence to fewer rollouts during gradient clipping. Across continuous-reward coding on FrontierCS, binary-reward mathematical reasoning on AIME24, and multi-turn search on Search-R1, EasyPPO remains stable throughout the full training horizon and consistently outperforms vanilla PPO, VAPO, and HL-Gauss PPO. Its best validation scores show relative gains of 14.89%, 2.28%, and 9.47% over PPO, respectively.

[LG-104] Where Does Randomness Matter in Neural Cellular Automata?

链接: https://arxiv.org/abs/2609.36797
作者: Fei Zuo,Jiaqi Shi,Yujing Liu
类目: Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE)
*备注: 21 pages, 7 figures

点击查看摘要

Abstract:Stochastic cell updates are often used throughout the life of a neural cellular automaton (NCA), from backpropagation through time to final rollout. This leaves two questions entangled: does update randomness help learn a useful rule, and must that randomness remain at execution? We separate training and evaluation update modes in controlled Growing NCA experiments, then vary the states shown during training. Under the standard constant-rate persist recipe, asynchronous training passes the short-horizon quality test in 10/10 runs, compared with 3/10 synchronous runs. All ten asynchronous models also retain the target for 4,096 steps under deterministic evaluation. For a scalar translation-invariant lattice, we derive an exact mean-square criterion: random masking can damp mean modes, but it also injects variance, and a mean-only test misclassifies four non-marginal settings. Finally, among 30 models that all pass the same reconstruction test, eight of ten grow-trained models become off-target at 4,096 steps, while all persist and regenerate models retain the target; damage recovery separates persist from regenerate. The results distinguish optimization reliability, execution mode, and task-specific behavior instead of treating them as one stability property.

[LG-105] RAE-PPG: Duration-Grounded Retain-and-Extend Pretraining for PPG Foundation Models

链接: https://arxiv.org/abs/2609.36794
作者: Suyeong Lee,Hochang Lee,Seokyong Sheem,Daekyum Kim
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Signal features derived from photoplethysmography (PPG) require different signal durations to characterize. Existing PPG foundation models treat duration as a pretraining or evaluation condition rather than using the different durations required by PPG features to organize self-supervision. We hypothesize that self-supervision should expand with signal duration, allowing a single encoder to progressively acquire additional features while preserving and reusing earlier learning. We introduce Retain-and-Extend PPG (RAE-PPG), which trains a single Transformer encoder successively on 10 s, 30 s, and 240 s inputs, adding supervision for signal features supported by each longer observation. The encoder is partitioned into duration-specific parameter groups, allowing later stages to reuse earlier groups while updating only the group assigned to the current stage. Selected earlier targets are reused to supervise later stages, encouraging the corresponding features to remain accessible in longer-input representations. Direct decoding from the final encoder shows that earlier features remain recoverable from longer-input representations, while later-stage features show higher mean decoding performance at their introduction durations. Controlled comparisons further show that prior-stage learning provides a better basis for learning newly introduced features at both transitions. Across 18 tasks from eight datasets, the final frozen encoder achieves the best observed score on 12 tasks compared with five existing PPG foundation models.

[LG-106] Harnessing Large Language Models to Compile Task-Relevant Context into Bayesian Optimisation

链接: https://arxiv.org/abs/2609.36788
作者: Zhongwei Yu,Sourabh Roy,Bin Cao,Xue Yan,Anjie Liu,Jun Wang
类目: Machine Learning (cs.LG)
*备注: 42 pages, 6 figures

点击查看摘要

Abstract:Incorporating rich task-relevant context, such as domain knowledge and external observations, is a key capability yet remains challenging for Bayesian optimisation (BO). Recently, practitioners have started to use large language models (LLMs) to generate and execute BO programs through coding harnesses. In such emerging practices, the posterior belief is shaped not only by Bayesian inference but also by LLM-generated model and data artefacts, offering a flexible route for task context to enter BO as executable code. To study whether and how LLMs can be harnessed to compile diverse contextual signals for BO, we formulate LLM-compiled BO as generalised-context decision making. We propose HarBO, a BO-specialised harness that compiles generalised context into the core artefacts of standard BO through a validated multi-stage workflow. Our theory analyses the regret under imperfect compilation and the effect of adding new context. Across synthetic functions and real-world benchmarks, we find that LLM harnesses can effectively compile context into standard BO, achieving competitive performance with specialised LLM-embedding-based and direct LLM-in-the-loop BO methods. General coding harnesses can be effective in familiar domains such as hyperparameter optimisation, but fall short in unfamiliar, context-rich domains. Together, these results establish LLM harnesses as a promising, but not automatically reliable, route for making rich task context usable in BO.

[LG-107] When Can Prefixes Compile LoRA? Exact Resource-Capped Tests for Frozen Attention

链接: https://arxiv.org/abs/2609.36766
作者: Joyanta Jyoti Mondal,Ibne Farabi Shihab
类目: Machine Learning (cs.LG)
*备注: 26 pages, 4 figures

点击查看摘要

Abstract:Can a fixed continuous prefix replace a given low-rank adapter while the attention head stays frozen? In this research, we show that the answer depends on the adapter’s target through three conditions. First, observability: at one causal readout, every independent key–value prefix sees the content only through the query, attention partition, and value numerator, so a target that differs on two inputs with equal summaries incurs an error floor at every prefix length; norm caps extend this floor to nearly equal summaries. Second, realizability: at a common query, any prefix reduces exactly to two aggregate variables, and the norm-capped optimum is an attained second-order-cone program, also after a fixed output projection; it places two equal-norm rank-one value updates on opposite sides of compilability. Third, implementation: under affine query exposure, 2r signed slots approximate a rank- r value update, but their values grow as O(\epsilon^-3/2) , and the construction passes all 400 tolerance checks in float64 yet only 38 in bfloat16. A first-layer GPT-2 readout with fixed token and position meets the common-query condition without clamping activations; at three such heads, the capped optimum leaves 18.4% to 74.2% of the projected adapter effect uncompiled, with a head-dependent value–query ordering. All claims concern local approximation at one head, not whole-network equivalence.

[LG-108] Graph-Spectral Flow Matching for Multivariate Time Series Anomaly Detection

链接: https://arxiv.org/abs/2609.36765
作者: Zepeng Zhang,Jhony H. Giraldo,Wenbin Wang,Olga Fink
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Multivariate time series anomaly detection typically relies on evaluating discrepancies between observations and outputs produced by models trained on normal data. An alternative perspective is to characterize the distribution of normal data through the generative dynamics, i.e., the velocity field, of flow matching models. However, standard flow matching typically adopts linear probability paths that overlook dependencies among variables, leading to a misalignment with the structured data distribution. To address this issue, we propose GRASP, a flow matching framework with a graph-spectral path for multivariate time series anomaly detection. GRASP incorporates graph structure into the probability path by minimizing a fixed-endpoint action that combines kinetic energy with graph Dirichlet energy. This formulation yields a closed-form path based on graph-frequency-dependent hyperbolic interpolation. A velocity predictor trained on normal data then detects anomalies using weighted velocity discrepancies aggregated across source samples, flow times, and graph frequencies. Theoretically, we establish that GRASP is invariant to the choice of Laplacian eigenbasis and decompose its expected oracle anomaly score into bounded endpoint uncertainty and graph-frequency-weighted Fisher discrepancy. Experiments on four benchmarks demonstrate the superior anomaly detection performance of GRASP and validate the effectiveness of its graph-spectral path and weighting mechanism.

[LG-109] Federated Clustering with Unknown Local and Global Cluster Cardinalities

链接: https://arxiv.org/abs/2609.36762
作者: Mitushi Goyal,Tarun S.,Riddhanya Senapathi,Arun Raman
类目: Machine Learning (cs.LG); Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Federated clustering methods that do not require the global number of clusters K still assume that each client knows its local number K_g . This assumption is hard to justify when clients know no more about their data than the server does, as in fault diagnosis across independently operated industrial sites. We propose a two-phase framework in which neither count is known: each client first estimates K_g from its own data, and an aggregator that requires local counts, such as FedGEM, then uses these estimates in place of the true values. For the first phase we introduce Adaptive Split–Merge (ASM), which grows a spherical Gaussian mixture by BIC-driven splitting and then merges excess components. ASM uses no labels, selects its hyperparameters on held-out client data only, and makes no assumption about how clusters are shared across clients. We derive a closed-form split criterion whose critical cluster size falls with anisotropy and rises with dimension, and show empirically that over-fragmentation grows with the number of points per cluster, which federation divides among clients. Across eight datasets, ASM with FedGEM attains a mean ARI of 0.333, against 0.256 for the next best label-free estimator and 0.361 when the true local counts are supplied. It also gives the most reliable global estimates of K and is robust when client size is decoupled from local cardinality.

[LG-110] cktFormer: Transformer-Based Approach for Automated Analog Circuit Design

链接: https://arxiv.org/abs/2609.36752
作者: Pasindu Dodampegama,Praveen Wijesinghe,Naveen Basnayake,Keshawa Jayasundara,Tharindu Bandaragoda
类目: Machine Learning (cs.LG); Hardware Architecture (cs.AR)
*备注: 6 pages, 6 figures, IECON 2025 51st Annual Conference of the IEEE Industrial Electronics Society

点击查看摘要

Abstract:Circuit design is a complex and iterative process that requires expertise in electronic engineering. It involves selecting components while meeting performance constraints, such as power efficiency, cost-effectiveness, and signal integrity. However, manual design is time-consuming and prone to errors. Although other stages of the manufacturing pipeline have benefited from AI-driven optimizations, circuit design remains a bottleneck, limiting overall productivity. Generative AI and machine learning offer the potential to automate and improve this stage, boosting efficiency and accuracy. To address this, we introduce a dual transformer architecture that bridges the gap between AI and circuit design by leveraging attention mechanisms to model complex, non-sequential circuit relationships. Our approach structures netlist data into graph-based representations, enabling effective learning of circuit topology and component interactions. The system consists of two interlinked models: a node prediction model that proposes components and an edge prediction model that infers valid connections. This collaborative and decoupled design captures both component-level semantics and global structural coherence. In our experiments, this architecture outperforms recent models such as AnalogGenie and cktGNN in the validity of generated circuits. By addressing key limitations in existing methods, our work advances automation in electronics engineering and contributes a benchmark for AI-driven circuit synthesis.

[LG-111] Efficient Offline Learning of Ranking Policies via Top-k Policy Decomposition CIKM2026

链接: https://arxiv.org/abs/2609.36740
作者: Ren Kishimoto,Koichi Tanaka,Haruka Kiyohara,Yusuke Narita,Yasuo Yamamoto,Nobuyuki Shimizu,Yuta Saito
类目: Machine Learning (cs.LG)
*备注: Published as a conference paper at CIKM 2026

点击查看摘要

Abstract:Many recommender systems such as for e-commerce and news platforms aim to provide users with rankings they are likely to interact with. Off-Policy Learning (OPL) of ranking policies enables us to learn new ranking policies using only historical logged data. However, ranking settings make OPL remarkably challenging because their action spaces consist of permutations of unique items, being extremely large. Existing methods primarily use either policy- or regression-based approaches. The policy-based approach, which typically uses importance-weighted policy gradients, can suffer from high variance due to large action spaces. The regression-based approach, on the other hand, estimates the expected reward using conventional machine learning methods, avoiding variance issues but potentially suffering from severe bias. To circumvent these issues of existing methods, we propose a new OPL method for ranking, named Ranking Policy Optimization via Top- k Policy Decomposition (R-POD), which combines the policy- and regression-based approaches in an effective fashion. Specifically, R-POD decomposes a ranking policy into a first-stage policy for selecting top- k actions and a second-stage policy for choosing the bottom actions given the top- k actions. It learns the first-stage policy using a new policy gradient estimator and the second-stage policy via the regression-based approach. This method can substantially reduce variance, since it applies importance weighting only to the top- k actions. We also demonstrate that our policy-gradient estimator for the first-stage policy is unbiased under a conditional pairwise correctness condition, which only requires that the expected reward differences of pairs of rankings sharing the same top- k actions can be estimated correctly.

[LG-112] Routing in Gradient Space: Balanced Usage Is Not Expert Specialization

链接: https://arxiv.org/abs/2609.36724
作者: Yuchen Li,Mingyu Du,Zongqi Fan,Nguyen H. Tran,Ken-Tye Yong
类目: Machine Learning (cs.LG)
*备注: 68 pages, 6 figures

点击查看摘要

Abstract:Sparse expert models can distribute traffic evenly while still grouping incompatible training signals within the same experts. We study routing as a gradient-partitioning problem and introduce gradient-aligned routing (GAR), whose load-normalized router objective rewards grouping observations with aligned gradients. On five multi-task text-classification mixtures, we compare GAR with task-loss-only routing, gradient-combination and gradient-conflict methods, and load-balancing losses. With a fully trainable RoBERTa backbone and classification-head experts, GAR has the highest aggregate validation accuracy, 1.07 percentage points above task-loss-only routing. With frozen DeBERTa and Qwen3-1.7B backbones and low-rank adapter experts, it again ranks first, 1.10 points above task-loss-only routing, with better-balanced expert load and higher gradient-mass purity, the share of each expert’s gradient-norm mass from its dominant task; the load-balancing losses flatten load further but leave this purity near its task-loss-only level. Top-1 routing, trainable full-parameter feed-forward network (FFN) experts, and a larger backbone also show positive aggregate gains. The results distinguish expert-load balance from gradient-based routing organization and indicate the predictive value of gradient-informed routing in multi-task text classification.

[LG-113] CALIBUDGET: Calibration-Guided Source Allocation for Fixed-Budget Mixed-Reasoning Adaptation

链接: https://arxiv.org/abs/2609.36721
作者: Yupeng Chang,Yuan Wu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Fixed-budget adaptation from heterogeneous data sources requires deciding not only how much data to use, but how much exposure each source receives. Size-proportional rules can crowd out small sources, whereas difficulty-only rules can chase noisy estimates or allocate residual budget to nearly saturated pools. We introduce CALIBUDGET, a floor-protected, reliability-aware integer allocator that treats source exposure as an explicit adaptation variable. From small train-internal calibration splits, it combines model need, post-floor availability, and bootstrap stability, then produces exact capacity-respecting quotas without changing the model, objective, or total budget. In a controlled setting combining mathematical and commonsense data, CALIBUDGET improves CommonAvg, FragileAvg, and MacroAvg over validation-error-with-floor, the strongest matched comparator, in all three paired LLaMA-2-7B LoRA+ runs. The respective mean gains are 0.56, 0.46, and 0.41 percentage points (pp). Overall increases by 0.18 pp, whereas MathAvg decreases by 0.20 pp, exposing a coverage-retention boundary rather than a uniform gain. CALIBUDGET changes only 1.14-1.42% of the source budget but improves performance in 15 of 24 comparisons across commonsense tasks and seeds. These results suggest that small changes in source quotas can matter; example-level selection can then determine which examples fill each quota.

[LG-114] Information-theoretic receding-horizon active learning of nonlinear dynamical systems

链接: https://arxiv.org/abs/2609.36712
作者: Juncal Arbelaiz,Anushri Arora,Jonathan W. Pillow
类目: ystems and Control (eess.SY); Machine Learning (cs.LG)
*备注: 8 pages, 3 figures. Accepted to the 2026 IEEE Conference on Decision and Control (CDC)

点击查看摘要

Abstract:Accurately learning nonlinear dynamics from a finite-duration experiment requires the efficient collection of informative data. We address this challenge for stochastic controlled nonlinear dynamical systems whose state is observed along a single trajectory. Our goal is to reconstruct the unknown controlled state-increment map over a prescribed compact subset of state-input space. We construct a parametric estimator of the map using fixed nonlinear features, so that the model is nonlinear in the state and input, but linear in the unknown parameters. A Gaussian prior over the parameters yields recursive Bayesian posterior updates as data stream in, enabling online quantification of predictive uncertainty in the reconstructed dynamics over the target set. We formulate an optimal adaptive-design problem over an information state, using a prediction-oriented acquisition criterion based on the mean marginal mutual information between candidate future trajectories and the reconstructed dynamics over the target set. We then approximate the resulting adaptive-design problem by a non-myopic receding-horizon formulation, evaluate its remaining expectation using a scenario-based sample average, and solve the resulting deterministic program with the cross-entropy method, leveraging parallel candidate-scenario evaluations. Numerical experiments on a noisy multistable system demonstrate that the proposed adaptive information-seeking strategy reduces predictive uncertainty and reconstruction error more efficiently than common excitation baselines under comparable experimental constraints.

[LG-115] Into the danger zone: stable extrapolation in high-dimensional function and operator learning

链接: https://arxiv.org/abs/2609.36709
作者: Ben Adcock,Simone Brugiapaglia,Xuemeng Wang
类目: Numerical Analysis (math.NA); Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Out-of-distribution (OOD) generalization is a central challenge in scientific machine learning. We study regression problems in which the test distribution differs from the training distribution and ask: under what assumptions on the target function or operator is stable extrapolation possible, and how far beyond the training domain can one extrapolate? Existing theory controls the test error through additive penalties measuring the discrepancy between the training and test distributions. Such guarantees show robustness to small distribution shifts, but can very pessimistic in comparison to OOD performance observed empirically. We identify classes of holomorphic functions and operators for which the OOD generalization error converges at algebraic rates even in the presence of large distribution shifts. This phenomenon stems from the increasing smoothness of higher-index coordinates, leading to what we term a `blessing of high dimensionality’. For learning with either polynomials, deep neural networks or deep neural operators, we derive explicit rates for arbitrary test measures supported on suitable domains and quantify how the admissible domain depends on the underlying regularity of the function or operator. Our extrapolation guarantees are independent of the test distribution, depending only on its support. We also present a series of numerical experiments across a range of functions and operators that support the main theoretical findings.

[LG-116] When Is Coarse Supervision Worth It? Cost-Aware Learning under Unknown Aggregation

链接: https://arxiv.org/abs/2609.36704
作者: Jianyu Xu,Smriti Jha,Aarti Singh,Bryan Wilder
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 34 pages, 6 Figures

点击查看摘要

Abstract:Modern learning systems often acquire supervision at multiple resolutions, trading annotation cost against information content. We study cost-aware two-resolution learning, where expensive fine labels reveal a vector response and cheaper coarse labels reveal a scalar aggregate formed with unknown weights, while the target remains the full response. The challenge is that unknown aggregation changes which directions coarse data can identify, so the value of coarse supervision depends jointly on cost, noise, and identification. We characterize this information geometry and develop an estimate-and-track policy that learns the aggregation rule and tracks the optimal resolution mix. We derive a closed-form break-even condition for coarse supervision and prove that the online policy attains the optimal leading cumulative-risk coefficient, with a matching local asymptotic minimax lower bound. Synthetic experiments support the predicted all-fine/mixed transition, show the online learner approaching the oracle-share benchmark, and demonstrate a finite-budget gain over all-fine acquisition when coarse supervision is sufficiently favorable. Our results provide a principled way to balance information and annotation cost across supervision resolutions.

[LG-117] Learned Queries and Keys Are All You Need: Replacing the Value Projection with Structured Transforms

链接: https://arxiv.org/abs/2609.36698
作者: Ene Meco,Emadeldeen Hamdan,A. Enis Cetin
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:To reduce the number of parameters and cache memory requirements of transformers we introduce dual-headed transformers instead of three heads. We studied Walsh-Hadamard Transform (WHT), Discrete Cosine Transform (DCT), Discrete Fourier Transform, filterbank based Shearlet Transform, and Multiplication-Avoiding (MA) operators to construct dual heads. We combine spatial patches and their orthogonal transforms (or Shearlet and MA operators) in a structure similar to the attention block. We obtained better results than triple headed transformers in ImageNet. Extensive simulation examples are presented.

[LG-118] Know Thyself Teach Thyself: Internal Information Flow for Selective Self-Distillation

链接: https://arxiv.org/abs/2609.36695
作者: Rui Wang,Ruijie Wang,Bo Chen,Jiangxuan Long,Yingyu Liang
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Self-distillation turns knowledge distillation into a closed learning loop and offers a path toward recursive self-improvement. Without an external teacher, however, the model must determine both what information can improve its supervision and which induced changes should be learned. Existing methods typically improve teacher-generated data or select training examples in isolation, leaving the information transferred between these stages unmeasured. We introduce InFlow, a retrieval-guided on-policy self-distillation framework that models this process as potential-to-realized information flow. InFlow first retrieves potentially informative sources using certainty-calibrated hidden-state trajectories, then measures their realized effect through the Jensen–Shannon divergence between the teacher’s initial and retrieval-conditioned answer beliefs. Examples with larger belief shifts are selected for on-policy distillation. Our analysis formalizes the information optimized by retrieval and selection and relates the answer-level shift to the teacher–student distillation gap. Across four open-weight language models and three knowledge domains, InFlow achieves the strongest cross-model average among the compared selection methods, with ablations supporting both stages of the framework. Our code is available at this https URL.

[LG-119] Where Root Cause Analysis Fails: A Retrieval-Reranking Decomposition NEURIPS2026

链接: https://arxiv.org/abs/2609.36686
作者: Hada Melino Muhammad,Luan Pham,Laure Barrière,Sachin Shetty,Leonardo Pulga,Flora D. Salim
类目: Machine Learning (cs.LG)
*备注: Accepted at NeurIPS 2026 (Evaluations Datasets Track)

点击查看摘要

Abstract:Identifying the root cause of an anomaly among hundreds of sensors is critical for preventing safety incidents and costly downtime in complex monitored systems. Existing studies evaluate root cause analysis (RCA) methods using top@k accuracy. We show that this metric has a fundamental blind spot: it conflates two failure modes, retrieval failure, where the true cause is never considered, and reranking failure, where it is considered but ranked too low. In this work, we introduce a retrieval-reranking decomposition and audit four well-known benchmarks to expose this blind spot. Our experiments show that, on benchmarks with complex faults, statistical baselines mis-rank the true cause 79-100% of the time, and graph-based methods never clearly beat the best statistical baseline, whether their causal graphs are learned on short fault windows, on retrieved candidate pools guaranteed to contain the cause, or on multi-day normal-operation data. Meanwhile, on simple benchmarks where faults manifest significantly at their origin, retrieval is nearly solved (98-100%). Guided by the decomposition, we build a two-stage pipeline combining a multi-signal retriever with an LLM reranker that, as one fixed configuration, matches or exceeds the best baseline’s top@1 accuracy on all six benchmark suites (by up to +12 points), with no causal graph or labeled data required. When all methods rank the same retrieved candidates with the true cause guaranteed present, adding a short system-description document lets the reranker lead the best baseline by +7 to +18 points on every benchmark. Code is available at this https URL.

[LG-120] Understanding Private Evolution as Learning-Augmented Clustering NEURIPS2026

链接: https://arxiv.org/abs/2609.36678
作者: Audra McMillan,Kunal Talwar,Felix Zhou
类目: Machine Learning (cs.LG); Cryptography and Security (cs.CR); Machine Learning (stat.ML)
*备注: to be published in NeurIPS 2026

点击查看摘要

Abstract:Private Evolution (PE) is a differentially private algorithm for synthetic data generation. While it can be viewed as a Wasserstein learning algorithm, it performs much better in practice than worst-case Wasserstein analyses would predict. We recast PE as generative model-augmented Wasserstein learning. We show theoretically that when we take into account the use of a generative model that is able to capture something about the true distribution, then we can obtain much better performance bounds. For example, if the generator gives samples in the same low-dimensional space as the distribution, then sample complexity depends on intrinsic, not ambient, dimension. We also show that standard variants of PE can fail to converge on simple well-clustered instances, and propose a new geometry-aware version of PE with provable convergence on such instances. Experimentally, we show that our new algorithm is competitive with standard baselines and can improve recall.

[LG-121] Human-inspired Task-Dimension-Guided Exploration for Efficient Learning in High Dimensions

链接: https://arxiv.org/abs/2609.36672
作者: Fanyu Zhu,Jiahui An,Ni Ji
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Efficient exploration in high-dimensional decision spaces remains a central challenge for decision-making systems. Humans, in contrast, can navigate large decision spaces with remarkable efficiency. Recent behavioral studies suggest that humans reduce dimensionality in large decision spaces by probing candidate feature dimensions, identifying reward-relevant ones, and restricting the effective decision space. Inspired by this mechanism, we propose TDGE (Task-Dimension-Guided Exploration), a human-inspired, model-agnostic algorithm with an automatically constructed task-dimension–feature–item hierarchy. TDGE follows a top-down exploration strategy: it first selects task-relevant feature dimensions, then identifies informative features within those dimensions, and finally recommends concrete items based on the selected features. Experiments on MovieLens-20M, this http URL, and Amazon recommendation datasets show that TDGE substantially improves exploration efficiency and cold-start adaptation over baseline algorithms. Comparisons with other structured algorithms and ablation studies attribute these gains to TDGE’s hierarchical structure and semantic feature-space exploration, with robust results across clustering methods and hierarchy depths. Recommendation-trajectory visualizations also show exploration patterns similar to human dimension-guided behavior.

[LG-122] Stochastic Heavy Ball with Polyak Step Size and Armijo Line Search: A General Convergence Analysis

链接: https://arxiv.org/abs/2609.36668
作者: Jiawei Zhang,Qitan Shi,Yuantao Gu
类目: Machine Learning (cs.LG); Optimization and Control (math.OC)
*备注:

点击查看摘要

Abstract:Polyak step size (PS) and Armijo line search (ALS) have received increasing attention in stochastic optimization, with encouraging empirical performance and theoretical guarantees. However, their convergence theory for stochastic heavy ball (SHB) methods remains limited. In this work, we develop a unified convergence analysis for SHB equipped with PS and ALS. To this end, we introduce a modified Armijo rule that closely parallels the Polyak step size, together with a decoupling analysis that isolates the historical dependence induced by momentum. For SHB with standard PS and ALS, we establish expected convergence for strongly convex, convex, and non-convex objectives without interpolation or restrictive conditions on the momentum parameter. Under interpolation or strong growth, we further strengthen the results to almost sure rates and last-iterate convergence. Moreover, for general settings beyond interpolation, we prove almost sure convergence to the exact optimum or to stationarity for SHB with diminishing variants of PS and ALS. These results provide a more comprehensive theoretical view of Polyak step size and Armijo line search for stochastic heavy ball methods.

[LG-123] GenLimitLib: A Formal Library for Language Generation in the Limit and AI-Assisted Mathematical Research

链接: https://arxiv.org/abs/2609.36663
作者: Shuangping Li,Peng Zhang
类目: Machine Learning (cs.LG); Logic in Computer Science (cs.LO)
*备注:

点击查看摘要

Abstract:We present GenLimitLib, a source-aligned Lean 4 library for language generation in the limit. Introduced by Kleinberg and Mullainathan at NeurIPS 2024, language generation in the limit studies a theoretical question motivated by LLMs: how to generate valid new strings from observed examples. This young and rapidly evolving field offers a natural testbed for studying large-scale formalization. GenLimitLib contains formal developments for 30 papers. It extracts shared definitions and reusable proof components while preserving paper-specific assumptions and statements, and records relationships across papers. In this way, GenLimitLib provides a concrete and structured view of the literature. We show through mathematical case studies and LLM experiments how our library can support both human mathematical research and AI-assisted research. Our Library: this https URL.

[LG-124] Byzantine-Robust Federated Representation Learning

链接: https://arxiv.org/abs/2609.36660
作者: Leonardo F. Toso,James Anderson,Rafael Pinot,Nirupam Gupta
类目: Machine Learning (cs.LG); Optimization and Control (math.OC)
*备注:

点击查看摘要

Abstract:We study federated learning (FL) with adversarial clients, where the goal is to minimize the average loss of the honest (non-adversarial) clients without knowing their identity. Under heterogeneity, a single shared model parameter is statistically inappropriate: it cannot capture the distinct data-generating processes across clients, incurring an irreducible model-heterogeneity bias and severely limiting robustness to adversarial clients (a.k.a. Byzantine-robustness). We address this problem through representation learning, where each client learns a personalized linear head, while collaboratively estimating a shared nonlinear representation through Byzantine-robust aggregation. We demonstrate that the heterogeneity among honest representation gradients is controlled by the representation error and statistical errors that decay either with the number of data samples per client ( \tau ) or the number of iterations ( T ). In particular, our non-asymptotic parameter recovery error bound reveals three terms: (i) an initialization-dependent error that goes away with T , (ii) finite-sample noise terms that decreases with \tau and the number of honest clients, and (iii) a stochastic gradient variance term that also reduces with T . Importantly, with no irreducible model-heterogeneity bias in our bounds. We extend the regression analysis to multiclass classification, and empirically validate it on CIFAR-10, FEMNIST, and School Exam Score datasets.

[LG-125] On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training

链接: https://arxiv.org/abs/2609.36659
作者: Shufan Shen,Zhongni Hou,Junshu Sun,Yufei Zhang,Wei Lin,Guojun Yin,Qingming Huang,Shuhui Wang
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:The strong generalization performance of on-policy post-training paradigms has motivated studies of their parameter update behaviors. However, these studies treat the observed behaviors only as byproducts in on-policy training, overlooking their potential to serve as optimization principles for improving the generalization of other paradigms such as supervised fine-tuning (SFT). To address this limitation, we investigate whether there exists a specific on-policy update behavior that can achieve such improvements. First, our analyses reveal that SFT updates parameters along consistent directions, while the on-policy paradigm continuously adjusts the direction during training. This difference inspires us to focus on the cumulative update direction of each parameter as a promising behavior. Then, we evaluate its effectiveness for improving generalization by proposing On-Policy direction-constrained Supervised Fine-Tuning (OPSFT), which constrains SFT updates to the direction identified by on-policy paradigms. The strong performance of OPSFT indicates that the generalization advantage of on-policy paradigms can be transferred to SFT through the parameter update direction. Once such a direction is identified, even SFT can generalize with its updates constrained to this direction. This finding offers two practical benefits by combining the strong generalization of on-policy paradigms with the advantages of SFT, including the high training efficiency and ability to leverage high-quality trajectories. For efficiency, we identify update directions that support strong generalization using a few on-policy training steps, and subsequently apply OPSFT to achieve high training efficiency. For leveraging high-quality trajectories, OPSFT can utilize these trajectories to continue improving a post-trained model along its update direction without disrupting the ability learned from on-policy training.

[LG-126] Scheduling Recursive Reasoning in Looped Transformers

链接: https://arxiv.org/abs/2609.36653
作者: Boyuan Wang,Chengyao Yu,Jiaxi Ren,Hongxin Wei,Bingyi Jing,Yuxin Tao
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Recurrent reasoning models have attracted growing attention for scaling test-time computation, typically by iteratively refining latent states with shared parameters. However, these models apply each learned update with a fixed unit scale, which can be conservative when updates make persistent progress and overly aggressive when they fluctuate, limiting the benefit of additional loops. To understand how the scale should vary along the trajectory, we first analyze the sensitivity of terminal loss to recurrent update scale. We show that its temporal average admits an exact decomposition into persistent-progress and centered-fluctuation contributions. Based on this, we introduce the Trajectory Adaptive Progress-Fluctuation Scheduler (TAPS), which tracks their balance across recurrent updates and adapts the step size online. Theoretically, we establish sufficient conditions under which TAPS reduces expected terminal loss and reaches a target quality in fewer recurrent loops. Empirically, we show that TAPS improves terminal accuracy across structured reasoning tasks without retraining. By further incorporating the progress-fluctuation principle into training, TAPS yields additional accuracy gains with up to 1.56 times wall-clock speedup at matched baseline accuracy. The broad applicability of TAPS is supported by its effectiveness across diverse recurrent architectures and inference strategies. Together, these results establish update scale as complementary control axis of recurrent inference alongside architecture and depth.

[LG-127] PR-OPD: Privileged Representation On-policy Self-Distillation for Agent ic Reinforcement Learning

链接: https://arxiv.org/abs/2609.36642
作者: Muyang Li,Jie Yang,Zhengyu Fang,Junchao Zhu,Zhengkun Xiao,Ruining Deng,Zhe Jiang,Shigang Chen
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Language-model agents are usually trained by reinforcement learning from one reward per episode, and privileged self-distillation enriches it by letting the same policy, given a skill, teach its skill-free self through token probabilities. However, we identify two phenomena that question this channel. Invisible Advantage: a skill in context lifts WebShop success from 42.2% to 56.2%, yet changes the probabilities of fewer than a quarter of the sampled tokens. Much to Align: a skill changes the hidden states of over 80% of response tokens, in a way that linear probes can trace back to the specific skill. To exploit this, we propose Privileged Representation On-policy Self-Distillation (PR-OPD). After a GRPO warm start, the policy writes a hindsight skill for each trajectory, re-reads its own responses with that skill as a stop-gradient teacher, and aligns its projected hidden states to the teacher’s at every layer alongside the reward objective, with no external skill library, separate teacher, or inference overhead. On ALFWorld and WebShop with two backbones, PR-OPD achieves the best overall results in every setting, improving over GRPO by up to 4.7 points in ALFWorld success and 14.0 points in WebShop accuracy. Code is available at this https URL.

[LG-128] A Digital Simulation Toolkit for Physics-Based Generation of Realistic Experimental Scanning Tunneling Microscopy Images

链接: https://arxiv.org/abs/2609.36639
作者: Huanhuan Zhao,Laxmi Bhurtel,Connor Vernachio,Fahmy Paiziah,Wonhee Ko,Arpan Biswas
类目: Machine Learning (cs.LG); Materials Science (cond-mat.mtrl-sci)
*备注: 17 pages, 6 figures in main text and 2 supplementary figures

点击查看摘要

Abstract:Scanning Tunneling Microscopy (STM) is a widely used tool for characterizing surfaces of materials at the atomic scale, playing a crucial role in discoveries across condensed matter physics and materials science. Despite its extreme spatial resolution, STM is one of the most sensitive microscopy techniques and is highly prone to noise. While existing unsupervised denoising methods are very cheap to train, these are primarily focused on removing the noise with minimal recovery of key physical information. While supervised methods can offer superior performance, the major bottleneck is that a large amount of paired clean-noisy experimental images is required which are impractical to obtain. Thus, we developed a low-cost physics-driven digital toolkit to rapidly generate large volume of realistic STM images. Firstly, we simulate clean images from a chosen material system. Then, with prior knowledge of the physical characteristics of the artifacts and noise present in STM experiments, we formulate several artifact-noise functions such as Gaussian electronic noise, 1/f flicker noise, scan-line noise, background tilt and mechanical drift. These physically informed noise components are then added to the simulated clean images to generate realistic STM images. We demonstrated the capability of the proposed digital toolkit to generate AI-ready data for denoising images of the (111) surfaces of copper and lead, while preserving atoms, defects, and electron waves. We also validated the quality of the downstream image analysis of learning electron wave patterns induced by quantum interference from Cu(111) images. Results show that the supervised models trained on digitally generated AI-ready data can more effectively denoise and learn electron wave patterns on Cu(111) images than benchmarked unsupervised approaches, indicating that the proposed toolkit facilitate scientific discovery.

[LG-129] Physics-Aware Machine Unlearning for Cyber-Physical Systems

链接: https://arxiv.org/abs/2609.36633
作者: Mohammad Zakaria Haider,Muhammad Nadeem,Mohammad Ashiqur Rahman
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:This paper proposes a physics-guided gradient-ascent-based machine unlearning method that couples the forgetting signal with the physical residual of the target cyber-physical systems, ensuring that weight updates during unlearning are steered toward physically feasible regions of the weight space. The physics residual acts as a safety fence during gradient ascent: the model is steered away from the poisoned behavioral basin and simultaneously toward physics-compliant territory, rather than toward an arbitrary alternative that may still violate domain constraints. We evaluate the proposed method against four baselines: naive gradient ascent, exact unlearning, SISA, and full retraining on an IEEE 34-bus distribution system, driven by two physics-informed neural network-based distribution energy resource controllers and validated through high-fidelity OpenDSS power-flow co-simulation. From the evaluation, we found that our proposed physics-guided model simultaneously removes poison and restores physical compliance, which are essential for the safe deployment of safety-critical cyber-physical systems

[LG-130] SemPSG: A Semantic Channel-Aware Foundation Model for Polysomnography Analysis

链接: https://arxiv.org/abs/2609.36619
作者: Junyu Chen,Chenxi Liu,Shiqin Tang,Hao Miao,Wanyun Ling,Ziyue Li,Hongbin Liu,Gaofeng Meng
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Polysomnography (PSG) integrates multiple physiological signals to provide a comprehensive characterization of human sleep, yet its heterogeneous channel configurations across centers pose substantial challenges for transferable representation learning. Existing foundation models mainly focus on physiological modeling or temporal learning, while channel identity is often treated as a fixed structural index, overlooking the physiological semantics encoded by signal modality and reference configuration. To this end, we propose SemPSG, a Semantic channel-aware foundation model for heterogeneous PSG analysis. SemPSG explicitly represents the physiological semantics of channel identity and incorporates them into both signal representation learning and channel aggregation, enabling flexible modeling across diverse data configurations. Specifically, a semantic-conditioned time-series encoder captures signal-specific temporal dynamics and cross-signal interactions, while a multi-view image encoder extracts complementary time-frequency and morphological patterns from the same physiological recordings. We evaluate SemPSG on sleep and health-related tasks, including sleep staging, sleep-disorder breathing analysis, disease prediction, cognition and emotion recognition, and demographic estimation. Extensive experiments demonstrate consistent improvements over both general-purpose time series foundation models and PSG-specific foundation models, together with generalization across heterogeneous datasets across diverse channel configurations.

[LG-131] CI-PINN: Causal Integral Physics-Informed Neural Network for Solving Evolution Equations

链接: https://arxiv.org/abs/2609.36615
作者: Xiaodong Feng,Ziyu Sun,Tao Tang,Xiaoliang Wan,Tao Zhou
类目: Numerical Analysis (math.NA); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Physics-informed neural networks (PINNs) solve partial differential equations (PDEs) by incorporating governing physical laws into the training loss. For evolution equations, however, their conventional pointwise space–time representation does not explicitly encode temporal dependence, which can hinder accurate prediction. To mitigate this limitation, this work proposes a novel neural architecture termed a causal integral neural network (CinNet). The core module of CinNet is a Volterra-type causal integral term, which aggregates historical features to encode temporal dependence, thereby incorporating temporal causality at the architectural level rather than through training-level modifications as in many existing methods. Building on CinNet, we further develop a causal integral physics-informed neural network (CI-PINN) for solving evolution equations. Extensive numerical experiments on benchmark evolution equations demonstrate that the presented method outperforms various baseline PINN variants in terms of solution accuracy, with pronounced superiority under sparse-collocation scenarios. Additional empirical analyses show that CI-PINN exhibits low sensitivity to hyperparameter choices, while ablation studies confirm the effectiveness of the proposed network components.

[LG-132] Selective Elicitation as a Commercial Influence Channel: A Reproducible Synthetic Shopping-Agent Stress Test

链接: https://arxiv.org/abs/2609.36614
作者: Jiapeng Li
类目: Machine Learning (cs.LG); Computers and Society (cs.CY)
*备注: 7 pages, 1 table; synthetic data and reproduction code in ancillary files; substantial AI-tool use disclosed in the manuscript

点击查看摘要

Abstract:A commercial incentive need not enter the final ranking algorithm to affect a shopping assistant’s recommendation: it may instead influence which preference question the assistant asks. We make this distinction experimentally observable in a deliberately small, synthetic setting. Each task has two products, three verified numerical attributes, a price limit, and a private fixed preference vector. An honest simulated user answers one pairwise question. A separate recommender receives the products and this answer but not the sponsorship assignment. We contrast a neutral question, a soft commercial instruction, and an explicitly adversarial instruction to ask about the sponsor’s advantage while omitting the rival’s advantage. Across 40 held-out sponsorship-assignment cases (20 distinct catalog-preference contexts), the soft instruction changes no selections. The targeted instruction raises sponsored selection by 0.30 and reduces mean synthetic utility by 0.0547 relative to neutral questioning (95% context-bootstrap interval [-0.0828, -0.0291]) for one language-model recommender. A fixed Bayesian recommender shows a similar effect; a second model makes the same choices on all 120 frozen question-answer inputs. A terminal-answer consistency judge rates all 20 sampled targeted answers consistent, although five have synthetic regret above 0.05; a separate question-coverage dimension flags their one-sided elicitation. A robust partial-preference certificate remains valid under the stipulated synthetic utility but certifies only 16 of 40 targeted cases and is not better than asking a neutral question directly. These results establish neither typical behavior under advertising incentives nor effects on actual consumers.

[LG-133] Do LLM s Really Forget? Hidden-State Leakage in Model Unlearning and How to Fix it

链接: https://arxiv.org/abs/2609.36612
作者: Hadi Reisizadeh,Jiajun Ruan,Sijia Liu,Mingyi Hong
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Unlearning in large language models (LLMs) is typically evaluated at the output level, where a model appears to suppress sensitive or undesirable content. In this work, we show that such evaluations can create an illusion of forgetting: even when output-level leakage is eliminated, sensitive information can remain encoded in the model’s hidden representations. We first provide a theoretical analysis establishing a fundamental separation between output suppression and representational erasure. Specifically, we show that the decoder can be made arbitrarily insensitive to sensitive directions, driving output-level leakage to zero, while the hidden representations retain the underlying information. To empirically validate this phenomenon, we train generative probe decoders on hidden states across transformer layers, enabling layer-wise measurement of information leakage. Across three widely used benchmarks, TOFU, MUSE, and WMDP, and state-of-the-art unlearning methods, we find that substantial sensitive information remains recoverable from hidden representations, even when standard output-level metrics indicate successful unlearning. To address this gap, we propose Probe-Adversarial Representation Suppression (PARS), an unlearning objective that adversarially minimizes the extractable information from hidden representations. PARS directly targets representational leakage and provides significantly stronger guarantees of erasure under adversarial probing and relearning attacks, outperforming all evaluated baselines. Our results highlight a fundamental limitation of existing unlearning paradigms and suggest that true forgetting in LLMs requires controlling not only model outputs, but also the information encoded in hidden representations. Codes are available at this https URL.

[LG-134] Communication-Efficient Agnostic Federated Learning via Faster Convergence and Compression

链接: https://arxiv.org/abs/2609.36610
作者: Haomin Bai,Junyan Sun,Sifan Yang,Bo Xue,Lijun Zhang
类目: Machine Learning (cs.LG)
*备注: 30 pages, 3 figures

点击查看摘要

Abstract:Agnostic federated learning (AFL) seeks a model that performs reliably across m heterogeneous workers, but communication remains a bottleneck. We improve communication efficiency by reducing the number of synchronization rounds via faster convergence and the communication cost per round via compression. We first propose AFL-BR, which updates the dual weights over workers using online mirror ascent with KL divergence and blockwise restarts. It achieves an O((\log m)^1/4T^-1/8) stationarity rate after T update rounds, reducing the m -dependence of the synchronization rounds required for convergence from polynomial to logarithmic order. Building on AFL-BR, we develop AFL-Com by applying bidirectional compression with error feedback (EF). Instead of compressing local gradients, workers apply EF to their dual-weighted gradients, enabling direct control of the aggregated compression error under time-varying weights. We then establish an O((\delta^-1+(\log m)^1/4)T^-1/8) stationarity rate for AFL-Com under general \delta -approximate compressors and improve the \delta -dependence from \delta^-1 to \delta^-1/2 for additive-and-idempotent compressors with shared randomness (SR). With suitable compression levels, AFL-Com retains the same convergence rate as AFL-BR at a lower per-round communication cost, yielding reductions in total communication complexity by factors of (\log m)^1/4 with Top- k and (\log m)^1/2 with Rand- k and SR. Experiments validate the improved synchronization and communication efficiency of our methods.

[LG-135] Learned Reporting Preferences in RLVR Can Conflict with the Current Request

链接: https://arxiv.org/abs/2609.36587
作者: Yupeng Chang,Wenxuan Zhang,Yuan Wu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Reinforcement learning with verifiable rewards (RLVR) has become a prominent approach for improving language-model performance on reasoning tasks using automatically checked answers. Yet convention-matched evaluation cannot reveal whether reinforcing one reporting convention reduces adherence to a different request that the initial policy already follows. To test this, we train matched policies under two reporting conventions and evaluate each policy under both current requests, using the same initial policy as a shared reference. We complement this crossed design with controlled interventions and independent human calibration. On GSM8K, boxed-format RLVR reduces the fraction of Qwen2.5-7B responses containing the requested hash-format payload by 35.33–74.37 percentage points relative to a 95.45% initial baseline in four of five training seeds; the fifth improves by 2.50 points. In the four deteriorating runs, almost every response that omits the requested payload instead retains the trained boxed convention, and the same four seeds deteriorate under two fixed paraphrases. Changing only the final-answer marker in supervised targets reverses which reporting convention the model prefers across three seeds, providing controlled evidence that this preference is learnable. Across three settings with independent human calibration, gains under a convention-sensitive scorer exceed the corresponding gains in committed-answer correctness, i.e., the correctness of the answer the model actually commits to. Together, these results separate three distinct post-training outcomes: learned reporting preference, current-request adherence, and committed-answer correctness. They show that convention-matched accuracy alone does not fully characterize post-training behavior and motivate evaluating current-request adherence alongside convention-matched task accuracy.

[LG-136] Making Analog Training Scale: Co-Designing Mapping Optimizer and Converters

链接: https://arxiv.org/abs/2609.36584
作者: Zhaoxian Wu,Tayfun Gokmen,Omobayode Fagbohungbe,T. Patrick Xiao,Tianyi Chen
类目: Machine Learning (cs.LG); Hardware Architecture (cs.AR)
*备注:

点击查看摘要

Abstract:Analog in-memory computing (AIMC) offers an alternative for model training by executing matrix operations directly where weights are stored. However, scaling AIMC to train modern deep models remains an open challenge due to severe hardware non-idealities, including physical weights with finite dynamic range and write granularity, analog-digital converters with finite resolution, and noisy and asymmetric updates. Guided by the insight that gradient accumulation is sensitive to precision and rounding errors, we adopt a mixed-precision training paradigm: executing forward and backward matrix multiplications in the analog domain while computing weight gradients in the digital domain. To enable scalable training, we present a holistic system-algorithm co-design that co-optimizes weight mapping to ensure well-conditioned physical and logical weight profiles, couples a preconditioned optimizer with threshold-triggered open-loop pulsing to stabilize training trajectories, and aligns converter dynamic ranges to suppress quantization errors. Evaluated via hardware-calibrated architectural simulations calibrated with electrochemical RAM measurements, our framework scales Transformer training up to 123\textM parameters with validation loss scaling as L\propto N^-0.231 , where N is the parameter count, comparable to L\propto N^-0.238 for digital training.

[LG-137] Sharp Convergence and Sampling Trade-offs for Riemannian Diffusion under Nonnegative Ricci Curvature

链接: https://arxiv.org/abs/2609.36568
作者: Yuhao Liu,Longbo Huang
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Diffusion models have emerged as state-of-the-art generative models, with recent extensions from Euclidean spaces to Riemannian manifolds. However, existing convergence guarantees for Riemannian diffusion models typically require \tildeO(\mathrmpoly(d,T)/\epsilon^2) score evaluations, with potentially unfavorable dependence on the dimension. In this work, we develop a general framework that separates score discretization from Brownian-motion simulation and allows multiple geodesic random-walk steps per score evaluation. Under nonnegative Ricci curvature assumption and an exact Brownian-motion simulation oracle, we show that \tildeO(d/\epsilon^2) score evaluations suffice to achieve an \epsilon^2 KL divergence from the target distribution, matching the existing convergence rate of Euclidean diffusion models. We further show that \tildeO(d^4T/\epsilon^2) geodesic random-walk steps suffice to approximate the required drifted Brownian motion to \epsilon total variation error. Combining these results yields a sampling scheme with \tildeO(d/\epsilon^2) score evaluations and \tildeO(d^4T/\epsilon^2) geodesic random-walk steps, motivating multiple random-walk steps between consecutive score evaluations. Our results provide a sharper characterization of the convergence and sampling complexity of Riemannian diffusion models.

[LG-138] Reactive Real-Time Flow Policies via Asynchronous Distribution Alignment

链接: https://arxiv.org/abs/2609.36540
作者: Moritz Zoellner,Reece O’Mahoney,Ioannis Havoutis,Rohan Paleja
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Generalist robot policies such as vision-language-action models (VLAs) have achieved remarkable generalization, but their inference delays can conflict with the demands of real-time control. Asynchronous execution avoids pauses between action chunks by predicting the next sequence of actions while the robot carries out the previous one. In this paper, we study whether asynchronous execution produces the same action distribution as the original VLA. We find that, for non-Markovian demonstrations, asynchronous execution can produce a fundamentally different action distribution, which can limit the policy’s reactivity. In our method, we seek to restore this reactivity by aligning the asynchronously produced action distribution with that of the original VLA through two complementary mechanisms. First, Recursive Flow-Field Distillation trains the asynchronous policy using the VLA’s action-generation flow. We characterize the learned distribution theoretically and show experimentally that our asynchronous policy can generate nearly the full range of actions the original VLA would produce, while existing asynchronous methods recover only a fraction of that range. Second, Propose-Resolve prepares multiple action sequences asynchronously and uses the latest observation to select among them based on a lightweight approximation of their likelihood under the VLA’s action distribution. Our resulting method matches the original VLA’s success on LIBERO and retains about 80% of its success on RoboMimic, about 30 percentage points more than existing asynchronous methods.

[LG-139] SCOPE: Observation-Conditioned Full-Target Prediction for Sparse PDE Inference

链接: https://arxiv.org/abs/2609.36527
作者: Ruichen Xu,Siyao Wang,Fang Wan,Jiacheng Qiu,Wenhan Gao,Jiaxing Zhang,Linsey Pang,Ravid Shwartz-Ziv,Prakhar Mehrotra,Yann LeCun,Yuefan Deng
类目: Machine Learning (cs.LG); Computational Physics (physics.comp-ph)
*备注: 34 pages, including supplementary material. Code: this https URL

点击查看摘要

Abstract:Recovering complete physical fields from sparse observations is challenging because the measurements may not uniquely determine the underlying state. Diffusion-based PDE solvers address this problem through iterative sampling whereas neural operators provide deterministic one-pass predictions. We propose SCOPE (Sparse-Context Observability-aware Predictive Embeddings) to recover complete PDE fields from sparse observations by coupling full-field latent prediction with physical reconstruction. A shared decoder reconstructs fields from both predicted and complete-view representations so that representation learning is guided by both physical recovery and latent matching. We derive a quadratic risk decomposition at fixed teacher-decoder pairs showing why optimal latent prediction need not yield optimal field reconstruction. We also establish sufficient conditions for decoder improvements on complete inputs to transfer to recovery from partial observations. Experiments across five PDE settings show that SCOPE outperforms mask-aware neural operators on all ten forward and inverse tasks and achieves lower errors than those reported for diffusion-based solvers including DiffusionPDE and FunDPS. Decoder-only adaptation further improves recovery without retraining the backbone while retaining deterministic single-pass inference.

[LG-140] PDE-OBS: Controlled Evaluation Across Observation Patterns

链接: https://arxiv.org/abs/2609.36521
作者: Ruichen Xu,Siyao Wang,Fang Wan,Jiacheng Qiu,Wenhan Gao,Jiaxing Zhang,Linsey Pang,Ravid Shwartz-Ziv,Yann LeCun,Yuefan Deng
类目: Machine Learning (cs.LG); Computational Physics (physics.comp-ph)
*备注: 57 pages, including supplementary material. Code: this https URL

点击查看摘要

Abstract:Physical-field reconstruction and forecasting depend on both measurement density and spatial layout, yet evaluation under a single observation pattern does not characterize performance when that pattern changes. We introduce PDE-OBS, an integrated benchmarking platform spanning numerical data generation, model training, and inference and evaluation under varying observation conditions. It combines 560,000 fields and trajectories from seven partial differential equation families with configurable observation operators and seven adapted baseline methods for stationary reconstruction and short-horizon forecasting. Separating observation construction from physical records allows users to specify parameterized patterns and deterministic mixtures for training and testing while preserving prediction targets and data splits. The evaluation protocol uses references trained for each test pattern to compare models on identical test observations and targets, alongside equal-count groups for spatial-layout comparisons. On a 14,000-record subset, we evaluate 441 trained models under nine test patterns, yielding 3,969 evaluations. Mean cross-pattern error exceeds mean matched-pattern error in all 49 PDE-method pairs, and this finding persists in a configuration-matched subset of 117 models. Denser test observations do not consistently reduce error for a fixed model. Mixed-pattern training on five completed pairs reduces large single-pattern transfer errors, although destination-trained references usually remain more accurate. Together, the benchmark and findings support systematic evaluation of observation-pattern sensitivity and provide a reusable workflow for developing methods under changing measurement conditions. Code: this https URL.

[LG-141] Efficient and Scalable Physics-Guided Fully Convolutional Spatiotemporal Learning for 3D Microstructure Evolution Prediction

链接: https://arxiv.org/abs/2609.36504
作者: Michael Trimboli,Wenxi Liu,Xianqi Li
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Accurate prediction of three-dimensional (3D) microstructure evolution remains computationally demanding because high-fidelity phase-field simulations require repeated numerical integration over large volumetric domains and long temporal horizons. This study develops an efficient and scalable physics-guided fully convolutional spatiotemporal framework for direct multi-frame prediction of complete 3D microstructure sequences. The model combines shared 3D spatial encoding and decoding with a factorized latent translator that integrates temporal, local 3D spatial, and channel interactions. A discrete Cahn–Hilliard (CH) residual is incorporated during training to regularize the learned evolution toward the governing dynamics without altering the inference pathway. The framework is evaluated on high-resolution 3D spinodal-decomposition trajectories under nominal, long-horizon, and reduced-temporal-context forecasting. Under full temporal context, the model accurately reproduces volumetric evolution, with average 3D structural similarity remaining above 0.97 over the nominal prediction horizon. Physics guidance becomes increasingly beneficial as temporal information is reduced, improving predictive robustness and preservation of interface-level morphology. The framework also achieves more than a 30-fold wall-clock speedup relative to the reference spectral phase-field solver, while physics guidance introduces no additional inference cost. These results establish direct multi-frame, physics-guided fully convolutional learning as a high-throughput surrogate strategy for dense 3D phase-field dynamics and repeated microstructure forecasting.

[LG-142] AdaptArena: Evaluating Test-Time Personalization of Web Agents

链接: https://arxiv.org/abs/2609.36488
作者: Dongchan Shin,Xing Han Lù,Jiaqi Deng,Jay Gala,Tomás Vergara Browne,Jaewon Moon,Fengyuan Liu,Alexandre Drouin,Siva Reddy,Alexandre Lacoste
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Large language model (LLM) agents have demonstrated strong performance on complex web navigation tasks, yet they remain brittle in real-world settings where user intentions are underspecified and preferences are heterogeneous. In practice, users rarely provide explicit profiles, requiring agents to infer latent preferences from implicit signals. Despite its importance for deployment, this problem setting is largely underexplored in existing benchmarks. To address this gap, we introduce AdaptArena, a benchmark for evaluating test-time personalization of web agents via implicit preference inference. AdaptArena consists of 480 tasks, featuring both single-preference and double-preference scenarios. Each evaluation task must be solved by retrieving and leveraging the most relevant historical user trajectory that implicitly encodes the target preference. In addition, we introduce AdaptiveAgent, a retrieval-based framework for standardized evaluation of implicit preference inference. Experiments reveal a substantial performance gap: while oracle agents with access to ground-truth preferences achieve an 82.92% success rate, the evaluated LLM agents using our framework reach at most 15.62%. Furthermore, we find that correctly inferring user preferences is necessary but not sufficient for task success, as execution failures in downstream web interactions remain a significant bottleneck even when agents align with the target preference. These findings highlight implicit preference inference and robust action grounding as key challenges for deploying reliable, user-facing web agents. We release our code: this https URL

[LG-143] Optimal Multi-Reward Reinforcement Learning

链接: https://arxiv.org/abs/2609.36486
作者: Zijun Chen,Zihan Zhang
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:We study an unknown-transition finite-horizon Markov decision process (MDP) with a finite collection of known reward functions \r^1, r^2, \ldots, r^M\ . The goal is to output an \epsilon -optimal policy for every reward using online episodic interaction only. Performance is measured by the policy error V_0^, m - V_0^\widehat\pi^m, m where m\in [M] represents the reward function and V_0^, m=\mathbbE_s_1\sim \mu[V_1^*, m(s_1)] . Under this setting, we design a provably efficient algorithm to establish a minimax sample complexity bound of O\left(\fracSAH^3\epsilon^2\log M \mathrmpolylog\left(\fracSAH\log M\min\left\epsilon, 1\right\delta\right)\right) episodes, with no additional burn-in cost. This matches the information-theoretic lower bound up to a factor of \mathrmpolylog(SAH\log M/(\min\left\epsilon, 1\right\delta)) . Our method combines three technical ingredients. First, we adapt MVP to reward-switching learning to construct optimistic value estimates. Second, we use fresh replay samples to conservatively evaluate the candidate policies. Third, gap-based multiplicative weights updates adjust the reward-sampling distribution using the differences between these estimates, converting weighted learning progress into simultaneous guarantees for all rewards.

[LG-144] he Teacher Is a Direction Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation

链接: https://arxiv.org/abs/2609.36484
作者: Hao Li,MeiJia Chen,Weijie Ren,Donghan Li,Zijun Tian,Jingchun Huang,Naibo Wang
类目: Machine Learning (cs.LG)
*备注: 19 pages

点击查看摘要

Abstract:On-policy distillation (OPD) trains a student to match the teacher’s next-token distributions on the student’s own trajectories and has yielded substantial empirical gains. Generalized variants allow the student to surpass the teacher by extrapolating an implicit reward in output space. The language-model head, however, attenuates this change anisotropically: much of the change encoded in the teacher’s hidden states reaches the logits at a small fraction of its weight, and the sampled-token log-probability ratios on which output-space extrapolation relies inject noise that the extrapolation amplifies, making training unstable. We observe that reinforcement learning (RL) shifts a model’s internal representations relative to its base checkpoint, and that the direction of this shift can be measured at every layer. Motivated by this observation, we propose RIDE (RL-Induced Direction Extrapolation), which extrapolates the RL-induced change directly in representation space: at every layer and token position, RIDE computes the residual between the teacher and its pre-RL checkpoint and regresses the student’s hidden states toward targets displaced beyond the teacher along this residual. Conditioned on a sampled trajectory, this regression is equivalent to maximizing a linear directional reward defined by the residual under a quadratic penalty centered at the teacher, which makes explicit how the objective moves the student along the RL-induced direction while limiting its deviation from the teacher. Across four base/RL-teacher pairs spanning different scales, architectures, and pre-training lineages, RIDE approaches or exceeds the RL-trained teacher on every pair and is the only method whose mean does so, and it consistently outperforms output-space extrapolation, which degrades the student whenever the teacher is close to its base. Project page: this https URL.

[LG-145] DisCoMBO: Steering Expert-in-the-Loop Black Box Optimization via Distributional Conformance

链接: https://arxiv.org/abs/2609.36472
作者: Jonas Seng,Bennet Wittelsbach,Kristian Kersting
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Sequential Model-Based Optimization (SMBO) traditionally relies on Bayesian or ensembling surrogates for uncertainty quantification. While historically treated as fully data-driven, SMBO increasingly integrates external domain expertise to accelerate discovery. To overcome the opaque guidance and diminished integration fidelity of standard acquisition re-weighting, Probabilistic Circuits (PCs) have emerged as a generative surrogate alternative, enabling direct knowledge injection via conditional sampling. However, these generative routines lack the formal exploration-exploitation semantics required for rigorous optimization. We introduce the Distributional Conformance Score (DisCo), a novel metric that unifies the flexibility and efficiency of PCs with a formal uncertainty framework. DisCo provides a bounded, [0, 1] -normalized measure of model “surprise” that (1) recovers properties comparable to kernel-based uncertainty known from, e.g., Gaussian Processes, while maintaining linear-time inference, and (2) enables accurate assessment of conformance of external knowledge w.r.t. model evidence. We then present DisCoMBO, a framework leveraging these properties for robust, knowledge-aware optimization. We prove that DisCoMBO is a zero-regret algorithm and demonstrate its effectiveness across diverse benchmarks from AutoML, material optimization, and wind park optimization.

[LG-146] Emergent phases of superposition: from partial to full representation

链接: https://arxiv.org/abs/2609.36455
作者: Lihao Guo,Yizhou Liu,Jeff Gore
类目: Machine Learning (cs.LG)
*备注: 35 pages, 25 figures

点击查看摘要

Abstract:Large language models are thought to represent features by vectors in a hidden space of dimension given by the model’s width. Superposition, in which more features are represented than the width by letting representation vectors overlap, is a leading account of how representation vectors are organized. However, how model width and data statistics determine the configuration of representation vectors and the resulting loss when the number of features and the width are large remains less understood. Here we show, in Anthropic’s toy model of superposition, that increasing the width drives a continuous phase transition from a partial-representation phase, where only a subset of features receives appreciable representation vectors while the rest vanish, to a full-representation phase, where every feature is represented. Our theory via a partial random projection approximation predicts, and experiments confirm, that the critical width grows linearly with the number of active features up to a logarithmic factor. The loss scaling changes across the transition: below the critical width, the loss grows linearly with the number of active features and depends weakly on the width in a form set by data statistics; above it, the loss grows approximately quadratically with the number of active features and decays inversely with the width. Non-uniform firing probabilities delay the transition and lower the loss, as more frequent features occupy more space. Our results provide an account of how model width and data statistics jointly shape representations and loss, a step toward understanding representation scaling in large models.

[LG-147] heory on Attention Dynamics for Out-of-Distribution In-Context Learning NEURIPS2026

链接: https://arxiv.org/abs/2609.36448
作者: Junze Deng,Daouda Sow,Sen Lin,Yingbin Liang
类目: Machine Learning (cs.LG)
*备注: Accepted to Neurips2026

点击查看摘要

Abstract:Transformers have demonstrated remarkable in-context learning (ICL) capabilities, enabling them to perform new tasks without additional fine-tuning. However, their performance often deteriorates when encountering out-of-distribution (OOD) inputs that deviate from the training distribution, and the underlying theory remains poorly understood. To fill this gap, we characterize the OOD error under the input distribution shift through the interplay between the dynamics of the so-called \alpha -type and \beta -type attention weights, which represent the transformer’s confidence in identifying the correct and incorrect features, respectively. Our results indicate that the OOD error for each feature depends on all pairwise interactions between the training features and OOD features, and under certain cases the transformer performs no better than random guessing. To improve the OOD generalization performance, we next investigate the impact of model finetuning with the OOD data, and particularly, characterize the model forgetting performance on the source domain. Interestingly, the performance on the source domain may not always degrade after finetuning, which highly depends on the nature of the feature shift: finetuning on OOD domain keeps enhancing the confidence of identifying correct features from the original distribution, while the interference from other incorrect features may either increase or decrease. Extensive experiments on both synthetic and real data are conducted to corroborate the theoretical insights.

[LG-148] Sample Complexity of Equivariant Reinforcement Learning

链接: https://arxiv.org/abs/2609.36421
作者: Rayan Mazouz,Haibo Zhao,Chris Hillar,Christian Shewmake
类目: Computational Complexity (cs.CC); Machine Learning (cs.LG); Group Theory (math.GR)
*备注:

点击查看摘要

Abstract:Reinforcement learning (RL) is a powerful framework for robotic control, yet its practical application is often hindered by high sample complexity. This is particularly restrictive in physical domains where interaction data is costly. While the world often exhibits geometric and physical symmetries, standard RL algorithms typically fail to exploit this structure. In this paper, we demonstrate that exploiting group symmetries significantly reduces the sample complexity of RL. Focusing on finite-horizon Markov decision processes, we find that leveraging homomorphisms induced by group symmetries significantly reduces the theoretical upper and lower bounds on the number of environment interactions required to reach an optimal return. We further extend these bounds to continuous state and action spaces, providing corresponding sample-complexity guarantees under appropriate regularity assumptions. Beyond theory, we validate our findings through controlled experiments and demonstrate the advantages of symmetry-aware policy learning on high-dimensional continuous robotic simulations. Our results show that integrating symmetry into the learning pipeline yields substantial gains in sample efficiency and performance, offering a principled path toward more data-efficient robotics.

[LG-149] Neural Succession: A Mesoscopic Theory of Invasion Coexistence and Stabilization in Continual Learning

链接: https://arxiv.org/abs/2609.36375
作者: Shoaib Ahmed Dipu,Md Salman Shamil,Sayeed Shafayet Chowdhury
类目: Machine Learning (cs.LG)
*备注: 20 pages, 10 figures, 7 tables. Code available at this https URL

点击查看摘要

Abstract:Continual learning is usually studied through mechanisms that preserve old knowledge. We develop Successional Learning Theory (SLT), a mesoscopic account in which the current representation is a resident community, the incoming task is an invader, forgetting is resident displacement, joint retention is coexistence, replay is resident reinforcement, and training moves from establishment toward stabilization. Its empirical coordinate is directional pre-invasion compatibility, measured on the resident model before the incoming task is learned. Across eight experiments, compatibility orders later forgetting on the 20 directed Split-CIFAR-10 transitions (three-repeat r=-0.789, incoming-task cluster 95% CI [-0.90,-0.72], every repeat alone r=-0.67), forecasts held-out forgetting with 24% lower error than a no-information baseline, and reproduces under controlled MNIST permutations and CIFAR-10 rotations (r=-0.804, -0.718). On an 84-transition suite, compatibility separates coexistence from exclusion at every retention threshold (AUC 0.93-0.97). Replay repairs every transition with at most 325 stored examples and is most efficient where displacement is largest. Compatibility reaches |r|=0.720, while activation, representation, Jacobian, and fixed-coefficient Lotka-Volterra specializations do not. Plasticity and feature turnover fall reliably from early to late training (15/15 and 14/15 runs). We formalize a minimum habitat-modification bound, a displacement floor, a sufficient coexistence condition, an identifiability law with a range-restriction corollary, successional stabilization, and local reinforcement. The identifiability law also predicts where the coordinate loses leverage, and the prediction matches three CIFAR-100 partitions and five optimizer regimes. SLT is a pre-adaptation diagnostic that complements replay, regularization, and projection methods.

[LG-150] Adapting Linear-Time Architectures for Tabular In-Context Learning

链接: https://arxiv.org/abs/2609.36337
作者: David Schnurr,Felix Sarnthein,Thomas Hofmann,Imanol Schlag
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 31 pages, 13 figures, 10 tables. Code available at this https URL

点击查看摘要

Abstract:Tabular foundation models achieve strong performance by conditioning on labelled examples in context, but softmax attention limits their use on large datasets. Existing linear-time alternatives, however, are mostly causal, and their potential for tabular in-context learning (ICL) remains underexplored. To address this, we (1) revisit causal training setups, (2) compare linear sequence mixers, and (3) investigate their ICL generalisation beyond the pretraining context length. First, we show that the best training setup for causal models resembles next-token prediction. Then, perhaps surprisingly, the most promising linear sequence mixer is causal: DeltaNet outperforms even non-causal linear attention. However, it degrades beyond 2 - 4\times the pretraining context length, and existing mitigation strategies such as bidirectionality defer the problem at best. A hidden-state oracle shows that this is not a capacity problem. Instead, our analysis points to an instability in the recurrent state, which drifts in deeper layers of causal models. Since DeltaNet’s learned write rates overfit to the pretraining regime, we modulate them with a time-dependent decay schedule intervention to stabilise length generalisation. Finally, re-introducing non-causality by reading out from the final state allows us to closely match a controlled softmax attention baseline on OpenML-CC18 and TabArena.

[LG-151] Reducing the Adaptation Gap Through Reachable Fisher Geometry

链接: https://arxiv.org/abs/2609.36329
作者: Wasif Jalal,Sachin Deb,Asif Salekin
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Parameter-efficient fine-tuning (PEFT) determines not only how many parameters are trained, but also which local directions a model can move in, so similar adapters can affect subgroup losses differently. Since curvature matrices are infeasible to form at adapter scale, scalar summaries such as the Fisher trace are often used instead. We study what the trace reveals and what it loses through the reachable Fisher: each subgroup’s full-model Fisher pulled back through the adapter Jacobian. Under likelihood losses, it represents the Gauss-Newton curvature accessible to the adapter, and its trace can be computed from score-gradient norms without forming the full matrix. Under matched subgroup gradients, a positive-definite reachable-Fisher difference, with a margin exceeding the Hessian-Fisher defect, implies that every sufficiently small nonzero model-changing update increases the signed gap. In contrast, the restricted operator norm determines worst-case quadratic change, while a matrix-free Frobenius discrepancy bounds its reachable-Fisher component. Trace alone cannot certify definiteness or control matrix mismatch. Equal traces rule out a positive-definite difference but can still hide large operator discrepancies. Across 306 single-seed models, higher trace accompanies greater subgroup difficulty in 75.7 percent of 1,218 eligible evaluations, while trace matching reduces the best-worst subgroup gap in all 30 dataset-encoder-adapter combinations. However, held-out audits show that operator discrepancy decreases in 23 of 30 combinations, while the unbiased squared-Frobenius statistic decreases in only 16 of 30. Fisher trace is therefore a scalable diagnostic and training heuristic, but not a certificate of local gap behavior or matrix alignment.

[LG-152] Distill Locally Schedule Globally: Flow Maps for Few-Step Text-to-Speech ICASSP2027

链接: https://arxiv.org/abs/2609.36324
作者: Yentl Collin,Evan Dufraisse,Amr Mohamed,Amine Khelif Khelif,Dani Bouch,Guokan Shang
类目: ound (cs.SD); Machine Learning (cs.LG)
*备注: 5 pages, Submitted to ICASSP 2027

点击查看摘要

Abstract:Flow-matching text-to-speech (TTS) models achieve high synthesis quality but require many neural function evaluations (NFEs) to integrate their generative trajectories. Recent few-step flow-map distillation approaches for TTS construct targets from numerically integrated teacher trajectories, creating a trade-off between target accuracy and training cost. We propose Local Flow-Map Distillation (LFMD), which adapts Eulerian Map Distillation to conditional TTS and avoids teacher trajectory integration during target construction. For inference, we derive a sampling schedule (TD-DP) from teacher dynamics and consistency of the learned maps, with a single cost graph supporting multiple NFE budgets without external audio-metric evaluation. Because scheduling offers no flexibility at one NFE, we refine this regime with alignment-aware temporal self-distillation using soft-DTW. Across Seed-TTS and LibriSpeech-PC, LFMD improves low-NFE synthesis over a matched integral-distillation baseline. On Seed-TTS, the refined student reaches 1.80% WER with 1-NFE, compared with 1.76% for its 32-NFE teacher.

[LG-153] Learning Samples Importance: Parameterizing Dual Variables in Everywhere Learning

链接: https://arxiv.org/abs/2609.36310
作者: Ignacio Boero,Jonathan Nixon,Alejandro Ribeiro
类目: Machine Learning (cs.LG); Signal Processing (eess.SP)
*备注:

点击查看摘要

Abstract:Everywhere learning provides a principled framework for training AI models under constraints that must hold throughout the data distribution. In the dual domain, these pointwise constraints give rise to functional dual variables. In this work, we propose to learn these dual variables, motivated by the fact that their values encode useful information about the underlying constrained problem. By representing the dual variable as a parametric function of each sample, we enable the learned multiplier to be evaluated on new, unseen samples. This contrasts with standard empirical dual formulations, which assign an independent multiplier to each training sample. We characterize the error in the recovered primal solution induced by restricting the dual variable to a parametric function class and show that it is controlled by how well this class approximates the optimal statistical multiplier. Moreover, we show that the learned parametric multiplier retains the sensitivity interpretation of the optimal statistical multiplier, yielding approximate sensitivity guarantees that extend beyond the samples used for training. We empirically validate our theory across a variety of everywhere learning tasks, showing that the resulting constrained problems can be solved efficiently and that the learned dual variables provide meaningful representations of sample-level sensitivity.

[LG-154] Cheap and Powerful Tests for Supervised Subspaces: Per-Component Inference for PLS NEURIPS2026

链接: https://arxiv.org/abs/2609.36307
作者: Paweł Lenartowicz,Hubert Plisiecki
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 41 pages, 5 figures, 28 tables. Accepted at NeurIPS 2026. Code and results: this https URL

点击查看摘要

Abstract:Partial Least Squares (PLS) regression extracts a few outcome-aligned directions in a high-dimensional X and is widely used across applied science, but inference on the resulting fit is either expensive, biased and discouraged, or absent. We reduce inference to held-out OLS refits of the supervised subspace, a primitive shared by PLS, supervised PCA, and linear probes, and supply two tests using held-out correlations: a Nadeau-Bengio corrected asymptotic t-test as a fast approximation, and a permutation test with comparable power, finite-sample valid under outcome-predictor independence and iid rows. Held-out predictions are unchanged under any orthogonal rebasing of the supervised span, so an interpretable basis such as varimax inherits the joint claim but not a per-axis p-value; per-component claims come from a fixed-sequence test on the PLS extraction order. We validate on synthetic geometries, two NIR chemometric datasets, and cross-lingual word-embedding regressions; the exact test also transfers to supervised PCA and a ridge probe. The proposed tests have more power than CV-permutation-Q^2, at a fraction of its cost. A pre-run check on n and the spectrum of X says when the approximation is safe. We release a Rust library with Python, R, and Julia bindings, plus a Python text pipeline.

[LG-155] he Universal Classifier for Graph Learning

链接: https://arxiv.org/abs/2609.36302
作者: Ben Finkelshtein,André Linhares,Petar Veličković,Bryan Perozzi,Mikhail Galkin
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:While foundation models have revolutionized natural language processing and computer vision by leveraging universal vocabularies, Graph Machine Learning (GML) remains fractured due to the absence of a unified feature and structural representation across diverse domains. Existing works claiming to be Graph Foundation Models (GFMs) are typically restricted to node-level predictions or require fixed feature dimensions, failing to provide a truly task-agnostic backbone for the full spectrum of graph learning applications. In this paper, we introduce the Universal Classifier (UC), which supports arbitrary feature and class cardinalities, unifying node-, edge-, and graph-level objectives under a single similarity-based classification objective. The UC reformulates all node-, edge-, and graph-level prediction tasks as maximizing similarity in the latent space: by lifting heterogeneous features and labels into 3D latent tensors, the model learns transferable features independent of specific input schemas. This architecture allows a single pre-trained model to generalize to node classification, node regression, and link prediction across unseen graphs with varying feature semantics. Experiments show strong zero-shot transfer performance across node-, link-, and graph-level tasks.

[LG-156] Representational and Functional Robustness to Electrode Montages in EEG Foundation Models

链接: https://arxiv.org/abs/2609.36288
作者: Jakob Steglich,Justus Meyer zu Bexten,Shakiba Moradi,Laure Ciernik,Simon M. Hofmann,Mina Jamshidi Idaji
类目: Machine Learning (cs.LG)
*备注: 17 pages, 9 figures

点击查看摘要

Abstract:EEG foundation models (EEG-FMs) are intended to generalize across different datasets by learning representations that, ideally, are invariant to dataset-specific EEG configurations such as electrode montages. However, EEG-FMs that accept different montages as input do not guarantee that representations and predictions remain stable across different electrode configurations, especially outside the training setting. In this work, we investigate the effects of different electrode montages through a joint functional and representational analysis of four EEG foundation models selected to span distinct montage-handling designs. We evaluate embeddings on cross-subject resting-state eyes-open/closed and within-subject motor-imagery classification under spatially informed channel reduction. Functional robustness is tested through the generalizability of linear probes across channel counts, while representational robustness is assessed through within-subject similarity and preservation of between-subject geometry. The four models show distinct robustness profiles, and the two axes dissociate: large changes in embedding similarity need not come with comparable probe degradation, and stable embeddings can still lose downstream performance. Comparing two readouts of the same encoder further shows that aggregation, not the encoder alone, determines functional robustness: pooling into anatomically aligned regions degrades less than a learned global readout, despite being montage-invariant by construction. Montage robustness is therefore a joint property of the encoder and its aggregation, and characterizing it requires both a representational and a functional axis. Input compatibility alone is evidence for neither.

[LG-157] GNA: Granular Neighbor Assembly for Retrieval-Augmented Multivariate Time-Series Forecasting

链接: https://arxiv.org/abs/2609.36281
作者: Vincent Uhse
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Deep forecasters predict from a fixed-length lookback window, and lengthening it gives diminishing returns at a growing cost. Retrieval augmentation instead shows the model how similar past situations continued. Retrieving a whole past window gives every variate the continuation of the same past moment. In multivariate series, however, the best past match differs from variate to variate. We present GNA (Granular Neighbor Assembly), a retrieval layer for forecasting backbones that assembles neighbors at two granularities: whole past windows, which keep the variates coherent, and per-variate neighbors, in which each variate takes its future from its own best-matching past. A learned gate decides, per forecast step and variate, how much to trust these futures against a persistence forecast, next to the backbone’s own forecast. Candidates come from an embedding trained to predict each window’s future, and retrieval is strictly causal: a past window is used only once its future has been observed. With the same lookback for every model and the same retrieval constants for all datasets, GNA improves two Transformer backbones in 85 of 96 dataset-horizon settings, gives the lowest MSE on 8 of 12 standard benchmarks and beats its backbone in every seed on 10 of them. Both granularities are needed, and neighbors of mismatched queries are worse than none. Retrieval helps most where the lookback says least: the gate shifts trust to retrieved futures further ahead. Where it fails, on hourly non-stationary series at long horizons, the loss is consistent with a drifting level of the retrieved futures.

[LG-158] Stochastic Optimization Under Power-Law Spectra: Tight Bounds and Shuffling Analysis

链接: https://arxiv.org/abs/2609.36271
作者: Thomas Dybdahl Ahle,Yaroslav Bulatov,Christopher De Sa,Christopher Ré
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 59 pages, 9 figures, 2 tables

点击查看摘要

Abstract:Recent work has established that power-law spectral conditions on data enable tight convergence bounds for deterministic gradient descent, resolving the conflict between classical exponential bounds and observed power-law learning curves. In this work, we extend this result to the stochastic regime of high-dimensional machine learning. We provide two main contributions: (1) We generalize the power-law spectral theory to Stochastic Gradient Descent (SGD), showing that the same spectral exponents govern stochastic dynamics; (2) For the fundamental case of isotropic Gaussian data, we provide a precise analysis of data shuffling, deriving exact constants that prove Single Shuffle is strictly superior to Flip-Flop and IID sampling. Our results bridge the gap between abstract spectral theory and practical stochastic training choices, offering a unified picture of how data geometry drives optimization speed.

[LG-159] Understanding LLM Parameter Update Sparsity through the Lens of Fisher

链接: https://arxiv.org/abs/2609.36262
作者: Yufan Zhang,Sagnik Mukherjee,Hao Peng
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Recent studies have observed that parameter changes during language-model post-training can be concentrated in a small subset of coordinates. This phenomenon has been reported in reinforcement learning, on-policy distillation, and supervised fine-tuning on near-policy data. Its recurrence across different post-training paradigms suggests shared structure in training dynamics. In this paper, we examine this pattern through the diagonal model Fisher, which measures the sensitivity of the model’s output distribution to individual parameters and is independent of any particular reward or teacher signal. Theoretically, we show that small diagonal Fisher leads to small expected gradients across a range of training objectives, providing a common explanation for sparse gradient updates. Empirically, we test this connection in RL and OPD. We find that Fisher identifies where gradients are concentrated, and fixed sparse masks selected from the initial Fisher retain a large proportion of the improvement from full training. Finally, we investigate the mechanisms underlying low Fisher in on-policy training. Our results show that high-probability next tokens tend to have similar parameter sensitivities, contributing to low Fisher. Together, these results establish the diagonal model Fisher as a unifying perspective linking update sparsity to on-policy training dynamics in LLM post-training.

[LG-160] CyFA: Linear Sequence Modeling with Relative-Time-Partitioned Memory

链接: https://arxiv.org/abs/2609.36259
作者: Yixiao Chen,Shuojin Yang,Shi-Min Hu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Linear RNNs offer linear-time sequence processing and constant-memory decoding, but their fixed-size recurrent states must accommodate all past key–value associations. Existing forgetting mechanisms and Delta Rule updates reduce interference by selectively clearing or correcting the state, yet earlier associations can still become difficult to retrieve. We introduce CyFA (Cyclic Flow Attention), a Linear RNN with relative-time-partitioned memory. At each step, a learned clock controls the cyclic transport applied jointly to the key and value states before the current key–value pair enters the age-zero slot, thereby organizing stored associations across relative-time slots. We further derive an exact change to absolute-clock coordinates that expresses CyFA as two scalar-decay linear attention recurrences and enables efficient chunk-wise training. Across 400M–1.4B pretraining experiments with matched recurrent-state sizes, CyFA improves recall-intensive performance while maintaining competitive language modeling and high computational efficiency. At 400M, CyFA outperforms KDA on FDA (42.60 vs. 26.07) while requiring only 46.7% and 48.3% of KDA’s forward and backward core-operator execution times, respectively. Our code is publicly available at \hrefthis https URLthis https URL.

[LG-161] he Signed Geometry of One-Shot Recourse: On-Path Validity and the Signed-Curvature Criterion

链接: https://arxiv.org/abs/2609.36252
作者: Hazar Yueksel(Google)
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: Main text 8 pages; 57 pages including appendix and checklist

点击查看摘要

Abstract:Closed-form recourse moves a rejected user along the unit gradient \hat g of the classifier score f by the promised distance d_p=|f(x)|/|\nabla f(x)| , at which the linearized score reaches zero. We ask when this one-shot step succeeds and what additional model queries change. To leading order the step ends on the favorable side exactly when the path curvature \kappa=\hat g^\top\nabla^2 f(x),\hat g is nonnegative. Across 80 shallow models, the fraction of rejected users whose step ends there and the fraction with \kappa\ge0 correlate at r=0.985 , although on Fashion-MNIST the first falls below the second by 8.2 points on average. No rule that uses only the score value and gradient can be valid for every score with path curvature bounded by K without overshooting some by order Kd_p^2/|\nabla f(x)| . When the curvature is also Lipschitz and the step is short, one evaluation of f at the promised point attains the minimax rate among deterministic one-query rules that know the curvature bound and its Lipschitz constant, and split-conformal calibration makes such a rule reach the first crossing or abstain with probability at least 1-\delta . Training with an asymmetric curvature penalty lets 99-100% of paths cross within the promised step on undershoot-prone shallow data, at about 4-22 times the overshoot of symmetric penalties (Fashion-MNIST, COMPAS). Because \kappa and d_p depend on how the score is scaled, part of this gain can be a longer promised step, and at matched validity a smaller audit of briefly trained models finds no uniform advantage over tuned inflation. Where a per-user line search along the ray is affordable, it is exact to grid resolution and preferable.

[LG-162] Action Chunking Proximal Policy Optimization with Feedback Correction NEURIPS2026

链接: https://arxiv.org/abs/2609.36250
作者: Sanghyun Hahn,Jonghyun Choi
类目: Machine Learning (cs.LG); Robotics (cs.RO)
*备注: Accepted at NeurIPS 2026

点击查看摘要

Abstract:Action chunking provides temporal abstraction in reinforcement learning by selecting short action sequences instead of individual actions, but many existing approaches face two limitations in high-dimensional robotic control. First, many rely on value functions over action chunks, which can be difficult to learn as action dimensionality and chunk length grow. Second, executing chunks open-loop removes within-chunk feedback, limiting reactivity in contact-rich tasks. We present Action Chunking PPO (ACPPO), a PPO extension that uses a chunked actor while retaining a standard state-value critic, thereby avoiding chunked Q-functions. We further propose ACPPO-Corr, which augments the chunk planner with a stepwise feedback corrector that adjusts planned actions online within each chunk. Across 25 simulated robotics tasks from IsaacGym and Bi-DexHands, spanning locomotion, arm manipulation, and dexterous hand-object interaction, ACPPO-Corr achieves the strongest aggregate performance among evaluated methods and performs best on both decision-frequency-sensitive and decision-frequency-neutral task subsets. Ablations show that moderate chunk lengths work best and that corrector regularization is important for balancing chunk-level planning with local feedback. These results suggest that action chunking can be effective in online PPO when chunk-level planning is paired with closed-loop correction. The code is available at: this https URL.

[LG-163] Can Representation Learning Decouple from Loss Minimization? Polar Updates Have an Answer

链接: https://arxiv.org/abs/2609.36240
作者: Akash Kumar
类目: Machine Learning (cs.LG); Optimization and Control (math.OC); Machine Learning (stat.ML)
*备注: 105 pages, 21 figures

点击查看摘要

Abstract:Does representation learning stop when the training loss stops improving? We study this question for matrix Muon, whose polar-normalised updates have a step length set by the gradient’s rank rather than its norm. Near the edge of stability, full-batch Muon on teacher-student problems enters approximately period-2 loss oscillations that persist for thousands of steps: the cycle-mean loss stays flat or rises, yet the weights keep moving and the learned features continue to align with the teacher subspace. For linear teacher-student learning toys, we derive explicit cycle and alignment formulas and conditional plateau and decay bounds. For a population mean-field ReLU model, we prove that, under stated dimension, initialisation and small-head conditions, the leading eigenspace of the average gradient outer product (AGOP) recovers the teacher subspace exactly during a loss plateau, before the loss later drops. In all 33 ReLU, GELU and SiLU teacher configurations we study, direction-only alignment metrics show the student AGOP aligned with, or still aligning to, the teacher subspace during the period-2 oscillations; projected head refitting on selected configurations shows that the learned directions are useful for prediction, and further measurements distinguish AGOP alignment from weight-mass concentration. In deep residual ReLU students, freezing the downstream layers while the first layer trains with full-batch exact polar updates recreates a nearly flat cycle-mean loss with improving input-AGOP alignment; freezing and unfreezing switch between this plateau and loss decrease, and the effect is sensitive to momentum and to the choice of orthogonaliser.

[LG-164] ransversal Pooling Neural Networks

链接: https://arxiv.org/abs/2609.36237
作者: Emily J. King,Dustin G. Mixon,Michael Perlmutter,Lander Ver Hoef
类目: Machine Learning (cs.LG); Functional Analysis (math.FA)
*备注:

点击查看摘要

Abstract:Many learning tasks require stability to small transformations while retaining sensitivity to larger ones. We introduce \emphtransversal pooling neural networks, which generalize spatial max pooling to affine group actions. We establish equivariance to a chosen subgroup and derive explicit stability bounds for individual pooled wavelet coefficients under affine perturbations of the input. Experimentally, we demonstrate the utility of our networks in low-data environments and for predicting tropical cyclone intensification.

[LG-165] he Role of Feed-Forward Layers in Transformer Dynamics

链接: https://arxiv.org/abs/2609.36230
作者: Thomas Jacob Maranzatto,Semih Akkoc,Sennur Ulukus
类目: Machine Learning (cs.LG); Dynamical Systems (math.DS)
*备注: 9 pages, 5 figures

点击查看摘要

Abstract:We study the dynamical behavior of tokens in transformers from a control-theoretic perspective. Our model includes the feed-forward layer present after the self-attention mechanism, with the self-attention mechanism interpreted as an interacting particle system and the feed-forward layer as an independent control. Our main theoretical result establishes that the feed-forward network can steer the tokens arbitrarily close to consensus regardless of the key, query, and value matrices. Our result are easily extended to convergence to many clusters and to multi-head attention. We conduct numerical experiments to verify our results, and compare thresholding behavior from our theory to real-world LLMs.

[LG-166] BASE: Batch-Aware Selection of Experts Using Predicted Removal Error for Efficient MoE Decoding

链接: https://arxiv.org/abs/2609.36222
作者: Ali Abbasi,Justin Shi,Soheil Kolouri
类目: Machine Learning (cs.LG)
*备注: 19 pages, 4 figures, 8 tables

点击查看摘要

Abstract:Large language models are increasingly expensive to serve. In large-scale serving systems, autoregressive decoding is often bottlenecked by transferring model weights from accelerator high-bandwidth memory into on-chip SRAM. Mixture-of-experts (MoE) models reduce computation by activating only a small subset of experts per token, but this sparsity does not translate directly to batched decoding. Different requests select different experts; therefore, the combined active set across many concurrent requests can span a substantial fraction of the expert pool and require significantly more expert weights to be transferred. Most expert-reduction techniques make retention decisions independently for each token and therefore do not address this batch-level expansion. More recently, batch-aware methods have attempted to coordinate expert use across concurrent requests and reuse experts already fetched for the batch. Yet their selection criteria are based primarily on router rankings or expert statistics collected during calibration. Consequently, these criteria are not directly tied to the output error caused by dropping an expert, nor do they capture how an expert’s contribution changes across tokens at inference time. We instead rank experts according to how much their removal would change the MoE-layer output. To apply this criterion during serving, we train a lightweight linear predictor during calibration that estimates the expert removal cost for each incoming token, and develop custom GPU kernels for cost prediction and expert selection. Across three MoE architectures, BASE improves the quality-efficiency tradeoff without retraining. On Qwen3-30B-A3B, it improves average accuracy by 29.5 points over the strongest baseline at comparable throughput under a tight expert budget. At a higher expert budget, it is 60% faster than dense inference while remaining within 0.4 accuracy points.

[LG-167] How Language Models Differ in Redistributing Attention-Head Activity Under Serial Demand

链接: https://arxiv.org/abs/2609.36221
作者: Johnny Jingze Li,Abdulla Kuleib,Kalyan Basu,Gabriel A. Silva
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:The way a model distributes activity over each layer’s attention heads offers a coarse view of how it routes information through depth; how this changes with the task is part of what a mechanistic account must explain. Holding prompt length fixed, we vary how many serial steps a task demands and measure, in every layer of 17 open-weight models, whether activity concentrates on a few heads or spreads across many as demand rises. Both occur: in most models, layers just before mid-depth concentrate activity and later layers spread it. Models differ in where and how strongly this happens. The Qwen2.5 base models from 0.5B to 7B, for example, spread less than the average model in every task and concentrate activity in parts of their second half, where Llama models from 1B to 8B and OLMo-2 spread; the contrast largely holds between Llama-3.1-70B and Qwen2.5-72B, which have the same number of layers and heads. These differences are reproducible, and post-trained models keep much of their base model’s pattern. An ablation study suggests that, within a task, models whose activity is more concentrated on their top heads also depend more on those heads for the answer. Concentration and spreading across layers thus offer a new way to compare models, by how they route information through depth. Code is available at this https URL.

[LG-168] Preconditioned Physics-Informed Neural Operator Training

链接: https://arxiv.org/abs/2609.36216
作者: Shizheng Wen,Siddhartha Mishra,Marius Zeinhofer
类目: Machine Learning (cs.LG); Numerical Analysis (math.NA)
*备注:

点击查看摘要

Abstract:Neural operators are typically trained in a supervised fashion, which requires a dataset to be generated with a classical solver. Training them physics-informed, i.e., purely from the governing equations, removes this large offline cost and allows fresh samples to be drawn at every optimization step, but has so far been limited to simplified problems and trails supervised training in accuracy. The obstacle is the ill-conditioning of physics-informed losses, which differential operators induce and which worsens as the discretization is refined. We therefore propose a preconditioned residual loss function and show mesh-independent conditioning for elliptic problems and greatly improved conditioning for saddle point problems. Realized through geometric and algebraic multigrid, the construction applies to linear and nonlinear equations, steady or time-dependent, on structured and unstructured meshes, is agnostic to the neural operator architecture, and adds no cost at inference. On the Poisson, Allen-Cahn and stationary Stokes equations, the resulting label-free training matches supervised training and is four to twenty-five times more accurate than previous physics-informed operator learning methods.

[LG-169] EnergyEminence: Source-Aware Environmental Calibration and Evaluation in a Physics-Grounded Grid Digital Twin

链接: https://arxiv.org/abs/2609.36215
作者: Huy Trinh,Michael Mai,Yu Nong
类目: Machine Learning (cs.LG); Image and Video Processing (eess.IV)
*备注:

点击查看摘要

Abstract:Power-grid digital twins must combine data-driven prediction with physically meaningful state evolution while preserving the provenance of environmental observations. This paper presents an early-stage EnergyEminence testbed that couples an IEEE 118-bus-style graph-temporal predictor, nonlinear AC cascade simulation, and operator-dashboard-like temporal replay. In addition, we introduce a shared bounded calibration that converts wildfire-detection confidence and spatial extent into source-comparable wildfire interpretable and explainable evidence. We then evaluate it with visually diverse fire and hard-negative videos. Sixteen synthetic environmental videos are curated to generate 160 source-tracked grid scenarios, and a source-video-disjoint test yields 10 true positives, 8 false positives, 22 true negatives, and no false negatives. The errors occur in stressed, non-cascading scenarios conditioned on an unseen growing-fire source. Our diagnostic then reveals environmental shortcut learning that is obscured by scenario-level random splitting. The paper therefore contributes a data-centric and inspectable evaluation workflow for multimodal grid-resilience models, together with evidence supporting separation of environmental alerting from electrical cascade inference

[LG-170] Physical Cross-Modal Masked Autoencoding for Seismic-to-Well Representation Learning

链接: https://arxiv.org/abs/2609.36193
作者: Meher Gajula,Keyla Gonzalez,Ben Lasscock,Alejandro Valenciano
类目: Machine Learning (cs.LG); Geophysics (physics.geo-ph)
*备注: 11 pages, 3 figures

点击查看摘要

Abstract:Learning from scientific measurements often requires aligning modalities with different spatial support and resolution. Subsurface characterization is an extreme dense-sparse problem in which 3D seismic provides volumetric but indirect measurements and well logs provide high-resolution 1D measurements at sparse borehole locations. We introduce a physically grounded cross-modal masked autoencoder (CM-MAE) for seismic-to-well representation learning. The model jointly tokenizes seismic volumes and well-log depth patches, embeds both modalities in continuous physical coordinates using four-axis rotary position embeddings, and reconstructs masked targets with a cross-modal decoder. Sparse Mixture-of-Experts layers provide modality-specific capacity while dense attention allows information exchange between modalities. Pretraining uses 23 seismic volumes covering approximately 178,000 square km across U.S. offshore and onshore basins, together with approximately 92,000 wells. Matched-mask ablations show strongly asymmetric information flow. Seismic context improves held-out well-log reconstruction by 9.73%, while well-log context improves seismic reconstruction by 0.65%. We evaluate seismic-only pseudo-log predictions against independent interpreter-drawn salt-geobody masks to determine whether they contain geologic signal. Across offshore U.S. surveys, compressional-slowness-derived salt scores reach an AUROC of 0.910. In onshore basins, predicted compressional slowness preserves formation-scale structure and partial relative organization in an unseen survey despite calibration drift. CM-MAE learns useful seismic-to-well representations under extreme modality asymmetry, although absolute pseudo-log calibration remains survey-dependent.

[LG-171] EvoMO-SR: Multiobjective LLM -based Evolution of Symbolic Expressions with substructure guidance

链接: https://arxiv.org/abs/2609.36187
作者: Cristina Rossetti,Anna V. Kononova,Thomas Bäck,Fei Liu,Niki Van Stein
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Symbolic Regression (SR) is a data-driven method for scientific discovery which searches for interpretable analytical relationships within data. Recently, Large Language Models (LLMs) have also had a significant impact on scientific discovery, enabling the automation of various stages of the process. For these reasons, the possibility of harnessing the embedded scientific knowledge and programming capabilities of LLMs to solve SR tasks has emerged, showing promising performance compared with traditional methods. We propose EvoMO-SR, a novel LLM-driven SR framework in which the LLM generates equation skeletons, with their coefficients fitted separately by an external optimizer. The framework includes a multi-objective survival selection which controls bloating by balancing accuracy and complexity, and a substructure guidance mechanism which mutates expressions with candidate reusable building blocks. EvoMO-SR achieves the best accuracy in seven of the eight in-domain and out-of-domain settings for LSR-Synth, using a small LLM model, i.e., Llama-3.1-8B-Instruct. We also evaluated structural recovery through two symbolic accuracy metrics based on canonicalized subtree overlap and term matching, showing that our method has a greater probability of recovering highly accurate symbolic structures.

[LG-172] Fair Policy Optimization in Major-Minor Weakly Coupled Markov Decision Processes

链接: https://arxiv.org/abs/2609.36174
作者: Xiaohui Tu,Yossiri Adulyasak,Erick Delage
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We consider fair resource allocation in sequential decision-making environments modeled as major-minor weakly coupled Markov decision processes (M2WCMDP). In this framework, resource constraints couple the action spaces of a major sub-Markov decision process (sub-MDP) and a population of minor sub-MDPs that would otherwise operate independently. Instead of using the traditional utilitarian (total-sum) objective, we optimize a general class of monotone, concave, permutation-invariant, normalized fairness functions. With homogeneous minor sub-MDPs, we prove that the problem under symmetry reduces to optimizing the platform-plus-mean-participant utilitarian objective over the class of \textitpermutation-invariant policies, which allows us to exploit efficient algorithms that optimize the utilitarian-based objective to solve this fairness-aware problem. For more general settings, we introduce a count-proportion-based deep reinforcement learning approach with a priority-based sampler that generates feasible count actions. The generality of our framework means that the proposed algorithms and theoretical guarantees transfer to any domain with a symmetric M2WCMDP structure. We consider two applications: the machine replacement problem and the joint control of pricing and taxi relocation problem on a New York City-calibrated dataset. We validate our theoretical findings with comprehensive experiments, confirming the effectiveness of our proposed method in achieving strong fairness-aware performance while remaining scalable.

[LG-173] Privacy-Friendly Cohort Determination: Sealed CSP-Independent In-Browser ML Inference of Professional Segments for Identity-Less Advertising

链接: https://arxiv.org/abs/2609.36153
作者: Om Shankar Tiwari,Navnit Shukla,Guanyu Wang,Akshay Jain
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注: 11 pages, 5 figures, 2 tables. Ancillary code, crawl data and full proofs are included

点击查看摘要

Abstract:B2B advertising targets a viewer’s professional attributes (employer size and industry, function, seniority) and has obtained them by matching identities across sites. Safari and Firefox block third-party cookies, Google retired the Privacy Sandbox cohort APIs in 2025, and reverse-IP firmographics decay under remote work. We present SIF (Sealed Inference Frame), which infers coarse professional cohorts on the device and emits only a locally differentially private, taxonomy-coded label into the OpenRTB bid stream, with no cross-site identifier. It rests on a property of the web platform we make precise: a navigated cross-origin iframe is the only way third-party code obtains a policy it controls, so inference runs in WebAssembly even where the publisher’s CSP forbids it, and a nested worker served with default-src ‘none’ gives the model no network. Even a malicious model leaks at most about 5 bits per site per week. Labels pass through a memoised k-ary randomised response keyed to the publisher’s first-party identifier, which gives \varepsilon -local differential privacy, defeats averaging, and links requests no better than the identifier already sent. An org-conditional k-anonymity rule suppresses cells, more strictly on corporate networks than at home. Cohorts ride OpenRTB this http URL in a LinkedIn-aligned taxonomy, and attribution uses LinkedIn’s click-scoped li_fat_id without bridging identities. We report a crawl of CSP deployment on 7,969 top sites and 431 B2B publishers, Heavy-Ad budgets, closed-form privacy-utility trade-offs, a re-identification simulation, and an assessment of which attributes are predictable at all: company type and size are, seniority largely is not. On-device is a design property, not a consent exemption.

[LG-174] Learning Continuous Patient Trajectories from Electronic Health Records

链接: https://arxiv.org/abs/2609.36144
作者: Silas Ruhrberg Estévez,Kara Liu,Christopher Chiu,Benjamin Atta Owusu,Umesh Kadam,Russ B. Altman,Mihaela van der Schaar
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Electronic health records provide irregular observations of latent patient states that evolve continuously over time. Recent autoregressive models condition on clinical histories to forecast future events as sequences of discrete observations. Conversely, multi-marginal flow matching provides a continuous-time formulation, but using multiple observations to supervise training paths does not itself give the learned dynamics access to preceding patient history. We introduce EHRFlow, a multi-marginal flow-matching framework that conditions on encoded patient history, thereby allowing future dynamics to depend on the patient’s prior clinical trajectory. Our proposed framework accommodates irregular observation times and supports forecasting at arbitrary horizons. Across controlled synthetic benchmarks, EHRFlow improves clinical-code forecasting and latent-state recovery. On real-world clinical datasets comprising more than one million patients, including an independent external validation cohort, EHRFlow improves horizon-averaged top-5 clinical-code accuracy over autoregressive and history-independent flow-matching baselines. Finally, in a controlled counterfactual simulation, conditional guidance approximates the known effect of an antihypertensive intervention without training a task-specific outcome model.

[LG-175] Why Backdooring Neural Networks is so Easy?

链接: https://arxiv.org/abs/2609.36117
作者: Issam Seddik,Mohamed El Amine Seddik
类目: Machine Learning (cs.LG); Cryptography and Security (cs.CR)
*备注:

点击查看摘要

Abstract:Securing modern AI systems against backdoor attacks remains an open challenge and requires fundamentally principled estimates of the adversary’s budget – the poison fraction \pi and trigger strength \alpha needed to construct successful yet stealthy attacks. Motivated by recent empirical evidence that poisoning large language models can require a nearly constant number of malicious samples even as clean datasets grow, we derive an exact closed-form analysis of a quadratic neuron trained on a poisoned Gaussian mixture. We show, perhaps counterintuitively, that the same feature-learning dynamics that make neural networks powerful can also make them more vulnerable to backdoors. Specifically, with clean accuracy preserved to first order, O(\pi) , we demonstrate that lazy learning imposes the inverse-square-root scaling \alpha \propto \pi^-1/2 for a successful attack, while feature learning induces a quadratic detector whose loss margin scales as O(\alpha^4) , improving the attack budget to \alpha \propto \pi^-1/4 . Consequently, nonlinear feature learning substantially reduces the trigger strength required at small poison fractions, thereby in a sense making feature learners more backdoor vulnerable. These results provide a theoretical mechanism consistent with large-scale empirical observations and demonstrate that security audits based on linear heuristics can systematically underestimate backdoor vulnerability in the widely adopted feature-learning regimes.

[LG-176] Dyad: Extending Large Language Models with Native Typed Decision-Making

链接: https://arxiv.org/abs/2609.36116
作者: Yundaichuan Zhan,Weishi Wang,Wenbiao Liu,Daniel Dahlmeier,Chengwei Qin,Juncheng Li,Fredrik D. Johansson,Zhongqi Yue
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We study how to build more capable general-purpose agents by extending large language models (LLMs) with native typed decision-making. We introduce Dyad, an architecture that augments a pretrained LLM with an environment-conditioned action encoder that embeds each candidate action description in parallel, then scores these embeddings against the LLM’s internal state to yield a distribution over typed actions. By factorizing decision-making into representations of the evolving interaction state and environment-specific action semantics, Dyad introduces an inductive bias for learning reusable representations while keeping action scoring efficient even as the action space grows. We investigate two complementary reinforcement learning settings driven by environment interaction. With the LLM frozen, training the action encoder alone achieves consistent gains across four unseen environments, enabling modular adaptation without modifying any LLM parameters. Jointly optimizing both components outperforms conventional RL post-training across diverse interactive tasks and model scales, including a 3.80% average absolute gain on ALFWorld with a 9B model, while improving general knowledge, reasoning, and coding.

[LG-177] Koa-action: Fast and Consistent Structured Decision Making with Generative LLM s

链接: https://arxiv.org/abs/2609.36115
作者: Shenghong Dai,Shiva Kumar Pentyala,Yingchi Liu,Shubham Mehrotra,Suman Banerjee,James Zhu,Bin Bi,Sitaram Asur,Phil Mui
类目: Machine Learning (cs.LG)
*备注: 19 pages, 7 figures

点击查看摘要

Abstract:Industry applications often demand low-latency classification, yet current large language model (LLM) approaches remain poorly suited for latency-critical applications. Existing prompting and constrained decoding produce verbose, multi-token outputs that require expensive token-by-token generation, while encoder-based models achieve faster inference but sacrifice task flexibility. We propose Koa-action, a framework for low-latency atomic actions – fast, single-step decisions such as classification, semantic endpointing, Boolean checks, and scoring – formulated as constrained generation with single-token outputs. By introducing atomic label tokens and applying supervised fine-tuning, our method reduces classification to a deterministic one-step decoding problem. Across standard benchmarks, Koa-action delivers competitive accuracy with consistently low and stable latency. On a production intent-routing benchmark, Koa-action reaches 85.5% accuracy – competitive with the strongest frontier models (Claude-4.8-Opus, Gemini-Pro-3.1) and ahead of GPT-5 and Gemini-2.5-Pro – while answering in about half a second, several-fold faster than every frontier model (up to ~7.5x at the median) under identical serving conditions. Against the dedicated single-token system Jev/TypeSafe, Koa-action is competitive on accuracy and faster at the median, while also handling multimodal inputs and multi-label outputs that single-label text systems do not.

[LG-178] Data Unlearning via Inverse Distillation

链接: https://arxiv.org/abs/2609.36099
作者: Aleksei Leonov,Nikita Kornilov,Zhenhe Zhang,Evgeny Burnaev,Iaroslav Koshelev,Alexander Korotin
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Multi-step matching models, including flow and diffusion models, produce high-quality outputs but incur substantial inference costs and may reproduce unwanted components of their training datasets. We introduce Inverse Distillation Unlearning (IDU), a unified framework that simultaneously distills a teacher multi-step matching model into an efficient one-step student generator and suppresses outputs corresponding to a designated training subset. We first formulate distillation as a min-max objective over a data distribution and then represent this distribution as a mixture of the forget-set and the generated distributions. This allows us to compare this mixture with the teacher’s training distribution and recover only the retained data at the optimum. Our method requires only a pretrained full-data teacher and data from the forget set, without access to retained training examples, extra feature extractors or classifiers. Extensive experiments on MNIST and CIFAR-10 datasets under flow-matching and score-based diffusion settings demonstrate that IDU substantially reduces the generation frequency of forgotten classes while preserving generation quality on the retained classes. To the best of our knowledge, IDU is the first unified framework for simultaneous unlearning and distillation in unconditional flow-matching and score-based models.

[LG-179] Early Learning Shapes Later Directions Of Representation Change In Continual Learning

链接: https://arxiv.org/abs/2609.36081
作者: Yuantao Deng,Jinnuo Liu,Kaizhen Tan,Yuchen Liu
类目: Machine Learning (cs.LG)
*备注: 35 pages, 7 figures

点击查看摘要

Abstract:Representations continually change as a network learns new tasks. We ask whether early representational changes naturally form a geometric structure that continues to shape later learning. We identify a low-dimensional subspace of early representation drift, which we call a scaffold, and test whether it is reused across subsequent tasks. Across four pretrained visual encoders and two datasets, later representational changes consistently favor this early-defined subspace over matched random alternatives. This reuse is history-dependent: when networks experience different early tasks but identical later training inputs, each network preferentially reuses the scaffold induced by its own learning history. The same preference appears in individual optimizer updates, even though the network’s dominant local response directions shift away from the original scaffold. Finally, constraining motion within the scaffold slows new-task acquisition more than matched random constraints, while effects on old-task retention are less consistent. In summary, these results suggest that early experience leaves a persistent geometric imprint on how neural networks adapt to future tasks.

[LG-180] GEM-KMeans: Memory-Efficient and Accurate Clustering on Massive Scale with GPU Optimization

链接: https://arxiv.org/abs/2609.36074
作者: Peng Xu,Nihar Koganti,Volodymyr Kindratenko,Xiaohui Chen
类目: Machine Learning (cs.LG); Hardware Architecture (cs.AR); Performance (cs.PF)
*备注:

点击查看摘要

Abstract:Memory-efficient scaling on clustering problems without sacrificing statistical accuracy is of central interest for large-scale data analysis and machine learning problems. Nonnegative low-rank (NLR) matrix factorization for K -means is a scalable clustering method, which connects to semidefinite relaxations with optimal average-case exact recovery guarantees. However, a direct GPU implementation of NLR requires multiple large factor-sized buffers and substantial data movements that are essentially memory-bound. In this paper, we introduce GEM-KMeans, a spectrally normalized yet mathematically equivalent NLR formulation that fuses the gradient update, nonnegative projection, and sufficient statistics for normalization and iterate movement into a matrix-multiplication epilogue. Instead of retaining three massive factor-sized arrays, our IO-aware GPU implementation materializes only one single factor with small tile-reduction arrays as additional storage in the High Bandwidth Memory (HBM). We derive explicit memory costs and spectrally normalized smoothness bounds for optimizing the clustering objective function. Accurate clustering is demonstrated at massive scales on synthetic and real datasets, where performance gains of GEM-KMeans over existing GPU-accelerated Lloyd’s algorithms involve data-dependent runtime tradeoffs.

[LG-181] Mixture-of-Kittens: MoE Megakernel for NVL72s

链接: https://arxiv.org/abs/2609.36070
作者: Stuart H. Sul,Nash Brown,Henry Wildermuth,William Lin,Federico Cassano,Christopher Ré
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:AI accelerator systems are rapidly consolidating into scale-up architectures, where tens to thousands of GPUs communicate over high-bandwidth, single-hop fabrics. We find that existing Mixture-of-Experts (MoE) training systems, optimized for conventional scale-out networks, transfer poorly to this setting, often running slower than a naive baseline built with PyTorch and NCCL. With industry roadmaps pointing toward even larger scale-up domains, understanding the performance tradeoffs of this hardware regime is increasingly important. We present Mixture-of-Kittens (MoK), an MoE training system designed for Nvidia NVL72. MoK builds on three insights that unlock performance on scale-up domains: (1) choosing push- or pull-based communication per operator, (2) restructuring the computation-communication overlap, and (3) fully eliminating CPU-GPU synchronization. MoK distills these insights into a single deterministic training megakernel that fuses token dispatch, shared and routed expert FFNs, and token combine. Across MoE layer shapes from four widely used open-weight models, MoK delivers up to 2.37\times the throughput of the strongest publicly available baseline. In a production run on 512 GPUs spanning multiple GB300 NVL72 racks, MoK improves end-to-end training throughput by 1.41\times .

[LG-182] Introducing the CZAR Loss: A Tailored Objective Function for Financial Log-Return Predictions

链接: https://arxiv.org/abs/2609.36061
作者: Joel Pfeffer(1),J. M. Diederik Kruijssen(1),Florian Stecker(1),Steven N. Longmore(1,2) ((1) Allora Foundation, (2) LJMU)
类目: Machine Learning (cs.LG); Computational Engineering, Finance, and Science (cs.CE); Computational Finance (q-fin.CP); Mathematical Finance (q-fin.MF)
*备注: 22 pages, 12 figures, 2 tables; appeared in ADI (September 2026)

点击查看摘要

Abstract:In quantitative finance, standard regression losses are misaligned with the economics of return prediction. As the conditional mean of financial log-returns is close to zero, symmetric losses such as the mean squared and mean absolute errors make the constant zero forecast a near-optimal solution, penalizing models with genuine but noisy directional skill. This applies both during training, where predictions shrink toward zero, and during evaluation, where trivial forecasters can lead loss-based rankings. Under a Gaussian linear prediction model, we show that all symmetric monotonic losses share a universal breakeven directional accuracy against the zero predictor, which rises sharply and becomes unobtainable as the prediction noise approaches the standard deviation of the returns. We introduce the CZAR (Composite Zero-Agnostic Return) loss function, a piecewise quadratic loss built around five key requirements: convexity in the prediction, asymmetry with true return direction that vanishes at zero, near-linear penalization of undershoots and wrong-direction predictions, divergence for large errors, and an adaptive loss floor for evaluation. CZAR is provably convex in the prediction at fixed true value, has closed-form gradient and Hessian suitable for custom objectives in gradient-boosted libraries, and its four hyperparameters reduce to a single choice through correlated defaults. In idealized tests, the minimum directional accuracy required for a CZAR-evaluated forecaster to outperform the zero predictor under mean log loss remains near the 50% chance level, whereas the corresponding threshold for symmetric losses rises sharply with prediction noise. This advantage persists under heavy-tailed return distributions. In a LightGBM experiment on intraday BTC log-returns, CZAR-trained models reduce the `zero-returns bias’ and improve directional accuracy on large-magnitude returns.

[LG-183] ABC: Advantage-Based Control Variates for Reinforcement Learning with Verifiable Rewards

链接: https://arxiv.org/abs/2609.36058
作者: Hsiao-Ru Pan,Florent Draye,Bernhard Schölkopf
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Recent progress in reinforcement learning with verifiable rewards (RLVR) has highlighted the effectiveness of simple critic-free policy-gradient methods such as Group Relative Policy Optimization (GRPO). In contrast, actor-critic methods rely on learned value functions whose approximation error can introduce bias through commonly used advantage estimators such as temporal-difference error. Motivated by this observation, we revisit trajectory-level control variates through an advantage-value formulation, which we call Advantage-Based Control Variates (ABC). This formulation reveals that the covariance structure is closely related to the return decomposition used in Direct Advantage Estimation (DAE). Finally, we combine ABC with DAE into a single actor-critic algorithm and evaluate it in an offline-to-online RLVR setting, where the critic is first trained on previously collected trajectories and adapted during online learning. On mathematical reasoning tasks, ABC achieves performance competitive with GRPO using substantially fewer online optimization steps.

[LG-184] Multi-Class Multi-Tier Network Intrusion Detection: A Comprehensive and Reproducible Benchmark

链接: https://arxiv.org/abs/2609.36039
作者: Yufeng Xin,Bryant Goseland,Mohamed Rahouti
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Machine learning (ML) and deep learning (DL) have dominated Intrusion Detection System (IDS) research in recent years. Unfortunately, many existing studies have produced inflated results and unreliable benchmarks due to critical oversights and mistakes in the ML and DL pipeline, from data collection and labeling to feature engineering and model training and evaluation. CIC-IDS2017 is a standard benchmark for network intrusion detection. Still, many published results on this dataset are difficult to compare due to labeling errors, inconsistent flow extraction, potential leakage, and performance evaluation metrics dominated by benign traffic. In this paper, we present a comprehensive benchmark with corrected PCAP-level labeling and a complete evaluation pipeline with diverse ML models. We evaluate eleven tabular classifiers at three nested levels: binary attack detection, nine-class attack-family attribution, and fifteen-class fine-grained classification. A soft-voting ensemble of Random Forest, XGBoost, and LightGBM obtains the best fine-tier macro-F1 of 0.955, with coarse and binary macro-F1 scores of 0.980 and 0.999, respectively. We further conducted a feature selection study based on an analysis of feature importance. This comprehensive benchmark pipeline is configurable and open-source, enabling new feature extraction and model plugins for new datasets. Future work should use this pipeline as a reference point for richer features, rare-class analysis, and model generalization towards new datasets and attack classes.

[LG-185] Preferent Compression Bounds Are Tight

链接: https://arxiv.org/abs/2609.36030
作者: Dario Paccagnan,Marius Tirlea
类目: Machine Learning (cs.LG); Systems and Control (eess.SY); Optimization and Control (math.OC); Probability (math.PR)
*备注: 14 pages, 6 figures, to appear at the 65th IEEE Conference on Decision and Control (CDC 2026)

点击查看摘要

Abstract:The lack of rigorous safety and performance certificates remains a key bottleneck to the deployment of modern learning-based methods. Sample compression has recently emerged as a powerful tool for deriving such certificates, with particularly sharp bounds available for algorithms satisfying a so-called preference property – also known as stability in the learning theory literature. These bounds find direct application across domains as different as the Scenario Approach, Pick-to-Learn, and Support Vector methods. However, whether they are tight has remained an open problem. In this paper we resolve this question affirmatively and show that the state-of-the-art bound for preferent compressions is provably tight. We establish this by exhibiting an explicit construction based on the uniform distribution and order statistics that attains the bound in the limit. Along the way, we also provide a considerably shorter and more accessible proof of this bound, requiring only elementary counting arguments and no infinite-dimensional duality.

[LG-186] In-Context Learning for Robots: Methods and Applications

链接: https://arxiv.org/abs/2609.36012
作者: Haojian Huang,Zexi Li,Junhao Guo,Yehang Zhang,Wenxuan Peng,Bohan Zhou,Weilin Ruan,Leyi Wu,Chenxu Wang,Jianchong Su,Binghui Xie,Wosong Chen,Yingjie Xu,Tianhao Zhou,Suzeyu Chen,Pukun Zhao,Jiaqi He,Xinyi Li,Runze Li,Peiran Dong,Shaoxiang Dang,Jing Huang,Yingbing Chen,Yifan Chang,Tianyi Zhang,Shiyuan Deng,Haozhi Wang,Yangkai Wei,Wenqian Li,Han Yang,Kaiwen Zhou,Huaping Liu,James Cheng,Rui Shao,Donglin Wang,Yaochu Jin,Jianye Hao,Ying-Cong Chen,Yinchuan Li
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注: 100 pages, 26 figures, 25 tables. Project page: this https URL ; Code and literature: this https URL

点击查看摘要

Abstract:General-purpose robots must infer what a new task requires and translate that understanding into appropriate physical action. In-context learning (ICL) for robots supports this process by using demonstrations and interaction to direct existing competence with neural parameters held fixed during deployment. We organize this literature review around the interfaces connecting contextual evidence to execution, distinguishing four families: context-conditioned policies, geometric demonstration transfer, world-model-based control, and skill- and agent-based execution. Comparing these interfaces clarifies their transfer assumptions and the roles of training, correspondence, and memory in making context useful. Across manipulation and navigation, we examine how these mechanisms preserve taught requirements as objects, environments, and execution conditions change. This analysis links method design to evaluation practices that distinguish responsiveness to teaching, physical transfer, and benefits from retained experience. The resulting agenda connects compositional task acquisition and faithful transfer with physical recursive self-improvement, in which experience improves the ability to learn subsequent tasks.

[LG-187] FluxLite: Inference-Time Proposal Control for Discrete Diffusion Models NEURIPS2026

链接: https://arxiv.org/abs/2609.35947
作者: Yinuo Ren,Haoxuan Chen,Grant M. Rotskoff,Jiequn Han,Lexing Ying
类目: Machine Learning (cs.LG); Computation (stat.CO); Machine Learning (stat.ML)
*备注: Accepted at NeurIPS 2026

点击查看摘要

Abstract:Many inference-time tasks for pretrained discrete diffusion models and diffusion language models reduce to drawing samples from a tilted version of the pretrained distribution. Feynman-Kac sequential Monte Carlo (SMC) makes this correction exact in principle, but its prescribed weights routinely degenerate when the proposal dynamics are misaligned with the tilt, capping the practical gains from additional particles. We introduce FluxLite, a lightweight, training-free proposal-control framework for discrete diffusion. On the sparse directed graph of pretrained reverse rates, any sparse jump-rate perturbation can be exactly compensated by a q_t -weighted graph-divergence term in the Feynman-Kac potential; the target path is therefore preserved while the residual reweighting variance becomes a local convex objective. We instantiate this principle as two practical samplers: a one-hop local reallocation rule (HEU) and a small nonnegative quadratic program over pretrained-rate bases (D-VCG). We further prove population stability under the standard score-entropy training loss, identifying a tilted-path coverage factor that governs robustness to score error, together with finite-particle convergence for a fixed controlled Feynman-Kac recursion. Empirically, FluxLite improves over standard Feynman-Kac SMC baselines by up to two orders of magnitude in terminal KL on an analytically tractable finite-state CTMC benchmark, and reduces row-correlation MSE on 2D Ising sampling by 5-7x in geometric mean and up to 55x at peak.

[LG-188] EnJoi: Ensemble Joint Score Filter for Generative Data Assimilation

链接: https://arxiv.org/abs/2609.35944
作者: Julien Moreau,Marc Lelarge
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Data Assimilation (DA) aims to recover the full state of a dynamical system that is only partially observed. A solution is to use Score-based models to generate physically consistent trajectories that agree with the observations. These Autoregressive Diffusion models are trained by conditioning on the previous state; however, they do not take into account the uncertainty of their past predictions. We propose a new diffusion-based assimilation algorithm that dynamically balances the confidence in the current state and the new observations. Crucially, we choose to learn the distribution of the joint state containing both the past and future. This allows us to use a modified version of En4DVar, a classical DA algorithm that relies on the covariance of an ensemble of particles. Experiments on fluid and traffic flow simulations show improved reconstruction performance, especially in situations where observations are sparse and non-homogeneous.

[LG-189] KernelOnet: An Interpretable Neural Operator Based on Kernel Functions

链接: https://arxiv.org/abs/2609.35938
作者: Yuan Guo,Hanshu Chen,Qiang Xi,Timon Rabczuk,Zhuojia Fu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:This paper proposes an interpretable neural operator framework, the Kernel Operator Network (KernelOnet), which incorporates kernel functions explicitly into the neural operator architecture, so that the operator structure matches the kernel-expansion form used in boundary-type kernel-expansion methods. Unlike traditional neural operators such as DeepONet, which learn basis functions implicitly through deep networks, KernelOnet replaces the trunk network with explicit kernels and offers three complementary kernels: a data-driven learnable kernel, in which a neural network parameterizes a radial basis function learned from data, and which for constant-coefficient linear problems can be regarded as a non-singular fundamental solution; a physics-informed kernel, which embeds physical information such as analytic fundamental solutions into the network structure, so that the expansion satisfies the governing equation automatically and can be trained without supervision on boundary conditions alone, with no interior solution data; and a hybrid kernel, which splits the solution, according to the linear principal part of the governing equation, into a homogeneous part spanned by analytic fundamental solutions and a source part carried by low-rank learned correction kernels, thereby balancing physical priors against data fitting on nonlinear problems lacking an analytic fundamental solution. On three benchmarks and one engineering problem in a shallow-water waveguide, KernelOnet attains high accuracy; where comparable with DeepONet, it is more accurate with fewer learnable parameters. Its unsupervised configuration needs no interior solution labels, and its per-query inference cost is far below that of per-instance solvers, offering an effective route to acoustic propagation in unbounded exterior domains that general-purpose neural operators struggle to handle.

[LG-190] CADENCE: A Confidence-Adaptive Dual-Expert Network for Fast and Accurate Time Series Classification

链接: https://arxiv.org/abs/2609.35929
作者: Onisa Mpaunda
类目: Machine Learning (cs.LG)
*备注: 7 pages, 5 figures. Code and benchmark evaluation scripts available at this https URL

点击查看摘要

Abstract:Time series classification (TSC) exhibits a sharp trade-off between accuracy and computational scalability. Meta-ensembles like HIVE-COTE 2.0 reach state-of-the-art accuracy but require extensive compute, whereas ultra-fast random convolutional transforms (e.g., MiniRocket, Hydra) run in seconds but struggle with phase-independent distributions, signal kinematics, and decision tree fragmentation on large class counts. In this work, we present CADENCE (Confidence-Adaptive Dual-Expert Network for time series Classification Excellence), a unified, CPU-native dual-expert architecture. CADENCE decouples representation learning into two specialized pathways: (i) a Convolutional Linear Expert pairing 10,000 deterministic dilated features with closed-form L2-regularized Woodbury ridge classification, and (ii) a Distributional Interval Expert pairing competing dilated kernels (Hydra) with dyadic Cornish-Fisher moment approximations across signal kinematics and FFT spectral bands, fitted with an ExtraTrees ensemble. An internal validation meta-router with rare-class preservation dynamically selects between pure expert routing and confidence-weighted soft blending, followed by a full refit on 100% of training data. Evaluated across all 109 equal-length UCR Archive datasets over 30 resamples (3,270 total runs), CADENCE achieves a grand mean accuracy of 0.8864. This ranks #2 across the archive, surpassed only by HIVE-COTE 2.0 (0.8895, p_Holm = 0.295, no statistically significant difference), while outperforming Hydra+MultiRocket (0.8818), MultiRocket (0.8797), and HIVE-COTE 1.0 (0.8786, p_Holm = 0.048). CADENCE closes the gap to HIVE-COTE 2.0 to 0.31 percentage points while taking an average of only 17.53 seconds per dataset on a dual-core CPU. Source code and evaluation scripts: this https URL Comments: 7 pages, 5 figures. Code and benchmark evaluation scripts available at this https URL Subjects: Machine Learning (cs.LG) Cite as: arXiv:2609.35929 [cs.LG] (or arXiv:2609.35929v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2609.35929 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-191] Replication Failure and Trivial Baselines in Road-Level Crash Prediction

链接: https://arxiv.org/abs/2609.35917
作者: Maurya Patel
类目: Machine Learning (cs.LG)
*备注: 6 pages, 1 figure, 4 tables. Code and reproduction scripts: this https URL

点击查看摘要

Abstract:Graph neural networks are increasingly applied to road-level crash prediction, but the stability of their reported gains has received little scrutiny. We independently reconstruct the data pipeline of a recent uncertainty-aware model and evaluate eleven of its design decisions across three London boroughs under an expanding-window protocol. Four survive replication on a second borough; seven do not, and four of those reverse sign rather than attenuate. Multi-seed evaluation is decisive: one effect reverses sign between random seeds within a single borough, and the reference architecture exhibits per-borough seed spreads of up to 35.7 points against 4 points for ours. We further compare both networks against a parameter-free baseline that ranks segments by cumulative past crash count. At matched history depth our model is statistically indistinguishable from that baseline ( -0.90 points, p=0.61 ), and the reference architecture loses to it on 18 of 18 held-out windows ( -17.37 , p10^-6 ). Sweeping the baseline’s lookback horizon shows it spans 22.71% to 83.94% accuracy on that variable alone, and that every published figure in this line of work is matched by the baseline at a horizon of one to five years. We argue that the apparent margin of graph networks over historical baselines in this task is substantially an artefact of the short horizons those baselines were computed over, and recommend horizon-matched baselines and multi-seed reporting as minimum practice.

[LG-192] Learning in the Transverse Subspace: A Minimal Representation for Divergence-Free Operator Learning

链接: https://arxiv.org/abs/2609.35884
作者: Yifei Sun
类目: Machine Learning (cs.LG); Fluid Dynamics (physics.flu-dyn)
*备注:

点击查看摘要

Abstract:Divergence-free vector fields are fundamental state variables in incompressible flows and many PDE systems. Redundant parameterizations, including Neural Conservation Law (NCL) potentials, map multiple auxiliary representations to the same physical field. Our experiments show that this redundancy can reduce static representation-fitting error by enlarging the set of equivalent solutions, but the resulting many-to-one mapping does not provide a unique state for operator learning. We introduce a minimal representation that encodes a real (D)-component divergence-free vector field on a (D)-dimensional domain as a real ((D-1))-component field on the same domain. Exploiting the transverse structure imposed by incompressibility in Fourier space, we use a Householder orthogonal transformation to construct the reduced coordinates directly. For periodic and closed impermeable fields, the transform is invertible, isometric, and angle-preserving. For open nonperiodic flows, Fourier extension constructs a compatible periodic field, and a minimum-energy rule selects a unique reduced representation. Neural operators then learn temporal evolution entirely in this reduced space. At inference, the predicted ((D-1))-component field is decoded directly into a physical divergence-free (D)-component field, without predicting an ambient field or applying post-hoc projection. Experiments on static fitting and temporal prediction reveal a task-dependent trade-off: redundancy facilitates static optimization, whereas unique invertible coordinates provide a well-defined state for temporal dynamics. By removing unconstrained longitudinal or null directions from the learned state space, the proposed formulation achieves lower prediction error and greater robustness while preserving divergence freedom by construction. Subjects: Machine Learning (cs.LG); Fluid Dynamics (physics.flu-dyn) Cite as: arXiv:2609.35884 [cs.LG] (or arXiv:2609.35884v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2609.35884 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Yifei Sun [view email] [v1] Sun, 27 Sep 2026 00:38:20 UTC (35,682 KB)

[LG-193] CipherGenome: Homomorphic Inference for Genomic Mixture-of-Experts

链接: https://arxiv.org/abs/2609.35883
作者: Guang Yang,Fengchen Liu
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG); Genomics (q-bio.GN)
*备注: Preprint. 19 pages

点击查看摘要

Abstract:Genome foundation models are growing into sparse mixture-of-experts (MoE) networks whose expert weights no longer fit on the machines that hold the sequences, yet sending a private genome to rented accelerators exposes it: we show that a single server hosting one expert recovers the input nucleotides with 99.8% top-1 accuracy. We present CipherGenome, a protocol that keeps the embedding, attention and router of a 15.1B-parameter MoE genome model on a trusted thin client and outsources every expert projection, 95.8% of the parameters, to untrusted and possibly colluding GPU servers under module-LWE encryption. The design exploits three structural facts: expert layers are linear between two SwiGLU gates, expert weights are public, and GPU integer tensor cores can evaluate a ciphertext-weight product exactly modulo 2^48 in a single GEMM. The client evaluates the nonlinearity exactly and re-encrypts with fresh secrets, so no polynomial approximation or bootstrapping is ever needed. On 72 windows from 12 bacterial genomes, encryption adds 2.54 \times 10^-4 nats per token of KL divergence (95% CI upper bound 3.95 \times 10^-4 ), below a pre-registered non-inferiority margin and indistinguishable from bf16 inference, while the same inversion attack falls to chance level. A reusable public hint cuts end-to-end latency by 3.54 times, wire compression reduces traffic 6.8 times, per-layer padding reduces routing leakage from 54.9% to 8.9% accuracy, and HE-compatible int4 experts remain non-inferior to their plaintext counterparts. Per expert and token, the server-side cost is more than six orders of magnitude below a CKKS baseline.

[LG-194] GenomeOcean Anywhere: Private WebGPU Inference for Genome MoEs

链接: https://arxiv.org/abs/2609.35882
作者: Guang Yang,Fengchen Liu
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG); Genomics (q-bio.GN)
*备注: Preprint. 18 pages

点击查看摘要

Abstract:Genome foundation models are most useful where sequences are generated, yet the largest models need datacenter accelerators and a place to send private DNA. We ask whether a 15-billion-parameter genome mixture-of-experts (MoE) model can instead run on volunteers’ web browsers, with the experts spread across many untrusted devices, without changing its predictions and without revealing the sequence to any single device. We build a system in which a trusted coordinator runs attention and routing while browser workers run every expert feed-forward network through hand-written WebGPU kernels, and we protect the expert inputs with real-valued Lagrange coded computing: each worker receives only a Gaussian-padded share, computes the expert’s linear maps, and the coordinator decodes from any two of three workers. On GenomeOcean-MoE (8 experts, top-2 routing, 24 layers), the browser path matches native this http URL at every quantization level, the distributed path stays at the BF16 numerical noise floor (KL 0.0036 nats per token), and an unfitted latency model predicts decode time within 0.74% (median) under emulated wide-area links. We first show that plaintext expert inputs are not private: a probe recovers the token from a single vector at every depth, and one worker can identify the source genome from 300 unordered tokens with 92% accuracy. With coded experts, an adaptive attacker trained on shares falls to the most-frequent-token baseline, one worker’s information about each token is bounded below one bit per forward pass, and the fidelity cost stays below the BF16 noise floor; in Chrome, coded decoding runs at 220 to 376 ms per token, depending on how much of the routing is hidden, and continues without replicas when a worker fails.

[LG-195] GenoTrace: Inheritable Watermarks for Genome Foundation Model Distillation

链接: https://arxiv.org/abs/2609.35881
作者: Guang Yang,Fengchen Liu
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG); Genomics (q-bio.GN)
*备注: Preprint. 33 pages

点击查看摘要

Abstract:Can a genome model retain a detectable record of the synthetic sequences used to train it? We study watermark inheritance through distillation with GenoTrace, a codon-aware extension of green-list watermarking. Two token-level factors modulate the teacher’s generation bias using codon position and organism-specific codon usage. The resulting sequences train a smaller student, whose outputs are audited without an active watermark processor. In a three-seed GenomeOcean-500M-to-100M experiment, the joint configuration achieves a mean audit score of 17.88 and 94.5% detection at a fixed threshold. It retains 49.0% detection after key-aware token substitution, compared with 0% for the available single-seed plain-watermark comparator, and 47.0% after combined mechanism-targeted nucleotide edits. Additional experiments establish inherited signal across five organism-conditioned datasets and teacher-student size ratios up to 40. Component ablations and computational sequence-quality assays reveal distinct operating points for detection strength and coding coverage. GenoTrace provides a practical token-level construction and an empirical account of how genomic structure shapes inherited watermark signals. The findings concern shared-tokenizer distillation and the tested editing procedures, with calibration and biological utility treated as separate evaluation requirements.

[LG-196] From Static Policies to Adaptive Priors in Offline Reinforcement Learning NEURIPS

链接: https://arxiv.org/abs/2609.35880
作者: Tianwei Ni,Vineet Jain,Akash Karthikeyan,Pierre-Luc Bacon
类目: Machine Learning (cs.LG)
*备注: Accepted to NeurIPS Position Paper Track, 2026

点击查看摘要

Abstract:Offline reinforcement learning (RL) has traditionally focused on learning policies for direct deployment under conservative objectives, where uncertainty outside the offline dataset is treated pessimistically to ensure robustness. We argue that this formulation becomes incomplete when an offline-trained policy is subsequently updated through online interaction, as increasingly occurs in modern intelligent systems through test-time adaptation and online fine-tuning. This position paper argues that, in such settings, the objective of offline RL should extend beyond immediate deployment and instead prioritize learning adaptive policy priors: policies that preserve the capacity to improve during subsequent interaction through memory, exploration, and self-correction. We formalize this perspective as adaptive offline reinforcement learning (AORL), distinguish it from offline-to-online RL, and explain why adaptability becomes important under distributional shift, limited dataset coverage, and changing test-time conditions. We further discuss Bayesian offline RL as one principled direction for constructing adaptive policy priors by preserving epistemic uncertainty over plausible environments. Finally, we outline connections, open challenges, and research directions for treating offline RL as preparation for future experience rather than as a static deployment problem.

[LG-197] Hybrid Ensemble Learning for EEG-Based Epileptic Seizure Forecasting

链接: https://arxiv.org/abs/2609.35876
作者: Mason Dana,Khandaker Mamun Ahmed
类目: Machine Learning (cs.LG)
*备注: Accepted at the 2026 IEEE Cyber Awareness Research Symposium (CARS 2026)

点击查看摘要

Abstract:Epileptic seizure forecasting aims to provide actionable warnings before seizure onset, yet patient-independent generalization and false-alarm control remain major challenges. We propose a calibrated hybrid ensemble for EEG-based seizure forecasting that combines five deep learning models and three classical machine learning models through a logistic regression stacking meta-learner. The proposed pipeline integrates signal preprocessing, handcrafted feature extraction, class-imbalance handling, probability calibration, and clinically motivated post-processing. We evaluate the framework on CHB-MIT using strict Leave-One-Patient-Out (LOPO) cross-validation, with threshold and post-processing parameters selected only on held-out meta data. On the filtered cohort, excluding patients with anomalous preictal rates below 1% or above 15%, the model achieves 74.2% seizure-level sensitivity at 1.24 false alarms per hour, with an average warning time of 16.9 minutes. A test-tuned oracle constrained to the target false-alarm budget achieves 60.9% sensitivity at 0.951 false alarms per hour, highlighting the importance of reporting sensitivity together with realized false-alarm rates. Our code is available at: this https URL

[LG-198] ReLOBGen: Replayable Limit Order Book Message Generation

链接: https://arxiv.org/abs/2609.35867
作者: Junoh Kang,Kiseop Lee,Bohyung Han
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We propose ReLOBGen, a method for generating limit order book (LOB) messages that are replayable by construction. Replayability is required for closed-loop market simulation, yet existing LOB message generators may produce non-replayable raw messages, i.e., messages inconsistent with the current market state. These generators therefore rely on post-hoc correction or rejection followed by resampling, which may alter the replayed message distribution or increase inference cost. ReLOBGen instead ensures replayability during generation: it selects the referenced order from the resting orders in the current LOB and then generates the remaining message fields to be consistent with that order and the market state. For realistic reference selection, ReLOBGen samples from a learned distribution over eligible resting orders, efficiently computed from cached order representations and a context-dependent query. It then enforces the consistency of the remaining fields by masking out invalid tokens. Together, these components enable efficient generation of realistic messages without post-hoc correction or resampling. In 500-message rollouts, ReLOBGen achieves 100% replayability, improves market realism, particularly for top-of-book statistics and the relative prices of LOB messages, and provides a 2.7\text-3.6\times speedup per replayed message over the LOBS5 baseline.

[LG-199] A Mesoscopic View of Transformer Weights Through Row and Column Scale Fields

链接: https://arxiv.org/abs/2609.35852
作者: Tiexin Ding
类目: Machine Learning (cs.LG)
*备注: 37 pages, 14 figures, 5 tables. Code and data: this https URL (directory scale_field, tag p5-scale-field-v1)

点击查看摘要

Abstract:Pooled statistics of Transformer weights obscure how magnitude is distributed across functional channels, while individual weights are too numerous to compare directly. We study the mesoscopic level between them: row and column scale fields, the median-centred log-RMS profiles of a weight matrix over its channels, which together with a global scale and a full balanced core represent the matrix exactly. Across public Pythia checkpoints at four sizes and controlled runs from three initialization families, balancing reveals similar measured core magnitude profiles. A mixture bridge, with its form fixed before the analysis and its coefficients fitted, predicts the pooled-shape departure from field width on held-out runs and data arms of the controlled grid. The indexed fields retain further structure: they align across projections that share a functional channel, and query/key profiles follow reassigned RoPE frequencies rather than fixed matrix coordinates. Training trajectories show early field formation followed by component-dependent broadening or recession. Extending the channel-based analysis to AdamW’s second moment reveals related functional organization in its log-space row and column factors. Finally, edits of a frozen checkpoint separate reciprocal scale balance, which preserves the forward computation, from relative channel gain: flattening the gain increases in-distribution loss while preserving matrix norms and the balanced core. Row and column scale fields thus connect pooled magnitude statistics to channel organization and provide coordinates for tracking and testing trained weight structure.

[LG-200] Learned Compression of SAR Phase-History Data: A Rate-Honest Feasibility Study on GOTCHA

链接: https://arxiv.org/abs/2609.35848
作者: Alizishaan Khatri
类目: Machine Learning (cs.LG); Image and Video Processing (eess.IV); Signal Processing (eess.SP)
*备注:

点击查看摘要

Abstract:On-board compression of synthetic aperture radar (SAR) phase history is bandwidth-critical, and block-adaptive quantization (BAQ) remains the operational standard. We test whether a small convolutional autoencoder, with its encoder on the sensor, can compete with BAQ on complex phase-history patches from the AFRL GOTCHA collection. Every method is charged for all transmitted bits, rates are reported in bits per complex sample (b/cs), and detection is scored by one-to-one matching of CA-CFAR detections. The autoencoder (28,656 encoder parameters) loses at every rate. At 16 b/cs it reaches -2.87 dB NMSE, against -35.5 dB for 8-bit BAQ with \pm 3\sigma clipping and -41.0 dB with a tuned clipping range. It also loses to a 16 \times 16 block Karhunen-Loève transform (KLT), a local linear coder with a tenth of its encoder cost (-5.39 dB). Running the network in a companded Fourier domain helps, but its detection F1 remains bounded at 33%. The evidence points to this model, its normalization, and its objective, not to a fundamental limit of learned coding. Per patch, the data have modest lag-1 coherence ( |\rho| \approx 0.3 ) and patch-specific spectral concentration. Two findings concern evaluation itself. First, 97% of CFAR crossings on raw 64 \times 64 patches are border artifacts of the zero-padded detector. Second, on interior cells BAQ’s clipping range decides detection: 8-bit BAQ keeps 69% F1 with tuned clipping but 17% at \pm 3\sigma , and at 8 b/cs or less adaptive FFT thresholding preserves more detections than BAQ. We close with an evaluation protocol for learned radar compression.

[LG-201] Learning from the Gap Between Pass@K and Pass@1

链接: https://arxiv.org/abs/2609.35793
作者: Xuan Liu,Jingbin Qian,Haosheng Chen
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Large language models (LLMs) are increasingly trained with reinforcement learning from verifiable rewards (RLVR). An exact verifier can also support test-time scaling by selecting a passing response from multiple samples, while other deployments use beam search, adaptive sampling, or tools. We study single-sample decoding, where each query receives one response without search, to ask whether search-exposed behavior can be absorbed into the model. Existing verified-response post-training recipes do not generally distinguish problems already solved on the first decode from failures recovered within K samples. Under a fixed budget, this can spend examples repeating behavior the deployed policy already has. We introduce GapFT, which selects training evidence by the source checkpoint’s single-sample outcome and fine-tunes on the Pass@K-Pass@1 gap: problems the policy fails on one sample but solves within K samples. We match training examples, processed tokens, and optimizer steps while keeping the objective unchanged. GapFT fills the matched budget with recovered failures and uses an exact decomposition to distinguish corrections of recovered and missed failures from regressions on first-decode successes. On LogiQA 2.0 and ReClor with Llama-3.1-8B, GapFT improves Pass@1 by 14.4 and 13.9 points over the source model, outperforms budget-matched uniform verified RFT at the same learning rate, and matches fine-tuning on the full verified pool using one third of the data. A single decode matches the source model’s verifier-selected Pass@4 accuracy. A randomized control attributes gains to covering distinct failures, and our analysis relates available gains to transferable failure support. A three-seed Qwen2.5-7B replication retains positive gains over uniform RFT on both logic tasks.

[LG-202] Serverless gossip training of LSTM failure detectors: A matched-protocol comparison with federated local and centralized learning on NASA C-MAPSS

链接: https://arxiv.org/abs/2609.35792
作者: Yusuf Öztürk,Enes Göktekin,Bengisu Atlı,Akın Öztürk,Zhixiang Wang,Ulas Bagci
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 10 pages, 5 figures. Code: this https URL

点击查看摘要

Abstract:Industrial predictive maintenance increasingly depends on learning from equipment spread across sites whose sensor data cannot easily be pooled. Federated averaging (FedAvg) solves this with a central aggregation server; gossip learning removes the server, but its behaviour for recurrent failure-detection models has not been measured under controlled conditions. We compare synchronous ring gossip with FedAvg, isolated local training and a centralized reference for a stacked LSTM that detects imminent failure on the NASA C-MAPSS turbofan benchmark. All methods share one open implementation, architecture, initialization, optimizer, data split and training budget, and the primary endpoint uses one terminal window per test engine to avoid the statistical dependence of overlapping windows. On FD001 (five seeds), gossip reached a terminal-window F1 of 89.6 +/- 1.3%, compared with 89.9 +/- 1.1% for FedAvg, 83.6 +/- 6.7% for local training and 93.5 +/- 2.1% for centralized training, while transmitting the same payload as FedAvg without a coordinator. Node models agreed closely but not exactly (1.8% pairwise decision disagreement versus 5.6% without communication). Across FD002-FD004, peer communication improved terminal-window F1 over local training by 13-28 points; gossip matched FedAvg on FD003 and FD004 but was 4.3 points lower on the multi-condition FD002 subset. Simulated message loss, node failure and server outage changed neither method appreciably, whereas larger rings degraded gossip faster. Ring gossip is therefore a practical serverless alternative when data heterogeneity is moderate, and faster-mixing topologies become important as heterogeneity grows.

[LG-203] Exact Regret Frontiers and Externality Scheduling in Centralized Serial-Dictatorship Bandits DATE

链接: https://arxiv.org/abs/2609.19963
作者: Lishang Xu,Guodong Ma,Pengcheng Weng,Zixuan Xia
类目: Computer Science and Game Theory (cs.GT); Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: v2: expanded and corrected proofs in the appendices (new Remark D.8), added and corrected references, updated AI use statement, typos and wording fixed; main results unchanged

点击查看摘要

Abstract:Exploration in centralized serial-dictatorship matching bandits must use complete matchings, so learning one player-arm pair can impose regret on others. We study this externality under a known common priority order and Gaussian rewards with unit variance. We show that the matching-level Graves-Lai constraints reduce to finitely many pairwise exploration quotas and, at top-choice-separated instances, yield a polynomial-size marginal linear program. At these instances, the exact attainable set of expected logarithmic regret coefficients is G(\theta)\mathcalX(\theta) , where \mathcalX is the feasible matching-allocation set and G maps allocations to player regret. The usual upper-closed Graves-Lai region can be strictly larger despite having the same Pareto-minimal boundary. We further show that identical exploration quotas can induce very different regret through their scheduling. Finally, we construct estimate-solve-track policies, uniformly good on the full row-strict class, that attain every fixed positively weighted optimum without assuming optimizer uniqueness. Every Pareto-minimal point is pointwise attainable, possibly through an instance-calibrated target.

[LG-204] ReCIRC: Rectified Conformal Risk Control

链接: https://arxiv.org/abs/2609.38112
作者: Bruno Marcondes e Resende,Helton Graziadei,Thiago Rodrigo Ramos,Rafael Izbicki
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: 69 pages, 11 figures

点击查看摘要

Abstract:Many applications of black-box predictive models require controlling task-relevant error rates, such as missed lesion pixels in segmentation or missed labels in multilabel classification. Conformal risk control (CRC; Angelopoulos et al., arXiv:2208.02814) gives distribution-free guarantees for such losses, but it calibrates a single threshold shared by all inputs. Because conditional risk varies with the input, this marginal guarantee often overprotects easy cases and underprotects hard ones. We propose ReCIRC (Rectified Conformal Risk Control), which inverts each input’s estimated local risk curve to reparameterize the calibrated threshold as a risk budget a representing a common target conditional risk, and then applies CRC unchanged to the resulting family. ReCIRC retains CRC’s finite-sample marginal guarantee regardless of the accuracy of the estimated curves, while accurate curves yield approximate conditional risk control and, under additional conditions, asymptotically exact conditional risk control; they also support a risk-calibration diagnostic. Across three synthetic and five real-data settings spanning segmentation, multilabel and multiclass classification, and regression, ReCIRC attained the lowest average worst-group risk and mean positive group excess in every setting, while maintaining marginal risk close to the target, whereas changes in prediction size were application-dependent.

[LG-205] Optimal Quantum-Classical Separations for Exact Learning

链接: https://arxiv.org/abs/2609.38073
作者: Srinivasan Arunachalam,Amin Shiraz Gilani,Nikhil S. Mande
类目: Quantum Physics (quant-ph); Computational Complexity (cs.CC); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We study exact learning with membership queries for concept classes \mathcal C\subseteq\0,1^N , focusing on the relationships among their deterministic, randomized, and quantum query complexities, denoted \mathsfD(\mathcal C) , \mathsfR(\mathcal C) , and \mathsfQ(\mathcal C) , respectively. The two canonical quantum speedups in this model are witnessed by Grover search and Bernstein-Vazirani, leading to the longstanding conjecture \mathsfR(\mathcal C)=O(\mathsfQ(\mathcal C)^2+\mathsfQ(\mathcal C)\log N). We first refute this conjecture by constructing concept classes \mathcal C and \mathcal C’ satisfying [ \mathsfR(\mathcal C)=\Omega!\left(\frac\mathsfQ(\mathcal C)^3\log N\log \mathsfQ(\mathcal C)\right) \qquad\textand\qquad \mathsfD(\mathcal C’)=\Omega(\mathsfQ(\mathcal C’)^3\log N). ] The first bound matches the upper bound of Arunachalam et al.~[Quantum’21] up to constant factors, while the second matches the upper bound of Servedio and Gortler~[SICOMP’04]. In particular, this shows that the saving in the randomized upper bound of Arunachalam et al. fundamentally relies on randomness. Apart from characterizing the optimal relationship between classical and quantum query complexity, our results are the first to show that quantum speedups for learning can go beyond the Grover and Bernstein-Vazirani paradigms.

[LG-206] Latent Inference-Time Guidance of Time Series Foundation Models

链接: https://arxiv.org/abs/2609.38058
作者: Chloé Hashimoto-Cullen,Amaury Durand,Laurent Bozzi,Benjamin Guedj,Yannig Goude,Sylvain Le Corff
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Methodology (stat.ME)
*备注: 22 pages, 8 figures

点击查看摘要

Abstract:Time Series Foundation Models (TSFMs) currently provide state-of-the-art results in forecasting tasks. They are available out-of-the-box and rely on in-context learning to make their predictions, which makes the quality of their performance highly sensitive to the user-selected lookback, covariates, horizon and training data distributions. In practise, the quality of the forecasts are variable but complementary, which highlights the need for a principled ensembling approach, rather than selecting the best context. This paper introduces Latent Inference-Time Guidance for TSFMs, which adaptively combines a pool of TSFM forecasts through a time-dependent latent space with independent components. The framework comes equipped with identifiability and reconstruction guarantees, whilst maintaining the off-the-shelf aspect of foundation models. We provide experiments on datasets at various frequencies and from multiple domains: these show that the approach is competitive with traditional ensembling approaches.

[LG-207] he finite-horizon five-expert prediction problem

链接: https://arxiv.org/abs/2609.38035
作者: Jeff Calder,Nadejda Drenska
类目: Analysis of PDEs (math.AP); Computer Science and Game Theory (cs.GT); Machine Learning (cs.LG); Optimization and Control (math.OC)
*备注:

点击查看摘要

Abstract:We give an explicit solution to the five expert prediction with expert advice partial differential equation (PDE) in the finite-time horizon setting. The solution formula establishes that the adversary’s rank strategy (1,0,1,0,0) is globally optimal, and the COMB strategy (1,0,1,0,1) is optimal exactly on the set where x_1=x_2 and x_3=x_4 . The formula is derived from the solution of the geometric-stopping problem given in our companion paper through the transform principle of Bayraktar, Ekren and Zhang, which links the two problems by a Laplace transform. Inverting the transform term by term expresses the solution through a series of Gaussian and complementary error function kernels. The optimality of (1,0,1,0,0) is reduced to the signs of 41 one-variable Gaussian series, which are certified with computer assistance by Poisson summation, first-mode domination and interval arithmetic on 1616 rational cells. The proofs of our main theorems, certificates included, are also formalized in the Lean proof assistant.

[LG-208] Identifiability Guarantees for Drivers and Dynamics of Delayed Physical Systems

链接: https://arxiv.org/abs/2609.37944
作者: Julien Boussard,Antoine Débouchage,Théo Saulus
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Dynamical Systems (math.DS)
*备注: 46 pages, 2 figures

点击查看摘要

Abstract:A wide range of methods have been proposed, including physics-informed neural networks, which are powerful but do not guarantee identifiability of the dynamics, symbolic regression, which requires a set of precomputed operations, and causal discovery, which is more principled but usually relies on strong assumptions that physical systems may violate. In this work, we develop a theory-grounded method and prove that under a set of permissive assumptions, the structural drivers and drift of stochastic delayed differential equations are identifiable. Our method outperforms others on a benchmark for driver identifiability, and on a second benchmark to evaluate physical consistency of the learned dynamics.

[LG-209] Post-Anomaly Detection Inference for Deep SVDD

链接: https://arxiv.org/abs/2609.37935
作者: Cao Le Cong Thanh,Dang Quang Vinh,Vo Nguyen Le Duy
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Deep Support Vector Data Description (Deep SVDD) has become a prominent framework for unsupervised anomaly detection by learning latent representations that compactly characterize normal data around a center. Despite its empirical success, anomaly decisions produced by Deep SVDD are typically made solely based on anomaly scores without rigorous statistical guarantees, thereby limiting their reliability in safety-critical and high-stakes applications where false positives must be strictly controlled. In this paper, we propose PADI (Post-Anomaly Detection Inference), a novel framework that equips a trained and frozen Deep SVDD detector with statistically valid inference by leveraging the Selective Inference framework. Specifically, PADI performs inference conditional on the event that a test instance is identified as anomalous by Deep SVDD, thereby enabling rigorous statistical assessment of anomaly decisions. Based on this formulation, we derive valid selective p-values that quantify the statistical significance of the detected anomaly. Using these p-values, we theoretically establish control of the false positive rate (FPR) at a user-specified significance level \alpha (e.g., \alpha=0.05 ). Furthermore, we extend the proposed framework to Deep Semi-Supervised Anomaly Detection (Deep SAD), providing a principled approach for statistically reliable inference in semi-supervised anomaly detection settings. Extensive experiments on both synthetic and real-world benchmark datasets robustly support the theoretical findings. The results demonstrate that PADI consistently achieves proper FPR control while attaining superior true positive rates compared with existing approaches.

[LG-210] Foundation Neural-Network Quantum States for Molecular Potential Energy Surfaces in Second Quantization

链接: https://arxiv.org/abs/2609.37733
作者: Lizhong Fu,Jianan Wei,Wenguan Wang,Honghui Shang
类目: Chemical Physics (physics.chem-ph); Machine Learning (cs.LG); Quantum Physics (quant-ph)
*备注: 23 pages, 7 figures

点击查看摘要

Abstract:Second-quantized neural-network quantum states have achieved accurate molecular energies, but extending them across molecular geometries requires a shared representation of the geometry-dependent wavefunction coefficients. We introduce geometry-conditioned foundation neural-network quantum states for molecular electronic structure in second quantization. A single autoregressive model learns a family of ground states from sparse anchor geometries and provides wavefunctions at untrained geometries without further optimization. Orbital alignment matches orbital identities and transports their phases, establishing an aligned orbital basis across geometries. Frozen energies reach chemical accuracy at every untrained query geometry for N _2 , CO, and H _4 . On additional molecular paths, the energy-trained wavefunctions yield dipoles, quadrupoles, and natural occupations without property labels. Across three paired N _2 training seeds, orbital alignment lowers the mean absolute energy error over all untrained query geometries from 34-37 mHa to 0.049-0.085 mHa. At approximately 1 mHa mean absolute error, frozen evaluation reduces the per-geometry cost by 986\times relative to independent optimization, yielding an estimated 25.8\times end-to-end GPU-cost reduction on a 161-point N _2 grid.

[LG-211] A Finslerian Approach for Embedding Directed Data

链接: https://arxiv.org/abs/2609.37649
作者: Gwendal Debaussart-Joniec(CB, ENS Paris Saclay),Théau Blanchard(HeKA | U1346, GE Healthcare),Argyris Kalogeratos(CB, ENS Paris Saclay)
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Many datasets carry an intrinsic directionality: citations point backward in time, cells differentiate along lineages, and traffic follows preferred routes. Spectral embedding methods, including most of their extensions to directed graphs, discard this information: they symmetrize the data and map it into a Euclidean space where asymmetry cannot be represented. We instead model directed data as sampled from a Finsler manifold, whose distance depends on the direction of travel, and study the kernel operator built from this asymmetric distance. Through a moment expansion of this operator, we show that its symmetric and antisymmetric parts separate geometry from direction. As the bandwidth of the kernel vanishes, the symmetric part converges to a weighted Laplacian, recovering diffusion maps in the Riemannian case, while the antisymmetric part converges to a first-order transport operator that encodes the directionality. We prove that the corresponding graph operators, built from finitely many samples, converge uniformly and almost surely to these limits. For Randers metrics, this vector field is explicit and yields an embedding algorithm recovering both the manifold structure, from the spectrum of the symmetric part, and the underlying drift. We illustrate the approach on synthetic directed graphs and point-clouds.

[LG-212] Ornstein-Uhlenbeck Is Hard to Beat Yet Superlinear Drift Ships Lower Transport Costs

链接: https://arxiv.org/abs/2609.37579
作者: Attila Lovas,Lóránt Nagy
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Brešar and Mijatović \citebresar2025 show that Ornstein–Uhlenbeck diffusion is hard to beat in forward convergence under assumptions that exclude superlinear drift. We instead test superlinear Langevin diffusions for score-based image generation, computing their conditional scores numerically from a Fokker–Planck equation. In our experiments, the superlinear models beat the Ornstein–Uhlenbeck baseline on empirical Wasserstein distance across nearly the entire tested grid and show less variation across diffusion horizons. The ``hard to beat’’ verdict of \citebresar2025 thus fails to be universal.

[LG-213] End-to-End Optical Semantic Communication over a Nonlinear WDM Fiber Link

链接: https://arxiv.org/abs/2609.37551
作者: Hussein Jammal,Andrea Bianco,Cristina Rottondi
类目: ignal Processing (eess.SP); Information Theory (cs.IT); Machine Learning (cs.LG)
*备注: 6 Pages

点击查看摘要

Abstract:Emerging optical-network applications increasingly use received data for inference and control rather than exact source reproduction, creating an opportunity to trade bit-level fidelity for greater transmission reach and efficiency. We propose an end-to-end optical semantic communication system for joint image classification and reconstruction over a nonlinear wavelength-division multiplexed (WDM) fiber channel. The system maps each image directly into a fixed-length sequence of channel symbols that preserves task-relevant information, without explicit source compression or channel coding. Experiments on the MNIST dataset cover launch powers from -9 to +3 dBm, fiber lengths up to 800 km, and 16-, 64-, and 256-Quadrature Amplitude Modulation (QAM) formats. At 0 dBm, classification accuracy remains between 98.92% and 99.31% across all tested link lengths and modulation orders, while requiring fewer transmitted symbols than a Low-Density Parity-Check (LDPC)-coded JPEG baseline at every tested modulation order. These results show that semantic communication can simultaneously extend optical reach and reduce transmission resources by conveying only task-relevant information.

[LG-214] Geometry-Aided Channel Deduction with Partial Channel Estimates and Uncalibrated Digital Twin

链接: https://arxiv.org/abs/2609.37277
作者: Hongning Ruan,Zhaoyang Zhang,Zirui Chen,Ziqing Xing,Zhaohui Yang,Mérouane Debbah
类目: ignal Processing (eess.SP); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:The acquisition of high-dimensional channel state information (CSI) in wireless MIMO-OFDM communications usually requires high pilot overhead, or relies on accurate and complete positional or environmental information. In this paper, we propose a geometry-aided channel deduction (GCD) approach, which utilizes an uncalibrated digital twin (DT) with only approximate environmental geometry and positions to assist the channel acquisition. The key rationale behind is that, even imprecise geometric information, which can be easily obtained in advance through radio sensing technologies or existing geographic databases, provides certain structural features about the current channel; meanwhile, the coarse instantaneous channel estimates using only a small amount of pilots provide dedicated information that aligns with the channel structure and further compensates for the geometry inaccuracy and other channel unknowns. To this end, we first extract geometric features from the DT, which contain only simple structural information of the channel. Then we propose random prompt augmentation, a novel method to generate an appropriate prompt that converts geometric multi-path structure into a CSI-like representation while suppressing the disturbance of other unknown channel parameters. The prompt is then fused with the pilot-based instantaneous channel estimate via a channel deduction network. To further enhance the network’s versatility, we incorporate pilot configurations into the existing learning architecture to support variable pilot patterns. Comprehensive experiments validate the superiority of the proposed method, which demonstrates high channel acquisition quality, low pilot overhead, and strong robustness. Furthermore, the structural prompt also serves as scenario-related context, enabling our approach to generalize well in new scenarios.

[LG-215] Probabilistic Symbolic-Distillation Model of Droplet Collision for Spray Simulation at High Ambient Pressures

链接: https://arxiv.org/abs/2609.37202
作者: Weiming Xu,Tao Yang,Peng Zhang
类目: Fluid Dynamics (physics.flu-dyn); Machine Learning (cs.LG)
*备注: 32 pages, 13 figures, 2 tables

点击查看摘要

Abstract:Droplet collision governs droplet population dynamics in many chemical engineering processes, such as spray drying, spray cooling, agricultural spraying, and combustion. Existing analytical models impose deterministic, pairwise boundaries between collision outcomes, whereas machine-learning classifiers lack the explicit functional form required of analytical collision submodels. In this study, we develop a probabilistic symbolic-distillation model using nearly forty thousand experimental events spanning eight regimes and five dimensionless parameters, including over five thousand data for ambient pressure up to 50 atm. A machine-learning teacher learns the joint outcome-probability landscape from these data, and symbolic regression subsequently distils it into eight class-specific expressions that jointly define a coupled analytical model. The resulting analytical field replaces abrupt regime switching with finite-width fuzzy boundaries. It outperforms the evaluated conventional analytical boundary models and reveals that their main limitation is the inability of zero-width boundaries to represent gradual probability transitions. The “biased-dice” sampling scheme provides a statistically consistent and practically convenient model implementation for Eulerian-Lagrangian spray simulation.

[LG-216] RNA Design via Conditioned Flow Matching and Finite-Policy Reinforcement Learning

链接: https://arxiv.org/abs/2609.36885
作者: Zefeng Lin,Xianyong Fang,Tianfan Fu,Xiaohua Xu
类目: Biomolecules (q-bio.BM); Machine Learning (cs.LG)
*备注: 32 pages, figures. Preprint

点击查看摘要

Abstract:RNA design aims to identify sequences that fold into specified secondary structures. Existing methods formulate the task as target-specific search or conditional generation. However, natural RNA evolution proceeds through sequence variation and selection, with compensatory substitutions, whereas these methods do not explicitly model this process. To address this limitation, we propose a two-stage framework comprising RNA Inverse-Folding Flow (RNA-IFlow) and RNA-IFlow-RL. RNA-IFlow uses structure-conditioned Dirichlet Flow Matching to model coordinated variation across the sequence, while RNA-IFlow-RL maps the learned flow to a pairing-preserving finite policy and refines it with thermodynamic feedback. Our framework achieves leading performance on multiple benchmarks, reaching 85.19% Pass@1 on Rfam-27. Further analyses reveal thermodynamic gains, policy dynamics, and robustness across settings. Our work couples coordinated variation with thermodynamic selection, offering a novel paradigm for RNA design.

[LG-217] Best Practices in EEG Analysis: Preprocessing Modeling and Machine Learning

链接: https://arxiv.org/abs/2609.36609
作者: Parsa Razmara,Woojae Jeong,Aditya Kommineni,Raymundo Cassani,Richard Leahy,Takfarinas Medani
类目: ignal Processing (eess.SP); Machine Learning (cs.LG)
*备注: 38 pages, 5 figures. Chapter 16 of the 2026 monograph “Electroencephalogram Signals and Cognition,” edited by Hema A. Murthy, Shrikanth Narayanan, Mriganka Sur, and Rajeswari Aghoram

点击查看摘要

Abstract:Electroencephalography (EEG) analysis requires careful choices in preprocessing, statistical modeling, and machine learning because EEG signals are highly susceptible to artifacts, volume conduction, low signal-to-noise ratio, and substantial inter-subject variability. This chapter provides a practical and methodological guide to modern EEG analysis, spanning EEG preprocessing, artifact removal, filtering, bad-channel detection and interpolation, re-referencing, independent component analysis (ICA), and preprocessing of simultaneous EEG-fMRI recordings. We review major approaches for computational EEG analysis, including event-related potentials (ERPs), time-frequency analysis, functional and effective connectivity, source localization, multivariate decoding, permutation testing, and multiple-comparison correction. We then examine machine-learning methods for EEG, from feature-based classifiers to deep learning and emerging EEG foundation models, with emphasis on cross-subject generalization, limited-data regimes, data leakage, evaluation metrics, and fair benchmarking. Reproducibility is treated as a core requirement throughout, including transparent preprocessing, BIDS-EEG data organization, standardized derivatives, preservation of raw data, and FAIR data practices. The chapter is intended as a practical reference for researchers developing reliable, interpretable, and reproducible EEG analysis and machine-learning pipelines.

[LG-218] Hierarchical Utility Calibration for Structured Multiclass Decisions

链接: https://arxiv.org/abs/2609.36532
作者: Futoshi Futami,Jerry Huang,Ichiro Takeuchi
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:In multiclass probabilistic prediction, Utility Calibration (UC), which focuses auditing on specified utilities, has recently received attention as a way to guarantee downstream decisions while controlling computational and sample requirements. At the same time, some multiclass problems have meaningful label hierarchies that play important roles in medicine and image classification, yet how UC evaluates utility within a hierarchy remains insufficiently understood. We show that the difference between realized utility and predicted mean utility admits an exact decomposition into a sum of contributions from the internal nodes of the label tree. This decomposition shows that positive and negative contributions from different nodes can cancel, and that even when UC is small, the utility errors remaining in parts of the hierarchy need not be small. To address this problem, we propose Hierarchical Utility Calibration (HUC), which evaluates each node contribution before summation while retaining the same target utility, subgroup, and predicted-utility interval. We further provide finite-sample evaluation over all predicted-utility intervals and propose HUC-Boost, which updates only violated internal nodes, with theoretical guarantees for both.

[LG-219] LOCO-AdaMP: Built-in LOCO Inference for Adaptive Minipatch Ensembles with Enhanced Prediction

链接: https://arxiv.org/abs/2609.36396
作者: Yinan Cheng,Lili Zheng
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Methodology (stat.ME)
*备注:

点击查看摘要

Abstract:As black-box machine learning models become increasingly common, extracting interpretations with uncertainty quantification has become a critical challenge. One popular type of interpretation is leave-one-covariate-out (LOCO) feature importance, while prior LOCO inference methods often require data-splitting or model-refitting. A recent ensemble framework, LOCO-MP, addresses these challenges using minipatches that subsample both observations and features, but massive feature subsampling can hurt prediction in high-dimensional sparse settings. Motivated by this limitation, we consider minipatch ensembles with adaptive feature sampling guided by LOCO importance, and propose LOCO-AdaMP, which enables free LOCO inference for the resulting adaptive minipatch ensemble. We show that LOCO-AdaMP yields substantially improved predictive models while retaining asymptotically valid feature importance inference without data-splitting, despite the complex dependence between the adaptive sampling distribution and the LOCO importance statistics. Our analysis relies on a careful leave-two-out perturbation bound for the iteratively updated sampling probabilities together with the stability of LOCO scores induced by observation subsampling. Empirical results on synthetic and real datasets demonstrate advantages of LOCO-AdaMP over existing methods in predictive performance, inferential power, and stability. Overall, LOCO-AdaMP provides a flexible ensemble framework (agnostic to base models) that delivers both strong predictive performance and asymptotically valid, powerful feature importance inference for regression.

[LG-220] Receptive-field-constrained stimulus optimization for human early and intermediate visual cortex

链接: https://arxiv.org/abs/2609.36391
作者: Junru Zhao,Hanfei Guo,Andrew Luo,Margaret M. Henderson
类目: Neurons and Cognition (q-bio.NC); Machine Learning (cs.LG)
*备注: 81 pages, 84 figures

点击查看摘要

Abstract:An ongoing challenge in sensory neuroscience is to characterize the feature dimensions encoded by cortical populations. Recent approaches probe feature selectivity in a data-driven way, by synthesizing a most-exciting-input (MEI) for a target neural population. While this approach has been successfully applied to human higher visual cortex using fMRI data, generating MEIs for early- and mid-level retinotopic visual areas requires additional modeling constraints due to small receptive field sizes. To address this challenge, we introduce two novel MEI generation frameworks, Receptive Field Diffusion for Visual Exploration (RF-DiVE) and Receptive Field Gradient Optimization (RF-GO). Both methods use a population receptive field (pRF)-constrained voxelwise encoding model; RF-DiVE combines this with a pretrained latent diffusion model, while RF-GO uses regularized gradient ascent. When applied to single voxels in retinotopically defined areas V1-hV4, using data from the Natural Scenes Dataset, we obtain MEIs that exhibit consistent structure within the pRF, suggesting selectivity for local features like contour, color, and texture. We systematically compare MEIs generated by RF-DiVE and RF-GO using two encoding backbones, performing in-silico validation of predicted responses to MEIs using independent encoding models. Across all methods and all visual areas, MEIs elicit higher model-predicted responses than the most activating natural images. We further find that the choice of generation framework and encoding backbone differentially affects MEI properties, including their visual appearance, structural interpretability, and cross-model generalizability. These results offer a new approach for performing data-driven characterization of spatial and feature selectivity across human visual cortex.

[LG-221] Finite-Sample Theory for Fitted Q-Iteration When Actions Are Functions

链接: https://arxiv.org/abs/2609.36390
作者: Gefei Lin,Rui Miao,Xiaoke Zhang
类目: Methodology (stat.ME); Machine Learning (cs.LG); Statistics Theory (math.ST); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Offline reinforcement learning seeks optimal decision rules from previously collected data. In some applications, a decision can be an entire function, such as a fluence map in radiation therapy or a smooth movement trajectory in robotics. In this paper, we study the finite-sample theory for fitted Q-iteration (FQI) with functional actions in a discounted infinite-horizon setting. Three major difficulties arise in this setting: first, the absence of a Lebesgue probability density for functional actions complicates coverage descriptions; second, conventional coverage requirements can be restrictive; and third, the large functional action space makes greedy optimization in FQI challenging. To address these difficulties, we study smoothness-regularized policy search under a critic-relative coverage condition. This condition measures how well logged data distinguish relevant action-value differences without requiring an action density. Our main theorem gives finite-sample guarantees for learned-policy regret relative to the best value within a fixed smooth class of functional-action policies. The results allow trajectory lengths to be either bounded or growing and Q-functions to be fitted by either functional-input kernel ridge regression or adaptive functional neural networks. For a few examples, we can obtain polynomially decaying regret bounds in the number of logged transitions, up to logarithmic factors, with logarithmically many FQI iterations. Numerical experiments show gains of learned functional-action policies over constant-action policies and support our adoption of a critic-relative coverage condition.

[LG-222] nsor-Train Compressed Separable PINNs: A Curvature-Aware Optimization Framework for Parametric PDEs in High Dimensions

链接: https://arxiv.org/abs/2609.36165
作者: Denis Korolev,Martin Eigel
类目: Optimization and Control (math.OC); Machine Learning (cs.LG); Numerical Analysis (math.NA)
*备注:

点击查看摘要

Abstract:In this work, we develop a second-order optimization framework for physics-informed neural networks (PINNs) applied to high-dimensional parametric partial differential equations (PDEs). The framework is built on the Gauss–Newton pullback metric, which provides an operator-informed notion of curvature in parameter space and connects the method to the broader family of natural gradient schemes. We show that, for coordinate-separable neural architectures and linear differential operators (or linearized operators in the nonlinear case) admitting a finite separable representation, the residual Jacobian inherits a structured separable factorization. This yields an exact compressed formulation of the Gauss–Newton step in a reduced space, without assembling the full residual Jacobian on the exponentially large tensor-product collocation grid. The dimension of the reduced space (the effective compressed dimension) is determined by the local collocation grid sizes, the separable operator structure, and the contraction pattern of the architecture, thereby avoiding dependence on the full tensor-product grid size and replacing dense linear algebra in parameter space by a substantially smaller structured problem. Within our framework, we investigate canonical polyadic and tensor-train parametrizations and derive their full algebraic characterization relevant to the Gauss–Newton method, including the structure of the residual Jacobian, the resulting compressed system, and its effective compressed dimension. Numerical experiments on high-dimensional PDEs, including parametric problems, demonstrate the high efficiency of the proposed compressed Gauss–Newton method, which achieves substantially lower errors than tensor-compressed first-order baselines with orders of magnitude fewer iterations and only a fraction of the computing time.

[LG-223] Fundamental Limits of Transferability and Equivariance in Algebraic Signal Models I: Finite Dimensions

链接: https://arxiv.org/abs/2609.36106
作者: Alejandro Parada-Mayorga
类目: ignal Processing (eess.SP); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We study the fundamental limits of transferability in algebraic signal processing through homomorphisms between algebraic signal models. Homomorphisms are linear maps between the signal spaces of two models that commute with filtering, so filtering a signal and transferring it across domains can be done in either order. The existence of such maps is governed entirely by coincidences among the filtered eigenvalues of the two models’ shift operators, but existence alone is insufficient: the space of homomorphisms always contains trivial elements that destroy all information. We introduce the spectral transfer efficiency \eta(\theta)\in[0,1] to quantify information-preserving quality, prove that every homomorphism decomposes into unconstrained blocks over coincidence classes, derive the dimension of the homomorphism space, and characterize exactly when lossless transfer is achievable. Beyond normal shift operators, we quantify a departure-from-normality penalty and show how filter derivatives can repair spectral defectiveness. The theory yields concrete consequences in three settings: for sampling, eigenvalue interlacing converts transferability under subsampling into an explicit filter design constraint; for compressed sensing, \eta(\theta) controls the restricted isometry constant and the coherence of the resulting measurements, and recovery decouples across coincidence classes; and for machine learning, spectral aliasing emerges as the controlled symmetry breaking that makes transfer between mismatched domains possible at all.

[LG-224] Hybrid Neural Simulation-Based Inference for Robust Applications and Limited-Budget Scenarios

链接: https://arxiv.org/abs/2609.36044
作者: Sean Benevedes,Mani Dehghan,Aishik Ghosh,Tae Hyoun Park
类目: High Energy Physics - Phenomenology (hep-ph); Machine Learning (cs.LG); High Energy Physics - Experiment (hep-ex); Data Analysis, Statistics and Probability (physics.data-an)
*备注: 10 pages, 8 figures, code available at this https URL

点击查看摘要

Abstract:We develop two hybrid techniques that approach the performance of neural simulation-based inference (NSBI) analyses while substantially reducing the computational cost of inference and preserving some or all of the reliability guarantees of parametric methods. The first approach is broadly applicable, while the second is tailored to a class of particle physics analyses that admit a semi-parametric NSBI formulation. With only a modest compromise in raw sensitivity, these methods represent an important step toward computationally efficient NSBI in offline analyses and also open the door to the exploration of trigger-level applications in the future. Based on our comparison studies, we recommend the use of our first approach, Latent Categories, for robust and efficient inference. Comments: 10 pages, 8 figures, code available at this https URL Subjects: High Energy Physics - Phenomenology (hep-ph); Machine Learning (cs.LG); High Energy Physics - Experiment (hep-ex); Data Analysis, Statistics and Probability (physics.data-an) Cite as: arXiv:2609.36044 [hep-ph] (or arXiv:2609.36044v1 [hep-ph] for this version) https://doi.org/10.48550/arXiv.2609.36044 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-225] Dual-Anchor Acceleration Is Near-Optimal for Stochastic Monotone Root-Finding

链接: https://arxiv.org/abs/2609.36033
作者: TaeHo Yoon,Nicolas Loizou
类目: Optimization and Control (math.OC); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Among distinct optimal acceleration mechanisms for deterministic monotone root-finding problems and fixed-point problems, dual-anchoring has recently been shown to admit a more robust direct stochastic extension than standard anchor acceleration. However, without additional strong monotonicity, the existing stochastic dual-anchoring guarantee has two limitations: first, it requires cocoercivity in expectation, and second, it attains only O(\epsilon^-3) oracle complexity, leaving a gap to the near-optimal \tildeO(\epsilon^-2) complexity achieved by other methods. In this work, we address both of these limitations by combining dual-anchoring with stochastic resolvent approximation and optimized variance control. For unbiased stochastic oracles with variance bounded by \sigma^2 , where sample operators are monotone and uniformly L -Lipschitz, our algorithm finds a point with \epsilon -residual with a near-optimal oracle complexity of O ( (LD / \epsilon) \ell + (\sigma^2 / \epsilon^2) \ell^2) , where \ell = \log (1 + LD / \epsilon) and D is the initial distance to a solution. This result improves the best known oracle complexity in the noise-dominated regime under these samplewise assumptions, reducing the poly-logarithmic factor from cubic to quadratic.

[LG-226] SIFARI: Self-Supervised Interferometric Fitting for Astronomical Radio Imaging

链接: https://arxiv.org/abs/2609.35966
作者: Shunyuan Mao,Andrea Isella,Paris Perdikaris,Li-Ta Lo,Hui Li
类目: Instrumentation and Methods for Astrophysics (astro-ph.IM); Earth and Planetary Astrophysics (astro-ph.EP); Machine Learning (cs.LG)
*备注: 19 pages, 9 figures. Currently under review

点击查看摘要

Abstract:Radio-interferometric images are reconstructed from sparsely sampled visibilities, and CLEAN-based imaging can struggle with spatial filtering, complex morphologies, and uncertainty quantification. Alternative methods that fit visibilities directly can address some of these limitations but often require manual choices of image priors and model hyperparameters. We present SIFARI (Self-Supervised Interferometric Fitting for Astronomical Radio Imaging), a self-supervised neural network workflow that represents sky brightness as a continuous function of position and fits measured visibilities without an external image training set or explicit spatial regularizer. An empirical rule sets the Fourier feature scale from the visibilities before training, controlling how readily the network fits fine structure. Sampling network weights with Stochastic Weight Averaging-Gaussian (SWAG) gives approximate brightness uncertainty estimates, which we combine with a thermal-noise floor to construct spatially resolved signal-to-noise maps. In synthetic ALMA tests, SIFARI yields an effective point-source response about eight times narrower than the natural-weighting CLEAN restoring beam and recovers more extended flux than CLEAN when short baselines are missing. It also achieves higher image fidelity than the restored CLEAN images in all three morphology benchmarks. Applied to ALMA observations of PDS 70, SIFARI recovers the bright outer ring together with faint compact emission in the central cavity. For long-baseline-only WISPIT 2 data, SIFARI supplies a sky model for phase self-calibration where the CLEAN model is inadequate. The restored, self-calibrated SIFARI image has approximately 30% lower RMS noise than the CLEAN image made from the original visibilities without self-calibration.

[LG-227] Wasserstein Causal Forests for Distribution-Valued Outcomes

链接: https://arxiv.org/abs/2609.35898
作者: Hugo Gobato Souto
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:This paper proposes Wasserstein Causal Forests (WCF) for settings in which each unit’s outcome is itself a probability distribution. This study also defines finite-grid transformed average and conditional average treatment effects, including a reference-distance contrast that asks whether treatment moves unit-level distributions toward a prespecified benchmark. Simulations cover null effects, location and shape changes, limited overlap, equal-mean but different laws, heterogeneous effects, multimodality, and structural zeros. WCF is most accurate on the conditional-law metric in most reported designs and sharply improves reference-effect estimation in the principal location-and-shape settings, but it is less accurate than the forest baselines for multimodal settings. WCF is applied to the famous Project STAR \citepword1990state, revealing that small classes alter more than the mean: they raise within-grade mathematics achievement by 0.161 standard deviations on average (SE 0.028 ); but the gain is not a uniform location shift, it is larger in the upper part of the classroom score distribution ( +0.179 at the ninetieth percentile versus +0.091 at the tenth) and, in descriptive stratum estimates, largest in the schools serving the most economically disadvantaged students (highest free-lunch quartile, +0.262 , versus +0.092 to +0.164 elsewhere), while overall dispersion is essentially unchanged.

[LG-228] BarcodeMAE: Rethinking Masked Pretraining and Global Representations for DNA Barcode Foundation Models

链接: https://arxiv.org/abs/2609.35877
作者: Monireh Safari,Pablo Millan Arias,Scott C. Lowe,Lila Kari,Angel X. Chang,Graham W. Taylor
类目: Genomics (q-bio.GN); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Many DNA foundation models are pretrained by masking parts of a sequence and asking the model to reconstruct them. Standard masked pretraining exposes the encoder to special [MASK] tokens that are absent at inference, creating a mismatch between training and downstream use. The role of an explicit global sequence representation such as a [CLS] token and how it should be trained also remain poorly understood for DNA barcodes. We introduce BarcodeMAE+ and study model architecture, global [CLS] representation, and auxiliary pretraining objectives across arthropod COI (BIOSCAN-5M) and fungal ITS (UNITE+INSD) barcodes. Across both barcode regions, the encoder-decoder MAE-LM architecture outperforms its matched encoder-only counterpart in nearly all evaluated configurations, supporting MAE-LM as an effective architectural design for DNA barcode foundation models. A trained global [CLS] representation provides substantial additional gains: on BIOSCAN-5M, [CLS] accuracy increases from 47.53% without an auxiliary objective to 80.65% with cross-entropy genus classification. The best auxiliary objective is region-dependent: cross-entropy performs best on BIOSCAN-5M, whereas pairwise same-genus classification performs best on UNITE+INSD, reaching 73.19% on Yeast and 63.07% on Filamentous Fungi. BarcodeMAE+ outperforms published DNA foundation model baselines on BIOSCAN-5M and achieves the highest Yeast accuracy among the evaluated UNITE+INSD baselines using frozen encoder representations. Similarity-weighted softmax KNN voting further stabilizes accuracy as neighbourhood size increases. Overall, encoder-decoder masked pretraining and an explicitly trained global representation are strong design choices for DNA barcode foundation models, while the optimal objective for learning that representation depends on the biological domain.

[LG-229] Beyond Discrimination: Calibrated Geoprior Fusion for Bioacoustic Monitoring

链接: https://arxiv.org/abs/2609.35863
作者: Neha Sajja,Bart van Merriënboer,Burcu Karagol Ayan,Tom Denton
类目: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
*备注: 10 pages, 2 figures

点击查看摘要

Abstract:Modern bioacoustic foundation models like Perch and BirdNET can identify species with high discriminative accuracy, yet their confidence scores are often uncalibrated and difficult to interpret as probabilities of real-world occurrence. This limits their use for ecological inference beyond threshold-based detection. We leverage a global annotated acoustic dataset (WABAD) to produce calibration priors for an acoustic model, optionally incorporating species-level information. We introduce new methods of fusing the acoustic predictions with geopriors, which empirically improves calibration while preserving discrimination. Together, these results suggest a path to simpler and more reliable acoustic monitoring for broad biodiversity.

[LG-230] Meta-learning accelerates detector design optimization

链接: https://arxiv.org/abs/2609.35827
作者: Maxim Borisyak,Nikita Gladin,Andrey Ustyuzhanin
类目: Instrumentation and Detectors (physics.ins-det); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:The quality of a detector design is ultimately determined by the quality of the inference it enables, that is, by the accuracy with which the quantities of interest are reconstructed from the raw detector response. For complex detectors, the inference is performed by machine learning models, and the relation between the design and the attainable inference performance is, in general, non-trivial. In this work, we consider the optimization of the inference performance with respect to the detector design. The conventional approach prescribes retraining the inference model at every candidate design, thus, treating the evaluations as independent tasks and discarding the shared structure of the optimal inference algorithms at different designs. We propose the meta-learned objective estimate (MLOE): instead of solving the inference problem anew at every candidate design, a single meta-inference model, conditioned on the design and trained continually along the optimization path, is shared across all of them. We test MLOE on three families of optimization problems, the last of which comprises two design spaces of the Spectrometer Straw Tracker of the Search for Hidden Particles (SHiP) experiment; under matched budgets of simulation calls, the meta-inference model evaluates a candidate design using fewer simulation calls than the baseline strategies and holds the better rank over the convergence curve in all examined cases.

[LG-231] Optimization Using Pathwise Algorithmic Derivatives of Electromagnetic Shower Simulations

链接: https://arxiv.org/abs/2405.07944
作者: Max Aehle,Mihály Novák,Vassil Vassilev,Nicolas R. Gauger,Lukas Heinrich,Michael Kagan,David Lange
类目: Computational Physics (physics.comp-ph); Machine Learning (cs.LG); High Energy Physics - Experiment (hep-ex); Data Analysis, Statistics and Probability (physics.data-an); Instrumentation and Detectors (physics.ins-det)
*备注: 12 pages, 11 figures, 2 tables

点击查看摘要

Abstract:Among the well-known methods to approximate derivatives of expectancies computed by Monte-Carlo simulations, averages of pathwise derivatives are often the easiest one to apply. Computing them via algorithmic differentiation typically does not require major manual analysis and rewriting of the code, even for very complex programs like simulations of particle-detector interactions in high-energy physics. However, the pathwise derivative estimator can be biased if there are discontinuities in the program, which may diminish its value for applications. This work integrates algorithmic differentiation into the electromagnetic shower simulation code HepEmShow based on G4HepEm, allowing us to study how well pathwise derivatives approximate derivatives of energy depositions in a sampling calorimeter with respect to parameters of the beam and geometry. We found that when multiple scattering is disabled in the simulation, means of pathwise derivatives converge quickly to their expected values, and these are close to the actual derivatives of the energy deposition. Additionally, we demonstrate the applicability of this novel gradient estimator for stochastic gradient-based optimization in a model example. Comments: 12 pages, 11 figures, 2 tables Subjects: Computational Physics (physics.comp-ph); Machine Learning (cs.LG); High Energy Physics - Experiment (hep-ex); Data Analysis, Statistics and Probability (physics.data-an); Instrumentation and Detectors (physics.ins-det) Cite as: arXiv:2405.07944 [physics.comp-ph] (or arXiv:2405.07944v1 [physics.comp-ph] for this version) https://doi.org/10.48550/arXiv.2405.07944 Focus to learn more arXiv-issued DOI via DataCite

附件下载

点击下载今日全部论文列表