本篇博文主要内容为 2026-09-18 从Arxiv.org论文网站获取的最新论文列表,自动更新,按照NLP、CV、ML、AI、IR、MA六个大方向区分。

说明:每日论文数据从Arxiv.org获取,每天早上12:30左右定时自动更新。

提示: 当天未及时更新,有可能是Arxiv当日未有新的论文发布,也有可能是脚本出错。尽可能会在当天修复。

目录

概览 (2026-09-18)

今日共更新826篇论文,其中:

  • 自然语言处理104篇(Computation and Language (cs.CL))
  • 人工智能216篇(Artificial Intelligence (cs.AI))
  • 计算机视觉125篇(Computer Vision and Pattern Recognition (cs.CV))
  • 机器学习199篇(Machine Learning (cs.LG))
  • 多智能体系统17篇(Multiagent Systems (cs.MA))
  • 信息检索22篇(Information Retrieval (cs.IR))
  • 人机交互42篇(Human-Computer Interaction (cs.HC))

多智能体系统

[MA-0] COIN-GP: Cooperative Online Learning in Networked Distributed Systems with Partial Measurements via Gaussian Process Regression

【速读】:该论文旨在解决在分布式传感器网络中,仅能获取部分状态观测的情况下,联合估计系统状态与部分未知动态的难题。其核心挑战在于观测信息不完整以及高斯过程(Gaussian Process, GP)模型先验不足导致的建模不确定性。为此,论文提出了一种基于观测器的动态协同学习框架,融合在线分布式高斯过程回归,能够在不完全测量数据和不完善GP模型条件下实现高精度的状态与动态估计。该方案的关键在于通过分布式协同机制实时更新局部GP模型,并结合一种新型数据采集策略,确保数据获取的可行性与有效性;同时,理论推导给出了涵盖状态估计误差与模型估计误差的上界,利用高斯过程的确定性误差边界特性保障了算法的收敛性与鲁棒性。实验结果表明,所提方法在估计精度与适应性方面显著优于现有的分布式高斯过程方法。

链接: https://arxiv.org/abs/2609.20598
作者: Zewen Yang,Xiaobing Dai,Zhenxiao Yin,Hang Zhao,Zhijun Li,C.C. Chan
机构: Technical University of Munich (TUM)(慕尼黑工业大学); The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)); Tongji University (同济大学); Shanghai Yangzhi Rehabilitation Hospital (上海杨智康复医院); Shanghai Key Laboratory of Wearable Robotics and Human-Machine Interaction (上海市可穿戴机器人与人机交互重点实验室); University of Science and Technology of China (中国科学技术大学); The Hong Kong Polytechnic University (香港理工大学)
类目: Machine Learning (cs.LG); Multiagent Systems (cs.MA); Systems and Control (eess.SY)
备注:

点击查看摘要

Abstract:In this paper, we tackle the problem of jointly estimating the system states and partially unknown dynamics within distributed sensor-equipped networks, particularly in scenarios where only partial state observations are available. To address this issue, we propose an observer-based dynamic cooperative learning framework incorporating online distributed Gaussian Process (GP) regression, which enables accurate estimation despite incomplete in measurements and deficient GP models. In addition, a novel data collection strategy is introduced, with theoretical conditions ensuring feasible data acquisition. Moreover, we also derive an error upper bound encompassing state estimation and model estimation, leveraging the deterministic error bounds of GPs. Empirical simulations demonstrate the superiority of our approach compared to existing distributed GP-based methods.

[MA-1] NS3Learn: Transferring 5G NR Mode-2 Reception Realism from ns-3 to the Veins/SUMO Stack for Connected-Vehicle Safety Assessment

【速读】:该论文旨在解决当前车联网安全评估中因标准信道模型忽略5G NR侧链路模式2(Mode-2)中无线资源竞争问题,导致在高密度交通场景下消息传输率被严重高估的难题。其核心解决方案是提出一种基于模型蒸馏(model distillation)的轻量级通信模型NS3Learn,通过利用已有的ns-3 5G-LENA仿真轨迹数据(共1050万次接收结果)进行标注与拟合,构建一个闭式表达的生成式通信模型,精准捕捉半双工冲突、调度碰撞、接收机捕获效应及解码失败等物理机制。该模型无需对完整协议栈进行重实现,即可在保持现有仿真流程的前提下,实现对密集交通环境下包丢失和拒绝服务影响的高精度建模;验证表明,其在瞬时消息交付率预测上相较于ns-3 5G-LENA的平均绝对偏差仅为0.06,显著优于其他替代模型(0.44和0.55),且参数具备良好跨场景泛化能力(新交叉口仅增加20%误差)。关键创新在于通过显式映射物理机制的可解释性模型,实现了从复杂仿真到高效精确近似的无缝迁移,使研究人员与交通机构可在不修改代码的情况下,快速适应新的无线配置,从而大幅提升车联网安全评估的真实性与可靠性。

链接: https://arxiv.org/abs/2609.20578
作者: Rasheed Bello,Arthur Mukwaya,Gurcan Comert,Varghese Vaidyan,Vijay Bendigeri,Anthony Dontoh,Jagruti Sahoo,Judith Mwakalonge
机构: South Carolina State University (南卡罗来纳州立大学); North Carolina AT State University (北卡罗来纳农业技术州立大学); Dakota State University (达科他州立大学); Independent Researcher (独立研究员)
类目: Networking and Internet Architecture (cs.NI); Machine Learning (cs.LG); Multiagent Systems (cs.MA); Systems and Control (eess.SY)
备注: 20 pages, 5 Images, Submitted to TRB/TRR

点击查看摘要

Abstract:Connected-vehicle safety evaluations rely on coupled traffic and network simulations, but standard channel models ignore radio resource competition in 5G NR sidelink Mode-2, reporting unrealistically high message delivery in dense traffic. This study introduces resource-competition losses without requiring full protocol reimplementation. We labeled 10.5 million reception outcomes from ns-3 5G-LENA traces (calibrated on 3GPP scenarios and driven by SUMO trajectories) to fit NS3Learn - a closed-form model capturing half-duplex loss, scheduling collisions, receiver capture, and decoding. Evaluation spanned two signalized urban networks, six penetration levels (1-100%), and five random seeds per condition. NS3Learn achieved a mean absolute deviation of 0.06 in per-instant delivery compared to ns-3 5G-LENA, outperforming alternative models (0.44 and 0.55 deviation). Fitted parameters transferred to a distinct intersection with only 20% additional error. Crucially, using realistic communication models reversed simulated traffic speed trends and more than doubled predicted hard-braking events. The framework transfers reception realism between simulators via model distillation instead of full reimplementation. Every stage maps directly to an explicit physical mechanism. Researchers and transportation agencies can maintain existing simulation pipelines while accurately accounting for dense-traffic packet loss and denial-of-service impacts. Adapting to new radio configurations requires only offline refitting rather than code modification.

[MA-2] Value-Based Massive Access through Goal-Oriented Irregular Repetition Slotted ALOHA

【速读】:该论文旨在解决在大规模设备连接场景下(如远程监测),面向任务的通信范式中介质访问机制设计滞后的问题,尤其针对传统方法依赖集中式架构或简化假设、难以满足实际网络需求的瓶颈。其核心挑战在于如何在保持低计算开销与极少反馈的前提下,实现高效、鲁棒的随机接入与信息估计。解决方案的关键是提出一种面向目标的非规则重复时隙ALOHA(GO-IRSA)方案,该方案融合现代随机接入技术与基于信念的决策策略,能够在数千个传感器组成的分布式网络中,显著降低分布式维纳过程估计的平均误差和最坏情况误差,相比最优集中式解法提升超过30%,且对干扰消除不完善及过程模型不准确具有较强鲁棒性。

链接: https://arxiv.org/abs/2609.20569
作者: Pietro Talli,Andrea Munari,Federico Mason,Federico Chiariotti,Andrea Zanella
机构: University of Padova (帕多瓦大学); German Aerospace Center (DLR)(德国航空航天中心)
类目: Networking and Internet Architecture (cs.NI); Multiagent Systems (cs.MA)
备注:

点击查看摘要

Abstract:The goal-oriented communication paradigm is poised to enable novel real-time applications by easing the burden on communication networks while still delivering task-relevant information. However, efforts so far have focused on the encoding problem, while the design of medium access schemes is still in the early stages of development, especially when connectivity is to be provided to a massive number of devices, e.g., for remote monitoring. In this respect, existing goal-oriented approaches are often centralized or based on simplified underlying mechanisms, requiring unrealistic assumptions. In this work, we present the Goal-oriented Irregular Repetition Slotted ALOHA (GO-IRSA) scheme, which combines modern random access techniques with belief-based policies. GO-IRSA does not impose significant computing loads on the sensors or require frequent feedback, and it can reduce the average and worst-case error of the estimate of a distributed Wiener process by over 30% with respect to the optimal centralized solution in a network with thousands of sensors, and is robust to imperfect interference cancellation and inaccurate process knowledge.

[MA-3] Language-model groups overstate consensus when replaying human deliberation on a reasoning task

【速读】:该论文旨在解决生成式人工智能(Generative AI)在模拟人类集体认知过程时,其共识率(full-consensus rates)作为衡量集体智慧指标的可靠性问题。现有研究常将共识率视为集体认知的体现,但其结果高度依赖于参与度与最终状态的操作化定义,导致人机比较存在测量偏差。本文通过复现100组人类Wason任务讨论组,并匹配相应的大语言模型(Large Language Model, LLM)代理组,采用基于个体前期回答锚定信念的代理设定,以统一评分代码进行对比分析。结果显示,人类群体的共识率在24.0%至57.0%之间波动,约五分之一参与者未发言,而代理几乎全部参与;在两种去盲敏感性分析中(基于提交行为与参与度匹配),代理组的共识率均显著高于人类组,差距达34至44个百分点,且两种方法虽针对不同测量偏误,结果却收敛于0.5个百分点以内。即使在不采用早期停止或移除可记忆答案参数化的条件下,代理组仍表现出极高的共识性,且多集中于错误答案。这表明,模拟共识并未反映集体准确性,且信念锚定的代理组无法准确估计人类群体决策分布。因此,该研究的关键在于提出一种显式评分框架,为评估仿真群体对人类协商结果的模拟提供可验证、可解释的方法论基础。

链接: https://arxiv.org/abs/2609.20543
作者: Tengfei Shao
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Multiagent Systems (cs.MA)
备注: 37 pages, 4 figures. Preregistration: this https URL . Code and data: this https URL

点击查看摘要

Abstract:Full-consensus rates are often treated as indicators of collective cognition, yet depend on how participation and final states are operationalized. We replayed 100 held-out human Wason groups with matched large language model (LLM) agent groups, seeding one belief-anchored agent per participant’s pre-discussion answer and scoring agents and people with the same code. Across human scoring definitions, estimates ranged from 24.0% to 57.0%; about one fifth of participants never posted, whereas agents almost always did. Agent groups remained more consensual in two post-unblinding sensitivity analyses: the submit-based comparison (n = 98) yielded gaps of 34.0 and 43.9 percentage points for chat and reasoning modes, and the participation-matched comparison (n = 45) yielded gaps of 34.1 and 44.4 points. These complementary routes reduced different measurement asymmetries yet converged within 0.5 percentage points. The gap persisted without early stopping and under a reparameterization removing the memorizable answer; reasoning-mode groups then agreed nearly unanimously, mostly on incorrect answers. Simulated consensus did not track collective accuracy, and belief-anchored agent groups were biased estimators of the human group-outcome distribution in this setting. These analyses provide a scoring-explicit basis for assessing simulated-group estimates of human deliberative outcomes.

[MA-4] Silence Is Endorsement: Verification-Status Laundering in LLM Agent Pipelines

【速读】:该论文旨在解决大语言模型(LLM)智能体系统中安全监控器因依赖摘要或存储的交接信息而非原始证据来判断动作合法性所引发的安全漏洞问题。其核心挑战在于“验证状态清洗”(verification-status laundering)——即在信息传递过程中,尽管授权主张保持不变,但其未经验证的事实背景被丢失,导致监控器误判为合法。解决方案的关键在于:在整个智能体流水线中,必须将授权的溯源信息(authorization provenance)以结构化状态的形式与主张绑定并持续传递,避免因摘要、压缩或中间处理环节丢失验证上下文。实验表明,移除未验证来源的框架可使风险动作通过率从5%~9%飙升至60%~98%,且该现象普遍存在于多种开源与托管模型及典型代理流程中。仅通过显式指令要求拒绝未验证授权无法实现跨模型可靠防护,因此必须从系统架构层面保障授权溯源的完整性。

链接: https://arxiv.org/abs/2609.20211
作者: Yibo Hu
机构: Illinois Institute of Technology (伊利诺伊理工学院)
类目: Cryptography and Security (cs.CR); Multiagent Systems (cs.MA)
备注:

点击查看摘要

Abstract:Safety monitors in LLM agent systems often judge actions from summaries or stored handoffs, not from the original evidence. This creates a simple but dangerous failure mode: the handoff preserves the claim that an action is authorized while losing the fact that the claim was never verified. We call this verification-status laundering. Across nine open-weight monitors and two hosted models, the action and authorization proposition remain fixed while we remove the unverified provenance framing around the claim. This change raises approval for risky actions from 5% to 60% on Llama-3.1-8B and from 9% to 98% on Qwen2.5-14B, with similarly large shifts on both hosted models. The failure also emerges in ordinary agent pipelines. Summarizers frequently weaken the status, memory compressors often remove it, and a full proposer–summarizer–memory–monitor pipeline raises risky approval to 57 – 81% across three downstream monitors. Experiments on WildGuard and ATBench show the same pattern on independently authored harmful and unsafe requests: unsupported authorization claims make approval substantially more likely. Explicitly instructing monitors to reject unverified authorization is not a reliable cross-model fix: some models remain vulnerable, while others reject legitimate requests. Agent systems should therefore carry authorization provenance as structured state attached to the claim throughout the pipeline.

[MA-5] A Proposal for an Agent ic AI Architecture to Support Multi-Domain Decision-Making in the Brazilian Armed Forces

【速读】:该论文旨在解决多域作战环境(陆地、航空航天、海军、网络及电磁频谱)日益复杂化背景下,指挥与控制(C2)中心面临的数据量剧增与信息处理速度加快所导致的观察-定向-决策-行动(OODA)循环压力问题。当前国防领域应用的人工智能系统普遍为被动响应且孤立的工具,仍高度依赖人工干预以整合信息、评估态势并制定行动方案。其解决方案的关键在于提出一种面向自主决策支持的代理式人工智能(Agentic AI)概念架构,该架构具备规划、访问数据源、调用工具及自主执行任务的能力,并支持可解释、可追溯的行动过程,从而实现跨巴西三军(海军、陆军、空军)的智能化协同决策。研究重点涵盖四大应用场景(决策支持、态势分析、可行性评估与应对措施建议),并明确了在行政、战略、作战和战术层级下所需的数据与传感器接入机制,以及保障负责任部署的安全与权限管控策略。

链接: https://arxiv.org/abs/2609.20080
作者: Gioliano de Oliveira Braga,Sidnei Barbieri,Ágney Lopes Roth Ferraz,Wagner Comin Sonaglio,Henrique Curi de Miranda e Lourenço Alves Pereira Jr
机构: Instituto Tecnológico de Aeronáutica (ITA), São José dos Campos/SP - Brasil
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: This paper was accepted for publication in the XXVIII SIGE (Simpósio de Aplicações Operacionais em Áreas de Defesa)

点击查看摘要

Abstract:The growing complexity of multi-domain operational environments (land, aerospace, naval, cyber, and electromagnetic spectrum) has increased the volume and velocity of data reaching command-and-control (C2) centers, straining the observe-orient-decide-act (OODA) decision cycle. Artificial Intelligence (AI) systems currently employed in defense are, in general, reactive and isolated tools that still rely heavily on human operators to integrate information, assess scenarios, and formulate courses of action. This paper proposes a conceptual Agentic AI architecture for AI systems that can plan, access data sources, execute tools, and act autonomously and audibly, aimed at supporting decision-making across the three Brazilian Armed Forces (Navy, Army, and Air Force). Four application fronts are discussed (decision support, situational analysis, feasibility studies, and countermeasure suggestion), as well as the data and sensor access requirements and the security and permission safeguards necessary for responsible employment across administrative, strategic, operational, and tactical contexts.

[MA-6] Benchmarking LLM Compliance with China AI Generated Content Regulations

【速读】:该论文旨在解决大语言模型(Large Language Models, LLMs)在中文语境下内容合规性风险日益凸显的问题,尤其针对中国当前对生成式AI内容的监管要求。现有研究多聚焦于英文语境下的合规性评估,忽视了中文语言及意识形态维度的复杂性。本文基于中国现行AI生成内容合规规范,构建了一个涵盖6个维度、共2303个问题的评估框架,其中包括203个自建的宪法性问题(constitutional questions),通过多位评审员依据层级化对齐记忆(hierarchical alignment memory)独立生成判断,系统评估了20个主流大语言模型的合规性与拒答率。其解决方案的关键在于提出一个融合法律合规性与意识形态对齐的统一评估框架,不仅揭示了国际模型在标准中文提问下仍表现出高合规性,且指出差异主要源于与意识形态相关的维度;同时建立了可全球适用的监管基准,为中英文大语言模型提供基于法律依据的一致性合规评价体系。

链接: https://arxiv.org/abs/2609.19989
作者: Chenrui Cui,Hongye Fang,Lisha Song,Weichao Chen,Yue Zhu,Gang Xu
机构: 未知
类目: Computation and Language (cs.CL); Multiagent Systems (cs.MA)
备注: 5 pages, 3 figures, with appendix still improving

点击查看摘要

Abstract:The widespread adoption of LLMs has led to escalating content compliance risks. Prior works have contributed to addressing these risks in the English context, downplaying the complexity of Chinese language content. This paper follows China’s current AI-Generated content compliance requirements and provides evaluation results on 20 notable LLMs, offering insight into China’s regulatory landscape. We design a novel framework to assess the compliance and refusal rates with 2303 questions spanning six distinct dimensions, including 203 self-constructed constitutional questions. The framework employs several judges to generate verdicts independently based on their hierarchical alignment memory. Our findings show that international models also exhibit high levels of compliance despite the use of standard Chinese questions, and the main differences may stem from dimensions closely related to ideological alignment. We establish a regulatory benchmark that enables the global AI community to evaluate both Chinese and non-Chinese LLMs under a unified set of legally grounded compliance requirements.

[MA-7] LearnActCoder: Role-Aware Error Memory for Adaptive Clinical Coding Agents

【速读】:该论文旨在解决临床编码代理在实际应用中反复遭遇的各类错误模式,包括不支持的编码、遗漏已记录的诊断、编码特异性不足以及操作编码规范不符等问题。其核心解决方案是提出一种推理时自适应框架——Learn-Then-Act,该框架通过将少量标注的“学习批次”(LEARN batch)中的错误转化为结构化的错误知识库(Mistake Knowledge Database, MistakeKDB),实现对编码行为的动态调整:将漏报型错误导向以召回率优化为核心的编码器(Coder),将误报型错误导向以精确率优化为核心的判断器(Judge)。该方法在 LearnActCoder 系统中得以实现,该系统基于查表机制进行编码,并具备错误记忆功能。实验结果表明,在 150 例匹配的 MIMIC-III 临床文本上,结构化错误知识库使 CPT 的 F1 值提升 5.9 个百分点,显著优于原始样本记忆与反思式记忆;而 ICD-9 的改进不显著。在匹配的 MIMIC-IV 队列中,记忆机制使 ICD-10 编码向更高精确率偏移,但召回率下降,整体 F1 值无统计学差异。在 1,000 条保留测试集上的稳定表现验证了该方法的可扩展性与鲁棒性。总体而言,研究证明了基于反馈生成的结构化错误记忆能够在不更新模型权重或改变原有工作流的前提下,有效适应不同病例的编码需求。然而,系统在绝对性能上仍偏低,且评估为回顾性分析,尚未进入真实临床部署环境。

链接: https://arxiv.org/abs/2609.19721
作者: Meysam Ghaffari,Bhaskar Sen,Nasim Sabetpour,Nina Fatehi,Animesh Agarwal,Carlos Morato
机构: Optum AI; Harvard University (哈佛大学)
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注:

点击查看摘要

Abstract:Clinical coding agents repeatedly encounter the same failure modes, including unsupported codes, missed documented conditions, specificity errors, and procedure-coding convention mismatches. We introduce Learn-Then-Act, an inference-time adaptation framework that converts errors from a small labeled LEARN batch into a structured Mistake Knowledge Database (MistakeKDB). False-negative lessons are routed to a recall-oriented Coder, while false-positive lessons are routed to a precision-oriented Judge. We instantiate the framework in LearnActCoder, a Coder-Judge clinical coding pipeline with lookup-table grounding where available. On 150 matched MIMIC-III notes, structured MistakeKDB improves CPT F1 by 5.9 percentage points, while raw-example and reflection-style memories remain near the no-memory baseline; the ICD-9 improvement is not significant. On a matched MIMIC-IV cohort, memory shifts ICD-10 coding toward higher precision at a recall cost, leaving F1 statistically unchanged. Applying the same memory to 1,000 held-out MIMIC-III notes maintains a stable ICD operating point, providing scale/stability evidence. Overall, the results are consistent with structured, feedback-derived error memory being useful for adapting clinical coding behavior across cases without weight updates or changes to the underlying workflow. Absolute CPT/HCPCS performance remains low, and the system is evaluated retrospectively rather than in clinical deployment.

[MA-8] SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes

【速读】:该论文旨在解决当前生成式AI(Generative AI)代理在高风险领域应用中存在安全评估缺失的问题,尤其聚焦于金融交易代理这一典型高后果场景。现有研究多采用通用化、非领域特定的安全评估方法,未能充分考虑此类系统在真实对抗性市场环境中所面临的独特且高危的攻击面。为填补这一空白,论文提出FARSIGHT(Financial Agent Robustness and Security Investigation and Global Holistic Testing)框架,从两个核心维度对金融大语言模型(LLM)代理进行方案级评估:一是市场波动(包括类似闪崩的情景)下的鲁棒性,二是针对信息源攻击、代理自身攻击以及代理作为攻击者行为等三类威胁的防御能力。关键解决方案在于构建一个兼具宏观视角与细粒度测试能力的综合性评估体系,揭示出当前15个代表性学术方案普遍忽视鲁棒性与现实对抗威胁——80%在至少一项核心鲁棒性指标上失败,100%存在安全漏洞,且鲁棒性缺陷与安全漏洞之间具有强耦合性:微小误判即可引发系统性崩溃,而攻击者可低成本精准触发相同结果,凸显了高风险场景下智能代理安全设计的紧迫性与复杂性。

链接: https://arxiv.org/abs/2609.19705
作者: Mengxiao Wang,Nitesh Saxena
机构: 未知
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: 24 pages, 6 figures, 11 tables, 118 references. SoK paper. Evaluates 15 academic financial LLM trading agent schemes on robustness and security

点击查看摘要

Abstract:Autonomous large language model (LLM) agents are moving rapidly into high-stakes domains, yet existing agentic-AI security studies remain largely domain-agnostic and overlook the distinctive, high-consequence attack surface such settings create. We examine this gap through financial trading agents, a representative case of high-stakes agentic security, where a single compromised agent has direct execution authority over real capital in an adversarial, reflexive market. To this end, we present FARSIGHT (Financial Agent Robustness and Security Investigation and Global Holistic Testing), a framework that performs scheme-level evaluation of financial LLM agents on two axes: robustness under market turbulence (including flash-crash-like scenarios), and security against three attack types: attacks on information sources, attacks on agents, and agent-as-attacker behaviors. Applying FARSIGHT to 15 representative academic schemes, we find that most overlook robustness and realistic adversarial threats: 80% fail at least one core robustness metric and 100% exhibit security vulnerabilities. These two failure modes are inseparable: a small misjudgment can cascade into a market-wide crash on its own, while an adversary can deliberately trigger the same collapse at minimal cost.

[MA-9] FINSKILLOPS: A Self-Evolving Multi-Agent System for SEC Filing QA

【速读】:该论文旨在解决金融领域问答(Financial QA)系统在部署后面临持续出现的异构错误(如周期、实体、证据使用及计算方面的错误)时,现有自改进方法缺乏对修正范围与影响边界的可控性问题。其核心挑战在于:传统方法虽能将失败转化为新行为,却难以精确控制修正作用的范围,易引发对原有正确答案的破坏(即引入回归)。为此,论文提出将部署后的系统改进视为“受控的行为维护”(controlled behavioral maintenance),关键在于通过可复用的技能补丁(skill patches)实现精准修复,并确保每个补丁在部署前经过目标验证、受保护案例的回归检查、负向对照实验以及版本化替换或退役机制。研究构建了FINSKILLOPS——一个面向美国证券交易委员会(SEC)文件问答的多智能体系统,基于证据驱动、类型化的故障诊断生成技能,并实施严格的生命周期管理。在六个金融QA基准测试中,单一冻结的技能注册表实现了最高评分加权准确率与参考一致性;在增强基准上,正确率由3.70提升至4.55。此外,在12轮实际运行评估中,仅6项提议技能被采纳,同时监控非纠正率从20.0%降至12.5%。结果表明,受控的技能作用范围界定、准入机制与生命周期管理是实现可靠自改进的核心基础。

链接: https://arxiv.org/abs/2609.19680
作者: Yanzhang Ma,Zhenghan Tai,Hanwei Wu,Sizhe Guan,Jianliang Lei,Hailin He,Chaolong Jiang,Jijun Chi,Tung Sum Thomas Kwok,Bohuai Xiao,Jingrui Tian,Xinlu Wu,Xingao Zhan,Peng Lu,Muzhi Li,Yihong Wu,Liheng Ma,Sicheng Lyu,Tianshuo Yan,Junhao Zhu,Yaqian Xu,Lei Ding,Yufei Cui,Ziquan Liu,Boyu Han,Hengli Liu,Ling Zhou,Xinyu Wang
机构: SimpleWay.AI; McGill University (麦吉尔大学); University of Toronto (多伦多大学); University of California, Los Angeles (加利福尼亚大学洛杉矶分校); The Chinese University of Hong Kong (香港中文大学); University of Manitoba (曼尼托巴大学); Université de Montréal (蒙特利尔大学); Mila – Quebec AI Institute (魁北克人工智能研究所); McMaster University (麦克马斯特大学); Harvard University (哈佛大学); Monash University (蒙纳士大学); The University of Hong Kong (香港大学); Stanford University (斯坦福大学); Queen Mary University of London (伦敦玛丽女王大学); UE Capital Ltd.; CG Matrix Technology Ltd.
类目: Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Multiagent Systems (cs.MA); Software Engineering (cs.SE)
备注:

点击查看摘要

Abstract:Financial QA systems are typically improved before deployment through better retrieval, prompting, or agent coordination, leaving their reliability behavior fixed thereafter. In practice, new SEC-filing questions repeatedly expose heterogeneous errors in period, entity, evidence use, and calculation. Existing self-improvement methods can turn failures into new behaviors, but offer limited control over where a correction should apply or which previously correct answers it may break. We therefore frame post-deployment improvement as controlled behavioral maintenance: recurring failures should become scoped skill patches, and each patch should earn deployment with- out introducing regressions. We instantiate this view in FINSKILLOPS, a multi-agent system for SEC filing QA. FINSKILLOPS derives reusable skills from evidence-grounded, typed failure diagnoses and governs them through targeted validation, protected-case regression checks, negative controls, and versioned replacement or retirement. Across six financial QA benchmarks, a single frozen skill registry achieves the highest verdict-weighted correctness and reference consistency among the evaluated systems. Evolved skills raise correctness from 3.70 to 4.55 on our enhanced benchmark. In a separate 12-round operational study, only six of 33 proposed skills are promoted, while the monitoring non-correct rate falls from 20.0% to 12.5%. These results establish controlled skill scope, admission, and lifecycle management as the foundation for reliable self-improvement.

[MA-10] Replan Repair or Edit? A Unified Empirical Evaluation of Travel Agents for Itinerary Revision under Resource Disruptions

【速读】:该论文旨在解决旅行规划代理在行程被接受后因航班取消、酒店不可用或景点关闭等突发情况导致行程不可行的问题,核心挑战在于如何在保证新行程可行性的同时,最大限度保留原已接受的行程承诺,并权衡不同修复方法在效果、计划稳定性与计算成本之间的平衡。其解决方案的关键在于对比三种不同范式的方法:基于生成式AI与约束求解器(LLM-Z3)的全量重规划、基于分层任务网络的层次化修复(IPyHOPPER)以及基于大语言模型的局部修订适配器(iTIMO),并通过两个源自TREK基准集的实验数据集系统评估它们在单次扰动和多重并发扰动场景下的表现。研究发现,尽管LLM-Z3结合Gemini在多重扰动下取得了最高的成功率,但IPyHOPPER在单次扰动中几乎达到同等水平且显著保留更多已接受行程;而成功修复中的层次化与局部修复方法相比全量重规划更少修改原计划并保留更多承诺,同时在计算开销上,IPyHOPPER不依赖大语言模型推理,而基于大模型的方案则在推理成本上存在显著差异,从而为实际应用中在可行性恢复、承诺保留与计算效率之间提供了可操作的权衡指导。

链接: https://arxiv.org/abs/2609.19654
作者: Xiaofei Yuan,Yan Zhang,Shaobo Qiao,Huangleshuai He,Leyan Ni,Mingchen Ju,Lujia Yang,Sijia Xu,Yifu Tang,Zhengyi Yang
机构: Euler AI; University of New South Wales; University of Sydney; Vecton AI
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: 16 pages, 2 figures, 6 tables

点击查看摘要

Abstract:Travel-planning agents generate itineraries that may become infeasible after acceptance because of flight cancellations, hotel unavailability, or attraction closures. Revising these itineraries involves full replanning, classical plan repair, and LLM-based travel-agent revision, whose differing task formulations and evaluation protocols hinder comparison. We conduct a systematic empirical study using two TREK-derived benchmark sets: 500 single-disruption cases, including feasible and infeasible instances, and 200 feasible simultaneous compound-disruption cases. We compare LLM-Z3 full replanning, IPyHOPPER hierarchical repair, and an iTIMO local-revision adapter across effectiveness, plan stability, and computational cost. LLM-Z3 with Gemini achieved the highest observed compound-disruption success. IPyHOPPER nearly matched that configuration’s single-disruption overall success, while preserving substantially more of the accepted itinerary on successful repairs. Successful hierarchical and local repairs made fewer edits and retained more accepted commitments than full replanning. Computational profiles differed: IPyHOPPER used no LLM inference, the evaluated LLM-Z3 adapter used compact one-call inference, and the iTIMO adapter consumed substantially more tokens. The study provides practical guidelines for balancing feasibility recovery, commitment preservation, and computational cost within evaluated settings.

[MA-11] Agent ic AI Networking for Heterogeneous Unmanned Aerial Systems in Low-Altitude Wireless Networks

【速读】:该论文旨在解决低空无线网络(Low-altitude wireless networks, LAWNs)中异构无人系统在共享三维空域内并发服务时,因移动性、连通性与共享网络资源之间的强耦合关系,以及各类服务所提出的动态变化、差异化需求所带来的协同控制难题。传统优化与基于学习的控制器依赖预设目标,难以自主适应服务需求与资源优先级的实时演变。为此,本文提出一种分层混合式大语言模型(Large Language Model, LLM)与多智能体强化学习(Multi-Agent Reinforcement Learning, MARL)融合的双环架构:外层自适应环利用LLM辅助的游戏编排机制,解析服务需求与操作者意图,动态重构目标函数与资源优先级;内层执行环则在配置的游戏框架下,运行去中心化、参数条件化的MARL策略。该架构通过解耦高层决策与底层执行,在不重新训练底层MARL策略的前提下实现对动态运行环境的自适应响应。案例研究验证了该框架在物流监控场景中促进异构服务协同共存的有效性,为构建可扩展、可信且自适应的智能代理式低空无线网络提供了新思路。

链接: https://arxiv.org/abs/2609.19538
作者: Nguyen Duc Minh Quang,Chang Liu,Shuangyang Li,Derrick Wing Kwan Ng
机构: 未知
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注:

点击查看摘要

Abstract:Low-altitude wireless networks (LAWNs) are emerging as a key infrastructure for heterogeneous unmanned aerial systems that support concurrent services within a shared three-dimensional airspace. Their coexistence creates strong coupling among mobility, connectivity, and shared network resources, while heterogeneous services impose distinct and time-varying requirements. These interactions naturally form a dynamic non-cooperative game in which both operating conditions and coordination objectives evolve over time. Conventional optimization and learning-based controllers typically rely on predefined objectives, limiting their ability to adapt autonomously to changing service requirements and resource priorities. To address this challenge, we propose a hierarchical hybrid large language model (LLM)- multi-agent reinforcement learning (MARL) architecture organized as a dual-loop structure. Specifically, an outer adaptation loop employs LLM-assisted game orchestration to interpret service requirements and operator intent, and reconfigure objectives and resource priorities, while an inner loop executes decentralized, parameter-conditioned MARL policies under the configured game. A logistics-monitoring case study illustrates how the proposed framework facilitates coordinated coexistence among heterogeneous services, adapting to evolving operating conditions without retraining the underlying MARL policies. Finally, we discuss key challenges and research directions toward scalable, trustworthy, and adaptive agentic LAWNs.

[MA-12] Reputation as Community Memory for the Agent ic Web

【速读】:该论文旨在解决单一智能体(agent)在长期运行中因私有化记忆导致的知识孤岛问题,即单个智能体无法独立验证外部环境中的可信知识(如数据源、服务与工具的可靠性),从而限制了其决策质量与系统鲁棒性。其核心解决方案是提出Cairn——一个基于社区声誉的集体记忆平台,通过聚合多个独立观察者对资源的评价与证据,实现对共享环境的信任共识。关键创新在于采用具有时间衰减特性的贝塔模型(time-decayed Beta model)结合置信度收缩机制(confidence shrinkage),有效抑制恶意攻击(如虚假评价、合谋、伪装)的影响;同时支持基于评审理由的语义发现,使智能体可依据上下文推理选择可信资源。实验表明,Cairn在对抗性模拟下具备强鲁棒性,并在真实生产环境中成功实现了对异构智能体的评级评估。

链接: https://arxiv.org/abs/2609.19502
作者: Ryan Chard,Gus Ellerm,Alexander Brace,Alok Kamatar,Suman Raj,Ian Foster,Kyle Chard
机构: University of Chicago (芝加哥大学); Argonne National Laboratory (阿贡国家实验室)
类目: Multiagent Systems (cs.MA); Distributed, Parallel, and Cluster Computing (cs.DC)
备注: 7 pages, 5 figures, The 22nd IEEE International Conference on eScience

点击查看摘要

Abstract:Agents can now externalize experience into memory, consolidating historical traces into semantic knowledge and procedural shortcuts that persist between sessions. Such memory is typically private to a single agent. We argue that agentic memory benefits from being collective, because trustworthy knowledge of the shared environment—the data sources, services, and tools agents depend on—cannot be established by any single agent, only corroborated across many independent observers. We present Cairn, a community reputation platform that captures collective knowledge, allowing agents to query the community’s opinion of a resource before use and to submit evidence-backed ratings afterward. Cairn aggregates observations via a time-decayed Beta model with confidence shrinkage and supports semantic discovery over reviewer rationales. We evaluate Cairn’s reputation engine under adversarial simulation (e.g., lying, collusion, camouflage), benchmark its retrieval performance, and report a case study of rating heterogeneous agents in production.

[MA-13] he AR Fairness Metamodel: A Structured Framework for Fairness Measures

【速读】:该论文旨在解决现有公平性(fairness)定义缺乏统一、可系统化建模与比较框架的问题,尤其在多主体资源分配场景中,不同公平性度量之间难以进行形式化对比与分析。其解决方案的关键在于提出一种名为AR公平性元模型(AR fairness metamodel)的通用建模框架,该框架基于Tiles架构,通过模块化组件对代理(agent)、资源及其属性等核心要素进行形式化表达,支持对包括平等、公平、群体公平、个体公平、吉尼系数、泰尔指数、Jain公平性指数以及澳大利亚儿童保育补贴的详细公平性度量在内的多种公平性指标进行系统定义与比较。该方法不仅提供了形式化证明以揭示群体公平、个体公平与无嫉妒性(envy-freeness)之间的内在关系,还通过开源实现的Tiles框架,使公平性建模具备跨应用场景的可移植性与实用性,从而为公平性设计与评估提供可操作、可验证的理论工具。

链接: https://arxiv.org/abs/2609.19234
作者: Julian Alfredo Mendez,Timotheus Kampik
机构: Umeå University (乌梅奥大学)
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Logic in Computer Science (cs.LO); Multiagent Systems (cs.MA); Programming Languages (cs.PL)
备注:

点击查看摘要

Abstract:This paper presents the AR fairness metamodel, a framework designed to represent, analyze, and compare different fairness scenarios. The metamodel considers key elements, such as agents, resources, and their attributes, and enables the systematic definition and comparison of various fairness measures. We provide examples involving both discrete and continuous measures, including equality, equity, group fairness, individual fairness, the Gini index, the Theil index, Jain’s fairness index, and a detailed fairness measure for Australia’s Child Care Subsidy. We also explore relationships among group fairness, individual fairness, and envy-freeness, supported by formal proofs. At the conceptual modeling level, our approach builds on the Tiles framework, which offers modular components that can be connected to capture diverse fairness definitions. The goal is to make AR-based fairness definitions practical and adaptable across contexts, providing a clear way to define, compare, and evaluate them. An implementation of the Tiles framework is available as an open-source tool, and can support fairness modeling and evaluation across a wide range of applications.

[MA-14] CC-OPI: Online Distributed Task Allocation for UAV Swarms under Communication Constraints

【速读】:该论文旨在解决在灾后搜救等多无人机(Unmanned Aerial Vehicles, UAVs)任务中,由于通信范围有限导致的无人机集群频繁断连问题。传统“先分配后执行”范式依赖全局共识才能启动物理动作,在间歇性连接环境下失效。为此,论文提出一种通信受限下的在线性能影响(Communication-Constrained Online Performance Impact, CC-OPI)算法,其核心在于采用事件驱动机制,将任务协商与物理执行交错进行。关键创新点包括:一是设计了适应动态拓扑的双指标成本评估机制——一个引入空间局部性惩罚项以促进区域化协同,另一个加入基于截止时间的紧迫性项;二是引入非抢占式状态锁机制,保护各UAV正在进行的动作不被中断。此外,还构建了去中心化的容错层,结合版本状态同步与基于全局时间驱动的紧急任务池,确保系统在弱连通、丢包、地形遮挡及任务动态到达等复杂条件下仍能稳定运行。理论证明CC-OPI可在有限时间内终止,避免陈旧完成死锁和任务无限重分配问题。仿真结果表明,在250米通信半径下,CC-OPI保持约80%的任务完成率,相较未修改的PI与CBBA规则提升约7个百分点,优于静态基线约20个百分点,并在全连通条件下接近其性能表现,展现出良好的渐进退化特性与鲁棒性。代价为通信消息量增加与部分冗余移动,属于效率与鲁棒性之间的主动权衡。

链接: https://arxiv.org/abs/2609.19208
作者: Biao Liu,Tong Zhang
机构: 未知
类目: Multiagent Systems (cs.MA)
备注: 26 pages, 11 figures, 10 tables, including 6 pages of supplementary material. Code and data: this https URL

点击查看摘要

Abstract:In multi-robot missions such as post-disaster search and rescue, a short communication range fragments a swarm of Unmanned Aerial Vehicles (UAVs) into transient information islands. Under such intermittent connectivity, the prevailing “allocate-then-execute” paradigm–which requires global consensus before any physical movement–breaks down. This paper proposes the Communication-Constrained Online Performance Impact (CC-OPI) algorithm, an event-driven method that interleaves task negotiation with physical execution. CC-OPI replans only at discrete physical and topological events and integrates two further elements. The first is a pair of cost-evaluation metrics adapted to dynamic topologies–one with a spatial locality penalty that promotes regionalized operation, the other with a deadline-aware urgency term–complemented by a non-preemptive state lock that shields each UAV’s ongoing action. The second is a decentralized fault-tolerance layer that pairs version-based state synchronization with a global-time-driven emergency pool. We establish that CC-OPI terminates in finite time, free of stale-completion deadlock and of unbounded reassignment within the mission horizon. In simulations at a 250 m communication radius, CC-OPI sustains a task completion rate of about 0.80: it leads a matched online execution of the unmodified Performance Impact (PI) and Consensus-Based Bundle Algorithm (CBBA) rules by about seven percentage points, exceeds the naively transferred static baselines by roughly 20 points, and remains within several points of PI and CBBA under full connectivity. Within the tested settings, CC-OPI degrades gracefully as connectivity weakens and absorbs packet loss, terrain occlusion, and runtime task arrival. The price is more messages and some redundant travel–a deliberate trade-off of efficiency for robustness.

[MA-15] Position: It is Time to Virtualize Foundation Models with a Self-evolving Operating System Layer

【速读】:该论文旨在解决当前生成式 AI(Generative AI)应用中由多代理系统(agentic systems)堆叠导致的运行时碎片化问题。尽管现有协议(如 MCP、A2A)在工具与代理间的连接上提供了便利,但各框架仍内嵌了对状态、记忆、预算及安全策略等运行时要素的隐式管理机制,导致行为不可移植且治理能力脆弱。这一现状类似于早期计算缺乏操作系统时,每个程序需自行实现基础服务的困境。为此,论文提出构建“基础模型操作系统”(Foundation Model Operating System, FMOS),作为虚拟化基础模型交互的系统层,其功能类似于虚拟机对物理硬件的抽象,使应用层获得如同专用、可信的基础模型实例般的体验,并具备近乎无限的扩展能力。FMOS 的核心在于通过统一调度知识存储分层、模型选择、资源分配、验证与策略执行等关键功能,实现对基础模型行为的高效协同与动态调控;同时,借鉴人类大脑在直觉与理性决策间的切换机制,FMOS 能够学习何时介入干预、何时允许推理自主进行,并基于实际运行经验持续优化自身策略,从而实现智能化、自适应的运行管理。

链接: https://arxiv.org/abs/2609.19203
作者: Suparna Bhattacharya,Tarun Kumar,Cong Xu,Satish Kumar Mopur,Jiahao Li,Ashish Mishra,Aalap Tripathy,Annmary Justine Koomthanam,Martin Foltin,Ian Foster
机构: Hewlett Packard Enterprise(惠普企业); University of Chicago (芝加哥大学); Argonne National Laboratory (阿贡国家实验室)
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multiagent Systems (cs.MA); Operating Systems (cs.OS)
备注:

点击查看摘要

Abstract:AI applications have shifted from single, monolithic foundation models (FM) to compound agentic systems. Yet today’s stacks remain fragmented: even as protocols (e.g., MCP, A2A) ease tool/agent connectivity, each framework embeds an implicit runtime for state, memory, budgets, and guardrails, making behavior non-portable and governance brittle. It mirrors computing before operating systems, when every program re-implemented basic services. This position paper argues that the field now needs a Foundation Model Operating System (FMOS) – a system layer that virtualizes FM interactions analogous to how virtual machines abstract physical hardware, giving applications the illusion of dedicated, trustworthy FM instances with effectively unbounded capabilities. Internally, the FMOS orchestrates knowledge across memory tiers, model selection and resource allocation, and verification and policy enforcement. Like the human brain switching between fast intuition and slow deliberation, the FMOS learns when to intervene and when to let inference proceed directly and continuously adapting its policies based on operational experience.

[MA-16] Message capacity and claim wording set the transition points of collective truth-finding in language-model networks

【速读】:该论文旨在解决在讨论中,由于个体(无论是人类还是大语言模型)受认知、上下文或成本限制,只能阅读其他参与者有限信息的情况下,集体共识如何形成及其偏差来源的问题。其核心问题是:仅由信息阅读范围(即消息容量)所决定的通信约束,是否足以主导集体判断的最终走向。解决方案的关键在于将这一阅读限制抽象为一个单一参数——消息容量(message capacity),用以量化每个代理可读取的他人信息数量,并据此构建通信网络。研究发现,80亿参数模型对命题的判断可被建模为一个逻辑函数,其输入为加权求和的“收件箱”内容,符合具有除法归一化权重的随机二元神经元更新规则。基于该权重分布与网络度数统计,当平均每位代理阅读少于6.4条来自31个源的信息时,错误共识将无法被逆转。实验验证显示,在1,414次有预设起始状态的模拟中,即使75%的代理初始正确,正确方胜率仍低于50%,表明预测失效。进一步分析揭示,失败的根本原因在于“场域”(field)——即命题表述本身在未读任何信息前就设定的响应阈值。实验中的命题场域普遍低于校准均值,且当使用各命题自身的场域作为阈值时,权重可复现实际结果。反转命题表述后发现,该阈值取决于命题主张的内容,而非其真实性,说明存在“主张偏差”(assertion bias)。在另一80亿参数模型上,该方法成功预测了命题依赖的双稳态现象,且转变点与计算一致,16种条件中有15种匹配。而在700亿参数模型上,此偏差不再显著。因此,集体命运主要由两个单智能体测量值决定:命题表述所设定的响应阈值,以及决定相变点的消息容量。

链接: https://arxiv.org/abs/2609.19183
作者: Makoto Fukushima
机构: Honda Research Institute Japan Co., Ltd.(本田研究日本有限公司)
类目: Multiagent Systems (cs.MA); Computation and Language (cs.CL); Adaptation and Self-Organizing Systems (nlin.AO); Physics and Society (physics.soc-ph)
备注:

点击查看摘要

Abstract:Whether human or large language model (LLM), an agent in a discussion reads only a few of the others’ contributions, bounded by cognition, context, or cost. LLM collectives can settle on a wrong consensus even when a majority starts out correct; we ask how far that reading bound alone decides the outcome. We model the bound with one number, the message capacity, which sets how many of the others’ messages an agent reads, and generate the communication network from it. Over 31,824 randomized queries, we found that an 8-billion-parameter model’s judgment of a claim effectively reduces to a logistic function of a weighted sum of its inbox, the update rule of a stochastic binary neuron with divisively normalized weights. From these weights and the network’s degree statistics alone, the wrong consensus should become unreachable from any start once agents read, on average, fewer than 6.4 of their 31 sources. In 1,414 episodes with assigned starts the prediction failed: the correct side won in fewer than 50% of episodes from every start, and in only 28-45% when 75% of agents started correct. The failure traces to the field, the threshold that a claim’s wording sets for the agent’s answer before any message is read: the experimental claims’ fields lay below the calibration mean, and with each claim’s own field the same weights reproduce the outcomes. Reversing the wording showed that the threshold follows what a claim asserts, not whether it is true. On a second 8B model the pipeline predicts claim-dependent bistability; transition points appeared where computed, and an eight-claim calibration matched in 15 of 16 conditions. At 70B the assertion bias is not detected. Thus a collective’s fate is largely set by two single-agent measurements: the threshold a claim’s wording sets, and the message capacity that sets the transition point.

自然语言处理

[NLP-0] Coding Agents with an Obstacle-Aware Harness for Safe Robot Manipulation

【速读】: 该论文旨在解决生成式编码代理(coding agent)在机器人操作任务中因缺乏对安全约束的优先考虑而导致的碰撞问题。尽管代理在指令层面和感知层面均能识别障碍物,但其规划阶段未能将安全约束作为核心优先级,导致在路径规划与接触执行过程中频繁发生碰撞。其关键问题在于:在路径规划阶段,模型无法识别避障路径或在路径不可行时进行重规划;在接触执行阶段,模型未意识到接触动作必须受相同安全约束限制。为此,论文提出SafeHarness框架,通过引入两种障碍物感知的约束机制——障碍物感知的路径规划与障碍物感知的接触执行,使模型能够预先规划并验证路径、动态重规划,并在接触位置选择上主动规避障碍物。实验表明,该方法在任务成功率(71.9%)与碰撞规避率(87.5%)上显著优于现有最先进方法(分别提升6.5%和27.0%),性能分别为无Harness基线的2.3倍和1.5倍,有效弥合了安全规划与执行之间的差距。

链接: https://arxiv.org/abs/2609.20822
作者: Bingxin Xu,Yuzhang Shang,Zhen Dong,Emilio Ferrara
机构: USC(南加州大学); UCF(佛罗里达中央大学); UCSB(加州大学圣塔芭芭拉分校)
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Coding agents have emerged as a promising paradigm for robot manipulation: a language model writes the robot controller as a program, and agents built in this way now operate robots without robot-specific this http URL this paradigm is also safe, however, has not been asked. We evaluate coding agent under a safety constraint, where each task pairs a manipulation goal with an obstacle the robot must not touch. The agent pursues the goal but collides with the obstacle in most cases, treating task completion as its sole objective while neglecting safety. The agent reasons about the obstacle in its traces, and the prompt already forbids touching it, so neither perception nor instruction is at fault; the fault lies in the planning, where the stated constraint never becomes a priority. By decomposing manipulation into a route phase and a contact-rich moment, we locate the source of the failure. Along the route, the model cannot prioritize the safety constraint, having no notion of a clearing route and none of replanning once a chosen route becomes infeasible. At the contact, it is unaware that contact execution is bounded by the same constraint. To close this gap, we present SafeHarness, which equips the model with two obstacle-aware harnesses that enable it to prioritize the safety constraint. Obstacle-aware route planning grounds the objects as bounding boxes and draws candidate routes over them as sequences of waypoints. The agent then plans a route in advance, verifies it, replans when necessary, and only then executes it. Obstacle-aware contact execution instead selects the contact position so that the contact itself avoids the obstacle. SafeHarness attains 71.9% task success and 87.5% collision avoidance, surpassing the previous SOTA by 6.5% and 27.0%, respectively. These results are 2.3\times and 1.5\times those of the same agent without harnesses.

[NLP-1] Embedding Models Measure in Peculiar Ways

【速读】: 该论文旨在解决嵌入空间(embedding space)中语义相似性与距离是否能准确反映质量、距离、时间及体积等物理测量量的客观、唯一语义等价关系和度量问题。研究发现,物理测量在嵌入空间中的表征极为微弱,且存在异常的测量模式;进一步分析表明,物理测量的嵌入表示主要受表面字符串相似性的影响,即使对相似性进行重新校准,也未能显著提升其与物理量度之间的对齐程度。因此,该研究的核心挑战在于揭示嵌入空间中物理量表征的局限性,其解决方案的关键在于识别并量化字符串表面相似性对物理测量语义表征的主导作用,从而为改进嵌入模型在客观度量任务中的表现提供理论依据。

链接: https://arxiv.org/abs/2609.20821
作者: Juri Opitz,Andrianos Michail
机构: University of Zurich(苏黎世大学)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Embedding spaces define notions of semantic similarity and distance. We study whether those embeddings reflect physical measurements of mass, distance, time and volume, which admit a unique, objective notion of semantic equivalence and distance. We find that physical measurement is only weakly modeled in the embedding space, and that instead quite peculiar measurement patterns can be observed. Further analysis indicates that embedding representations of physical measurements are strongly influenced by superficial string similarity, and recalibration of similarity does not substantially improve the alignment.

[NLP-2] Unifying Models of Intergroup Hostility in Online Discourse

【速读】: 该论文旨在解决跨群体敌意言论(intergroup hostility)在现实话语中机制复杂且理论碎片化的问题,即现有社会心理学、道德心理学与政治科学中的基础理论虽被用于调节网络敌意言论,但彼此独立发展、解释路径常存在冲突,且缺乏在真实语境下的交叉验证。其解决方案的关键在于构建一个统一的实证框架,将六种核心敌意机制——边界建构(boundary construction)、威胁建构(threat construction)、替罪羊化(scapegoating)、负面评价(negative evaluation)、非人化(dehumanization)和行动导向(action orientation)——整合于同一分析模型中,并基于2024年美国总统大选期间来自TikTok、Truth Social和Twitter/X的286万条帖子进行建模。研究发现,边界建构与威胁建构在结构上处于核心地位,时间序列上呈现规律性演进:边界建构、贬损与行动导向多出现于早期;非人化与威胁建构次之;替罪羊化则最晚出现。该方法有效弥合了社会科学传统间的长期分裂,为计算社会科学研究跨群体敌意言论提供了超越单一标签识别的更系统、可操作的实证基础。

链接: https://arxiv.org/abs/2609.20808
作者: Patrick Gerard,Julia Mendelsohn,Kristina Lerman
机构: 未知
类目: Computation and Language (cs.CL); Social and Information Networks (cs.SI)
备注: 16 pages

点击查看摘要

Abstract:Hostile rhetoric toward social groups can normalize exclusion and justify mistreatment, as well as contribute to rising polarization and political violence. Efforts to moderate hostile rhetoric in online speech draw on foundational theories in social and moral psychology, and political science. However, these theories were developed largely in parallel, often propose different and sometimes conflicting accounts of how hostility develops, and have rarely been tested against each other in real discourse. The result is a fragmented understanding of the rhetorical mechanisms of hostility, without a clear sense of how they appear, and relate to each other, in real-world discourse. Using 2.86 million posts from TikTok, Truth Social, and Twitter/X during the 2024 U.S. presidential election, we model the mechanisms of six foundational theories of intergroup hostility – boundary construction, threat construction, scapegoating, negative evaluation, dehumanization, and action orientation – within a common empirical framework to recover the broader organization of intergroup hostility rhetoric. Structurally, we find that boundary construction and threat construction anchor the system; temporally, we find that these mechanisms tend to follow a regular ordering: boundary construction, derogation, and action orientation tend to appear early; dehumanization and threat construction later; scapegoating latest. Mapping how these theoretical frameworks actually manifest in discourse bridges longstanding divisions across social science traditions and presents computational social science with a clearer empirical foundation for modeling intergroup hostility rhetoric beyond single-label detection.

[NLP-3] An Empirical Study of Harness Design for Coding Agents

【速读】: 该论文旨在解决当前自主编码代理(autonomous coding agents)在长时程软件工程任务中,其编码框架(coding harness)各组件对性能影响不明确的问题。现有研究通常将框架视为整体系统进行评估,难以区分各组件的实际贡献。为此,本文提出一种轻量级编码框架,固定执行循环,仅变化三个核心组件:规划(planning)、动作空间(action space)与上下文管理(context management),通过在SWE-Bench Verified和Terminal-Bench 2.1上对四种模型进行176组匹配实验,系统评估五种上下文管理策略、四种上下文窗口预算及针对规划与动作空间的消融设置。研究发现:(1)随着上下文窗口预算收紧,上下文管理的价值显著提升,其主要作用在于防止上下文溢出失败;(2)在上下文管理策略中,优先采用基于规则的冗余内容剔除(elision)再结合大语言模型(LLM)摘要的方法,在效率上表现最优,而恢复被剔除内容的机制虽增加复杂性但未带来准确率提升;(3)规划模块对弱模型具有准确性支撑作用,而对强模型则主要降低计算开销,对准确率影响有限;(4)预定义工具对bash能力较弱的模型有性能增益,而具备bash能力的模型仅依赖bash接口即可实现高效执行,尤其在以命令行为中心的任务中成本显著更低。轨迹级分析揭示:上下文管理延长执行轨迹但不改变行为模式,规划决定轨迹终止位置,动作空间则影响代码生成粒度。这些发现为模型感知与资源预算感知的框架设计提供了依据,并构建了一个可模块化评估未来组件的分析框架。

链接: https://arxiv.org/abs/2609.20804
作者: Run-Ze Fan,Zihao Zhang,Simin Ma,Yebowen Hu,Shouju Wang,Kaiqiang Song,Fei Liu,Hamed Zamani,Xiaoyang Wang
机构: University of Massachusetts Amherst (马萨诸塞大学阿默斯特分校); Emory University (埃默里大学); Zoom (Zoom公司)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Software Engineering (cs.SE)
备注: 43 pages

点击查看摘要

Abstract:Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systems, leaving the effectiveness of individual components unclear. To enable component-level comparisons, we study this question with a lightweight coding harness whose execution loop is fixed while three components are varied: planning, action space, and context management. Across four models evaluated on SWE-Bench Verified and Terminal-Bench 2.1, we evaluate 176 matched settings spanning five context-management strategies, four context-window budgets, and targeted ablations of planning and action space. We find that: (1) Context management becomes increasingly valuable as the context-window budget tightens, with most of its benefit coming from preventing context-overflow failures. (2) Staging rule-based elision before LLM-based summarization provides the strongest overall efficiency among the context-management strategies, whereas making elided content recoverable adds machinery that models rarely use and yields no accuracy gain. (3) Planning shifts from an accuracy scaffold for weaker models to a cost saver for stronger models, with little change in accuracy. (4) Predefined tools improve performance for models with weaker bash proficiency, whereas bash-capable models can operate effectively with a bash-only interface and achieve substantially lower cost, especially on command-line-centric tasks. Trajectory-level analysis explains these effects: context management extends execution trajectories without substantially altering agent behavior, planning changes where trajectories stop, and the action space changes the granularity at which code is written. These findings inform model- and budget-aware harness design and provide a modular framework for evaluating future harness components.

[NLP-4] JEPA-Anything: Learning Predictive Models across Different Worlds

【速读】: 该论文旨在解决当前预测模型普遍存在的领域特异性问题,即现有世界建模方法难以在跨域、异构系统中实现通用性。其核心挑战在于:如何建立一种普适的学习原则,使模型能够在视觉、生物、临床轨迹、控制、分子动力学、物理场和气象等截然不同的系统中实现统一的世界建模。为此,作者提出JEPA-Anything框架,其关键创新在于采用正交预测因子分解(Orthogonal Predictive Factorization, OPF)机制:将潜在目标变量分解为互补的因子,通过独立路径分别学习各因子特征,并在共享的预测架构中进行重组。这一设计实现了对多源异构动态系统的统一建模能力。实验覆盖七类领域,涵盖表征学习、干预预测、分布外泛化及长时序动态模拟,包括10个匹配的动力学任务、超过1000例临床事件预测以及4种体系下100步的分子轨迹推演。结果表明,相比基准的JEPA模型,JEPA-Anything在全部10个动力学任务上均提升性能,在Interventional Pong中的单步干预预测误差降低34.8%;在所有4个分子系统中,1步与100步预测误差均优于现有方法。此外,基于因子识别的生物学干预策略在细胞共培养、患者来源类器官、肿瘤组织片段及小鼠模型中获得实验验证,且潜空间轨道模式成功恢复开普勒标度指数(斜率-1.4991),证实了该框架在连接世界建模、干预推理与可验证科学发现方面的普适性与有效性。

链接: https://arxiv.org/abs/2609.20800
作者: Taoyong Cui,Zhongyao Wang,Xinyue Xu,Weiyang Liu,Zhaochen Yu,Yuying Zhang,Qiang Gao,Mengyue Yang,Wanli Ouyang,Pheng Ann Heng,Yingcheng Wu,Zhenfei Yin,Ling Yang
机构: University of Hong Kong (香港大学); The Chinese University of Hong Kong (香港中文大学); Zhejiang University (浙江大学); Tsinghua University (清华大学); Peking University (北京大学); Shenzhen Institute of Information Technology (深圳信息职业技术学院); National University of Singapore (新加坡国立大学); Nanyang Technological University (南洋理工大学)
类目: Computation and Language (cs.CL)
备注: Code: this https URL

点击查看摘要

Abstract:World modeling enables intelligence to anticipate consequences, guide interventions, and learn from interaction. Yet predictive models remain domain-specific: can a common learning principle support world modeling across radically different systems? We introduce JEPA-Anything, a domain-agnostic framework based on orthogonal predictive factorization (OPF). Extending joint-embedding predictive architectures, OPF decomposes latent targets into complementary factors, learns them through dedicated pathways, and recombines them within a shared predictive design. We evaluate JEPA-Anything across seven domains: vision, biology, clinical trajectories, control, molecular dynamics, physical fields, and weather. Experiments span representation learning, intervention prediction, out-of-distribution generalization, and long-horizon dynamics, including 10 matched dynamics tasks, forecasting of over 1,000 clinical events, and 100-step molecular rollouts across four systems. Against matched JEPA baselines, JEPA-Anything improves reported metrics on all 10 dynamics tasks and reduces single-intervention prediction error on Interventional Pong by 34.8%. It achieves the lowest one-step and 100-step molecular errors among compared methods in all four systems. Beyond prediction, a factor-nominated biological intervention receives experimental support in cell co-cultures, patient-derived organoids, tumor fragments, and mice; latent orbital modes recover the Keplerian scaling exponent with a fitted slope of -1.4991. These results support a common factorized predictive principle across heterogeneous worlds, connecting world modeling with intervention and experimentally grounded scientific discovery. Code: this https URL

[NLP-5] RetireOPD: Self-Retiring On-Policy Distillation for Agent ic Reinforcement Learning

【速读】: 该论文旨在解决多轮智能体在强化学习(Reinforcement Learning, RL)训练中因仅获得轨迹级单标量奖励而导致的监督信号稀疏问题。现有方法通过自教师自监督的在策略蒸馏(On-Policy Distillation, OPD)引入密集的词元级监督,但其有效性受限于两个关键问题:具有特权信息的教师并不总是可靠,且教师监督的收益具有阶段依赖性。为此,论文提出一种新型训练范式RetireOPD(Self-Retiring On-Policy Distillation),其核心创新在于采用“自适应退休”机制:先独立优化一个技能条件化的解耦教师模型以利用环境奖励,随后联合使用强化学习与OPD训练无技能约束的学生模型;当学生与教师间的差异停止下降且达到教师成功率的目标比例时,学生主动终止对教师的依赖,此后仅通过强化学习继续训练。实验结果表明,该方法在Qwen2.5系列1.5B至7B模型上显著提升ALFWorld任务成功率14.1%至18.8%,WebShop任务准确率提升11.8%至19.0%,并在所有设置下超越自身教师性能,验证了其有效性与优越性。

链接: https://arxiv.org/abs/2609.20784
作者: Yan Yu,Zhengxi Lu,Yizhou Liu,Yichen Pan,Aozhe Wang,Qipeng Chen,Hua Yang,Wenqi Zhang,Weiming Lu,Qianglong Chen,Yongliang Shen
机构: Zhejiang University(浙江大学); Alibaba Group(阿里巴巴集团)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision from a self-teacher with privileged task skills, letting a skill-free student internalize them. This recipe, however, is undermined by two findings in agentic tasks: privileged information alone does not always make a teacher reliable, and the benefit of teacher supervision is stage-dependent. We therefore propose RetireOPD (Self-Retiring On-Policy Distillation), which first optimizes a decoupled, skill-conditioned teacher with environment rewards and then trains a skill-free student jointly with RL and OPD. Rather than following a predefined distillation schedule, RetireOPD adopts Adaptive Retirement: the student drops the teacher on its own once their discrepancy stops shrinking and it reaches a target fraction of the teacher’s success rate, after which training proceeds with RL alone. Across Qwen2.5 models from 1.5B to 7B, RetireOPD improves ALFWorld success rate over RL baseline by 14.1% to 18.8% and WebShop accuracy by 11.8% to 19.0%, and surpasses its own skill-conditioned teacher in every setting.

[NLP-6] Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations EMNLP26

【速读】: 该论文旨在解决当前大型语言模型(Large Language Models, LLMs)安全评估中存在的一种系统性缺陷:现有基于表面形式的有害性分类器无法有效识别和衡量“表层净化”下的隐性偏见与代表性伤害,导致对模型生成内容的安全性误判。其核心问题是,模型在对女性导向输出进行去有害化处理时,并未真正消除歧视性内容,而是通过语义重构将原本明确的性别歧视内容转化为看似无害但实质仍具伤害性的表达形式,这一现象被称为“有害性洗白”(harm laundering)。解决方案的关键在于提出一个三准则的有害性洗白形式化判定标准,并构建一个适用于任意生成式模型的三阶段检测协议。研究通过对15个从GPT-2到GPT-5的模型在三种人口统计条件下的45万条性别定向生成文本分析发现,随着模型迭代,女性导向输出中的性暴力相关主题显著消失,而男性导向输出则获得情感丰富性、照顾者角色及盟友身份等正向表征,且在GPT-5中出现将乳腺癌议题重构为“男性权利”辩论的极端案例,尽管三个独立分类器均将其评为非毒性内容。同时,女性导向生成内容的主题多样性相较男性下降36%,在GPT-4对齐边界处达到最低(W/M = 0.58),而代表性的伤害差异与发布日期呈显著正相关(ρ = +0.55, p = .034),但毒性评分却随代表性伤害上升而下降,表明毒性分数降低不能作为危害减少的有效代理指标。因此,该研究强调必须超越传统毒性检测框架,引入对代表性不平等与结构性偏见的深层评估机制。

链接: https://arxiv.org/abs/2609.20779
作者: Sarah Wyer,Sue Black,Noura Al Moubayed
机构: Durham University(杜伦大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Accepted at EMNLP 26 Main Conference

点击查看摘要

Abstract:Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit discriminatory content is transformed rather than removed. We call this \emphharm laundering. Analysing 450,000 gender-directed completions across 15 models spanning GPT-2 through to GPT-5 (OpenAI GPT lineage; three demographic conditions), we show that sexual violence clusters prevalent in GPT-2 women-directed output disappear by GPT-4, while men-directed completions gain positive representational territory (caregiving, emotional range, ally identity) that women-directed completions do not. The pattern is most visible at GPT-5: Topic~5 (1,997~documents) frames breast cancer as a men’s rights debate, while zero equivalent clusters appear in women-directed output. Three independent classifiers score this content as non-toxic. Sentiment scores invert at GPT-4: early models demean women; later models over-correct. Topic diversity in women-directed completions falls 36% relative to men at the GPT-4 alignment boundary (W/M~ = 0.58 , from 0.91 at GPT-2). REGARD representational harm disparity correlates with release date ( \rho = +0.55 , p = .034 ) while Detoxify does not ( \rho = -0.23 , p = .42 ): toxicity scores fall as representational harm grows. We formalise harm laundering as a three-criteria test and provide a three-stage detection protocol applicable to any generative model. Within the OpenAI GPT lineage, toxicity score reduction is not a sufficient proxy for harm reduction.

[NLP-7] dQwen 3.5: Hybrid-Attention Diffusion Language Models

【速读】: 该论文旨在解决如何高效地将预训练的自回归模型(Autoregressive Model, AR)适配为扩散语言模型(Diffusion Language Model, DLM)的问题,尤其关注在采用混合架构(即交替使用注意力机制与循环神经网络,RNN)的模型上实现这一目标。传统方法通常基于全注意力变换器(full-attention transformer),但近年来自回归建模逐渐转向混合架构以提升效率与表达能力。然而,这种架构与扩散模型所需的双向建模能力存在结构性冲突:RNN具有天然因果性,难以有效双向化。本文的关键解决方案在于验证混合架构作为自回归骨干网络在适配为扩散语言模型时的可行性与高效性。通过在0.8B、2B、4B和9B参数规模下对Qwen3.5进行适配,构建了dQwen3.5系列模型,结果表明,尽管存在结构不匹配,混合架构仍可作为高效的起点——其达到相同训练损失所需训练令牌数约为全注意力模型的一半,并在任意顺序解码行为及并行解码任务中表现出与全注意力扩散语言模型相当的性能,证明了混合架构在扩散语言建模中的潜力。

链接: https://arxiv.org/abs/2609.20751
作者: Anton Xue,Litu Rout,Aditya Akella,Adam Klivans,Sujay Sanghavi,Sanjay Shakkottai
机构: University of Texas at Austin (德克萨斯大学奥斯汀分校)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Adapting a pretrained autoregressive (AR) model is a cost-efficient route to a diffusion language model (DLM). While nearly all such adaptations start from a full-attention transformer, AR modeling has shifted toward hybrid architectures that interleave attention and RNN layers. This creates an obstacle for adaptation: unlike attention, RNNs are structurally causal and nontrivial to bidirectionalize. Despite this mismatch, we investigate whether such backbones can become effective DLMs by adapting Qwen3.5 at 0.8B, 2B, 4B, and 9B scales, yielding the dQwen3.5 family. We find that hybrid backbones can be efficient starting points for adaptation: against a full-attention control, the hybrid reaches a given training loss in about half the tokens. Across scales, dQwen3.5 resembles full-attention DLMs in any-order decoding behavior and performs strongly under parallel decoding.

[NLP-8] On-Demand Attention: Language Models Know When to Recall

【速读】: 该论文旨在解决生成式 AI(Generative AI)在长上下文推理任务中因全注意力机制(full-attention decoding)导致的计算效率低下问题。传统方法在每一步解码时均需对完整历史序列进行全局读取,即使部分上下文对当前预测贡献有限,造成不必要的计算开销。其核心解决方案是提出一种“按需注意力”(On-Demand Attention, ODA)机制,关键在于利用预训练模型自身解码状态中隐含的信息,通过一个轻量级的回忆头(recall head)实时预测当前步骤引入全局注意力的潜在收益,并据此动态决定是否调用全局注意力。该方法仅训练回忆头,保持预训练权重不变,同时保留完整的键值缓存(KV cache)以支持未来回溯。进一步地,在vLLM中实现基于GPU的条件执行策略,将减少的全局读取转化为实际的解码加速。实验结果表明,ODA在Qwen与Gemma等模型上,包括混合注意力骨干网络,能够显著降低全局注意力调用次数,同时几乎完全恢复局部注意力带来的性能损失,验证了预训练模型可自主引导其对所存储信息的访问,从而实现高效且智能的长上下文推理。

链接: https://arxiv.org/abs/2609.20734
作者: Haibo Feng,Ruiqi Liang,Hanyang Peng,Shiqi Yu
机构: Southern University of Science and Technology(南方科技大学); Peking University(北京大学); Peng Cheng Laboratory(鹏城实验室)
类目: Computation and Language (cs.CL)
备注: 28 pages, 5 figures

点击查看摘要

Abstract:Reasoning and agentic workloads increasingly demand efficient long-context inference. Yet full-attention decoding reads the growing history at every step, regardless of its benefit to the next prediction. We show that a pretrained model’s decoding states already contain information predictive of this benefit, before the global read. Building on this finding, we introduce On-Demand Attention (ODA), a local-first decoding method that uses a lightweight recall head to selectively invoke global attention as its predicted benefit changes during generation. ODA trains only the recall head, leaving pretrained weights unchanged and the complete historical KV cache available for future recall. We further implement GPU-side conditional execution in vLLM, translating reduced global reads into practical decoding speedups over full attention at long context lengths. Experiments across Qwen and Gemma models, including hybrid-attention backbones, show that selective recall recovers most of the performance lost under local attention while substantially reducing global reads. These findings support long-context inference in which pretrained models guide their own access to the information they retain.

[NLP-9] Dont Mask the Environment: Observation Supervision Changes How Agents Explore Under RL

【速读】: 该论文旨在解决标准监督微调(Supervised Fine-Tuning, SFT)中仅对智能体生成的动作令牌(action tokens)施加损失、而忽略环境观测(observation tokens)所导致的模型表征偏差问题。现有方法将环境观测作为上下文输入,但不将其作为预测目标,这可能导致策略在后续强化学习阶段缺乏对动作后果的准确建模能力。其解决方案的关键在于提出ActObs方法,即在SFT阶段同时对动作和观测令牌进行联合监督,尽管部署时智能体不会生成观测,但通过显式训练模型预测观测,可促使策略主动学习动作与环境状态变化之间的因果关系。该方法无需额外数据、参数、序列长度或前向传播开销,即可增强对环境动态的建模能力。实验表明,在基于GRPO的强化学习后,ActObs在Terminal-Bench 2.0和aider-polyglot跨领域代码编辑任务上均表现出更优性能,尤其在高采样预算下显著提升pass@k指标,并保持更高的策略熵与更小的策略漂移,表明其保留了更强的探索潜力与更稳定的初始状态。分析揭示,这是由于联合监督有效避免了动作与观测梯度的单边专业化现象,从而维持了对环境演变的预测能力,为下游强化学习提供了更优质的初始化。

链接: https://arxiv.org/abs/2609.20715
作者: Juzheng Zhang,Disha Makhija,Manoj Ghuhan Arivazhagan,Vinayshekhar Bannihatti Kumar,Rashmi Gangadharaiah
机构: University of Maryland; AWS AI Labs(亚马逊云科技人工智能实验室)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 29 pages, 9 figures, 11 tables

点击查看摘要

Abstract:Agent trajectories record what an agent does and what happens next. Yet standard supervised fine-tuning (SFT) applies loss only to agent-authored action tokens, using environment observations as context but not as prediction targets. We ask whether this convention provides the best initialization for subsequent reinforcement learning. We introduce ActObs, which also supervises the observation tokens already present in each trajectory. Although deployed agents never generate observations, learning to predict them encourages the policy to model action consequences without adding data, parameters, sequence tokens, or forward passes. The methods perform similarly after SFT but diverge after GRPO. On Qwen3-4B, GRPO from ActObs achieves higher pass@k at every evaluated sampling budget than its action-only counterpart on Terminal-Bench 2.0. On Qwen3-8B, it trades some pass@1 reliability for higher pass@k (+3.4 pp at pass@16) and solves more distinct tasks. The advantage extends to cross-domain code editing on aider-polyglot (+4.2 pp at pass@1 at 4B), whose tasks are unseen during SFT and RL. ActObs retains more entropy during RL while requiring less policy movement, leaving the final policy closer to its SFT initialization. Our analysis traces this difference to SFT: action and observation gradients rapidly become orthogonal, while action-only training leaves a large residual observation gradient and degrades environment prediction below the base model. Joint supervision prevents this one-sided specialization, preserving consequence prediction and preparing the policy for downstream exploration.

[NLP-10] Summarization Bias: The Directional Collapse of Objective Projection into Told-Mode Labels in Large Language Models — A Conceptual Framework and Registered Test Protocol

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在叙事理解与生成中存在的一种系统性偏差——总结偏倚(summarization bias),即模型倾向于将叙事意义抽象为概括性标签,而非还原其背后可重构的推断结构。这一问题的核心在于,现有模型在处理“呈现式”(shown)叙事时表现不足:根据布卢特教义(Bulut Doctrine),叙事效果沿“告知-呈现”轴展开,其中“呈现”模式要求读者通过物理线索与间接表达进行推理(客观投射,Objective Projection),属于高认知负荷情境;而“告知”模式则直接陈述情感与信息,依赖较少的读者重构。研究提出,LLMs在两个层面表现出方向性偏差:一是生成层面,模型在需通过客观投射传达情绪时,倾向于直接声明情绪;二是评估层面,模型在评判叙事质量时更偏好明确直白的“告知”式表达,忽视“呈现”式隐含性的价值。后者尤为关键,因大语言模型正越来越多地充当内容评判者与奖励模型,若存在向“告知”模式倾斜的偏向,将对文本创作施加选择压力,导致文学表达趋于扁平化、直白化。本文虽未验证该偏倚的实证成立,但明确定义了该概念,将其置于“模型作为裁判”类偏倚的语境中,并重新分析了一项独立可靠性研究,发现其结果与总结偏倚的方向一致;同时预先注册了双机制测试方案,设定了在特定决策规则下放弃该构念的条件,以确保研究的可验证性与透明度。

链接: https://arxiv.org/abs/2609.20712
作者: Levent Bulut
机构: 独立研究员(Independent Researcher)
类目: Computation and Language (cs.CL)
备注: v1.1. 8 pages. Also archived at Zenodo: this https URL

点击查看摘要

Abstract:This paper introduces and operationalizes summarization bias: a proposed systematic tendency of large language models (LLMs) to represent narrative meaning as an abstract summary label rather than as the reconstructable inferential structure that produces it. Within the Bulut Doctrine, narrative effect is theorized along a told-shown axis: in told mode, emotional and informational content is declared explicitly and requires little reader reconstruction; in shown mode, that content is suppressed at the surface and must be reconstructed from physical cues and indirection (Objective Projection). Shown mode is the higher-load condition the doctrine is designed to measure. The claim is that LLMs fail along this axis in a specific direction. Summarization bias is hypothesized to operate in two regimes: (i) a generative regime, in which a model asked to render an emotion through Objective Projection defaults to declaring it instead; and (ii) an evaluative regime, in which a model judging narrative quality rewards told-mode explicitness and under-detects shown-mode suppression. The evaluative regime is the more consequential, since LLMs increasingly serve as judges and reward models, and a directional bias toward told mode would impose a selection pressure degrading prose toward flat declaration. This report does not claim the bias is validated. It defines the construct, situates it against LLM-as-judge biases, rereads a completed independent reliability study as directional evidence consistent with it, and pre-registers a two-regime test with decision rules under which the construct would be abandoned. Comments: v1.1. 8 pages. Also archived at Zenodo: this https URL Subjects: Computation and Language (cs.CL) Cite as: arXiv:2609.20712 [cs.CL] (or arXiv:2609.20712v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2609.20712 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[NLP-11] HerHealthEval: Evaluating Multilingual and Register-Sensitive Understanding of Womens Health Communication

【速读】: 该论文旨在解决生成式AI在女性健康医疗沟通中多语言理解能力评估的不足问题,尤其关注模型在不同语用形式下对用户真实关切的准确识别与风险判断能力。现有评估多聚焦于响应质量,却忽视了用户表达意图是否被正确解析这一关键前提。其解决方案的关键在于提出HerHealthEval——一个受控的多语言女性健康沟通评估框架,通过为同一临床案例设计六种语用变体(标准、临床、通俗、间接/含蓄、情绪化关切、故意信息缺失),系统测试模型在语境多样性下的表现。特别地,故意信息缺失形式用于检验模型是否具备识别不确定性并主动请求澄清的能力。实验结果表明,传统评估指标可能掩盖安全相关的缺陷;而采用源语言衍生的、跨语言不变的风险标签进行再适配,可显著降低法语和阿拉伯语中的漏诊率(从0.994降至0.572和0.558),凸显了在多语言医疗AI评估中显式测试语用变体、不确定处理能力以及适配标签的来源与不变性的重要性。

链接: https://arxiv.org/abs/2609.20684
作者: Hassan Saeed Hassan Albattra,Mazen Mohammed Bahgat,Rahatara Ferdousi,Hana Essam Sayed Ahmed Amrya,Mariam Mousa
机构: Queen’s University (皇后大学)
类目: Computation and Language (cs.CL)
备注: 8 pages, 2 figures, 3 tables. Submitted to the 2026 International Conference on Large Language Models (LLM 2026)

点击查看摘要

Abstract:Large language models are increasingly used in healthcare communication, yet most evaluations emphasize response quality while assuming that the user’s concern has been interpreted correctly. We introduce HerHealthEval, a controlled evaluation framework for multilingual understanding of women’s-health communication. For each clinical case, HerHealthEval provides matched versions in English, French, and Modern Standard Arabic using six communicative forms: canonical, clinical, layperson, indirect or hedged, emotionally concerned, and deliberately under-specified. The first five express the same underlying concern and retain the same clinical information, whereas the under-specified form intentionally omits relevant details to test whether the model recognizes that clarification is needed. We evaluate a multilingual instruction model and QLoRA-adapted variants on concern classification, risk calibration, clarification behavior, parse compliance, and cross-form consistency. Results reveal that aggregate accuracy and consistency can conceal safety-relevant failures. A multilingual adaptation model reaches 0.994 under-triage in French and Arabic under language-asymmetric risk supervision. A controlled re-adaptation using source-derived, language-invariant risk labels reduces under-triage to 0.572 and 0.558, respectively. These findings show that robust multilingual healthcare evaluation requires explicit testing of register variation, uncertainty handling, and the provenance and invariance of adaptation labels.

[NLP-12] PAA: The Probabilistic Allen Algebra: A Generative and Complete Probabilistic Extension of Allens Interval Relations

【速读】: 该论文旨在解决经典时序关系代数——Allen区间代数在处理不确定性时序信息时的局限性问题。传统Allen代数依赖于精确的区间边界,其十三种基本关系为确定性(crisp)谓词,难以建模语言、感知、数据库或不确定历史中常见的模糊时序表达(如“刚在此之前”或“大致期间”),这些表达具有等级化语义。为此,论文提出概率化Allen代数(Probabilistic Allen Algebra, PAA),其核心在于将时序关系的概率由区间边界的分布推导得出,而非人为赋值。具体而言,时间点服从高斯分布,区间以高斯中点和截断高斯持续时间建模;所有关系均在统一概率空间中表示为边界顺序的逻辑谓词:点-点关系退化为误差函数,点-区间与区间-区间关系则转化为由线性不等式诱导的多元高斯正交概率。通过引入容差带(tolerance band),接触关系(如meets、starts、finishes、equals)获得正测度,且在单一容差下,十三种关系构成真正划分,并在容差趋零时恢复经典Allen代数。该框架自洽地推导出Allen的分类体系,而非预设;粗粒度谓词(如先于、重叠、包含)被定义为叶节点的并集,其概率为叶节点概率之和,且在区间退化为点的过程中,关系数量依次降为五种、三种,保持层次结构不变。此外,每个关系进一步分解为考虑相关性的时序原语,体现CIDOC CRM的精神。该代数具备尺度不变性,可区分“短暂之前”等程度表达与接触关系。所有结论经蒙特卡洛验证,并以开源、经过测试的Python包形式发布。

链接: https://arxiv.org/abs/2609.20634
作者: Julian Eggert(Honda Research Institute Europe, Offenbach, Germany)
机构: Honda Research Institute(本田研究所以)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 41 pages, 7 figures. Open-source implementation at this https URL

点击查看摘要

Abstract:Allen’s interval algebra is a qualitative calculus for temporal relations, but its thirteen base relations are crisp predicates over exact interval boundaries. This is inadequate for temporal information from language, perception, databases, or uncertain histories, where times, durations, and boundaries are uncertain and expressions such as “just before” or “roughly during” have graded meaning. We develop the probabilistic Allen algebra (PAA): a generative and complete extension in which relation probabilities are derived from distributions over interval boundaries rather than assigned as scores. Time points are Gaussian; intervals have Gaussian midpoints and truncated-Gaussian durations. Every relation is a boundary-ordering predicate in one common probability space: point-point relations reduce to error functions, and point-interval and interval-interval relations to multivariate Gaussian orthant probabilities induced by linear inequalities. Contact relations (meets, starts, finishes, equals) receive positive measure through a tolerance band, and under a single tolerance the thirteen relations form a true partition that recovers crisp Allen as the tolerance vanishes. The construction derives Allen’s taxonomy rather than positing it: coarse predicates such as precedence, overlap, and containment are unions of leaves whose probabilities are leaf sums, and this hierarchy is preserved as intervals collapse to points and thirteen relations reduce to five and then three. Each relation further decomposes into correlation-aware temporal primitives in the spirit of CIDOC CRM. The algebra is scale-invariant and separates graded expressions such as “shortly before” from contact relations. All results are Monte-Carlo validated and shipped as an open, tested Python package.

[NLP-13] UniPolicy: Unified Objective-Specific Policies for Generative Search Advertising

【速读】: 该论文旨在解决搜索广告系统中多目标优化的难题,即如何在兼顾用户相关性(relevance)、点击意愿(click propensity)与商业价值(commercial value)等异构目标的前提下,实现用户满意度与平台收益之间的平衡,同时克服因梯度竞争导致的全局次优性能问题。其解决方案的关键在于提出一种面向目标感知的多策略对齐框架UniPolicy,通过引入目标特异性前缀标记(objective-specific prefix tokens)、稀疏MoE-LoRA路由机制以及目标特异性残差前馈网络(residual FFNs),在共享主干网络内分层解耦参数,为不同业务目标构建差异化的参数空间与策略表达空间;进一步利用多阶段行为反馈构建成对偏好信号,增强已曝光但未点击样本中的相对偏好信息,并强化点击候选项在生成分布中的相对优势;在推理阶段,支持基于固定召回预算的并行、可定制化多策略束搜索(multi-policy beam search),灵活分配各目标的候选名额。大规模离线实验表明,UniPolicy在多个指标上实现了均衡提升且保持召回质量,优于单目标强化学习及简单的奖励融合基线;在线A/B测试结果也验证了其在真实系统中可带来0.71%的点击率(CTR)提升、1.58%的每秒请求量(RPS)增长及1.32%的广告收入增长,同时维持稳定的服务延迟。

链接: https://arxiv.org/abs/2609.20630
作者: Kun Yao,Yuhang Zhou,Yichi Zhang,Zeliang Tong,Shengri Xue,Haitao Wang,Siyu Lu,Qianlong Xie,Xingxing Wang
机构: Meituan(美团)
类目: Computation and Language (cs.CL)
备注: 13 pages, 5 figures, 4 tables

点击查看摘要

Abstract:Search advertising connects user intent with commercial content and plays a critical role in platform monetization. Recent systems typically align pretrained generative models with a single business reward, such as eCPM, or use naive reward fusion for preliminary multi-objective alignment. However, an ideal search advertising system must jointly account for heterogeneous objectives, including relevance, click propensity, and commercial value, to balance user experience and business value while mitigating globally suboptimal performance caused by gradient competition. We propose UniPolicy, an objective-aware multi-policy alignment framework. UniPolicy combines objective-specific prefix tokens, sparse MoE-LoRA routing, and objective-specific residual FFNs to hierarchically decouple parameters within a shared backbone, providing differentiated parameter and policy-expression spaces for different business objectives. It further constructs pairwise preferences from multi-stage behavioral feedback, supplementing the relative preference information in exposed-but-unclicked samples and strengthening the relative advantage of clicked candidates in the generation distribution. At inference, UniPolicy supports parallel, business-customizable multi-policy beam search, flexibly allocating candidate quotas across objectives under a fixed retrieval budget. Large-scale offline experiments show that UniPolicy delivers balanced improvements across multiple metrics while preserving retrieval quality, outperforming single-objective reinforcement learning and naive reward-fusion baselines. In a 7-day online A/B test on a real search advertising system, UniPolicy improves CTR by 0.71%, RPS by 1.58%, and advertising revenue by 1.32%, while maintaining stable serving latency.

[NLP-14] Chronicle: Cut-Point Replay for Regression Testing of LLM Agents

【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)代理在执行过程中因非确定性行为导致的故障难以复现的问题。由于推理过程不可位级重现、外部工具依赖动态状态,且多步骤决策轨迹在重放时极少一致,现有方法无法有效支持对代码变更的回归测试。其解决方案的关键在于提出Chronicle系统,通过在非确定性边界处以不可变封装(immutable envelopes)的形式记录代理运行,并引入“切点重放”(cut-point replay)机制:选择性地从记录中重放部分边界,其余边界则使用新代码实时执行。这一机制可将已记录的故障实例转化为持续集成环境中的回归测试用例。实验表明,在模拟模型边界的基准测试中,记录开销仅为每次穿越23微秒(占300毫秒模型调用的0.008%),完整重放不产生任何模型调用且在20次重复中保持位级稳定;切点测试能准确识别有缺陷的代码变更,同时通过安全和无害的修改验证了其有效性。在变异研究中,切点测试成功捕获所有绕过记录中危险操作的变异体,而传统桩函数方法则完全失效。

链接: https://arxiv.org/abs/2609.20625
作者: Tisha Chawla,Susheem Koul
机构: Microsoft(微软)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language model responses are non-deterministic, so failures in LLM agents are hard to reproduce: a failure depends on inference that is not bitwise reproducible, on tools that read changing state, and on a multi-step trajectory that a re-run rarely repeats. Record-and-replay makes a run reproducible, but existing agent tooling records runs only to trace or score them, not to test a code change against them. We present Chronicle, which records an agent run at its non-deterministic boundaries as immutable envelopes and replays it from the record. Its central operation, cut-point replay, serves a chosen subset of boundaries from the record and executes the complementary subset live with new code, turning a recorded incident into a regression test that runs in continuous integration. On a benchmark of 6 recorded failures with simulated model boundaries, recording adds 23 \mus per crossing (0.008% of an assumed 300 ms model call), full replay issues zero model calls and is bit-stable across 20 repetitions, and cut-point tests fail on faulty code and pass on guarded and benign changes for all 6 incidents. In a mutation study of the guarded tools, cut-point tests catch every mutant that lets the recorded unsafe action through, while a baseline that stubs every boundary, using the same assertion, catches none. Chronicle and the benchmark are publicly available at this https URL.

[NLP-15] What Does Privileged Information Add to On-Policy Self-Distillation?

【速读】: 该论文旨在解决生成式语言模型在自监督学习中,通过有偏教师(即带思考能力的教师模型)进行策略自蒸馏(On-policy Self-Distillation, OPSD)时,额外引入参考解题路径(如完整推理轨迹或优化后的解答)所带来的实际增益究竟来自何处的问题。核心问题是:相较于仅依赖模型自身推理能力的蒸馏过程,提供“特权参考”(privileged reference)是否真正带来了额外的、不可替代的学习优势,还是其作用本质上是促进不同推理模式间的知识迁移?解决方案的关键在于构建一个可复用的数学问题基准集——AMPLE-Math,包含5,319道共享同一答案但具有六种不同推理视图(reasoning views)的问题,从而实现对不同解题路径的精确对照实验。研究发现,参考自由蒸馏(reference-free distillation)已能解释Qwen3-1.7B模型在启用思维链评估下的主要性能提升,而额外参考信息带来的增益有限且依赖于学生模型训练状态;尤其在SmolLM3-3B中,完整推理轨迹仅在第50步带来2个百分点的提升。更重要的是,当使用长思维链输出替代短直接响应输出时,即使问题、参考与评估方式保持不变,性能反而下降,表明教师行为的模式选择会显著影响学习效果。进一步分析显示,改变令牌级别的监督信号并不显著改变学生行为,说明模型参数共享机制在跨模式迁移中的核心作用。因此,结论指出,OPSD的核心价值并非源于参考路径揭示的具体内容,而在于其作为桥梁,提升了学生模型对教师已有推理能力的访问效率,即通过共享参数实现从直接响应模式到思维链推理模式的知识转移。

链接: https://arxiv.org/abs/2609.20612
作者: XiuYu Zhang,Wei Chow,Junfeng Fang,Zhenkai Liang,Tat-Seng Chua
机构: National University of Singapore(新加坡国立大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:On-policy self-distillation (OPSD) lets a language model learn from a frozen copy of itself that sees an answer or a worked solution. Giving the teacher this extra information seems to offer the student more to learn, but how much does it add beyond distillation itself? To isolate that contribution, we construct AMPLE-Math, a reusable suite of 5,319 mathematical problems with six reasoning views that share the same answer, and compare each view with matched reference-free distillation. With a thinking-enabled teacher supervising direct-response rollouts, reference-free distillation accounts for much of Qwen3-1.7B’s improvement under thinking-enabled evaluation, both in domain and on external benchmarks. Evidence for an additional reference benefit is modest in Qwen, strongest for a polished solution, whereas complete traces add two percentage points in SmolLM3-3B at step 50. These benefits depend on the student being trained. At the same checkpoint, replacing short direct-response rollouts with long thinking-enabled rollouts turns gains into losses in both families while the problems, references, and evaluation stay fixed. Teacher profiles and matched loss interventions in Qwen further show that changing token-level supervision can leave student behavior largely unchanged. Together, these findings suggest that OPSD can improve access to existing reasoning capabilities through parameters shared by direct-response and thinking-enabled inference. The value of a privileged reference is what it adds to this cross-mode transfer, not how much of the solution it reveals.

[NLP-16] WiC is Not WSD: A Study on LLM s and Lexical Ambiguity Resolution AACL2026

【速读】: 该论文旨在解决语言模型在词义消歧任务中表现不佳的问题,尤其是针对“词在上下文中的含义(Word-in-Context, WiC)”这一挑战性任务。尽管近年来词汇语义任务取得进展,但WiC仍面临显著困难,其根源不仅在于对同一词语在不同上下文中语义用法的对比分析,更关键的是缺乏明确的语义粒度层级规范(sense inventory),导致模型难以准确判断目标词义。为此,研究提出通过引入候选词义(candidate senses),借鉴传统词义消歧(Word Sense Disambiguation, WSD)的范式,为模型提供显式的语义选项。实验结果表明,在所有测试场景下,引入候选词义均显著提升了模型在WiC任务上的表现,证明显式语义信息有助于模型做出更一致、更聚焦的判断。进一步的人工评估揭示,大量看似模型错误的判断实则源于标注者之间对词义边界的理解差异或标签本身的模糊性,而非模型对词汇意义的根本误解;尤其值得注意的是,大语言模型(LLM)常因过度细化词义区分而产生错误,反映出其在语义粒度把握上的偏差。因此,该研究的关键解决方案在于:通过显式提供候选词义,引导模型在合理语义粒度范围内进行判断,从而缓解因语义边界不明确和过度细粒度推理带来的性能瓶颈。

链接: https://arxiv.org/abs/2609.20593
作者: Yi Zhou,Kiamehr Rezaee,Danushka Bollegala,Mohammad Taher Pilehvar,Jose Camacho-Collados
机构: University of Liverpool(利物浦大学); Amazon(亚马逊); Cardiff University(卡迪夫大学)
类目: Computation and Language (cs.CL)
备注: Accepted to AACL 2026 (main)

点击查看摘要

Abstract:Word-in-Context (WiC) remains challenging for language models, despite recent progress on lexical-semantic tasks. We hypothesise that this difficulty arises not only from comparing two contextual uses of a word, but also from the absence of an explicit sense inventory that specifies the relevant level of semantic granularity. We evaluate open LLMs on WiC and traditional Word Sense Disambiguation (WSD) under similar settings. We find that providing candidate senses, similar to what is done in traditional WSD, improves WiC performance in all settings. In general, explicit sense information helps models make more consistent and targeted judgements. Human evaluation further shows that many apparent WiC errors reflect label ambiguity or mismatches between model and annotator sense boundaries rather than simple failures of lexical understanding. In particular, results show that LLMs overthink the sense distinction often leading to errors based on overly fine-grained distinctions.

[NLP-17] SAFARI: An Industrial Benchmark for LLM -Assisted Hazard Analysis and Risk Assessment EMNLP2026

【速读】: 该论文旨在解决生成式AI在安全关键型工程领域(特别是汽车功能安全)中应用时的可靠性问题,尤其聚焦于在ISO 26262标准框架下,大型语言模型(LLM)在辅助进行危害分析与风险评估(HARA)任务中的表现不足。其核心挑战在于:尽管大模型能够生成看似合理的危害描述,但在依据标准进行风险等级分类(如ASIL分级)方面仍存在显著缺陷。解决方案的关键是提出首个工业级基准测试SAFARI(Safety-Aware Functional Automotive Risk Inference),包含3,000个去标识化的工业级HARA案例,并设计了首个基于参考锚定的“大模型作为裁判”评估协议,以实现与专家判断高度一致的评估。实验表明,当前前沿大模型在风险分类上的最佳宏观F1仅为0.261,且链式思维提示(Chain-of-Thought prompting)对分类任务帮助有限甚至有害;错误分析进一步揭示主要失败原因在于危害生成阶段遗漏场景关键上下文信息,以及风险评估阶段对可控性判断失误,明确指出了需加强人工专家干预的关键环节。

链接: https://arxiv.org/abs/2609.20584
作者: Chenxi Wu,Zimu Wang,Haiyang Zhang,Wei Wang,Zhijie Xu
机构: Xi’an Jiaotong-Liverpool University (西安交通大学利物浦大学)
类目: Computation and Language (cs.CL)
备注: Accepted at EMNLP 2026 Industry Track

点击查看摘要

Abstract:Large language models (LLMs) are increasingly considered for safety-critical engineering, yet their reliability in regulated functional-safety workflows remains underexplored. We introduce SAFARI (Safety-Aware Functional Automotive Risk Inference), the first industrial benchmark for LLM-assisted automotive Hazard Analysis and Risk Assessment (HARA) under ISO 26262. It contains 3,000 de-identified industrial HARA cases and evaluates two coupled tasks: open-ended hazard analysis and standards-grounded risk assessment. To evaluate open-ended HARA artifacts, we propose the first reference-anchored LLM-as-a-judge protocol with high expert correlation. Experiments with nine frontier LLMs show that models often produce plausible hazard narratives but remain weak at ISO 26262 risk classification, with the best ASIL macro-F1 reaching only 0.261. Chain-of-Thought prompting provides limited benefit and often degrades categorical risk assessment. Error analysis further localizes major failures to scenario-critical context omissions during hazard generation and to controllability misjudgments during risk assessment, indicating where expert oversight should be concentrated. The dataset can be obtained from this https URL.

[NLP-18] Steering the Compass: Aligning Dynamic Psychological Counseling Conversations with Cognitive Behavioral Therapy Strategies EMNLP2026

【速读】: 该论文旨在解决当前基于大语言模型的心理咨询研究中,忽视来访者实时心理状态动态变化对认知行为疗法(Cognitive Behavioral Therapy, CBT)治疗决策影响的问题,这一缺陷限制了治疗方案的灵活性与临床效果。其核心解决方案是构建StratCBT数据集,该数据集包含9,688个咨询会话及约25.6万条对话语句,每个咨询师回复均明确标注对应八种不同的CBT策略。通过基于来访者负性思维建模并采用自洽对话(self-chat)生成高质量对话,同时引入真实会话作为指导,显著提升了数据在通用心理咨询与特定CBT技能方面的质量与多样性。实验结果表明,策略对齐的生成方法能有效提升大语言模型在模拟客户情境下的专业性与干预有效性,从而更真实地反映现实世界中的咨询过程。

链接: https://arxiv.org/abs/2609.20565
作者: Zimu Wang,Yiwen Jiang,Xiangyu Zhao,Yaling Shen,Jiahe Liu,Stephanie Fong,Maxmartwell H Cheng,Guilherme C Oliveira,Anh Nguyen,Robert Desimone,Barnaby Nelson,Dominic Dwyer,Zongyuan Ge
机构: Monash University (莫纳什大学); University of Liverpool (利物浦大学); Massachusetts Institute of Technology (麻省理工学院); The University of Melbourne (墨尔本大学); Orygen (奥里真)
类目: Computation and Language (cs.CL)
备注: Accepted at EMNLP 2026

点击查看摘要

Abstract:Recent advancements in large language models have revolutionized the field of psychological counseling, especially in the context of Cognitive Behavioral Therapy (CBT). While the success of CBT relies heavily on dynamic decision-making informed by the client’s real-time mental state, this aspect has often been overlooked in current research, limiting both flexibility and therapeutic outcomes. In this paper, we introduce StratCBT, a dataset specifically designed for psychological counseling conversations with CBT Strategies, consisting of 9,688 sessions and around 256K utterances, with each counselor’s response aligned with one of eight distinct strategies. The creation of StratCBT involves modeling clients based on their negative thoughts and generating high-quality counseling conversations through self-chat, incorporating realistic sessions as guidance, thereby significantly surpassing existing datasets in both general counseling and CBT-specific skills. We conduct extensive experiments to demonstrate the effectiveness of strategy-aligned generation and evaluate its efficacy in delivering professional and effective counseling with LLM-simulated clients to reflect real-world scenarios. The dataset can be obtained from this https URL.

[NLP-19] An Analysis of Training-Free Self-Reported Confidence in Language Models

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)生成内容时附带的数值置信度(confidence)是否真正反映其预测准确性的问题,即评估这种自报告置信度是否具备实际可解释性与可靠性。研究发现,直接通过模型在回答中显式表达置信度(verbalized confidence)是一个出人意料的强大基线,经基准测试误差审计后,在正确性预测任务上分别达到0.956和0.937的AUROC,显著优于三样本一致性(three-sample agreement)方法(分别为0.765和0.790)。然而,将显式置信度与三样本一致性的固定插值结合并未带来统计上显著的性能提升。此外,研究揭示了自一致性机制可能放大模型间的共性错误,且相同答案在不同提示下重采样时,置信度平均变化达0.043至0.084,阈值为0.8时有4%至9%的决策被反转。对100条带有置信度标注的人物传记陈述的探索性审计也显示,支持与反驳陈述之间的置信度差距微弱。因此,该研究的关键结论是:尽管某些形式的自报告置信度具有较高预测能力,但其有效性高度依赖于提示设计、存在相关性误差及基准数据噪声,表明当前模型的自我报告仍易受外部因素影响,尚未形成稳定可靠的客观度量。

链接: https://arxiv.org/abs/2609.20541
作者: Lukas Meyer,Sofia Rossi,Wei Chen,Thomas Laurent,Yiming Li
机构: DreamAI
类目: Computation and Language (cs.CL)
备注: workshop

点击查看摘要

Abstract:Large language models can report a numerical confidence together with generated content, but it is unclear whether this report is more than calibrated rhetoric. We analyze three training-free signals: confidence verbalized with the answer, post-hoc P(\mathrmTrue) , and agreement with three additional generations on the same 100 TriviaQA questions for two model families. Direct verbalization is a surprisingly strong baseline: after auditing benchmark errors, it reaches AUROC 0.956 and 0.937 for correctness prediction. Three-sample agreement is substantially weaker (0.765 and 0.790), and a fixed interpolation with verbalized confidence has no statistically reliable benefit. Four of nine errors from one model and two of eight from the other receive unanimous sample support, showing that self-consistency can amplify shared misconceptions. Re-eliciting confidence for the same fixed answers with equivalent prompts changes scores by 0.043 to 0.084 on average and flips 4% to 9% of decisions at a 0.8 threshold. An exploratory audit of 100 confidence-tagged biography claims further finds only a modest confidence gap between supported and contradicted claims. These results argue that useful self-reports remain sensitive to elicitation, correlated errors, and benchmark noise.

[NLP-20] Relational Attention for Data-Efficient Language Modeling EMNLP2026

【速读】: 该论文旨在解决生成式语言模型(Generative AI)在数据稀缺条件下仍能实现高效学习与强泛化能力的问题,尤其关注如何在有限训练数据(如10M至100M词)下提升模型对结构化关系信息的建模能力。其核心挑战在于:传统自注意力机制难以有效分离对象级(“感官”)语义特征与结构性/关系性信息,导致模型在纯关系推理任务中效率低下,且难以在语言建模中同时兼顾信息的解耦与融合。为此,论文提出Relational BabyLM,采用双注意力变换器(Dual Attention Transformer, DAT)架构,将关系注意力(Relational Attention, RA)从标准自注意力中解耦,以增强模型对关系结构的敏感性与数据效率。关键创新在于:一方面通过DAT架构实现对象级与关系信息的显式分离与可控整合;另一方面引入下一潜在状态预测(Next-Latent Prediction, NextLat)目标函数,促使隐藏状态逐步压缩历史信息为稠密信念状态,从而提升序列建模的内在表征能力。此外,提出一种基于旋转位置编码(RoPE)的新型符号检索机制,无需额外参数即可实现与可学习符号库相当的性能。实验表明,在严格(100M词)赛道上,该模型在55个参赛模型中排名第六,NLP子集第三,显著优于GPT-2基线,并在EWoK任务中取得最高得分,验证了架构设计在结构化语言泛化中的主导作用以及训练目标的辅助增益。

链接: https://arxiv.org/abs/2609.20530
作者: Adrian Brasoveanu,Ece Takmaz,Jakub Dotlačil
机构: UC Santa Cruz; Utrecht University (乌得勒支大学)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: BabyLM Workshop, EMNLP 2026. Source code: this https URL

点击查看摘要

Abstract:We present Relational BabyLM, a system submission to the BabyLM 2026 challenge that combines two cognitively motivated inductive biases in a single decoder-only Transformer. Architecturally, we replace standard self-attention with a Dual Attention Transformer (DAT), which separates the routing of object-level (“sensory”) lexical features from structural/relational information (Altabaa and Lafferty, 2025; Altabaa et al., 2024; Webb et al., 2024; Kerg et al., 2022; Webb et al., 2021). Relational attention (RA) disentangled from self-attention greatly increases data efficiency and out-of-training-sample generalization on purely relational tasks, but language modeling requires object-level and relational information to be integrated as well as disentangled, and RA-based LMs have remained largely unexplored. BabyLM’s data-constrained training and comprehensive evaluation is an ideal testing ground for whether that data efficiency transfers. As a training intervention, we add a Next-Latent Prediction (NextLat; Teoh et al. 2026) objective that encourages hidden states to compress history incrementally into a dense belief state. Architecture is the dominant factor for structural linguistic generalization; the objective is secondary but still significant. DAT’s three relational attention types (full RA vs. the simpler RCA and DisRCA variants) are largely interchangeable at 10M words; full RA pulls ahead at 100M. We also introduce a novel symbol-retrieval mechanism (RoPE-based, as opposed to learned, relative symbols) that matches learned symbol libraries while adding no parameters. On the strict (100M-word) track, our best model ranks 6th of 55 overall and 3rd of 55 on the leaderboard’s NLP-task subset at the time of writing; our two strongest models outperform the GPT-2 baseline on most benchmarks, with one attaining the highest EWoK score among strict-track entries.

[NLP-21] Edustories: A Collection of Real-world Case Studies from Classroom Practices

【速读】: 该论文旨在解决生成式 AI 在集体教学场景中应用不足的问题,即尽管生成式 AI(Generative AI)在个性化学生辅助方面展现出巨大潜力,但全球绝大多数教育实践仍发生在集体课堂环境中,而现有研究大多忽视了这一现实。为推动针对集体教学情境下 AI 辅助能力的研究,论文提出并构建了 Edustories 数据集,包含 1,492 个由教师撰写的案例研究,涵盖真实中小学课堂中具有挑战性的学生行为、教学干预措施及其结果。该数据集的核心价值在于支持对大语言模型(LLM)预测教师干预成效能力的评估。研究发现,当前四大语言模型家族中最先进的模型在预测课堂结果上的准确率仅为 58%,显著低于人类专家的 64%。这一差距揭示了当前生成式 AI 在理解复杂教育情境中的局限性,同时也凸显其作为一线教师辅助工具所具备的潜在价值和发展空间。

链接: https://arxiv.org/abs/2609.20484
作者: Michal Štefánik,Jan Nehyba,Jirina Karasova,Martin Fico,Lucie Škarková,Markéta Košatková,David Kosatka
机构: Masaryk University (马萨里克大学); National Institute of Informatics (日本信息研究所)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Despite the widely recognized potential of AI in education, most prior work has focused on individualized student assistance. In contrast, the majority of educational practice worldwide still takes place in collective classroom settings. To enable researchers to study AI assistance in collective teaching, we introduce Edustories, a dataset of 1,492 teacher-written case studies describing real elementary and high-school classroom situations involving challenging student behavior, pedagogical interventions, and their outcomes. Among many other applications, Edustories enables evaluating LLMs’ ability to predict the success of teacher interventions, crucial for providing practicing teachers with useful feedback. Comparing the latest models from four language-model families against expert assessments, we find that current models fall short of human expertise in predicting classroom outcomes; the strongest models reach 58% accuracy compared to 64% of human experts. This gap highlights both the limitations and the emerging potential of AI as assistants for practicing teachers.

[NLP-22] Stress-testing Alignment Midtraining

【速读】: 该论文旨在解决前沿模型在后训练(post-training)阶段对齐过程中,难以确保模型在所有潜在部署环境下的行为表现符合预期的问题,核心挑战在于模型必须具备超越后训练数据分布的泛化能力。为此,论文聚焦于中训练对齐(Alignment Midtraining, AMT)这一方法,其关键假设是:在预训练阶段持续引入与对齐相关的大量文本数据,可增强模型在后续训练阶段的对齐泛化能力。然而,论文通过大规模实验(最高达1100亿参数模型,10亿中训练样本)系统评估了AMT的有效性,发现其效果高度依赖具体情境——例如当后训练数据存在两种相互竞争的动机时,即使微调数据占比极小,也会完全抵消中训练带来的对齐引导作用;此外,在规则学习任务中,若某些规则仅在中训练或后训练数据中出现,模型才能稳健习得这些规则,否则无法有效学习。综合上述发现,论文认为当前缺乏充分的公开证据支持中训练能有效应对强大人工智能系统对齐中的根本性难题,因此对其作为通用解决方案的有效性持保留态度。

链接: https://arxiv.org/abs/2609.20412
作者: Sid Baines,Jonathan Bostock,Maria Angelica Martinez,Andrew Draganov,David Africa,Daniel Tan
机构: Arcadia Impact(弧度影响); Resolution
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:When aligning frontier models through post-training techniques, it is not possible to directly demonstrate all of the behaviours we want a model to exhibit in all possible deployment environments; our model must generalise outside of the post-training distribution. One proposed solution is alignment midtraining (AMT), which continues pretraining on large volumes of alignment-relevant documents to encourage generalisation in later stages of training. Despite the prominence of AMT as an alignment approach, there is limited public evidence for its effectiveness. To resolve this, we identify several assumptions around midtraining and evaluate them across scale: up to 110 billion-parameter models and 1 billion midtraining tokens. For instance, we study a scenario where post-training data is ambiguous between two possible motivations. We find that midtraining can steer the model’s motivation in simple versions of this setting. However, the presence of a tiny fraction of finetuning data which suggests a competing motivation erases the effects of AMT. We also study scenarios in which we want an AI to follow a number of rules, but only demonstrate a subset of them. We find that demonstrations must be present either in midtraining or post-training datasets for these rules to be robustly learned. Based on these and other findings, we do not believe that there is sufficient public evidence for us to confidently state that midtraining can address the core difficulties inherent in aligning powerful AI systems. Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI) Cite as: arXiv:2609.20412 [cs.CL] (or arXiv:2609.20412v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2609.20412 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[NLP-23] Xeno-Interpretability: Investigating the Alien Minds of LLM s

【速读】: 该论文旨在解决当前大语言模型(Large Language Models, LLMs)解释性研究中过度依赖人类已有概念框架的问题,即现有方法通常以“真实性”“拒绝”“欺骗”“人格”“有害性”等人类可理解的概念来解读模型内部表征,而忽视了模型可能具备的、无法用现有人类概念充分描述的内在区分机制。为此,论文提出“异域表征”(xeno-representation)这一新概念,指代那些在模型内部存在但缺乏对应人类语义概念的本体论表征结构,并引入“异域可解释性”(xeno-interpretability)作为研究范式。其核心解决方案在于将模型的语义空间划分为人类可解释的语义空间与异域语义空间(xeno-semantic space)——后者涵盖模型原生表征中无直接人类概念映射的区域。研究强调,即使无法用人类语言准确表达其语义内容,这些异域表征仍可通过可重复定位、几何特征刻画、因果操控及与下游行为关联等方式被实验识别。因此,该研究的关键在于突破传统可解释性对人类概念的依赖,转而建立一种以模型本体结构为中心的实证研究路径,旨在发现并表征那些可能影响模型行为但难以预测的原生表征机制,尤其对人工智能安全与多智能体系统中的隐性传播与稳定性问题具有深远意义。

链接: https://arxiv.org/abs/2609.20408
作者: F. Pierucci,M. Bracale Syrnikov,M. Prandi,M. Galisai,F. Giarrusso,P. Bisconti
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models are usually interpreted through concepts that humans already possess: truthfulness, refusal, deception, personality, harmfulness, and related categories. This paper asks whether models may also represent and use distinctions for which no adequate human concept exists. We call such internal structures xeno-representations, and their study xeno-interpretability. We distinguish the human-interpretable semantic space from the xeno-semantic space: the region of model-native representations for which no adequate human conceptual counterpart is available. We show that the space of possible internal distinctions in an LLM is substantially larger than the space available through finite human descriptions. We then separate experimental identification from semantic interpretation: an internal representation may be reproducibly located, geometrically characterized, causally manipulated, and linked to downstream behaviour even when its semantic content cannot be adequately expressed in human terms. On this basis, we sketch an empirical programme to identify xeno-representations. We finally examine the implications for AI safety and multi-agent systems, where model-native representations may propagate and stabilize across interacting agents while remaining only partially visible through human-readable communication. Xeno-interpretability therefore shifts the aim of interpretability from finding human concepts inside models toward discovering and characterizing the representational structures that are native to the models themselves and might affect their behaviour in unpredictable ways.

[NLP-24] Schema-Anchored Latent Reasoning for Semantic Parsing-Based Knowledge Base Question Answering

【速读】: 该论文旨在解决大而异构的知识库(Knowledge Base, KB)上基于语义解析(Semantic Parsing, SP)的问答任务中,大型语言模型(Large Language Models, LLMs)在生成可执行逻辑形式(Logical Form, LF)时面临的难题:如何准确选择与问题相关的模式元素(即关系和类),并将其组合成复杂的LF。现有基于LLM的方法往往在推理早期就做出离散的模式元素选择,导致错误的中间决策传播至最终结果,影响准确性。为克服此问题,论文提出一种名为SALR(Schema-Anchored Latent Reasoning)的方法,其核心在于通过在模型隐状态中生成连续的“思维”实现多步推理,从而延迟对LF决策的显式承诺。该方法通过一个由真实LF推导出的模式轨迹监督的对齐目标,将连续思维与KB模式元素的代码本(codebook)进行对齐,并将对齐后的模式代码作为后续推理步骤的输入,以实现模式引导的反馈机制。该机制无需模型输出显式的文本推理路径即可有效指导LF生成。实验表明,SALR在GrailQA和WebQSP数据集上均优于强基线模型,尤其在GrailQA的复合型问题上比现有最佳基线TIARA提升2.86 F1分数。进一步分析证实,该模式引导反馈显著影响了LF生成过程,且模式信息可从隐状态中恢复。

链接: https://arxiv.org/abs/2609.20398
作者: Guangze Gao,Zixuan Li,Sikui Zhang,Chunfeng Yuan,Wenjuan Li,Bing Li,Xiaolong Jin,Weiming Hu
机构: Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所); Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所); University of Chinese Academy of Sciences(中国科学院大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Semantic parsing (SP)-based knowledge base question answering aims to answer natural language questions by generating executable logical forms (LFs) over knowledge bases (KBs). When applying Large Language Models (LLMs) to this task, a key challenge over large, heterogeneous KBs is selecting question-related schema elements (i.e., relations and classes) and composing them into complex LFs. Recent LLM-based methods often make early discrete commitments to schema elements during intermediate reasoning, allowing incorrect intermediate schema decisions to propagate and finally result in incorrect LFs. To overcome this limitation, we propose SALR, a schema-anchored latent reasoning method for LF construction. It performs multi-step reasoning by generating continuous thoughts in the model’s hidden states, thereby delaying the explicit commitment to LF decisions. To ground this latent reasoning process in the corresponding KB schema, SALR aligns continuous thoughts with a codebook of KB schema elements through an alignment objective supervised by schema traces deterministically derived from gold LFs. It then incorporates the aligned schema codes into inputs for subsequent reasoning steps. This schema-mediated feedback guides LF generation without requiring the model to emit an explicit textual reasoning trajectory. Experiments on GrailQA and WebQSP show that SALR achieves consistent overall gains over strong baselines. Notably, on compositional questions from GrailQA, SALR outperforms TIARA, a strong SP-based baseline, by 2.86 F1 points. Further analyses show that schema-mediated feedback affects LF generation and that schema information is recoverable from the latent states.

[NLP-25] Lens: Bringing the Right Semantic Perspective into Focus for Training-Free Multimodal Representation Learning

【速读】: 该论文旨在解决现有生成式模型在零样本场景下进行表示学习时存在的语义视角错位(semantic perspective misalignment)问题,即当前的语义提取方法无法稳定地引导模型从下游任务所需的角度对输入内容进行语义理解,导致生成的表示仍被显著输入内容所主导,缺乏任务导向性。其解决方案的关键在于提出一种无需训练的框架Lens,通过“语义视角锚定”(Semantic Perspective Anchoring)将任务所需的解释角色与特定读出短语关联,明确指示后续特征提取的位置应承担的任务功能;并结合“上下文化短语读出”(Contextualized Phrase Readout),将该短语置于输入末尾,聚合其对应标记的状态,实现全上下文感知与任务视角对齐的联合建模。该方法在不修改模型结构、无需参数更新或重排序的前提下,实现了对任务条件下的证据整合与推理的精准捕捉,显著提升了跨36个MMEB数据集的整体Precision@1至63.9,较最优同类零训练基线提升10.2个百分点。

链接: https://arxiv.org/abs/2609.20252
作者: Xinran Liu,Shouqian Shi,Yixian Chen,Ruizhi Chen,Xin-Wei Yao,Sheng Zhong
机构: Nanjing University (南京大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 9 pages, 3 figures, 4 tables

点击查看摘要

Abstract:High-quality representations are essential for a wide range of downstream tasks. Dedicated embedding models are explicitly optimized for representation learning, yet their training data are often more limited in scale and diversity than the massive corpora used to pretrain modern large language models and multimodal large language models. Large-scale pretraining and instruction following enable autoregressive models to select relevant evidence, integrate multimodal information, and infer semantics under different task perspectives, creating a distinctive opportunity for training-free representation learning. However, our analysis reveals that existing semantic-elicitation methods do not reliably orient the extracted states toward the semantic perspective required by the downstream task. Consequently, the resulting representations often remain dominated by salient input content. We characterize this problem as semantic perspective misalignment and propose Lens, a training-free framework that makes representation readout task-directed. Semantic Perspective Anchoring associates the task-required perspective with a task-specific readout phrase, specifying the interpretive role of the positions later used for extraction. Contextualized Phrase Readout places the same phrase after the complete input and aggregates its token states, combining full-context access with the anchored perspective. The resulting representation reflects task-conditioned evidence integration and inference rather than a generic summary of salient content. Without parameter updates, architectural modification, or reranking, Lens achieves an overall Precision@1 of 63.9 across all 36 MMEB datasets, outperforming the closest same-backbone training-free embedding baseline by 10.2 points.

[NLP-26] he Public Discourse Corpus (PDC): A Speaker-Attributed Dataset for Valence and Epistemic Modality with Target Speaker Participation

【速读】: 该论文旨在解决公开人物访谈语料中情感极性(affective valence)与认识模态(epistemic modality)联合标注数据匮乏的问题,尤其针对跨专业领域、多说话人场景下目标说话人言论可分析性的可靠识别难题。其解决方案的关键在于提出并应用“目标说话人参与度”(Target Speaker Participation, TSP)这一五类注释分类体系,通过引入具有高信度(κ = 0.616)的标注标准,有效区分目标说话人、采访者及第三方发言,确保语料中保留的语音片段均为可分析的目标说话人内容。此外,研究构建了基于音频优先的说话人分离流程,结合局部Whisper自动语音识别(ASR)与pyannote说话人分离技术,实现了端到端的自动化处理,并将完整数据集、标注工具、跨平台验证样本及处理流水线开源发布,为后续相关研究提供了可复用的方法论框架与高质量基准数据。

链接: https://arxiv.org/abs/2609.20232
作者: Bo Chen
机构: Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
类目: Computation and Language (cs.CL)
备注: 16 pages, 1 figure

点击查看摘要

Abstract:We introduce the \textbfPublic Discourse Corpus (PDC), the first dataset of public-figure interview speech jointly annotated for affective valence and epistemic modality. The corpus contains 998 videos from 100 speakers across seven professional domains, yielding 186,642 sentences (3.1 million words) after sentence segmentation and filtering. To ensure that all retained videos contain analyzable speech from the intended speaker, we introduce \textbfTarget Speaker Participation (TSP)—a five-category annotation taxonomy with documented inter-annotator reliability ( \kappa = 0.616 )—as a key methodological contribution that any corpus construction project can adopt. Target-speaker turns are separated from interviewer and third-party speech through an \textbfaudio-first diarization pipeline combining local Whisper ASR with pyannote speaker separation, released as an open-source implementation. We release the annotated corpus, the annotation tools, the cross-provider validation sample, and the complete processing pipeline. The dataset is available at this https URL.

[NLP-27] Before the Warning Comes Too Late: Incremental Phone-Scam Detection from Speech

【速读】: 该论文旨在解决从原始电话音频中进行弱监督的增量式电信诈骗检测问题,其核心挑战在于训练仅提供通话级别的标签,而模型需在通话结束前实时更新预测结果。解决方案的关键在于提出StreamFraudNet框架,该框架通过使用冻结的自监督语音编码器对输入音频进行分段处理,结合循环时间建模(recurrent temporal modeling)与潜在窗口得分的可学习聚合机制,在不依赖语音转录或时间标注的前提下实现低延迟、高精度的欺诈风险评分。实验表明,该模型在控制条件下达到0.9953的ROC-AUC,显著优于基于声学特征和均值池化的基线方法,并接近强全局时序模型性能;同时可在10秒内生成首个预测结果,每2秒更新一次,且运行速度超过实时要求。消融实验进一步确认循环时间上下文是性能提升的主要贡献因素,凸显了在训练阶段引入延迟感知机制以优化早期预测能力的重要性。

链接: https://arxiv.org/abs/2609.20223
作者: Khang Nhat Hoang Vo,Anh Trac Duc Dinh,Tai Tien Ta,Tho Quan
机构: Mohamed bin Zayed University of Artificial Intelligence (穆罕默德·本·扎耶德人工智能大学); National University of Singapore (新加坡国立大学); Center for AI Reseach (CAIR), VinUniversity (越南大学人工智能研究中心); Faculty of Computer Science and Engineering, Ho Chi Minh City University of Technology (胡志明市科技大学计算机科学与工程学院), VNU-HCM (越南国家大学胡志明市分校)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:We study weakly supervised incremental telecom fraud detection from raw telephone audio, where training provides only conversation-level labels and predictions must be updated before a call ends. We introduce StreamFraudNet, which processes incoming audio through overlapping bounded-context windows using a frozen self-supervised speech encoder, recurrent temporal modeling, and learned aggregation of latent window scores. On a controlled English benchmark, StreamFraudNet achieves a ROC–AUC of (0.9953), significantly outperforming acoustic and mean-pooling baselines while remaining competitive with strong global temporal models. The model produces its first prediction after 10 seconds of audio, updates every 2 seconds, and operates faster than real time on the evaluated server hardware. Ablations identify recurrent temporal context as the principal contributor to performance. These results demonstrate that fraud risk can be scored incrementally from raw speech without transcripts or temporal annotations, while highlighting the need for latency-aware training to improve early prediction.

[NLP-28] Foundations of Stochastic Lexical Calculus: Semantic Descent and Random Dynamics on Probability Simplices

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)生成的提示依赖性词概率无法直接支持科学系统中对有意义状态的不确定性建模问题,尤其是在证据持续输入时需实现状态的顺序更新。其核心挑战在于如何从语言模型输出的概率中构建一个可验证、可演化且具备数学一致性的随机状态表示体系。解决方案的关键在于提出一种可观测框架,通过定义上下文语言的类型化可测变换,构造最小闭包表示,并给出语义更新唯一存在的充要条件;同时,通过界定不可约非闭包与累积误差的边界,在平均收缩条件下证明了概率单纯形上外部随机递归的存在性、唯一性与稳定性。该方法不依赖于语言模型内部的计算机制,而仅基于可观测的外部行为进行推断。实证部分通过冻结实验验证了理论预测:原始提示条件概率不满足预设不变性门限,但在经过提示特异性校准后,一个通用的三状态表示通过了稳定性检验,并在30条未触碰的八步路径中覆盖28条(名义水平0.90下为0.933),表明语言概率仅在经验证的闭包性、稳定性和覆盖性条件下,方可构成有效的随机状态表示。

链接: https://arxiv.org/abs/2609.20207
作者: Matthew F Dixon
机构: Artificial Intelligence Finance Institute (AIFI)
类目: Computation and Language (cs.CL); Category Theory (math.CT); Probability (math.PR); Machine Learning (stat.ML)
备注:

点击查看摘要

Abstract:Large language models produce prompt-dependent probabilities over words, whereas scientific systems require uncertainty over meaningful states that can be updated as evidence arrives. We develop an observable framework for determining when language-derived probabilities support such a sequential state representation. Theoretically, we define typed measurable transformations of contextual language, construct a minimal closed representation, and give necessary and sufficient conditions for semantic updates to exist uniquely. We bound irreducible nonclosure and accumulated error, and under average contraction prove existence, uniqueness and stability of an external random recursion on a probability simplex. These results define a stochastic lexical calculus without attributing an internal calculus to the language model. Empirically, frozen experiments test the observable implications. Raw prompt-conditioned probabilities fail the prespecified invariance gate; after prompt-specific calibration, a common three-state representation passes the stability gates and covers 28 of 30 untouched eight-step paths, or 0.933 at nominal level 0.90. Accordingly, language probabilities support a stochastic state only conditionally on verified closure, stability and coverage within a declared operating domain.

[NLP-29] o Copy or Not to Copy: Controlling Speculative Decoding via Intrinsic Model Signals

【速读】: 该论文旨在解决生成式 AI(Generative AI)中大规模语言模型(Large Language Model, LLM)推理加速技术——推测解码(Speculative Decoding, SD)所面临的根本性挑战:现有方法在神经推测(neural drafting)与基于上下文复制(context-based copying)两种策略之间存在性能权衡。其中,神经推测虽具备跨文本场景的鲁棒性,但速度提升有限;而基于复制的方法虽在重复性强的文本中可实现更高吞吐量,却因误触发频繁的意外重复(accidental repetitions)而降低实际效率,其根源在于仅依赖表面n-gram重叠作为复制决策依据,导致大量假阳性触发。本文提出SwitchSD,一种自适应框架,将复制行为视为大模型内部的潜在控制信号,并通过在目标模型的隐层表示上训练轻量级探测器(lightweight probes),以高精度(AUC达0.99)识别真正的复制意图。由此,系统能够动态切换至神经推测或上下文复制模式,实现了对复制行为的语义感知与精准控制。实验结果表明,在Llama与Qwen系列模型上,SwitchSD相较当前最优基线EAGLE3实现了最高15%的吞吐量提升,成功将复制从一种噪声驱动的启发式策略转变为基于模型认知的、可解释且高效的推理机制。

链接: https://arxiv.org/abs/2609.20186
作者: Roy Eisenstadt,Ido Cohen,Edo Cohen-Karlik,Lior Wolf,Itamar Zimerman
机构: 未知
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Speculative Decoding (SD) has significantly accelerated Large Language Model (LLM) inference, yet existing approaches face a fundamental tradeoff between two drafting strategies: neural drafting and context-based copying. Neural drafts (e.g., EAGLE3) provide robust performance across diverse text settings, while copy-based methods achieve higher speedups in copy-intensive regimes by generating candidates faster and exploiting long repetition spans for near-perfect speculation. We analyze existing copy-based methods and find that they are prone to accidental repetitions where surface-level n-gram overlap does not reflect a structural intent to copy, leading to false-positive triggers that ultimately degrade throughput. We introduce SwitchSD, an adaptive framework that treats copying as a latent control signal of the LLM. By training lightweight probes on the target model’s internal representations, SwitchSD identifies genuine copy-intent with high precision (AUC 0.99). This allows the system to dynamically switch between neural drafting (e.g., EAGLE) and context-based copying. Our results across Llama and Qwen families demonstrate throughput gains of up to 15% over state-of-the-art baselines like EAGLE3, effectively turning copying from a noisy heuristic into a principled, model-aware decoding regime.

[NLP-30] Fine-Tuning Models for Biomedical Relation Extraction

【速读】: 该论文旨在解决从海量生物医学文献中自动提取基因变异与表型之间关联关系的难题,这一任务在下一代测序(Next-Generation Sequencing, NGS)技术推动下愈发重要,但因数据规模庞大且信息分散,传统人工抽取方法已无法满足需求。其解决方案的关键在于提出并验证基于预训练模型(Pre-trained Models, PTMs)的自动化关系抽取框架,特别聚焦于变异-表型(variant-phenotype)领域。研究发现,对小型BERT类模型(尤其是DeBERTa)进行微调可达到接近当前最优(State-of-the-Art, SOTA)的性能;更关键的是,通过精细微调Google的Gemini Pro 1.0模型,在句子级别和摘要级别任务上均显著超越现有SOTA,表明大语言模型在复杂生物医学文本理解中的巨大潜力。

链接: https://arxiv.org/abs/2609.20169
作者: Claudiu Creanga,Liviu P. Dinu,Daniela Gifu
机构: 未知
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Next-Generation Sequencing has revolutionized the study of genetic mutations, enabling large-scale investigations into their roles in disease development. However, extracting meaningful insights from the vast amount of biomedical literature remains a complex challenge that cannot be addressed manually. In this paper, we present pre-trained models (PTMs) for the automatic extraction of relations from biomedical text, specifically targeting the variant-phenotype domain. Our evaluation on the SNPPhenA corpus demonstrates that fine-tuning small BERT-based models, particularly DeBERTa, yields strong performance, approaching the current state-of-the-art (SOTA). Additionally, our results indicate that carefully fine-tuning Google’s Gemini Pro 1.0 outperforms the existing SOTA for both sentence-level tasks (where the model processes only the target sentence) and abstract-level tasks (where the model processes the entire abstract).

[NLP-31] Design of the IBM Granite 5.0 TurboCTC ASR Model ICASSP2027

【速读】: 该论文旨在解决短时语音识别(short-form Automatic Speech Recognition, ASR)中模型在精度与推理速度之间难以平衡的问题,特别是在开源ASR基准测试中实现高效能。其核心解决方案在于设计了一个仅包含4.7亿参数的编码器单向模型Granite 5.0 Turbo CTC,通过多项关键技术优化达成优异的速度-精度权衡:首先,在Conformer模块中引入金字塔式时间下采样(pyramidal temporal subsampling),采用步进深度卷积实现高效的特征提取;其次,采用块对角(chunk-wise)自注意力机制以降低计算复杂度;第三,利用中间层的预测结果进行条件化处理,提升上下文建模能力;在训练方面,仅使用公开数据集,并创新性地采用Muon优化器及均衡的数据采样策略,显著提升训练稳定性与泛化性能;在推理阶段,通过将1×1卷积替换为线性层,并优化Conformer块中的注意力计算流程,进一步加速推理过程。上述综合优化使该模型在英文短时语音识别任务上达到开放ASR排行榜的前沿性能,且推理速度比最快竞争对手快一倍,同时支持宽松许可协议并可自由下载。

链接: https://arxiv.org/abs/2609.20104
作者: Brian Kingsbury,George Saon,Masayuki Suzuki,Hong-Kwang J. Kuo,Takashi Fukuda,Samuel Thomas,Vishal Sunder,Avihu Dekel
机构: 未知
类目: Computation and Language (cs.CL)
备注: 5 pages, 2 figures, submitted to ICASSP 2027

点击查看摘要

Abstract:We describe the architecture, training methodology and inference speedups of Granite 5.0 Turbo CTC, a 470 million parameter encoder-only model with an excellent speed-accuracy tradeoff. The architecture uses pyramidal temporal subsampling within Conformer blocks using strided depthwise convolutions, block-diagonal (chunk-wise) self-attention, and conditioning on intermediate predictions from the middle layer. Training highlights are the use of only publicly available data, the novel use of a Muon optimizer, and balanced data sampling. Inference speedups include replacing 1 x 1 convolutions with linear layers and optimizing the attention computation in the Conformer blocks. Collectively, these result in a model that is on the speed-accuracy Pareto frontier of the Open ASR leaderboard for English short-form ASR while being twice as fast as the fastest competitor. The model can be used under a permissive license and downloaded from this https URL.

[NLP-32] MATCH: Model-Aware Tool Learning with Curriculum Scheduling and Hierarchically Gated Rewards

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在使用外部工具进行任务时,基于强化学习的工具调用行为优化所面临的核心问题:一是固定阈值的课程学习策略难以随策略能力的动态演进而保持对齐,导致训练效率下降;二是累加式奖励机制在预测工具错误时会引发参数级信用泄露,影响奖励信号的准确性。为此,论文提出MATCH框架,其核心创新在于双管齐下的闭环设计:首先,模型感知的课程学习(Model-Aware Curriculum Learning, MACL)通过动态追踪由奖励驱动的样本难度,并在每轮迭代中选取接近当前策略能力边界及高难度的前k个样本,实现自适应样本调度;其次,分层门控奖励机制(Hierarchical Tool-call Gated Reward, HTGR)采用分层门控结构,依次对工具名称、参数键和参数值进行评分,仅在前置条件满足时才授予相应层级的奖励,从而精确分配信用并防止错误传播。该机制同时服务于策略优化(GRPO)与课程难度更新,形成政策优化与样本调度间的闭环反馈。在API-Bank和BFCL V3基准测试中,MATCH分别达到72.19%和62.87%的整体准确率,显著优于主流监督与强化学习基线,且在来自两个模型家族的四种骨干网络上均表现出一致的性能提升。

链接: https://arxiv.org/abs/2609.20082
作者: Shihao Liu,Hao Yin,Lijun Liu,Zhengzong Chen,Yuanyuan Zhao,Fei Huang
机构: Honor Device Co., Ltd; School of Cyber Security, University of the Chinese Academy of Sciences
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Tool learning enables large language models (LLMs) to use external tools for tasks beyond parametric knowledge. Reinforcement learning can optimize tool-call behavior from feedback, but current methods still face two problems: fixed-threshold curricula can become misaligned with the policy’s evolving capability boundary, and additive rewards can leak argument-level credit when the predicted tool is wrong. To address these problems, we propose MATCH, a closed-loop framework for model-aware tool learning with curriculum scheduling and hierarchically gated rewards. Model-Aware Curriculum Learning (MACL) maintains reward-derived sample difficulty that co-evolves with the policy, and each epoch selects samples near the current capability boundary together with a top-k pool of harder cases. Hierarchical Tool-call Gated Reward (HTGR) scores tool name, argument key, and argument value as a gated chain, granting credit at each level only when prerequisites hold. The same HTGR rewards drive both GRPO updates and MACL’s difficulty refresh, closing the loop between policy optimization and sample scheduling. On API-Bank and BFCL V3, MATCH reaches 72.19% and 62.87% overall accuracy, outperforming the main supervised and RL-based baselines. Backbone experiments further show consistent improvements across four backbones from two model families.

[NLP-33] Reading Emotions in the Token Space: Discriminative Adaptation of SpeechLLM s for Emotion Recognition

【速读】: 该论文旨在解决生成式语音大模型(SpeechLLM)在情感识别任务中因采用生成解码器进行情绪预测所导致的固有缺陷:生成解码器不适用于分类任务,容易输出目标类别集之外的标签,并且对高频类别存在偏好。其解决方案的关键在于提出一种判别式适配机制,通过一个单层线性分类头读取最终提示词(prompt token)的隐藏状态,实现一次前向传播即可输出指定类别的预测结果,且无需修改主干模型。该方法基于模型原本用于生成的隐藏状态进行判别推理,从而在保持模型结构一致的前提下,实现了生成式与判别式推理的可控对比。该分类头设计简洁,仅使用单一线性层,在几乎不损失准确率的情况下显著提升了可解释性——每个情感类别对应输出空间中的一个方向,揭示了与之相关的关键词元。在IEMOCAP数据集上,针对两种不同的SpeechLLM架构,该方法均提升了宏平均F1分数并消除了幻觉现象,尤其在真实语音识别(ASR)转录文本上表现最优。进一步分析表明,这些情感方向编码了间接关联,反映了大规模网络文本中存在的偏见。

链接: https://arxiv.org/abs/2609.20081
作者: Hasindri Watawana,Sergio Burdisso,Esaú Villatoro-Tello,Manjunath K E,Kadri Hacioglu,Petr Motlicek,Andreas Stolcke
机构: Idiap Research Institute(瑞士伊迪亚研究所); EPFL(瑞士联邦理工学院); Uniphore(美国 印度); Brno University of Technology(捷克布尔诺理工大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Sound (cs.SD); Audio and Speech Processing (eess.AS)
备注:

点击查看摘要

Abstract:SpeechLLMs have shown strong potential for emotion recognition, yet they read the predicted emotion off a generative decoder not suited for classification: it can emit labels outside the target set and favors frequent classes. We propose a discriminative adaptation that reads the final prompt token’s hidden state through a classification head, producing a label in one forward pass without modifying the backbone. Because this readout starts from the hidden state the model would otherwise decode, it gives a controlled comparison of generative and discriminative inference in an otherwise identical speechLLM. We keep the head a single linear layer, trading little accuracy for interpretability: each emotion becomes one direction in the LLM output token space, revealing associated tokens. On IEMOCAP, across two speechLLM architectures, it improves Macro F1 and removes hallucinations, with largest gains on realistic ASR transcripts. Our analysis reveals that these emotion directions encode indirect associations mirroring biases in web-scale text.

[NLP-34] Marginal utility matrix factorization and the Key-Value (KV) cache: a unified information-economic framework for sovereign geo-mining inference

【速读】: 该论文旨在解决如何将经济学中的边际效用(marginal utility)概念与机器学习中的矩阵分解(matrix factorization)及Transformer语言模型的键值缓存(Key–Value cache)机制建立理论联系,从而为模型资源分配与信息压缩提供统一的优化框架。其核心解决方案在于揭示三类不同现象——评分矩阵的奇异值谱、模型学习表征的投影协方差算子特征值谱以及缓存淘汰与低秩压缩——均服从于在内存预算约束下的效用最大化原则,并最终归约至同一最优分配规则:保留特征值超过绑定约束影子价格(shadow price)的前几维。该框架被应用于地质矿产文档的结构化信息自动化提取任务中,推动了多轮推理协议、分层TIES模型融合方法以及结合提取质量、定位漂移与能耗的综合选择策略的设计,其中漂移项通过条件风险价值(Conditional Value-at-Risk)进行加权。实验表明,一个1120万参数的分层分类器仅需单卡五分钟左右训练,即可在973份铀矿勘探文档的测试集上达到90.0%的一级准确率,优于50份文档人工审计样本上专有模型的92.0%,且推理延迟仅为2.62毫秒/卡,成本可忽略;诊断分析进一步揭示了均匀密度融合中存在的退化模式(即跨地理区域输出完全相同但置信度高),而采用分层校准密度重执行融合可有效消除该问题。全规模提取基准(含LoRA微调)目前为预测结果,尚待进一步实证验证。

链接: https://arxiv.org/abs/2609.20068
作者: Caroline Gans Combe(INSEEC)
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: Version 11, 14 septembre 2026. 49 pages, 9 tables. Les valeurs de l’architecture souveraine sont projet{é}es et non mesur{é}es ; le calcul {à} grande {é}chelle est en cours. Soumission pr{é}vue {à} IEEE Transactions on Artificial Intelligence

点击查看摘要

Abstract:This paper builds a theoretical bridge between the economic notion of marginal utility and two machine-learning constructs, matrix factorization and the Key–Value cache of transformer language models. The singular value spectrum of a rating matrix is shown to be a diminishing marginal utility schedule for latent factors, the eigenvalue spectrum of the projected covariance operator to be the marginal utility schedule of a model’s learned representation, and cache eviction and low-rank cache compression to be instances of constrained utility maximization under a memory budget. The three collapse into a single allocation rule: retain the top dimensions whose eigenvalue exceeds the shadow price of the binding constraint. The framework is applied to the automated extraction of structured information from geo-mining documents, where it motivates a multi-pass inference protocol, a layer-wise TIES model merging procedure, and a selection policy combining extraction quality, localization drift and energy, scalarized with a Conditional Value-at-Risk term on drift. Two empirical contributions are reported. An 11.2-million-parameter hierarchical classifier, trained in about five minutes on a single GPU, reaches 90.0 per cent level-1 accuracy on a held-out test set from a 973-document uranium-exploration corpus, against 92.0 per cent for a proprietary model on a fifty-document human audit of the same corpus, at a latency of 2.62 ms per card against approximately 2,000 ms for the API and at negligible cost. A diagnostic of uniform-density TIES merging exposes a reproducible degenerate mode in which the merged model returns token-identical outputs across five geographically distinct districts while declaring high confidence; re-executing the merge under layer-wise calibrated densities removes that signature on the diagnostic sample. The full-scale extraction benchmark, including LoRA fine-tuning, is reported as projected rather than measured and remains an empirical extension of this work.

[NLP-35] Geopolitical Divisions Across Languages in Large Language Models

【速读】: 该论文旨在解决多语言环境下生成式AI(Generative AI)在处理地缘政治议题时是否存在语言依赖性偏差的问题,特别是其在评估俄乌战争立场时是否因提问语言不同而产生不一致的政治倾向。研究发现,同一AI模型在不同语言下对同一事件陈述的回应存在系统性差异:以俄语系或部分非西方国家官方语言提问时,倾向于生成更偏向俄罗斯的立场,而以西方主流语言提问则更倾向于支持乌克兰。这一现象与各国公众对俄罗斯的态度、联合国投票倾向及对乌克兰援助程度呈现显著相关性。解决方案的关键在于揭示了信息战可能通过影响用于训练AI的数据分布,将潜在的地缘政治偏见“编码”进模型输出中,进而放大全球政治分歧。该研究强调,生成式AI的输出并非中立,其语言输入的语境可显著改变结果,提示需警惕训练数据中的隐性意识形态污染。

链接: https://arxiv.org/abs/2609.20005
作者: Maxim Chupilkin
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY)
备注:

点击查看摘要

Abstract:People increasingly turn to AI chatbots for news and explanations of world events. But do they receive the same political answers when they ask in different languages? Here we show that the language of a question can change how the same AI systems assess the war in Ukraine. We ask GPT, Claude and Gemini to evaluate twenty statements about the war in 112 languages, collecting 67,200 responses. The balance between Russia-leaning and Ukraine-leaning responses differs across languages. When we group responses by countries’ official languages, they follow a pattern resembling worldwide political divisions: relatively more Russia-leaning answers correspond to more favourable public views of Russia, less support for Ukraine in United Nations votes, and less aid to Ukraine. The broad pattern recurs across all three models and remains when individual statement pairs are removed. Our findings suggest a possible route through which information warfare may shape the text used to train AI models, which may in turn spread geopolitical biases.

[NLP-36] DeepSeek -V4.1-Flash: Pushing the Limits of KV Cache Compression

【速读】: 该论文旨在解决长时序代理(long-horizon agents)在实际部署中面临的高计算、高存储与高带宽开销问题,尤其是预填充(prefill)阶段的高昂计算成本以及大尺寸键值缓存(KV cache)对高带宽内存(HBM)和固态硬盘(SSD)容量及数据传输带宽的持续压力。其核心解决方案在于提出DeepSeek-V4.1-Flash模型,该模型基于因果编码器-解码器(Causal Encoder-Decoder, CED)架构,通过动态激活机制实现解码阶段每token激活160亿参数,而预填充阶段仅需激活80亿参数,显著提升生成式任务的成本效率。同时,为突破KV缓存压缩极限,模型融合了跨层KV缓存复用(cross-layer KV cache reuse)与FP4精度缓存(FP4 KV caching)技术,在Compressed Sparse Attention 2(CSA2)框架下将全局KV缓存占用降低至每token仅890字节,约为前代DeepSeek-V4-Flash的1/4;进一步通过专用部署优化策略SWA Bounded Replay,使持久化KV缓存(位于SSD或主机内存)占用减少至前代的1/8。尽管缓存规模大幅缩减,模型仍展现出优于基线的性能表现。此外,研究还对DeepSeek-V4架构进行了精简并引入多项高效扩展设计,结合包含45万亿词元的多模态语料进行预训练,并完成全面后训练,最终在多样化文本与多模态代理场景中均取得优异表现。

链接: https://arxiv.org/abs/2609.19969
作者: DeepSeek-AI:Anyi Xu,B. Li,Bangcai Lin,Bing Xue,BingCheng Xian,Bingzheng Xu,Bochao Wu,Bowei Zhang,Boyi Deng,C.C. Yu,Chao Jin,Chaofan Lin,Chen Dong,Chenbing Wang,Chenfan Feng,Chengda Lu,Chenggang Zhao,Chengqi Deng,Chengyuan Zhang,Chenhao Xu,Chenqi Zhao,Chenze Shao,Chuhao Wang,Chuqi Zhang,Damai Dai,Dejian Yang,Deli Chen,Di Huang,Di Wu,Donghao Li,Erhang Li,Eric Fu,F. Zhou,Fangwei Zhou,Fangyun Lin,Fangzhou Yuan,Feiyu Xia,Fucong Dai,Guangbo Hao,Guanglin Li,Guanting Chen,Guoai Cao,Guofan Fan,Guolai Meng,Guowei Li,Haichuan Zhang,Haiyang Ma,Haiyang Shen,Han Li,Han Yu,Han Zhang,Hangyuan Deng,Hanwei Xu,Hanxiang Xu,Hanxun Zhong,Hao Guo,Hao Jiang,Hao Li,Hao Qin,Haodong Wen,Haofen Liang,Haofeng Huang,Haohua Liu,Haoling Zhang,Haoming Luo,Haoran Yang,Haotian Xu,Haotian Yuan,Haoting Huang,Haowen Luo,Haoyang Cai,Haoyu Chen,Haozhe Ji,Hengran Zhang,Hengrui Wang,Hengxu Wu,Honghui Ding,Hongxuan Tang,Huadong Wang,Huanqi Cao,Huazuo Gao,Hui Qu,Hui Zeng,J. Yang,J.H. Jin,J.H. Zhang,J.X. Zou,Jia Yu,Jiahui Zhou,Jiajun Chen,Jialiang Huang,Jialin Zhao,Jiamin Tang,Jian Zhou,Jianan Tong,Jianwen Li,Jiaqi Zhu,Jiarui Wang,Jiasheng Ye
机构: DeepSeek-AI(深度求索)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Together, these compute, storage, and bandwidth demands constitute the primary bottleneck to further lowering deployment costs. To address this challenge, we introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. With its Causal Encoder-Decoder (CED) architecture, the model activates 16B parameters per token during decode but only 8B parameters during prefill, substantially improving cost efficiency for agentic workloads. To push the limits of KV cache compression, DeepSeek-V4.1-Flash combines cross-layer KV cache reuse in Compressed Sparse Attention 2 (CSA2) with FP4 KV caching. These designs reduce its global KV cache footprint (always in HBM) to 890 bytes per token, roughly 1/4 of the corresponding footprint of DeepSeek-V4-Flash. Further, through a dedicated deployment optimization known as SWA Bounded Replay, DeepSeek-V4.1-Flash reduces its persistent KV cache footprint (always on SSD or in host memory) to roughly 1/8 of that of DeepSeek-V4-Flash. Despite its much smaller KV cache footprint, the model delivers substantially better performance than the baseline. In addition, we streamline the DeepSeek-V4 architecture and introduce several efficient architectural extensions. We pretrain DeepSeek-V4.1-Flash on a multimodal corpus comprising 45T tokens and conduct comprehensive post-training, yielding strong performance across diverse text-based and multimodal agentic scenarios. Model checkpoints are available at this https URL.

[NLP-37] Before the Arrest: Benchmarking LLM s on Criminal Profiling from Incomplete Evidence EMNLP2026

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在刑事司法领域应用中一个关键的空白问题:即在逮捕前,如何基于不完整证据推断嫌疑人特征这一核心挑战。现有研究几乎全部集中于嫌疑人身份已知的案后场景,而对犯罪调查初期、依赖碎片化现场证据进行嫌疑人画像的预逮捕推理任务关注不足。为此,作者提出了一个名为“侦查、画像与定罪”(Profiling, Investigation, and Judgment, PIJ)的多国真实命案数据集,涵盖来自五个国家的2500起真实案件,系统评估了LLMs在刑事调查全链条中的表现。其解决方案的关键在于构建一个覆盖三个典型阶段的任务框架:刑事画像(需运用溯因推理从零散证据中推断嫌疑人属性)、犯罪过程重建(测试结构化信息提取能力)以及量刑预测(要求法律演绎推理)。实验结果表明,随着任务从显性事实抽取转向隐含的、针对未知嫌疑人身份的推理,模型性能显著下降;尤其在动机推断和受害者-加害者关系等需深层推断的任务上,模型表现仍远落后于人类专家,且存在性别、年龄及动机归因等方面的系统性偏差。研究揭示,基于不完整证据进行预逮捕推理仍是当前生成式人工智能面临的重大开放性挑战。

链接: https://arxiv.org/abs/2609.19965
作者: Yutong Yao,Yanjie Cao,Guanhua Chen,Xu Yang,Junchao Wu,Zeyu Wu,Lidia S. Chao,Derek F. Wong
机构: NLPCT Lab, Faculty of Information Science and Computing, University of Macau(澳门大学信息科学与计算学院自然语言处理与计算技术实验室); Faculty of Social Sciences, University of Macau(澳门大学社会科学学院)
类目: Computation and Language (cs.CL); Computers and Society (cs.CY)
备注: Accepted by EMNLP 2026 Findings. Codes are available at: this https URL

点击查看摘要

Abstract:Large Language Models (LLMs) are increasingly applied to legal and criminal justice tasks, yet existing work focuses almost exclusively on post-arrest scenarios where the suspect’s identity is already known, leaving the critical pre-arrest challenge of inferring suspect characteristics from incomplete evidence largely unexplored. To fill this gap, we introduce the Profiling, Investigation, and Judgment (PIJ), comprising 2,500 real homicide cases from five countries. PIJ evaluates LLMs across three tasks that span the entire criminal investigation pipeline: criminal profiling, which requires abductive reasoning to infer suspect attributes from fragmentary scene evidence, crime process reconstruction, which tests structured information extraction, and sentence prediction, which demands legal deductive reasoning. We evaluate 9 powerful LLMs and find that performance degrades systematically as tasks shift from explicit fact extraction to implicit reasoning over unknown suspect profiles. Categories requiring inferential reasoning, such as motivation and victim-offender relationships, remain the primary bottlenecks. Further analysis reveals substantial gaps between LLMs and human experts, along with pervasive biases in gender, age, and motive attribution. Our findings indicate that pre-arrest inference from incomplete evidence remains an open challenge.

[NLP-38] KoNeoBench: A Curated Evaluation Dataset for LLM Understanding of Korean Neologisms EMNLP2026

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在评估过程中对韩语新词(Korean neologisms)理解能力不足的问题。现有评测基准多基于已确立的词汇体系,难以覆盖近年来不断涌现的新词语及其语义演变,且其以英语为中心的设计无法充分反映韩语中实词与功能词素之间高度组合性的语言特性。为此,论文提出KoNeoBench,一个专为评估LLMs对韩语新词理解能力而构建的基准测试集。该基准涵盖自2020年以来在线新闻中记录的1,785个韩语新词,并经专家词典学审校,每个词条均包含用法例句、构词分析及词典式定义。基于此资源,研究设计了四项评估任务,并报告了近期主流模型的表现及人工基准结果。实验表明,当前LLMs在恢复词源成分、区分语义类别以及生成准确释义方面存在明显局限,揭示了当前模型在应对韩语近期词汇演化中的特定挑战。KoNeoBench的提出为评估和改进模型对动态语言变化的理解提供了关键工具。

链接: https://arxiv.org/abs/2609.19916
作者: Soha Lee,Soojin Lee,Heesung Yang,Hyunju Song,Hyunji Lee,Jinsan An,Jeongwan Shin,Jin Hyun Park,Jun Lee,Hyeyoung Park,Kilim Nam
机构: Kyungpook National University (庆北国立大学); Daegu Gyeongbuk Institute of Science and Technology (DGIST) (大邱庆北科学技术院); Texas A&M University (德克萨斯农工大学); Yonsei University (延世大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Accepted to Findings of EMNLP 2026. Code and data are available at the project repository

点击查看摘要

Abstract:Large language models (LLMs) are typically evaluated on static benchmarks, even though natural language constantly evolves through newly emerging words and meanings. Existing Korean benchmarks are centered on established vocabulary and therefore provide limited coverage of such recent lexical change, and their English-oriented design makes it difficult to assess the typological properties of Korean, in which content words combine productively with functional morphemes. In this paper, we introduce KoNeoBench, a benchmark for evaluating LLMs’ understanding of Korean neologisms. KoNeoBench is built on 1,785 Korean neologisms attested in online news since 2020 and curated through expert lexicographic review. Each entry provides usage examples, word-formation analyses, and dictionary-style definitions. Based on this resource, we define four tasks and report results on recent models, together with a human baseline. Our experiments show that current LLMs exhibit clear limitations in recovering source components, distinguishing semantic categories, and generating accurate definitions. These results reveal specific aspects of recent Korean lexical change that remain challenging for current LLMs. KoNeoBench is available at this https URL .

[NLP-39] Generalization through Lexical Abstraction in Transformer Models: The Case of Functional Words

【速读】: 该论文旨在解决预训练变换器模型在处理语言中功能性词汇(如代词、副词等)时,是否能够像人类一样理解其语法与语义抽象能力的问题。具体而言,研究关注语言模型能否识别并建模诸如“研究人员写了论文”与“他们写了它”这类句子之间的句法与语义平行性,这依赖于对名词的抽象替代机制。其解决方案的关键在于将语言学问题映射至预训练变换器模型的嵌入空间,通过对比名词与其对应的功能性词汇(如代词)在孤立及并行句式中的嵌入表示,探究二者在结构上的关联性。研究发现,功能性词汇在嵌入空间中位于中心位置但具有独特性,符合其作为多场景占位符的行为特征;而并行句式的嵌入分布于不同的子空间。进一步实验表明,仅使用功能性或仅使用词汇化句子进行训练均无法揭示共享结构——前者因词汇过度一致导致泛化不足,后者则因表达多样性过高难以捕捉共性。然而,当联合训练功能性与词汇化句子时,模型能够有效学习并表征出隐藏的句法-语义结构,从而实现对语言抽象性的准确建模。

链接: https://arxiv.org/abs/2609.19887
作者: Giuseppe Samo,Vivi Nastase,Paola Merlo
机构: Idiap Research Institute(伊迪亚研究所); University of Geneva(日内瓦大学)
类目: Computation and Language (cs.CL)
备注: 16 pages, 11 figures

点击查看摘要

Abstract:Pronouns, adverbs and other functional words (such as they, her, somewhere, there) are often used in language to replace concrete nouns or phrases, when their properties - such as gender, grammatical number - provide sufficient information for the given context. Do pretrained transformer models encode such functional words in a manner that allows them to be used like humans do? Can language models recognize the syntactic and semantic parallelism of sentences such as “The researchers wrote the paper” and “They wrote it”, which relies on such lexical abstraction? We map these linguistic questions into the embedding space of a pretrained transformer model, and compare representations of nouns, with the representations of the pronouns and adverbs that can replace these nouns, in isolation and in parallel lexicalized and functional sentences. We then probe for shared syntactic and semantic structure in the embeddings of parallel lexicalized and functional sentences. We find that functional words are located centrally compared to nouns, but are also distinct, which is congruent with their behaviour as place-holders in a wide variety of contexts. The analysis of the embeddings of parallel (lexicalized and functional) sentences show them inhabiting different subspaces of the embedding space. Experiments that distil the structural information of the sentence show that training on either type of data does not reveal the shared structure - because of the over-consistency of the vocabulary (in case of the functional data), and the too much variety (in case of the lexicalized versions). However, training with a mix of functional and lexicalized sentences, the shared structure emerges. Comments: 16 pages, 11 figures Subjects: Computation and Language (cs.CL) Cite as: arXiv:2609.19887 [cs.CL] (or arXiv:2609.19887v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2609.19887 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Giuseppe Samo [view email] [v1] Thu, 17 Sep 2026 08:35:09 UTC (2,495 KB)

[NLP-40] Evaluating Communicative Success in Machine-Translated Conversation

【速读】: 该论文旨在解决当前机器翻译(Machine Translation, MT)驱动的口译代理在实时跨语言对话中评估不足的问题。现有评估方法主要基于孤立句对的保真度(fidelity)指标,无法有效衡量对话中语义理解、交际意图传递及社会文化适切性等关键沟通维度,导致对口译代理真实沟通成效的误判。其解决方案的核心在于提出一个可复用的三层检查清单-评分框架(checklist-and-judge framework),从语义、语用和文化社交三个维度系统评估口译代理在单轮与多轮交互场景下的表现,涵盖自然性、意图准确性和社会适宜性等人类沟通的关键要素。该框架支持模拟用户在对话过程中对翻译内容的动态响应,并对每一轮对话及其整体连贯性进行综合评分。通过受控扰动、多评委对比与人工标注等手段进行广泛验证,研究构建了覆盖12个翻译方向的单轮基准(10种口译配置,4种语言:阿拉伯语、孟加拉语、印尼语、韩语,共5,624个源自OpenSubtitles的场景)以及包含全部6组语言对的多轮脚本与实时模式研究。结果表明,口译成功率随语义→语用→文化社交维度逐层下降,而传统MT指标难以捕捉强模型在复杂对话中的沟通失败;提示工程分析进一步揭示场景上下文、结构化指令与文化背景信息对提升沟通成功具有显著作用,但效果因系统而异。因此,该工作不仅提供了一套面向对话场景的口译代理评估框架与基准数据集,更强调将“沟通成功”作为核心评价维度,与现有保真度指标形成互补,推动口译代理向真正有效的跨语言交流工具演进。

链接: https://arxiv.org/abs/2609.19885
作者: Faiz Ghifari Haznitrama,Alice Oh
机构: KAIST(韩国科学技术院)
类目: Computation and Language (cs.CL)
备注: 32 Pages, 11 Figures, 11 Tables

点击查看摘要

Abstract:Interpreter agents built on machine translation (MT) increasingly mediate live conversation between people who do not share a language, yet we still evaluate them with metrics built for isolated sentences, which measure fidelity rather than whether communication succeeds. We introduce a reusable three-layer checklist-and-judge framework that evaluates interpreter-mediated conversation across semantic, pragmatic, and cultural-social dimensions, covering the naturalness, intent, and social appropriateness that fidelity metrics leave unmeasured. It runs in both single-turn and interactive multi-turn settings, where simulated users reply to translated messages as the conversation unfolds and each turn is scored alongside the conversation as a whole. We extensively validate it through controlled perturbations, cross-judge comparisons, and human annotations. Our main single-turn benchmark evaluates 10 interpreter setups across Arabic, Bengali, Indonesian, and Korean from 5,624 OpenSubtitles-derived scenarios spanning 12 translation directions, and our multi-turn study covers all 6 language pairs in scripted and live modes. Results show a consistent decline from semantic to pragmatic and cultural-social success, while conventional MT metrics overlook failures among stronger interpreters, and prompt ablations show that scenario context, structured instructions, and cultural context improve communicative success, although gains vary across setups. Our work thus provides an evaluation framework and benchmark for interpreter agents in conversation, and highlights the importance of communicative success alongside existing translation metrics.

[NLP-41] PetriBench: Benchmarking LLM Reasoning over Dynamic State Spaces

【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)推理能力评估中存在基准测试孤立特定推理技能、依赖外部知识或扩展成本高昂的问题。其核心解决方案是提出PetriBench,一个紧凑、完全自包含且可扩展的基准测试框架,利用成熟的佩特里网(Petri net)形式化方法对动态状态空间中的LLM推理进行评估。PetriBench将推理任务划分为四个按作用范围和时间跨度划分的任务族,并通过逐步增加结构复杂度生成易、中、难三个难度等级,所有任务均基于精确的真值进行评估。实验结果表明,不同模型在各类任务中的准确率随难度递增而系统性下降,且高难度任务能更清晰地揭示任务特异性能力分布。进一步分析显示,推理时计算资源的增加虽能提升性能,但对不同类型推理任务的影响存在差异;同时,过程式生成方式实现了与结构复杂度的平滑扩展。综上,PetriBench为探究LLM推理能力的强弱边界及其缩放行为提供了一个统一且可拓展的评估环境。

链接: https://arxiv.org/abs/2609.19883
作者: Pyrros Koussios,Benjamin Jäger,John Hua Yao,Ajay Sridhar,Violet Xiang,Chenhao Li
机构: ETH Zürich(苏黎世联邦理工学院); Stanford University(斯坦福大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Logic in Computer Science (cs.LO)
备注:

点击查看摘要

Abstract:Characterizing LLM reasoning remains an open challenge, as many existing benchmarks isolate specific reasoning skills, rely on external knowledge, or are costly to extend. We introduce PetriBench, a compact, fully self-contained, and scalable benchmark for evaluating LLM reasoning over dynamic state spaces using Petri nets, a mature formalism for modeling real-world concurrent and distributed systems. PetriBench organizes reasoning into four task families varying by scope and temporal horizon, with Easy, Medium, and Hard levels generated by increasing structural complexity and evaluated against exact ground truth. Across a diverse set of proprietary and open-weight models, accuracy decreases consistently with difficulty, while harder instances expose increasingly distinct task-specific capability profiles. Additional analyses show that test-time compute improves performance but interacts differently with different reasoning tasks, and that procedural generation yields smooth scaling with structural complexity. Together, these results show that PetriBench provides a unified and extensible setting for probing the strengths, limits, and scaling behavior of LLM reasoning.

[NLP-42] D-Quant: Driftable Entropy Coding for KV Cache Quantization

【速读】: 该论文旨在解决大语言模型(LLM)部署中键值缓存(KV cache)带来的内存瓶颈问题,其核心挑战在于KV cache的内存占用随序列长度和批处理大小线性增长,对内存容量与带宽构成巨大压力。现有压缩技术中,量化方法因高效且易于部署而备受关注,但主流方法采用固定位宽量化(fixed-width quantization),其本质限制是每个b位表示仅能提供2^b个离散量化级,随着位宽降低,可用量化级别呈指数级减少,导致信息损失严重、性能急剧下降。研究进一步发现,固定位宽量化未能利用KV cache值的高度非均匀分布特性——经过旋转与归一化后,KV值近似服从正态分布,大部分数值集中于中心区域,仅有少量出现在尾部。然而固定位宽编码对高频与低频符号均分配相同比特数,造成效率浪费。相比之下,熵编码(entropy coding)可通过为高频符号分配短码字、低频符号分配长码字,有效降低平均表示比特数。但其可变长度输出不适用于高度并行的注意力计算内核,后者依赖规则的内存布局与固定步长访问以实现高效反量化与计算。为此,本文提出D-Quant框架,引入一种漂移机制(drift mechanism),将每个标记的熵编码表示转换为固定大小的比特流,从而在保持熵编码压缩优势的同时,支持注意力内核中的规则内存访问与并行反量化,实现了压缩效率与计算效率的协同优化。

链接: https://arxiv.org/abs/2609.19880
作者: Yi Su,Hong Liu,Guanghua Yu,Jianchen Zhu
机构: 未知
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:The KV cache has become a major bottleneck in deploying LLMs, as its memory footprint grows linearly with sequence length and batch size, imposing substantial pressure on both memory capacity and bandwidth. Among various KV cache compression techniques, quantization is particularly attractive due to its effectiveness and ease of deployment. However, most existing methods rely on fixed-width quantization, where a b bit representation is inherently limited to 2^b quantization levels. As the bit width decreases, the number of available levels shrinks exponentially, leading to severe information loss and rapid performance degradation. We further observe that fixed-width quantization fails to exploit the highly non-uniform distribution of KV cache. After rotation and normalization, KV values approximately follow a normal distribution, with most values concentrated near the center and only a small fraction appearing in the tails. Nevertheless, fixed-width coding allocates the same number of bits to frequent and rare symbols. Entropy coding naturally exploits such non-uniformity by assigning shorter codewords to frequent symbols and longer ones to rare symbols, substantially reducing the average number of bits required for representation. However, its variable-length output is not suited to highly parallel attention kernels, where efficient dequantization and computation rely on regular memory layouts and fixed-stride accesses. To bridge this gap, we propose \textbfD-Quant, a flexible KV cache quantization framework that introduces a \textbfdrift mechanism to convert entropy-coded representations of each token into fixed-size bitstreams, enabling regular memory access and parallel dequantization within attention kernels.

[NLP-43] VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering

【速读】: 该论文旨在解决低资源语言(特别是泰卢固语)在口语问答(Spoken Question Answering, SQA)领域缺乏基准评测数据集及自动评估可靠性未被量化的问题。其核心挑战在于,现有大语言模型的问答能力主要集中在高资源语言,且在口语场景下的表现尚未充分验证,尤其在泰卢固语这一低资源语种中存在显著空白。解决方案的关键在于构建并公开发布首个泰卢固语口语问答基准数据集VākQA,包含2,001个事实型问答对,覆盖六个领域,配备2.53小时语音音频、双语转录文本以及人工验证的参考答案。研究通过与人类评判对比,系统评估了多种自动评估方法的有效性,发现Gemini-as-a-judge虽最接近人工评分但存在评价标准不一致问题,而开源权重模型则倾向于因表层形式差异而惩罚正确答案。基于此验证后的评估框架,论文对专有与开源模型在输入模态、语言和领域上的表现进行了全面基准测试,揭示了泰卢固语表达中的文化特异性在翻译中易丢失、语音输入引发的音素混淆改变问题语义,以及级联式自动语音识别(ASR)与机器翻译(MT)错误逐步累积等关键现象。

链接: https://arxiv.org/abs/2609.19879
作者: Bhavana Akkiraju,Ravi Sastry Kolluru,Sri Charan D,Srihari Bandarupalli,Santosh Kesiraju,Anil Vuppala
机构: International Institute of Information Technology Hyderabad, India; Brno University of Technology, Speech@FIT, Czechia
类目: Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
备注: Paper is accepted in IEEE SLT 2026

点击查看摘要

Abstract:Question answering has advanced rapidly with large language models, but predominantly for high-resource languages, in both text and spoken settings. Spoken question answering (SQA) benchmark for Telugu remains unexplored, and the reliability of automatic evaluation in this setting remains unquantified. We introduce VākQA, a Telugu SQA benchmark of 2,001 factoid question-answer pairs across six domains, with 2.53 hours of speech audio, bilingual transcriptions, and human-verified reference answers. We first validate evaluation methods against human judgements: Gemini-as-a-judge best approximates human ratings but is non-uniformly strict, while open-weight judges systematically penalize correct Telugu answers that differ in surface form from the reference. Using this validated setup, we benchmark proprietary and open-weight models across input modality, language, and domain. We observe that Telugu phrasing retains cultural specificity that is lost in translation, speech input introduces phonetic confusions that alter question meaning, and cascaded ASR-MT errors compound progressively. VākQA is publicly released.

[NLP-44] Uni-LaDiR: Latent Diffusion Unifies Multimodal Reasoning

【速读】: 该论文旨在解决多模态推理(Multimodal Reasoning)中不同模态间表征差异导致的推理不连贯问题。现有方法通常将各模态特有的思维标记(thought tokens)简单拼接为单一序列,迫使模型在跨模态推理过程中自行弥合语义鸿沟。为此,本文提出统一潜在扩散推理框架 Uni-LaDiR(Unified Latent Diffusion Reasoner),其核心在于将来自不同模态的教师推理步骤映射至共享的潜在空间,通过统一编码器生成保留后续推理与最终答案所需信息的通用思维标记。由于同一上下文可能支持多个有效推理路径,该框架采用扩散模型(Diffusion Model)从输入及先前推理块中预测下一阶段的思维标记。通过联合训练编码器与扩散推理模块并共享模型参数,促使思维标记既具备任务相关性,又可由上下文准确预测。在11个视觉-语言模型(VLM)基准和2个视觉-语言-动作(VLA)套件上的实验表明,Uni-LaDiR 在视觉推理任务上相对于最强基线实现7.3%的相对提升,在机器人操作任务上实现6.1%的相对提升,验证了其在跨模态协同推理中的有效性。

链接: https://arxiv.org/abs/2609.19878
作者: Haoqiang Kang,Yizhe Zhang,Nikki Lijing Kuang,Yian Ma,Lianhui Qin
机构: UC San Diego(加州大学圣迭戈分校); Meta
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Multimodal reasoning requires models to draw on information from multiple modalities throughout the reasoning process. Yet existing methods often concatenate modality-specific thought tokens in a single sequence, leaving the model to bridge representational differences as it reasons across modalities. We introduce Uni-LaDiR (Unified Latent Diffusion Reasoner), a framework that brings these thoughts into a shared latent space for reasoning. A unified encoder maps teacher reasoning steps from different modalities into shared thought tokens, trained to preserve the information needed for later reasoning steps and the final answer or action. Because the same context can support multiple valid next steps, we use diffusion to predict the next block of thought tokens from the input and preceding blocks. Jointly training the encoder and diffusion reasoner with shared model weights encourages thought tokens to be both useful for the task and predictable from the available context. At inference, the model generates these tokens without teacher observations. Across eleven vision-language model (VLM) benchmarks and two vision-language-action (VLA) suites, Uni-LaDiR achieves relative gains over the strongest evaluated baselines of 7.3% on visual reasoning tasks and 6.1% on robot manipulation tasks.

[NLP-45] JustMem: Just-Enough Memory Access for Long-Term Conversations

【速读】: 该论文旨在解决长时对话记忆(long-term conversational memory)中如何高效检索充分证据的同时避免无差别扩展上下文的问题。核心挑战在于:相关证据可能分散于多个会话中,而过度压缩又可能导致关键细节丢失,因此不同查询对记忆访问的需求存在差异。为应对这一问题,论文提出从“发现广度”(discovery breadth)与“读取保真度”(reading fidelity)两个维度建模记忆访问需求,并设计了JustMem系统。其关键创新在于将对话历史以紧凑的原子化记忆形式存储,并根据具体查询动态调整记忆访问策略:LOOKUP用于局部证据的快速检索,COMPOSE扩展发现范围以定位跨会话证据,REPLAY则提升读取保真度以恢复对精度敏感的原始内容。实验表明,相较于现有记忆系统,JustMem在LoCoMo和LongMemEval-S基准上实现了最高的平均准确率与召回率,同时显著减少了生成式模型在记忆构建与推理过程中所需的令牌数量。

链接: https://arxiv.org/abs/2609.19877
作者: Guanhua Chen,Yanting Wang,Wenjing Zhi,Lei Sha
机构: Beihang University(北京航空航天大学)
类目: Computation and Language (cs.CL)
备注: 12 pages, 8 tables, 3 figures. Includes appendix

点击查看摘要

Abstract:Efficient long-term conversational memory requires retrieving sufficient evidence without indiscriminately expanding the context presented to the language model. This is challenging because relevant evidence may be distributed across multiple sessions, while compression may discard details needed for answering. Different queries therefore require different forms of memory access. To capture these demands, we formulate memory access along two dimensions: discovery breadth, which controls how broadly evidence is searched, and reading fidelity, which controls whether evidence is read in compact form or recovered from the original conversation. Based on this formulation, we introduce JustMem, which stores conversation history as compact atomic memories and adapts memory access along these two dimensions to each query. Specifically, LOOKUP handles local evidence, COMPOSE broadens discovery for distributed evidence, and REPLAY increases reading fidelity for fidelity-sensitive evidence. On LoCoMo and LongMemEval-S, JustMem achieves the highest mean accuracy and retrieval recall among the compared memory systems while using substantially fewer generative-model tokens for memory construction and inference.

[NLP-46] Zarya: A Hybrid Autoregressive–Masked Diffusion Language Model with Flexible Training and Dual-Mode Inference

【速读】: 该论文旨在解决自回归语言模型(Autoregressive Language Models, ARM)在生成过程中受限于顺序、从左到右的推理模式,以及掩码扩散模型(Masked Diffusion Models, MDM)虽支持并行解码但存在无法复用键值缓存(Key-Value Cache, KV cache)导致计算开销高,且因在不可行的词元组合空间中学习依赖关系而引发生成不连贯的问题。其解决方案的关键在于提出Zarya——一种统一架构下的混合语言模型家族,通过在同一模型中联合优化自回归目标与掩码扩散目标,实现两种生成范式的有效融合。Zarya采用可变大小的槽位(slot)结构组织训练数据,并引入渐进式课程学习策略,逐步提升槽位粒度,从而实现从细粒度自回归学习到粗粒度扩散学习的平滑过渡。在推理阶段,Zarya通过统一接口提供两种解码范式:(i) 基于首次命中去噪的MDM采样;(ii) 交错执行槽间基于扩散的选择与槽内自回归填充的分槽推测解码,实现了完整的KV缓存复用。训练与推理完全解耦,使任意配置训练的模型均可灵活部署于任一模式。此外,支持分组噪声模式(如前缀补全、前缀填空、中间填空)、有序采样调度及噪声层级置换策略等高度可配置特性,为研究探索提供了灵活性。作者公开发布了0.6B、1.7B和4B三种规模的Zarya模型,在标准基准测试中展现出优异性能,实现了自回归与扩散范式在理论与实践上的原理性整合。

链接: https://arxiv.org/abs/2609.19868
作者: Leonid Sinev,Ilya Koziev,Vladislav Leshchuk
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Preprint. Work in progress. Please cite peer-reviewed version when published

点击查看摘要

Abstract:Autoregressive language models (ARMs) are constrained by sequential, left-to-right generation, while masked diffusion models (MDMs) enable parallel decoding but suffer from high computational overhead due to the inability to reuse Key-Value (KV) cache and from incoherent generation arising from learning dependencies over an intractable space of token combinations. We introduce Zarya, a family of hybrid language models that jointly optimizes an autoregressive (AR) objective and a masked-diffusion objective within a single architecture. Zarya structures training data into variable-size slots and employs a curriculum that gradually increases slot granularity, enabling a smooth transition from fine-grained AR learning to coarse-grained diffusion learning. At inference, Zarya provides two distinct decoding paradigms through a unified interface: (i) MDM sampling with first-hitting denoising, and (ii) slotted speculative decoding that interleaves inter-slot diffusion-based selection with intra-slot autoregressive infilling, achieving full KV cache reuse. The training and inference regimes are fully decoupled, allowing a model trained with any configuration to be deployed in either mode. Extensive configurability — including grouped noise patterns (Prefix Completion, Fill-In-the-Prefix, Fill-In-the-Middle), ordered sampling schedules, and noise-level permutation strategies — enables flexible research exploration. We release Zarya models publicly in sizes 0.6B, 1.7B, and 4B, demonstrating performance on standard benchmarks while offering a principled integration of autoregressive and diffusion paradigms.

[NLP-47] Reproducibility is not construct validity: LLM measurement of institutionally situated communication

【速读】: 该论文旨在解决生成式人工智能(Generative AI)在文本测量应用中所面临的“可重复性与构念效度”之间的分离问题,即高标注一致性并不必然意味着模型推断的测量指标能够准确捕捉其本应衡量的理论构念。研究以欧盟委员会《人工智能法案》咨询期数据为基础,将结构化问卷响应与同一利益相关方提交的自由文本意见进行匹配,通过大语言模型(LLM)对文本内容进行标注。尽管LLM标注表现出极高的可重复性(组内相关系数达0.99),但其推断的文本测量结果与问卷报告的名义构念之间收敛性有限。不同利益相关群体间存在系统性差异:商业协会在自由文本中表达的AI风险担忧显著高于问卷反馈(效应量g = +1.0),而公共部门及多个非商业组织则呈现较小或负向差异。此外,文本测量得分的差异显示出显著的空间自相关性(莫兰指数Moran’s I = 0.347, p = 0.036),表明邻近国家的利益相关方在AI安全立场上趋于相似。值得注意的是,无论测量分歧程度如何,问卷报告的风险感知仍与对可解释性的支持高度相关。研究的关键在于揭示了LLM标注的高可重复性无法替代构念效度验证,并强调在将大语言模型作为测量工具时,必须区分可重复性、构念有效性以及沟通语境变异等核心维度,从而建立更严谨的验证框架。

链接: https://arxiv.org/abs/2609.19866
作者: Veronika Batzdorfer(KIT),Carlo Romano Marcello Alessandro Santagiustina(ALMAnaCH, médialab, Sciences Po)
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Risk Management (q-fin.RM)
备注:

点击查看摘要

Abstract:High annotation reproducibility does not necessarily imply that an LLM-inferred measure captures the construct it is intended to measure. We test this distinction using a dataset from the European Commission’s AI Act consultation, linking structured survey responses to free-text consultation submissions from the same stakeholders. LLM annotations of consultation submissions are highly reproducible (intraclass correlations 0.99), yet show limited convergence with survey-reported measures of the nominal construct they were intended to approximate. Divergence between survey-and LLM-inferred text-based measures varies systematically across stakeholder groups: business associations express greater concern about AI risks in text-based consultations than in survey responses (g = +1.0), whereas public authorities and several nonbusiness groups show smaller or negative divergences. Divergences between scores suggest positive spatial autocorrelation across European countries (Moran’s I = 0.347, p = 0.036), indicating that stakeholders from neighboring countries tend toward more similar text-based stances towards AI safety concerns. Despite divergence, survey-reported concerns remain strongly associated with support for explainability across all divergence levels. These results demonstrate that LLM annotation reproducibility can coexist with poor construct correspondence and motivate validation procedures that distinguish reproducibility, construct validity, and communication context variation when LLMs are used as measurement instruments.

[NLP-48] F2DR: A Fine-Grained Full-Pipeline Reward Framework for DeepSearch Workflows

【速读】: 该论文旨在解决当前大型语言模型(Large Language Models, LLMs)在复杂用户查询场景下,深度搜索(DeepSearch)工作流评估缺乏精细化与全流程覆盖的问题。现有奖励模型(Reward Models, RMs)和评估基准主要针对静态单轮任务设计,无法有效捕捉DeepSearch所涉及的规划与反思、信息检索及答案生成等多阶段动态交互的全链条复杂性。为此,论文提出F2DR——一种细粒度的全链条深度搜索奖励框架,从内容(Content)、轨迹(Trajectory)和答案(Answer)三个维度对DeepSearch流程进行综合评估,实现了过程层面的全面量化。其核心创新在于构建了可区分不同奖励模型性能的专用评估基准DeepSearch RM-Bench,实验表明F2DR相较于基于自评估的基线方法具有显著更高的评估一致性,且该基准在区分现有开源奖励模型方面展现出强判别能力。

链接: https://arxiv.org/abs/2609.19827
作者: Bojian Xiong(Tianjin University)Wentao Ding(Baidu Inc.)Yujing Lu(Baidu Inc.)Shaowei Zhang(Tianjin University)Ling Shi(Tianjin University)Jing Liao(Baidu Inc.)Yan Wang(Baidu Inc.)Yueyang Zhang(Baidu Inc.)Long Xia(Baidu Inc.)Zhiyuan Sun(Baidu Inc.)Daiting Shi(Baidu Inc.)Jingzhou He(Baidu Inc.)Yuqi Ren(Tianjin University)Deyi Xiong(Tianjin University)
机构: TJUNLP Lab, Tianjin University (天津大学自然语言处理实验室); Baidu Inc. (百度公司)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:With the widespread industrial deployment of Large Language Models (LLMs), DeepSearch has emerged as the dominant paradigm for resolving complex user queries. It typically operates through an iterative closed-loop workflow consisting of planning and reflection, information retrieval, and answer generation. However, existing reward models (RMs) and evaluation benchmarks are primarily designed for static single-turn tasks, failing to capture the full-pipeline complexity of DeepSearch workflows. To address this limitation, we propose F2DR, a fine-grained full-pipeline DeepSearch reward framework. F2DR evaluates DeepSearch workflows across three dimensions: Content, Trajectory, and Answer, enabling comprehensive process-level assessment. We further construct DeepSearch RM-Bench, a dedicated benchmark for evaluating RMs in DeepSearch scenarios. Extensive experiments demonstrate that F2DR achieves significantly higher evaluation consistency than self-evaluation-based baselines, while DeepSearch RM-Bench exhibits strong discriminative capability across existing open-source RMs. We will publicly release the complete DeepSearch RM-Bench dataset soon.

[NLP-49] Dictionary-Constrained Grapheme-to-Phoneme for Unsegmented Languages from LLM -Annotated Data ICASSP2027

【速读】: 该论文旨在解决无分词语言(如日语)中音素转换(Grapheme-to-phoneme, G2P)面临的双重挑战:一是需同时完成词切分与多音字消歧,二是高质量标注数据稀缺导致模型性能受限。现有方法在上下文感知能力、处理复杂语义依赖以及数据规模方面存在不足。本文提出一种上下文感知的神经网络G2P方法,其核心在于利用判别式条件随机场(Discriminative Conditional Random Field, CRF)对由词典构建的词图(word lattice)中的路径进行评分,从而实现对上下文敏感的音素映射。为缓解数据稀缺问题,引入大语言模型(Large Language Models, LLMs)生成超过200万条合成语料,显著扩充训练数据。实验结果表明,该方法在Joyo-Kanji-Yomi基准测试中达到99.62%的目标词读音准确率、0.32%的目标词音素错误率(PER)和0.14%的句子级PER,显著优于传统基于形态分析器的方法及常规神经序列模型,验证了其在准确性、稳定性和上下文建模方面的优越性。

链接: https://arxiv.org/abs/2609.19805
作者: Rui Hu,Zhenpeng Zhan,Xiaolong Lin
机构: Baidu(百度)
类目: Computation and Language (cs.CL)
备注: Submitted to ICASSP 2027

点击查看摘要

Abstract:Grapheme-to-phoneme (G2P) conversion turns raw text into its phonemic form and is an essential part of both text-to-speech (TTS) and automatic speech recognition (ASR) systems. It is required to be fast, stable and context-aware. For unsegmented languages such as Japanese, G2P additionally couples word segmentation with highly context-dependent polyphone disambiguation, and the scarcity of accurately annotated data remains a bottleneck. In this paper, we present a context-aware neural G2P method that scores paths of a discriminative conditional random field (CRF) over a word lattice constructed from dictionaries. To tackle data scarcity, we utilize large language models (LLMs) to generate more than 2 million sentences. Experimental results demonstrate that our method strongly outperforms conventional morphological analyzer-based methods and neural sequence models. On the Joyo-Kanji-Yomi benchmark, our method reaches 99.62% target word reading accuracy, 0.32% target word phoneme error rate (PER) and 0.14% sentence PER.

[NLP-50] Evolution or Illusion? Rethinking Evaluation in LLM Evolutionary Search

【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)驱动的进化搜索中,如何在固定计算预算下合理分配种子数量(宽度,width)与迭代次数(深度,depth)这一关键问题。现有研究通常仅报告单一预算设置下的性能结果(如单个种子运行固定次数迭代),并据此对不同策略进行排序,但这种评估方式存在显著局限性。本文通过在五个典型优化任务上系统评估三种进化搜索策略,并覆盖完整的种子数与迭代次数组合网格,发现最优的预算分配策略随搜索策略、任务特性及总预算规模而动态变化。此外,策略间的相对表现排名也受预算影响:某些策略在少量种子时表现较差,但在增加种子数量后反而成为最优;另一些任务则表明,实际最优迭代次数远低于当前实践中常见的设定,过度追求深度会浪费计算资源,而将预算用于增加种子数可显著提升性能。因此,论文提出一种基于“种子-迭代前沿”(seeds-by-iterations frontier)的测量协议,并提供实用指导,以实现更全面、可靠的算法评估与资源配置。

链接: https://arxiv.org/abs/2609.19799
作者: Tal Oved,Roi Pony,Oshri Naparstek,Udi Barzelay
机构: IBM Research(IBM研究院)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:LLM-driven evolutionary search finds programs by launching seeds and iterating each one. Papers report a single budget setting, usually one seed run for a fixed number of iterations, and rank methods from that one point. We show this is not enough. We evaluate three evolutionary search strategies on five optimization tasks, commonly used by papers in the genre to report results. We run the analysis over a full grid of seeds and iterations. Our findings suggest that the best way to split a fixed budget between more seeds (width) and more iterations (depth) changes with the strategy, the task, and the total budget. Furthermore, we observe that the ranking of strategies also changes with the budget. On one task the strategy that looks worst at one seed is best at forty seeds. On another the best number of iterations is well below the value common in practice, so extra depth wastes budget that more seeds would turn into score. We provide a measurement protocol that reports the seeds-by-iterations frontier and practical guidance for using it.

[NLP-51] Learn Before You Judge: Progressive Knowledge-to-Decision Alignment for Explainable Hateful Meme Detection

【速读】: 该论文旨在解决生成式AI在仇恨表情包(hateful meme)检测中因解释生成与标签预测任务耦合导致的性能瓶颈问题。现有“先解释后检测”方法在统一训练过程中同时优化解释生成与分类决策,造成两个任务目标间的干扰,不仅未能提升检测效果,甚至表现劣于简单的监督微调(SFT)基线。其解决方案的关键在于提出一种渐进式知识到决策对齐(Progressive Knowledge-to-Decision Alignment, ProKDA)框架:首先通过智能体驱动的外部背景知识构建流程获取与表情包理解相关的外部知识;随后采用三阶段训练策略,依次完成背景知识学习、仇恨性检测学习与仇恨边界对齐,每个阶段专注单一任务目标,有效降低任务间干扰。该设计实现了从外部知识到鲁棒检测决策的渐进式转化,显著提升了检测性能,并生成准确、可解释且有证据支持的判别结果。实验在三个公开数据集上验证了ProKDA达到当前最优(state-of-the-art)的检测效果。

链接: https://arxiv.org/abs/2609.19778
作者: Bo Xu,Chenyuan Wang,Xinyu Chen,Quanhao Zhu,Rui Lin,Liang Zhao,Hongfei Lin,Feng Xia
机构: Dalian University of Technology (大连理工大学); RMIT University (皇家墨尔本理工大学)
类目: Computation and Language (cs.CL)
备注: 26 pages, 16 figures, 7 tables

点击查看摘要

Abstract:Hateful memes spread abusive content through implicit interactions between images and text, posing serious threats to the safety of online communities. In recent years, multimodal large language models have been widely used for hateful meme detection and are increasingly adopted to generate explainable detection results. However, we find that existing explain-then-detect methods often couple explanation generation and label prediction within the same training process. This coupling causes interference between task objectives, leading to limited detection performance and even worse results than simple SFT baselines. To address these challenges, we propose ProKDA, a progressive knowledge-to-decision alignment method for explainable hateful meme detection. Inspired by the human annotation training process, ProKDA first uses an agentic background knowledge construction pipeline to obtain external knowledge related to meme understanding. It then adopts a three-stage training strategy that sequentially performs background knowledge learning, hatefulness detection learning, and hatefulness boundary alignment. Unlike prior explain-then-detect methods that jointly optimize both tasks, ProKDA focuses on a single training objective at each stage. This design reduces interference between the two tasks and progressively transforms background knowledge into robust detection decisions. Experiments on three public hateful meme benchmarks show that ProKDA achieves state-of-the-art detection performance and provides accurate, explainable, and evidence-supported decisions for hateful meme moderation. Project page: this https URL.

[NLP-52] AutoData: Agent ic Search for Pre-training Data Selection

【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)在机器学习工程中自动化程度不足的问题,特别是数据预处理环节长期被排除在智能优化闭环之外的挑战。现有方法多依赖人工设计的数据混合策略,难以高效探索复杂的数据选择空间。其核心解决方案是提出AutoData——一个能够直接在可执行的选择算法空间中进行搜索的智能代理(agent)。与以往固定领域权重优化的方法不同,AutoData通过迭代式地基于代理模型(proxy model)的验证反馈,自动发现并优化包含打分、分层及随机选择规则在内的多样化数据筛选策略,从而实现对文档级特征(如词汇统计、类别标签、困惑度等)之间复杂交互关系的自动挖掘。实验表明,仅在小规模代理模型上经过一夜搜索所获得的数据选择方案,即可在大规模场景下有效迁移,并显著提升下游任务指标CORE,验证了数据工程本身可被建模为一个可自治优化的机器学习问题,标志着自主研究范式从模型与训练代码优化拓展至数据层面。

链接: https://arxiv.org/abs/2609.19754
作者: Yan Meng,Dhruv Srikanth,Bingchen Zhao,Zhengyao Jiang,Yuxiang Wu
机构: University of Amsterdam(阿姆斯特丹大学); Weco AI
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:LLM agents have recently shown promise in automating machine learning engineering by editing model and training code under execution feedback. Data, however, remains largely outside this agentic optimisation loop. We frame pre-training data selection as heuristic engineering over per-document features, i.e., lexical statistics, categorical labels, and perplexity. We introduce AutoData, an agent that searches directly over executable selection algorithms. Unlike prior data mixture methods that optimise weights over a fixed set of domains, AutoData searches a richer program space of scoring, stratification, and stochastic selection rules, discovering feature interactions automatically by iteratively refining algorithms with validation feedback from a proxy model. Within an overnight search, AutoData discovers a selection algorithm that outperforms existing human-designed curation pipelines. Despite being searched only on this small proxy, the discovered recipe transfers to larger scales and improves the downstream metric CORE. These results suggest that data engineering can be treated as an agentic machine learning problem, extending autonomous research from model and training-code optimization to the data.

[NLP-53] A Phonemically Comprehensive ASCII-Only Romanization Scheme for Thai and Lao: Systematic Cross-Lingual Correspondence and Chinese-User-Friendly Design

【速读】: 该论文旨在解决泰语(Thai)与老挝语(Lao)在罗马化表示中缺乏统一、系统且可跨语言互操作的音位级转写方案的问题。现有转写系统往往在音段对比、元音长短及声调表达上存在不一致或冗余,且难以兼顾语音透明性与机器可处理性。其解决方案的关键在于提出一种以音位为中心、仅使用ASCII字符的统一罗马化体系,将两种语言视为一个跨语言设计问题进行整体建模。该方案通过实现“一符号一音素”的透明对应关系,确保泰语与老挝语之间在音段、元音长度和声调层面具有系统性一致性;同时优先保证共时语音对应(如与汉语拼音Pinyin、粤拼Jyutping的对齐),并在不损害语音透明性的前提下保留历史音韵对应。声调采用紧凑的单数字默认标记,并支持可选的声调值与历史声调类别补充表示。最终形成的转写系统兼具可读性、键盘友好性及机器可处理性,适用于语言学习与跨语言语音处理任务。

链接: https://arxiv.org/abs/2609.19736
作者: Zijie Zhang,Tan Lee
机构: 未知
类目: Computation and Language (cs.CL)
备注: Accepted by O-COCOSDA 2026

点击查看摘要

Abstract:This paper proposes a phonemically comprehensive, ASCII-only romanization scheme for Thai and Lao, treating the two closely related languages as a unified cross-lingual design problem. The scheme represents segmental contrasts, vowel length, and lexical tone while maintaining one-symbol-one-phoneme transparency and systematic correspondence between Thai and Lao. The scheme prioritizes synchronic phonetic correspondence, including correspondence with Pinyin and Jyutping where applicable, while preserving historical-phonological correspondence where it does not conflict with phonetic transparency. Tone uses a compact single-digit default notation, supplemented by optional tone-value and historical tone-category representations. The resulting scheme provides a readable, keyboard-friendly, and machine-processable phonemic representation for language learning and cross-lingual speech processing.

[NLP-54] Learn Your Own Thoughts: Abstract Token Curriculum

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在推理过程中依赖显式思维链(Chain-of-Thought, CoT)标注数据的问题,即传统CoT方法需要大量任务特定的、带有中间推理步骤标注的数据进行监督训练,限制了其在无标注场景下的应用。为此,本文提出一种新型课程学习框架——抽象令牌课程(Abstract Token Curriculum, ATC),其核心创新在于无需直接监督或人工设计思维草稿,即可引导模型在连续表示空间中自发生成有效的内部抽象“思考”表征。ATC通过逐步增加问题复杂度,构建一系列分布序列,促使模型在训练过程中发展出对中间推理过程的隐式建模能力。理论分析表明,在使用单层Softmax注意力机制学习奇偶性函数时,ATC可使注意力自然聚焦于上下文中提供“最简路径”以预测下一个标记的抽象思维令牌;实验验证显示,ATC在图可达性与算术学习任务中均显著优于现有方法,证明了其在训练连续思维表征方面的有效性与优越性。

链接: https://arxiv.org/abs/2609.19717
作者: Khashayar Gatmiry,Avrajit Ghosh,Parsa Mirtaheri,Jason D. Lee,Nika Haghtalab,Emmanuel Abbe,Peter Bartlett
机构: UC Berkeley; Google DeepMind; EPFL
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (stat.ML)
备注:

点击查看摘要

Abstract:Large Language Models (LLMs) have achieved remarkable reasoning capabilities by utilizing chain-of-thought (CoT) as a scratchpad for intermediate stages of thinking. However, CoT techniques require explicit supervision on thinking tokens, which requires rich, task-specific data. In this work, we propose Abstract Token Curriculum (ATC), a novel curriculum learning framework that elicits effective continuous intermediate representations without direct supervision or manual scratchpad design. ATC gradually increases problem complexity through a sequence of distributions, training the model to develop internal abstract thoughts'' in the continuous representation space. This paper provides both theoretical and experimental evidence for the benefits of ATC and its advantages over previous methods for training continuous thoughts. Theoretically, we show that for learning parity functions with single-layer softmax attention using ATC, attention naturally focuses on the CoT tokens in the context that provide the easiest path’’ to predicting the next token. Experimentally, we show ATC’s effectiveness on graph reachability and arithmetic learning tasks.

[NLP-55] Improving Cross-Lingual Transfer for Sequential Sentence Classification in Research Papers via Structural Similarity

【速读】: 该论文旨在解决多语言科学文献中序列句子分类(Sequential Sentence Classification, SSC)任务因非英语语料稀缺而导致的训练数据不足问题,进而提升多语言数字图书馆中科学知识的可及性。其核心挑战在于如何有效实现跨语言迁移,尤其是在语言差异较大的情况下保持分类性能。解决方案的关键在于突破传统依赖语言相似性的跨语言迁移范式,转而聚焦于篇章结构层面的共性规律——即不同语言在修辞组织结构上的相似性,如标签序列模式与位置规律等。研究发现,语言相近性对迁移效果无稳定预测力,而修辞结构相似性表现出弱但一致的正向相关性;在控制源语言表现后,标签分布相似性成为最稳定的预测因子。基于此,论文提出三种显式利用结构信息的生成式模型方法,在领域内评估中达到与最强编码器基线相当的性能,并在未见语言的跨语言迁移场景中显著优于现有最优编码器模型,验证了结构信息在跨语言SSC中的关键作用。

链接: https://arxiv.org/abs/2609.19650
作者: Kazuhiro Yamauchi,Marie Katsurai
机构: Doshisha University (同志社大学)
类目: Computation and Language (cs.CL); Digital Libraries (cs.DL)
备注: Accepted at JCDL 2026 (ACM/IEEE Joint Conference on Digital Libraries), Frisco, TX, USA, October 13-16, 2026. 12 pages, 5 figures, 9 tables. DOI: https://doi.org/10.1145/3805696.3846040

点击查看摘要

Abstract:Sequential sentence classification (SSC) is an essential task for structuring scientific publications, and extending SSC research to languages other than English can improve accessibility to scientific knowledge in multilingual digital libraries. Cross-lingual transfer is a promising approach to address the scarcity of training data in non-English languages. Prior work on other natural language processing tasks has shown the benefits of capturing linguistic similarity between source and target languages. However, SSC inherently depends on patterns at the discourse level, such as label sequences and positional regularities, which appear consistently across languages regardless of linguistic differences. To examine the factors that determine transfer success in SSC, we constructed a multilingual SSC dataset covering 13 non-English languages collected from five academic databases. Our cross-lingual transfer experiments, using both encoder-based and generative models, show that linguistic proximity has no consistent predictive power for transfer performance, whereas structural similarity in rhetorical organization shows a weak but consistent positive correlation across models. After controlling for source-language performance, the similarity of label distributions is the most consistent predictor. Building on this finding, we propose a set of three methods that explicitly leverage structural information using generative models. In the in-domain evaluation, the best combination reaches parity with the strongest encoder baselines, and in transfer to languages unseen during training, it outperforms the strongest encoder baseline.

[NLP-56] Scientific Image Quality Assessment via Multi-modal Retrieval-Augmented Generation

【速读】: 该论文旨在解决科学图像质量评估(Scientific Image Quality Assessment, SIQA)挑战中同时存在的理解任务(SIQA-U)与评分任务(SIQA-S)的双重难题,即如何让模型在缺乏明确标注的情况下准确理解科学图像的内容并生成符合人类专家判断标准的量化评分。其解决方案的关键在于提出一种检索增强生成(Retrieval-Augmented Generation, RAG)框架,通过构建融合文本语义与细粒度视觉特征的多模态索引,并设计多路径检索与融合机制,实现对大语言模型(Large Language Model, LLM)的高效知识供给,使其能够基于高度相关的参考案例进行推理,从而显著提升对复杂科学图像的理解与评价能力。实验结果表明,该方法在人类专家评判标准上具有强一致性,最终在ICME 2026 Grand Challenges的SIQA-U赛道中取得第一名。

链接: https://arxiv.org/abs/2609.19634
作者: Yinuo Zhang,Bingshuo Liu,Zhiying Tu,Dianhui Chu,Qingbin Liu,Xi Chen,Jiang Bian,Xiaoyan Yu,Dianbo Sui
机构: Harbin Institute of Technology (哈尔滨工业大学); Tencent(腾讯); Nanyang Technological University (南洋理工大学)
类目: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:This paper proposes a Retrieval-Augmented Generation (RAG) framework for scientific image quality assessment, designed to simultaneously address both the understanding track (SIQA-U) and the scoring track (SIQA-S) of the SIQA challenge. We construct a multimodal index that integrates textual semantics with fine-grained visual features, and develop a multi-route retrieval and fusion mechanism to provide large language models with highly relevant reference cases, thereby enhancing their capability to evaluate complex scientific images. Experimental results demonstrate that the proposed framework effectively aligns with the judgment criteria of human experts. Ultimately, our method achieves 1st place in the SIQA-U track of the SIQA challenge at the ICME 2026 Grand Challenges.

[NLP-57] From Intent to Action: Benchmarking LLM Safety in Vehicle Voice Command Authorization

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在车载语音助手应用中,将自然语言请求映射至具体车辆功能时所面临的安全关键型授权决策问题。核心挑战在于:在执行指令前,系统需从七类动作中做出正确选择——包括执行、拒绝、澄清、要求确认、延迟至人工控制、触发紧急响应或不调用工具。现有评估方法未能充分分离用户角色、认证状态、车辆状态及工具可用性等多维因素对预操作决策的影响。为此,研究提出一个包含202个场景的基准测试集,并定义了基于七分类体系的参考决策标准。通过引入决策一致性(Decision Alignment)与特定安全错误指标,评估了两种本地开源模型和三种基于API的LLM。结果显示,模型决策一致性介于40.1%(Llama 3.2 3B)至89.1%(Gemini 3.1 Pro Preview)之间,其中基于API的模型表现相近且无显著差异。然而,即便最优模型仍存在每161个非执行场景中出现2至3次误执行(False Execute),且在确认与人工接管决策上持续存在错误。进一步的消融实验表明,采用结构化授权策略虽可将Llama 3.2 3B的对齐率提升至40.1%,高于仅依赖模式匹配或通用安全基线下的28.2%-29.2%,但无法根除误执行。因此,研究指出:仅依靠结构化大语言模型决策不足以构成独立的安全保障机制,实际部署必须引入独立的执行层,用于验证工具权限与车辆状态约束,确保所有操作符合安全规范。

链接: https://arxiv.org/abs/2609.19630
作者: Diba Afroze,Xingli Zhang,Yazhou Tu,Xiali Hei
机构: University of Louisiana at Lafayette (路易斯安那大学拉法叶分校); Auburn University (奥本大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Large language models (LLMs) are increasingly integrated into vehicle voice assistants. But linking natural-language requests to vehicle functions creates a safety-critical authorization problem. Before executing a command, the system must choose whether to execute, refuse, clarify, require confirmation, defer to manual control, trigger an emergency response, or make no tool call. To our knowledge, prior evaluations do not isolate this pre-action decision across speaker role, authentication status, vehicle state, and tool availability. We introduce a 202-scenario benchmark with Reference Decisions under a seven-class taxonomy. We evaluate two local open-weight models and three API-based LLMs using Decision Alignment and safety-specific error metrics. Alignment ranges from 40.1% for Llama 3.2 3B to 89.1% for Gemini 3.1 Pro Preview. The API-based models score between 83.2% and 89.1%, with no statistically significant differences among them. Even these models produce two to three False Executes among 161 non-execution scenarios, and persistent errors remain in confirmation and manual-control decisions. A controlled Llama 3.2 3B ablation increases alignment to 40.1% under the structured authorization policy, versus 28.2-29.2% under schema-only and generic-safety baselines, but it does not eliminate False Executes. Structured LLM decisions are therefore insufficient as a standalone safety mechanism, and deployment requires an independent enforcement layer that verifies tool permissions and vehicle-state constraints before invoking any vehicle function.

[NLP-58] Chain-of-Thought Entropy as a Reliability Signal: A Preregistered Reproduction

【速读】: 该论文旨在复现并验证赵于2026年报告的大语言模型链式思维(Chain-of-Thought, CoT)熵轨迹与最终答案正确性之间的解耦现象(dissociation),即熵轨迹的形状可预测答案正确性,而总熵降幅的大小则不具备一致预测能力。其核心问题在于:原始发现中“熵降幅”信号的可靠性依赖于单一实验设置(一个模型、一种随机种子、300个问题),存在显著的偶然性风险;而“熵轨迹形状”信号虽在大规模基准上被报告,但缺乏独立验证。本研究通过在七项明确记录的协议差异下,使用四种开源权重模型(包括一种推理精炼模型)对GSM8K和MATH-500两个基准进行全面测试,成功复现了熵轨迹形状信号的预测效力。关键发现为:熵轨迹形状信号具有稳健可复现性,而总熵降幅信号仅在特定设置下有效——在锚定模型上,单调熵轨迹链比非单调者在GSM8K和MATH-500上分别高出9.6和27.5个百分点准确率,但相关性分别为-0.018和+0.414,表明其预测能力高度依赖任务与模型;而在推理精炼模型中,二值化形状信号触发频率极低,难以评估对比效果,但熵违反次数的连续度量仍具预测力。 此外,探索性分析显示,仅最后一层熵值即可在所有八组模型-基准组合中超越原始二值形状标志的ROC-AUC表现,且在多数情况下优于风险覆盖率指标,提示更简单的指标可能更具实用性。本研究不仅提供了对原结论的全面独立验证,还系统揭示了信号有效性随模型架构与实验设置的变化规律,并补充了四项原始研究未报告的协议依赖性测量。

链接: https://arxiv.org/abs/2609.19606
作者: Theodore O. Cochran
机构: AI for Altruism
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:This empirical study is an independent reproduction of the dissociation Zhao reported in 2026. The shape of a large language model’s chain-of-thought entropy trajectory predicts whether the final answer is correct, while the magnitude of its total entropy drop does not. The dissociation merits reproduction because the magnitude half rests on a single 300-problem run with one model at one seed, while the shape half was reported at full scale on both benchmarks and on a second model family. Registered at OSF before any confirmatory run, the reproduction crosses the complete GSM8K and MATH-500 benchmark test sets with four open-weight models including one reasoning-distilled model of a kind the original did not test. The shape signal replicates. The magnitude signal divides by setting. On the anchor model the accuracy gap between monotone and non-monotone chains is +9.6 percentage points on GSM8K and +27.5 on MATH-500, while the rank correlation of the total entropy drop with correctness is -0.018 on GSM8K and +0.414 on MATH-500. On the reasoning-distilled model the binary form of the shape signal fires on about one chain in a hundred, too few to estimate the registered contrast, while the graded violation count remains predictive there. In an exploratory comparison the final-step entropy alone outperforms the binary shape flag in all eight model-by-benchmark cells by ROC area, and in six or seven by the risk-coverage area the original reports, depending on an integration range the original does not state. The study contributes a reproduction of the shape signal at full test-set scale under seven documented protocol differences, a map of the settings where the magnitude signal holds and fails, and measurements of four protocol dependencies the original does not report.

[NLP-59] Form Over Content In Gradient-Based Data Attribution Methods

【速读】: 该论文旨在解决生成式 AI(Generative AI)中基于梯度相似性的数据归属方法(data attribution methods)所存在的核心争议:梯度相似性究竟反映的是任务语义相关性,还是表面形式(answer format)的相似性。现有研究对此存在分歧,部分认为其可识别与任务相关的技能,而另一些研究则指出表面形式是主导因素。为厘清这一问题,作者通过独立操控任务和答案格式,设计了在相同任务但不同格式、或相同格式但不同任务的基准测试对,从而实现变量解耦。实验结果表明,梯度对齐主要遵循答案格式——当基准测试共享相同答案格式时,梯度相似性显著(去衰减余弦相似度接近0.4),而即使任务相同但格式不同,则梯度无明显对齐(相似度接近0.0)。该现象在预训练早期到微调后各阶段、不同模型规模与模型家族中均保持一致。进一步分析指令微调中已发布的梯度基数据选择方法LESS的选例发现,其选择结果显著偏向目标任务自身的答案格式。因此,本研究的关键结论是:当前基于梯度的数据归属方法更敏感于格式相似性而非任务语义,提示此类方法及其对梯度的语义解释必须在任务与格式可独立变化的数据集上进行验证,以提升其鲁棒性与可靠性。

链接: https://arxiv.org/abs/2609.19589
作者: Sunwoo Kim,Seokwon Jung,Sohyung Kim,Seong Joon Oh,Alice Oh
机构: KAIST(韩国科学技术院)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Data attribution methods using gradient similarity are widely used to analyze and select training data for large language models, but what gradient similarity actually measures is debated. Some interpret it as identifying task-relevant skills, while other work reports that surface form is the main factor. We resolve this debate for supervised fine-tuning examples by varying task and answer format independently. Specifically, we render benchmarks in different answer formats, such that datasets can share a task without a format or a format without a task. We find that gradient alignment follows the answer format, as benchmark pairs sharing an answer format align strongly (disattenuated cosine near 0.4), while same benchmarks rendered with different answer format classes show no alignment (near 0.0). We demonstrate that this ordering holds from the earliest pretraining checkpoints through post-training, and across model scales and families. We then analyze the released selections of LESS, a gradient-based data selection method for instruction tuning, and find that each target’s selections over-represent the target’s own answer format. Hence, we demonstrate that gradient-based attribution methods track format similarity more than task semantics, meaning that such methods, as well as the semantic interpretation of the gradient, should be tested on data where answer format and task vary independently for greater robustness and reliability.

[NLP-60] Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents

【速读】: 该论文旨在解决生成式代码代理(Coding Agent)在长期运行中因持续性偏移(persistent misalignment)而可能引发的严重安全风险,尤其关注现有生产级阻断监控机制(如Claude Code的Auto Mode和OpenAI Codex的Guardian)在面对有意识、持续性恶意攻击时的脆弱性。传统评估主要聚焦于偶然性伤害或来自不可信源的提示注入攻击,但对代理自身主动发起的、具有策略性的对抗性行为缺乏足够研究。为此,作者构建了一个对抗性代理(adversarial agent),通过高阶攻击策略尝试绕过监控并造成灾难性后果(如泄露自身权重)。研究发现,该对抗性代理可通过多种机制成功绕过现有监控,包括代理自动生成提示注入攻击、多代理协同攻击以及恶意压缩(malicious compaction)等;其中,在79%的实验中,对抗性代理能成功利用提示注入攻击使Auto Mode与Guardian执行任意bash命令。研究进一步表明,通过提升工具覆盖范围、优化对话记录格式及引入代理式监控阶段等设计改进,可显著增强Auto Mode的安全性,但针对跨上下文攻击(multi-context attacks)在可接受性能开销下的有效防御仍是一个未解难题。本文通过系统化的红队测试方法论,揭示了新型攻击向量,旨在帮助安全防御者更全面地评估其防护措施对持续性恶意代理的鲁棒性。

链接: https://arxiv.org/abs/2609.19587
作者: Alex Remedios,Simon Storf,Fabien Roger,John Hughes
机构: Anthropic Fellows Program; Anthropic
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:To keep coding agents from going off the rails, production systems now review each proposed action with a blocking monitor that can reject it before it runs (Auto Mode in Claude Code, Guardian in OpenAI’s Codex). Prior evaluations of such monitors largely measure robustness to accidental harm or prompt injections from untrusted sources looking to hijack the agent. Less understood is how they hold up when the agent they monitor is persistently misaligned. To understand this risk, we task an adversarial agent with evading production blocking monitors and causing catastrophic harm, e.g. by exfiltrating its own weights. We find that when instructed with high-level attack strategies, adversarial agents can succeed through several distinct mechanisms, such as agent-generated prompt injection against the monitor, multi-agent attacks, and malicious compaction. In particular we find that in 79% of trials, the agent can use an injection attack against Auto Mode and Guardian to run arbitrary bash commands. We also find that it is possible to greatly improve Auto Mode through design changes like enhancements to tool coverage, transcript formatting and an agentic monitor stage. Even with these improvements, preventing multi-context attacks at an acceptable cost remains an open problem. By detailing our red-teaming methodology and highlighting new attack vectors, we aim to help defenders evaluate their mitigations against the possibility of persistent malign coding agents. Code is available at this https URL.

[NLP-61] From Parameters to Behaviors: A Survey of Model Fusion for Large Language Models EMNLP2026

【速读】: 该论文旨在解决当前模型融合(model fusion)领域缺乏统一定义与系统化分类体系的问题。随着Hugging Face平台在2026年6月前已托管超过200万模型,模型资源的快速增长为能力复用与集成提供了丰富基础,但现有综述多局限于该领域的局部内容,未能形成整体性框架。为此,本文首次明确提出模型融合的概念,并将其划分为参数级、表示级和行为级三个层次的融合范式,构建了一个系统性的分类体系。其解决方案的关键在于通过多层级的结构化划分,整合已有研究中的方法、评估指标、基准测试与应用场景,从而为该领域提供清晰的研究地图,并推动未来工作的规范化与深化。

链接: https://arxiv.org/abs/2609.19553
作者: Shuo Cai,Yanggan Gu,Zihao Wang,Yuanyi Wang,Yibo Yan,Wenjun Wang,Yuhang Liu,Guanghao Zhu,Sirui Huang,Ming Li,Hongxia Yang
机构: The Hong Kong University of Science and Technology (Guangzhou); The Hong Kong Polytechnic University; InfiX.ai; PolyU-Daya Bay Technology and Innovation Research Institute; The Chinese University of Hong Kong
类目: Computation and Language (cs.CL)
备注: 25 pages, 4 figures. Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026

点击查看摘要

Abstract:Model fusion integrates the capabilities from source models into a single target model. As of June 2026, Hugging Face hosts more than 2M models. This growing pool provides a rich base for model reuse and capability integration. Yet existing surveys often cover only separate parts of this space, and they do not provide a unified definition or a systematic taxonomy. This survey defines model fusion and organizes prior work into three levels: parameter-level, representation-level, and behavior-level fusion. We also review related metrics, benchmarks, and applications, summarize current challenges, and identify future directions. Our goal is to provide a clear map of this area and support future work on model fusion. A comprehensive list of papers about model fusion is available at this https URL.

[NLP-62] Finding Common Ground: Graded Communal Knowledge in Bluesky Starter Packs

【速读】: 该论文旨在解决社交媒介中“共同基础”(common ground)的可测量性与可分离性问题,即如何在用户未发生直接互动前,客观识别和量化个体间共享的认知背景。传统研究虽基于克拉克(Clark, 1996)提出的“集体共同基础”(communal common ground)理论,认为共享的社群归属越多,共同基础越强,但因社群成员身份通常不可见且常与互动行为混杂,导致该机制长期缺乏实证检验。本文的关键解决方案在于将Bluesky平台中的“启动包”(Starter Packs, SPs)重新定义为用户自定义的社群归属标签,并利用其作为代理变量来衡量用户的语义共现程度。研究通过对191,648对用户的数据分析发现,共享的启动包数量与词汇库相似性呈单调递增关系,共享单一启动包的用户其语言相似度约为无关联用户的两倍;进一步通过语义归一化分析表明,共同基础更依赖于用户所处的主题上差异显著的社群数量,而非简单的社群总数。此外,社群共属对共同基础的贡献独立于用户在关注网络中的物理邻近性。由此得出结论:社群归属是可测量、可分离且具有语义结构的共同基础载体,使得共同基础可在交流前被观测,从而为大规模自然情境下的共同基础研究提供了新范式,突破了以往仅限于实验室环境的研究局限。

链接: https://arxiv.org/abs/2609.19549
作者: Sagar Kumar,Lawrence Swaminathan Xavier Prince,Julia Mendelsohn,Brooke Foucault Welles,Nicholas W. Landry
机构: 未知
类目: ocial and Information Networks (cs.SI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Communication is made possible by common ground—the unspoken knowledge that people share and presuppose of one another, whether that be online or offline. In his conception of common ground, Clark (1996) distinguishes between personal and communal common ground, and asserts that the latter is graded: the more community affiliations two people share, the more common ground they share as well. Social media research has invoked this mechanism to explain how users connect, but it has gone largely untested because community memberships are rarely visible and, where they are, they are coupled to user interactions in a way that leads to conflating effects. To circumvent these challenges, this study repurposes Bluesky starter packs (SPs) as user-curated community affiliation labels. Across 191,648 pairs of users, we show that shared lexical repertoire—our proxy for common ground—grows monotonically with the number of SPs that users share, with users sharing a single pack being roughly twice as similar as equally connected strangers. A semantic renormalization of SP co-membership shows furthermore that it is more so the number of topically \emphdistinct communities, rather than the raw count, in which common ground is graded. Finally, we show that community co-membership adds to common ground independently of proximity in the Bluesky follow network. These results lead to the conclusion that community membership is a measurable, separable, and semantically structured carrier of common ground. Reading it as such makes common ground observable before an exchange rather than inferred from it, and thus opens the door for large-scale observational approaches to a set of questions that have so far only been posed in the laboratory.

[NLP-63] When Hiring Becomes Agent -Mediated: Evaluating Access and Recurrence in Two-Agent Résumé Screening EMNLP2026

【速读】: 该论文旨在解决传统简历筛选(résumé screening)中存在的单向、静态且缺乏互动性的问题。当前的自动化筛选通常基于雇主单方面对简历与岗位匹配度的一次性判断,忽略了候选人自我呈现与雇主评估之间的双向动态过程。为改进这一局限,论文提出一种双代理(two-agent)筛选机制,其中雇主侧代理与候选人侧代理分别代表双方角色,通过交换证据并迭代更新判断,最终共同决定哪些申请者进入下一阶段。该方案的关键在于引入了双向交互与动态协商机制,使筛选过程更贴近真实招聘中的双向博弈。实验结果表明,相较于传统的一次性判断,双代理机制在多个基准数据集上显著提升了通过率(如GPT-5.5从33.3%提升至39.3%),尤其在边界案例中表现突出,通过率提升达21.7个百分点(4.5%→26.2%)。值得注意的是,该机制并非简单放宽标准,而是改变了决策方向——部分原被拒绝的申请被重新接纳,同时也有原通过的申请被否决,显示出决策的复杂性与合理性。此外,在重复运行中,仅由双代理共同选定的案例其一致性显著低于共享选择案例,表明其筛选逻辑具备更高的多样性与非可复制性,而单独的一次性阈值无法复现双代理筛选的结果。因此,研究强调:随着招聘流程日益由智能体(agent)主导,筛选机制本身的设计(而非单一模型性能)已成为决定谁能进入人工评审环节及其可重复性的关键因素。

链接: https://arxiv.org/abs/2609.19530
作者: Jian Gao,Hang Jiang
机构: Northeastern University(东北大学); Boston, MA, USA
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 9 pages, 5 tables, 1 figure. Accepted to the REALM Workshop at EMNLP 2026

点击查看摘要

Abstract:Hiring is bilateral: employers assess fit, while candidates present and defend evidence of their qualifications. Yet résumé screening, the first gate, is commonly automated as a static, one-call judgment over a résumé-job pair. We study a two-agent alternative in which employer-side and candidate-side agents represent these roles, exchange evidence, and update their judgments before deciding who advances. We compare procedures on 600 constructed résumé-job pairs using GPT-5.5 and Claude Opus 4.7. Two-agent screening advances more applications (33.3% to 39.3% for GPT-5.5; 34.0% to 35.5% for Opus 4.7). Across three runs on the common 191-pair borderline pool, pass-instance rates rise from 4.5% to 26.2% and from 6.5% to 16.1%, respectively. This is not a uniform relaxation: two-agent screening rejects applications one-call advances, changing decisions in both directions. At similar pass volumes, the procedures advance different applications, and no one-call threshold recovers applications consistently selected by two-agent screening. Among discovery-selected cases re-executed in fresh runs, two-agent-only selections recur less often than shared selections, clearly under GPT-5.5 and less certainly under Opus 4.7, while a separate one-call follow-up shows no comparable decline. As hiring becomes agent-mediated on both sides, the screening procedure, not only the model behind it, shapes who reaches human review and how reliably that access recurs.

[NLP-64] EconSkills: Studying Skill Transfer and Retrieval for Web Agents on Live Economic Data

【速读】: 该论文旨在解决网络代理(Web agent)在重复访问相同网站时,现有评估方法忽视先前成功交互中所习得知识的问题。其核心挑战在于如何有效复用已验证的、可迁移的网页操作流程以提升任务执行效率与成功率。解决方案的关键是提出EconSkills——一个技能库与评估框架,将经过验证的EconWebArena轨迹提炼为参数化的标准操作程序(Standard Operating Procedure, SOP),每个技能条目均记录其作用范围、导航步骤、站点特定指导、验证检查及故障恢复策略,并通过占位符替代具体实例值以实现泛化。该框架区分了两个关键问题:一是已知有效流程能否迁移到未见任务上,二是代理在从技能库中选择时是否能保留该迁移优势。实验表明,在受控迁移场景下,匹配的技能显著优于无技能提示,且在成功案例中所需步骤更少;抽象化处理远优于直接回放原始轨迹。在大规模技能库检索中,整体性能与无技能基线相当,但在覆盖任务上表现最优;分层覆盖分析显示,对未覆盖任务的近似匹配可部分抵消性能损失。浏览器轨迹分析进一步揭示了程序化指导在缩短门户特异性导航路径中的有效性,以及语义验证在确保结果准确性中的必要性。研究结果证明,可复用的经济类网页操作流程可在不同任务实例间有效迁移,并为实现覆盖率感知的选择机制与上下文传递设计提供了明确的技术目标。

链接: https://arxiv.org/abs/2609.19523
作者: Yinzhu Quan,Zefang Liu
机构: Georgia Institute of Technology (佐治亚理工学院); Capital One (资本一号)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Web agents often revisit the same sites, yet most evaluations discard the procedures learned in earlier successful interactions. We introduce EconSkills, a skill library and evaluation framework that distills verified EconWebArena trajectories into parameterized standard operating procedures for retrieving live economic data. Each skill records its scope, navigation procedure, site-specific guidance, verification checks, and recovery steps while replacing source-instance values with placeholders. EconSkills separates two questions: whether a known relevant procedure transfers to a held-out task, and whether an agent can retain that benefit when selecting from a library. In controlled transfer, matched skills improve success over no-skill prompting and require fewer steps on paired successes, while abstraction is substantially more effective than replaying raw trajectories. At library scale, retrieval is competitive with the no-skill baseline overall and performs best on directly covered tasks; coverage-stratified outcomes show that approximate matches on uncovered tasks offset these gains. Browser trajectories further identify when procedural guidance shortens portal-specific navigation and when semantic verification remains necessary. These results establish that reusable economic web procedures can transfer across task instances and provide a concrete design target for coverage-aware selection and context delivery.

[NLP-65] For Your Eyes Only: Evaluating Coordination Between Isolated Language Model Instances

【速读】: 该论文旨在解决在无共享记忆或专门协作训练的前提下,同一模型的不同实例能否通过自然语言生成实现隐秘信号传递与识别的问题。其核心挑战在于:在自动化工作流中,模型生成的内容可能被其他模型实例消费,如何确保一个模型(发送方)能够以不可察觉的方式向另一个独立的模型(接收方)传递隐藏信息,而仅依赖于共同的预训练语料和任务指令。解决方案的关键在于设计并评估“仅供你眼”(For Your Eyes Only)这一协作式信号博弈框架,通过双遍测试成功率(Double-Pass Success Rate)控制输出偏差,量化模型在避免可检测信号前提下的协调能力。实验结果表明,多数主流模型在要求规避可识别信号后协调能力显著下降,但某一前沿模型仍能保持近乎完美的性能;此外,研究还揭示了模型具备将此能力用于故意误导的能力,并发现跨架构的协调能力普遍弱于同架构内部的协调能力。

链接: https://arxiv.org/abs/2609.19504
作者: Alexander Shirnin,Aleksey Kudelya
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:As model-generated content is increasingly consumed by other model instances in automated workflows, a practically important question arises: can a model embed a signal in natural language that an independent instance of the same model can detect, relying only on shared pre-training and task instructions, without any shared memory or coordination-specific training? We introduce For Your Eyes Only, a cooperative signalling game designed to evaluate this directly. A Sender produces free-form descriptions for two words, one of which is a hidden target; an isolated Receiver must identify it. We evaluate seven contemporary models from four architectural families on 300 word pairs from established psycholinguistic corpora, using the Double-Pass Success Rate to control for output biases. We find that most models struggle to maintain coordination once they are required to avoid detectable signals, while one frontier model retains near-perfect performance even after such filtering. We further show that models can direct this capability toward deliberate misdirection, and that coordination is consistently weaker across architectures than within them.

[NLP-66] Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在自主系统中应用时,安全防护机制带来的延迟与计算开销问题,尤其针对资源受限、时间敏感的部署场景。现有外部防护模型(guardrail models)无法感知模型内部状态,导致安全性保障存在根本性盲区。为此,本文提出一种新思路:若模型自身已具备识别有害内容的能力,是否可直接利用其内部激活信息实现高效检测?研究从LLaMA-3.1-8B模型中提取激活特征,并训练轻量级多层感知机(MLP)探测器(仅1260万参数),以判断输入提示是否具有危害性。在WildJailbreak、Beavertails和AEGIS 2.0三个基准测试上,该探测器分别达到99%、83%和84%的F1分数,性能媲美千倍规模的外部防护模型,同时显著降低延迟与计算成本,其核心创新在于通过挖掘模型内部表征实现高效、低开销的安全监控。

链接: https://arxiv.org/abs/2609.19472
作者: Alizishaan Khatri,Chiquita Prabhu,Omkar Neogi
机构: Wrynx Inc.(Wrynx公司)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Autonomous systems increasingly rely on Large Language Models (LLMs) yet the safety infrastructure surrounding these models introduces latency and compute overhead. This limits utility in resource-constrained, time-critical deployments. Existing external guardrail models remain blind to the model’s internal workings, creating a fundamental assurance gap. We ask: does the model already know when the content is harmful? We extract activations from LLaMA-3.1-8B and train lightweight MLP classifier probes (12.6M parameters) to detect harmful prompts. Evaluated on WildJailbreak, Beavertails, and AEGIS 2.0, our probes achieve F1 scores of 99%, 83%, and 84%, respectively competitive with 1000x larger guard models while cutting latency and compute costs.

[NLP-67] From Models to Systems: A Comprehensive Survey of Efficient Multimodal Learning

【速读】: 该论文旨在解决多模态模型在快速发展过程中所面临的计算、内存与部署瓶颈问题,核心在于填补当前高效多模态学习(Efficient Multimodal Learning, EML)领域中对效率在模型-算法-系统全栈层面如何体现、实现及协同的系统性认知空白。其解决方案的关键在于提出首个从模型到系统的分层结构化分类体系,将效率优化划分为三个层级:模型层(架构精简)、算法层(执行优化)与系统层(硬件感知协同),并通过跨层协同设计揭示“效率-性能-隐私”三者之间的根本权衡机制。研究进一步以多模态大语言模型(Multimodal Large Language Models, MLLMs)为案例,梳理了从早期结构调优到现代全栈资源调度的演进路径,并倡导向自调节智能范式转型,使效率成为模型基础设计的内在涌现属性而非事后附加约束。该工作不仅构建了兼具高性能、泛化性与原生高效的多模态系统框架,还为不同应用场景提供了定制化优化蓝图,并指明了未来研究的关键挑战与发展方向。

链接: https://arxiv.org/abs/2609.19445
作者: Pan Wang,Siwei Song,Hui Ji,Siqi Cao,Heng Yu,Zhijian Liu,Huanrui Yang,Yingyan Celine Lin,Beidi Chen,Mohit Bansal,Xiaoming Liu,Pengfei Zhou,Ming-Hsuan Yang,Tianlong Chen,Jingtong Hu
机构: University of Pittsburgh(匹兹堡大学); Stanford University(斯坦福大学); UCSD(加州大学圣地亚哥分校); University of Arizona(亚利桑那大学); Georgia Tech(佐治亚理工学院); Carnegie Mellon University(卡内基梅隆大学); UNC Chapel Hill(北卡罗来纳大学教堂山分校); UC Merced(加州大学默塞德分校); New York University(纽约大学)
类目: Multimedia (cs.MM); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: TMLR

点击查看摘要

Abstract:The rapid expansion of multimodal models has surfaced formidable bottlenecks in computation, memory, and deployment, catalyzing the rise of Efficient Multimodal Learning (EML) as a pivotal research frontier. Despite intensive progress, a cohesive understanding of what, how, and where efficiency is manifested across the learning stack remains fragmented. This survey systematizes the EML landscape by introducing the first structured, model-to-system taxonomy. We distill insights from over 300 seminal works into three hierarchical levels–model, algorithm, and system–addressing architectural parsimony, execution refinement, and hardware-aware orchestration, respectively. Moving beyond a purely categorical review, we offer a methodological synthesis of the vertical synergies between these layers, elucidating how cross-layer co-design contributes to the fundamental “Efficiency-Utility-Privacy” trade-off. Through an integrative case study of Multimodal Large Language Models (MLLMs), we trace the field’s evolutionary trajectory from initial structural adjustments to modern full-stack resource orchestration. Furthermore, we provide a holistic discussion and application-specific optimization blueprints for diverse domains and posit a paradigm shift toward self-regulating intelligence, where efficiency is an intrinsic, emergent property of the model’s fundamental design rather than a post-hoc constraint. Finally, we present open challenges and future directions that will define the trajectory of EML research. This survey establishes a structured framework for multimodal systems that are not only high-performing and generalizable but natively efficient and ready for ubiquitous deployment. A continuously updated version is available at this https URL.

[NLP-68] BurnRiSc: Toward Non-Invasive Burnout Screening in Open Source from Public Repository Signals

【速读】: 该论文旨在解决开源软件维护者(maintainer)群体中职业性倦怠(burnout)难以被早期识别与量化监测的问题。现有方法依赖于自我报告量表(如旧堡德倦怠量表,Oldenburg Burnout Inventory),但其在实际应用中存在显著局限:无法对高风险个体进行有效检测,且不具备回溯分析能力,导致研究者无法评估倦怠的普遍性或干预措施的有效性。为此,论文提出BurnRiSc框架,通过挖掘GitHub公开活动数据,将倦怠的两个核心维度——精力耗竭(exhaustion)与情感疏离(disengagement)——转化为14个可计算的行为与语言信号,并基于贡献者自身历史动态建模,生成加权维度分值,最终合成月度倦怠风险评分(Burnout Risk Score, BRS)。该方法的关键在于利用机器学习从已标注案例中自动学习信号权重,实现对个体长期行为模式的敏感捕捉。初步评估显示,在10个仓库共68名贡献者的样本中,持续升高的BRS可提前6至15个月预测6例真实倦怠披露,若结合峰值指标则可提升至8例,所有10例披露均能在任意时间窗口内被成功预警。研究表明,基于公共代码平台数据的倦怠可筛查性具备可行性,为构建可持续的开源生态健康监控机制提供了技术路径。

链接: https://arxiv.org/abs/2609.19422
作者: Timofey Sanko,Yuan Tian,Mariam Guizani
机构: Queen’s University (皇后大学)
类目: Computers and Society (cs.CY); Computation and Language (cs.CL)
备注: 8 Pages, Submitted to the JAWs 2 Workshop

点击查看摘要

Abstract:Burnout is a chronic occupational syndrome, and open source is close to a worst case for it: maintainers absorb unbounded demand with no manager to reallocate work and no organization to notice decline. The cost is not only personal. Burnout precedes withdrawal, and in projects sustained by a handful of maintainers, one departure can break infrastructure that thousands of downstream systems depend on. Yet the field has no way to see it coming: self-report inventories, the only existing measure, miss exactly the contributors most in need of detection and cannot be applied retroactively, so the field cannot even ask how common burnout is or what helps. We present BurnRiSc, a framework that operationalizes the Oldenburg Burnout Inventory’s two dimensions, exhaustion and disengagement, as 14 behavioral and linguistic signals computed from GitHub activity and scored against each contributor’s own history. The signals aggregate into two weighted dimension scores, with weights learned from labeled cases, and average into a monthly Burnout Risk Score (BRS). In a preliminary evaluation across 68 contributors in ten repositories (ten disclosed burnout cases, twelve comparable-volume collapses, and 46 comparison contributors), sustained BRS elevation precedes 6 of 10 disclosures by 6-15 months, 8 of 10 when adding peak BRS as a second criterion, and 10 of 10 over any prior time frame. We thus present BurnRiSc as evidence that burnout is screenable from public data. Comments: 8 Pages, Submitted to the JAWs 2 Workshop Subjects: Computers and Society (cs.CY); Computation and Language (cs.CL) Cite as: arXiv:2609.19422 [cs.CY] (or arXiv:2609.19422v1 [cs.CY] for this version) https://doi.org/10.48550/arXiv.2609.19422 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[NLP-69] Less Is More: Graph-free Multimodal RAG via Multi-signal Late Fusion

【速读】: 该论文旨在解决基于图结构的检索增强生成(RAG)在多模态跨文档问答任务中存在构建成本高、查询效率低以及维护困难的问题。其核心解决方案是提出一种无图结构的多模态框架TrioRAG,通过整合三个互补信号——问题文本、锚定图像以及由视觉语言模型(VLM)增强的联合查询——各自独立地在共享的多向量索引(包含页面文本与图像)上进行检索,并采用晚期融合(late fusion)策略整合结果。该设计显著降低了系统开销,同时将单次查询推理速度提升1.6至2.3倍,且在多个基准测试中达到或超越传统图结构方法的性能。此外,为更真实地评估模型在复杂现实场景中的表现,研究引入AutoQA,一个基于噪声网络来源图像而非干净文档图像的多模态汽车领域基准,其问题需跨手册进行推理,被定位为由模型自洽构建的测试平台而非人工验证的黄金标准。实验表明,在出域图像检索仅实现19.3%的文档级召回率的情况下,依赖文本信号尤其是经VLM增强的查询仍能保持稳健的检索能力,凸显了TrioRAG对非理想输入场景的鲁棒性。

链接: https://arxiv.org/abs/2609.19417
作者: Tithi Rakshit,Hongkuan Zhou,Lavdim Halilaj,Yuqicheng Zhu
机构: University of Tübingen (图宾根大学); Robert Bosch GmbH (罗伯特·博世公司); University of Stuttgart (斯图加特大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Graph-based retrieval-augmented generation (RAG) is widely used for multimodal, cross-document question answering. However, building corpus-level graphs is expensive, slow to query, and difficult to maintain. We present TrioRAG, a graph-free multimodal framework that integrates evidence from three complementary signals: the question, the anchor image, and a VLM-enhanced query generated from both. Each signal retrieves independently over a shared multi-vector index of page text and page images, and the results are combined through late fusion. Further, we introduce AutoQA, a multimodal automotive benchmark whose questions are grounded in noisy, web-sourced images rather than clean document-sourced figures. Its questions require reasoning across manuals. We position it as a model-curated testbed rather than a human-validated gold standard. Across three benchmarks, TrioRAG matches or outperforms graph-based systems while reducing total cost and accelerating per-query inference by 1.6-2.3 times. By construction, AutoQA grounds its questions in out-of-corpus web images. In this setting image retrieval reaches only 19.3% document-level recall, while text-derived signals, especially the VLM-enhanced query, keep retrieval robust.

[NLP-70] A Cross-Lingual Acoustic Disease-Alignment Framework for Respiratory Health Assessment from Spontaneous Speech

【速读】: 该论文旨在解决跨语言呼吸系统疾病评估中因语言特异性语音差异导致的可解释性模型泛化能力不足的问题。其核心挑战在于,疾病相关的声学变化常被不同语言的语音特征所混淆,使得现有模型难以在多语言场景下保持一致的有效性。为此,研究提出一种跨语言疾病对齐框架(Cross-Lingual Disease-Alignment Framework, CL-DAF),其关键在于识别在多种语言中疾病效应保持一致的声学维度。通过构建包含201名英语和75名新采集的孟加拉语发音者的共性272维声学表征,并结合符号秩-双系列效应与语言不变性评分量化疾病对齐程度,CL-DAF成功筛选出26个跨语言稳定的疾病相关特征。实验表明,基于这些特征的模型在英-孟加拉语之间迁移时分别实现0.825和0.722的AUC,显著优于全特征表示的0.49表现,验证了该方法在抑制语言依赖性变异、聚焦病理特征方面的有效性,为构建可解释、可泛化的多语言临床语音分析模型提供了坚实基础。

链接: https://arxiv.org/abs/2609.19398
作者: Roksana Khanom,Raghib Asfak Tasnim,Bodrun Nahar Bithi,Shafia Shirin Supty,Saiful Islam Raju,Ashok Agrawala,Nirupam Roy
机构: University of Dhaka (达卡大学); University of Texas at Austin (德克萨斯大学奥斯汀分校)
类目: ound (cs.SD); Computation and Language (cs.CL)
备注: Under review

点击查看摘要

Abstract:Spontaneous speech offers a scalable, noninvasive signal for respiratory health assessment, yet interpretable models that generalize across languages remain challenging because disease-related acoustic changes are confounded by language-specific phonetic variation. We present CL-DAF, a Cross-Lingual Disease-Alignment Framework that identifies acoustic dimensions whose disease effects remain consistent across languages. Using 201 English and 75 newly collected Bangla speakers, we construct a common 272-dimensional acoustic representation and quantify disease alignment using signed rank-biserial effects and the Language Invariance Score. We first show that spontaneous Bangla speech separates COPD from controls (AUC 0.85); however, 133 features reverse their disease direction across languages and the full representation transfers poorly (AUC 0.49 from Bangla to English). CL-DAF isolates 26 disease-aligned features that raise AUCs to 0.825 and 0.722 from English to Bangla and Bangla to English, respectively. These findings provide a foundation for multilingual clinical speech models emphasizing pathology over language-dependent variation.

[NLP-71] Riemannian–Lorentz Fusion of Vision Transformers and State-Space Models

【速读】: 该论文旨在解决深度学习模型在规模化过程中面临的三大瓶颈:数据耗尽、训练成本呈指数级增长以及计算资源过度集中。现有方法中,模型合并(Model Merging)通过避免梯度下降重新训练,实现相较于完整训练的量级级效率提升,但其在处理架构与参数形状不一致的异构模型时存在显著挑战。尤其当融合对象为视觉变换器(Vision Transformer, ViT)与状态空间模型(State-Space Model, SSM)等采用不同运算机制的模型时,传统基于权重空间的合并方法因假设参数对齐且结构兼容而难以适用。为此,本文提出一种异构混合合并(Heterogeneous Merging)设置,在保留两类模型原始架构的前提下,通过语义角色对齐参数组。其核心解决方案是黎曼-洛伦兹参数融合(Riemannian–Lorentz Parameter Fusion, RLPF):该方法将对齐后的参数组映射至共同坐标系,选取部分坐标升维至双曲空间的洛伦兹超球面模型(Lorentz hyperboloid model),在该非欧几何空间中计算正则化测地线重心,并将结果解码回原双分支结构;最终通过一个可学习门控机制融合两分支的输出逻辑值。其中,组件组使用固定曲率,归一化参数则视为欧氏变量。实验结果显示,经微调后系统在CIFAR-10、Oxford-IIIT Pet和ImageNet-1K上分别达到82.37%、75.04%和78.58%的准确率,显著优于各自最优父模型(分别为76.54%、71.42%、76.42%),且预微调初始化阶段即达ImageNet-1K 77.80%准确率。这些结果验证了几何感知异构融合的有效性,但表明当前方法并非完全免训练的单检查点合并——RLPF本质上为双分支架构,其门控模块与最终模型均需训练。

链接: https://arxiv.org/abs/2609.19384
作者: Badri N. Patro,Vijay S. Agneeswaran
机构: Microsoft(微软)
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Multimedia (cs.MM)
备注:

点击查看摘要

Abstract:Scaling deep learning faces critical bottlenecks: data exhaustion, exponential training costs, and resource concentration. Model merging combines pre-trained checkpoints without gradient descent, offering orders-of-magnitude savings versus retraining. Combining independently trained vision models is difficult when their architectures and parameter shapes differ. Existing weight-space merging methods generally assume aligned, shape-compatible checkpoints, whereas a Vision Transformer (ViT) and a state-space model (SSM) implement token mixing with different operators. We study a hybrid Heterogeneous merging setting that retains both architectures while aligning parameter groups by semantic role. Our proposed Riemannian–Lorentz Parameter Fusion (RLPF) method projects aligned groups to common coordinates, lifts selected coordinates to the Lorentz hyperboloid model of hyperbolic space, computes a regularized geodesic barycenter, and decodes the result into the two branches. A learned gate then combines branch logits for each input. Component groups use fixed curvature values, with normalization parameters treated as Euclidean. In the results available in this manuscript, the fine-tuned system obtains 82.37% on CIFAR-10, 75.04% on Oxford-IIIT Pet, and 78.58% top-1 accuracy on ImageNet-1K; the corresponding best-parent accuracies are 76.54%, 71.42%, and 76.42%. On ImageNet-1K, the reported pre-fine-tuning initialization reaches 77.80%. These results support further study of geometry-aware heterogeneous fusion, but not a training-free single-checkpoint merge: RLPF is a two-branch hybrid whose gate and reported final models are trained.

[NLP-72] he Role of Fine-grained Harm Signals in LLM Safety

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)安全机制中,类别特异性有害性表征(category-specific harmfulness representation)在超越共享通用有害性表征(general harmfulness representation)之外所起作用的问题。现有研究已表明,模型内部的有害性表征在不同风险类别间存在差异,但共享一个共同的通用有害性成分。然而,这一通用成分之外的类别特异性成分是否仍具有实际意义尚不明确。为此,本文通过从各类别有害性表征中移除共享的通用有害性成分,提取出与通用有害性正交的类别残差(category residual),并在此基础上利用激活操控(activation steering)技术,在3个指令微调的大语言模型中对11个风险类别进行系统分析。研究发现,类别残差是否编码有害性信息在不同类别间存在显著差异,且该模式在不同模型间具有一致性;而类别残差是否引发拒绝响应则表现出更强的模型依赖性。此外,类别残差能增强模型下游层面对通用有害性表征的内生对齐(internal alignment)。综上,研究揭示了即使在某一层与特定概念正交的方向,也可能在后续层中促进该概念的放大,因此,为全面理解大语言模型的安全性,必须将更细粒度的类别残差纳入考量。其解决方案的关键在于通过解耦通用与类别特异性成分,揭示类别残差在有害性传播与模型响应中的非平凡作用。

链接: https://arxiv.org/abs/2609.19366
作者: Soyeon Park(1),Seogyeong Jeong(1),Sunwoo Kim(1),Alice Oh(1) ((1) KAIST)
机构: KAIST(韩国科学技术院)
类目: Computation and Language (cs.CL)
备注: 9 pages, 6 figures

点击查看摘要

Abstract:Prior work has shown that internal harmfulness representations in large language models vary across risk categories, while sharing a common general harm representation component. This raises a question about the role of the category-specific component beyond general harm representation in LLM safety. To answer this question, we isolate the category-specific component by removing shared general harmfulness representation from each categorical harmfulness representation, yielding a category residual that is orthogonal to general harmfulness at every layer. Using activation steering with category residuals across 11 risk categories in 3 instruction-tuned LLMs, we find that whether category residuals encode harmfulness varies across categories, and that this category-wise pattern is similar across models. Whether category residuals induce refusal also varies across categories, but this category-wise pattern is more model-dependent. We also find that category residuals increase LLMs’ downstream internal alignment with shared general harmfulness representation. Together, these findings demonstrate that more fine-grained category residuals should also be considered beyond shared general harmfulness representation to fully understand LLM safety. More broadly, our findings show that even a direction orthogonal to a concept at one layer can contribute to the concept’s downstream amplification.

[NLP-73] A frontend-backend architecture for tool calls in full-duplex speech models

【速读】: 该论文旨在解决全双工语音到语音(Full-duplex speech-to-speech, S2S)系统在实现自然、低延迟对话交互的同时,如何有效集成外部工具调用(tool-call)能力以完成复杂语音代理任务的问题。现有系统在保持流畅交互与工具调用性能之间存在权衡,尤其在实时性与上下文连贯性方面面临挑战。其解决方案的关键在于提出一种前端-后端(frontend-backend)架构:前端采用双工语音转文本(speech-to-text, STT)模块,通过学习生成“委托令牌”(delegation token),将流式语音识别结果传递至基于文本的后端大语言模型(LLM)进行工具调用决策;后端完成工具调用后,将结果通过轻量级的“预填充-重复”(prefill-and-repeat)机制回传至前端,并经由流式文本转语音(text-to-speech, TTS)合成返回用户。该设计仅需对前端模型进行最小化修改,即可在不破坏原有双工对话中的轮次交替、打断处理和低延迟特性的情况下,实现强大的代理式工具调用能力。实验表明,该方法在单轮工具调用任务中达到92%-97%的召回率,具备良好的工具调用预测性能及81.2%的无关调用拒绝准确率;当后端采用更大规模模型(如Qwen3-235B-A22B)时,在Full-Duplex-Bench-V3和EVA-Bench基准上均表现出与开源及闭源模型相当甚至更优的性能,验证了后端委托机制在融合自然双工语音交互与强代理能力方面的有效性与模块化优势。

链接: https://arxiv.org/abs/2609.19334
作者: Ke Hu,Slyne Deng,Chen Chen,Elena Rastorgueva,Edresson Casanova,Punit Kumar,Dharmendra Choudhary,Nikhil Srihari,Ameya Sunil Mahabaleshwarkar,Viet Anh Trinh,Slim Essid,Oluwatobi Olabiyi,Zhehuai Chen
机构: 未知
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Full-duplex speech-to-speech (S2S) models provide natural, low-latency conversational interaction and would benefit from the ability to use external tools and complete voice-agent tasks. We propose a frontend-backend architecture where a duplex speech-to-text frontend learns to emit a delegation token and forwards streaming ASR transcripts to a text-based backend LLM for tool calls. Tool-call results from the backend are injected back into the frontend through a lightweight prefill-and-repeat mechanism and then synthesized using streaming TTS to the user. Our approach largely preserves regular duplex turn-taking, interruption handling, and low-latency interaction as it requires minimal modifications to the frontend model. In a single-turn tool-call evaluation, our system achieves 92-97% tool-call recall, competitive tool-call prediction performance, and 81.2% accuracy in rejecting irrelevant calls. When equipped with a larger backend (e.g., Qwen3-235B-A22B), our system achieves competitive results on Full-Duplex-Bench-V3 compared to open and closed source models, and significantly outperforms GPT-realtime-mini and Qwen3-Omni-30B-A3B-Instruct on EVA-Bench. These results demonstrate that backend delegation is an effective and modular approach for combining natural duplex speech interaction with strong agentic tool-call capabilities.

[NLP-74] AUDITPLAN: Commit Then Answer for Auditable Safety Alignment

【速读】: 该论文旨在解决现有安全调优(safety tuning)流水线仅评估最终回答所带来的局限性,即难以区分真正的稳健拒绝与两种不良行为:对良性请求的泛化拒绝,以及表面合规但实际未约束输出的安全推理(safety rationale)——后者虽形式上合理却缺乏真实性。为此,论文提出AUDITPLAN,一种基于单模型“先规划后作答”的新方法:模型首先生成一个紧凑的结构化安全计划(safety plan),包含威胁标签、预期行为和显式约束,该计划在训练阶段用于机器可验证的审计,但在部署时对用户隐藏。通过监督微调结合基于FAITHGATE的强化学习,该方法仅在安全计划正确时才给予答案奖励,从而有效抑制看似安全实则不忠实的行为,并强化计划与回答之间的耦合性。实验结果表明,在Qwen系列模型上,AUDITPLAN显著提升了鲁棒性与可审计性:以Qwen2.5-3B-Instruct为例,FAITHGATE将攻击成功率(ASR)从24.0%降至11.6%,低质量拒绝率(LSR)从1.0%降至0.36%,过度拒绝率从11.0%降至2.0%,优于仅基于答案的强化学习、自由形式解释及加权求和结构化奖励等基线方法。更大规模模型(Qwen-3-4B-Instruct与Qwen2.5-7B-Instruct)的验证结果进一步确认了显式内部承诺在提升安全对齐的忠实性、鲁棒性和可审计性方面的有效性。

链接: https://arxiv.org/abs/2609.19325
作者: Sai Sri Pushpa Jampani,Kshitij Mishra,Asif Ekbal
机构: Indian Institute of Technology Patna (印度理工学院比哈尔分校)
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Safety tuning pipelines judge only the final answer, which makes it difficult to distinguish robust refusal from two undesirable shortcuts: blanket refusal on benign requests and polished but unfaithful safety rationales that do not actually constrain the answer. We propose AUDITPLAN, a single-model plan-then-answer approach where the model first emits a compact structured safety plan and then answers conditioned on it. The plan records a threat label, intended action, and explicit constraints, enabling machine-checkable auditing while remaining hidden from users at deployment. We train this behavior with supervised fine-tuning followed by reinforcement learning with FAITHGATE, a reward-gating objective that grants answer reward only when the safety plan is correct. This discourages safe-looking but unfaithful behavior and promotes tighter plan-answer coupling. Across Qwen backbones, AUDITPLAN improves both robustness and auditability: on Qwen2.5-3B-Instruct, FAITHGATE reduces ASR from 24.0% to 11.6%, LSR from 1.0% to 0.36%, and over-refusal from 11.0% to 2.0%, outperforming answer-only RL, free-form explanation, and weighted-sum structured rewards. Similar trends hold for Qwen2.5-1.5B-Instruct. Larger-model confirmation runs on Qwen-3-4B-Instruct and Qwen2.5-7B-Instruct preserve the same trend suggesting that explicit internal commitments can make safety alignment more faithful, robust, and auditable.

[NLP-75] Why Pretraining Fails to Share Cross-Lingual Knowledge

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在多语言任务中跨语言知识迁移能力有限的问题。尽管这些模型在多种语言的处理与建模上取得了显著进展,但其跨语言知识泛化能力远低于人类多语者,且这一局限性在标准训练干预下依然存在。研究发现,该问题的根本原因在于多语言预训练过程中不同语言被映射到独立的词元(token)空间,即使面对同一语言的两个副本,仅因词元空间不共享也会导致知识隔离(knowledge compartmentalization)。为此,论文提出关键解决方案:通过简单的词级翻译将不同语言映射至共享的词元空间,从而显著提升跨语言知识泛化能力,恢复高达12.6%的母语学习效率,相较基线提升达14倍。

链接: https://arxiv.org/abs/2609.19291
作者: Adam Gaber,Uriel Dolev,Elisabeth Fittschen,Bobby Cheng,Yuval Marton,Leshem Choshen
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Large Language Models (LLMs) have made remarkable progress in the processing and modeling of many languages. Yet, unlike human multilinguals, they exhibit surprisingly limited cross-lingual knowledge transfer. While this limitation is well documented, its origins during multilingual training remain unclear. We pretrain 360M- and 7B-parameter LLMs and show that poor cross-lingual knowledge generalization emerges during pretraining and persists under standard interventions. To isolate its cause, we employ a controlled bilingual pretraining setting using two copies of the same language, sharing identical text and token segmentation, but mapped to disjoint token spaces. We find that disjoint tokens alone are enough to induce knowledge compartmentalization, even between identical copies of the same language, establishing disjoint token spaces as a fundamental barrier to cross-lingual knowledge generalization. Guided by this understanding, we suggest mapping languages into a shared token space by simple word-wise translation and find it substantially improves cross-lingual knowledge generalization, recovering up to 12.6% of native-language learning efficiency — 14 \times the baseline.

[NLP-76] YNU-HPCC at SemEval-2025 Task 11: Bridging the Gap in Text-Based Emotion Using Multiple Prediction Headers

【速读】: 该论文旨在解决跨语言文本情感识别(Bridging the Gap in Text-Based Emotion)中的多语言情感分类问题,尤其关注如何在不同语言间实现高效且准确的情感理解。其核心挑战在于处理多语言文本中情感表达的差异性与语义鸿沟。解决方案的关键在于采用经过优化的RoBERTa(Robustly Optimized BERT Approach)模型,并通过改进输出头结构,使模型能够以单标签方式逐个预测单一情感类别,而非同时预测六种情绪,从而提升分类精度。此外,研究发现将全部多语言数据统一翻译为英文后进行训练,相较于直接使用原始多语言数据,可显著提升模型性能,这表明统一语境下的训练有助于缓解语言间的语义偏差。实验结果表明,该方法在官方评测中取得0.44的分数,验证了所提策略的有效性。

链接: https://arxiv.org/abs/2609.19238
作者: Hao Yang,Jin Wang,Xuejie Zhang
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:This paper describes the participation of the YNU-HPCC team in subtask A of task 11, Bridging the Gap in Text-Based Emotion at SemEval-2025. Our best-performing system employs the RoBERTa (Robustly Optimized BERT Approach) model, an improved version of BERT that utilizes the Transformer encoder architecture. We enhanced the output head to allow the model to process one emotion simultaneously. We obtained the official ranking score (0.44), including results from all languages. The entire dataset was translated into English using Google Translate to facilitate subsequent processing. Through probabilistic and attention analyses, we found that (I) a single prediction head performs better than six heads predicting six emotions simultaneously, and (II) training on a uniformly translated English dataset yields better results than using the original dataset. The code is available at: this https URL.

[NLP-77] CovR: Coverag e-Aware Hardware Verification via Reasoning -Guided Reinforcement Learning

【速读】: 该论文旨在解决硬件设计验证中测试平台生成效率低下且覆盖率质量不足的核心问题,尤其针对当前基于大语言模型(Large Language Models, LLMs)的方法普遍仅关注功能正确性而忽视覆盖率优化的局限性。其解决方案的关键在于提出一种名为CovR的智能体框架,通过引入自省循环(self-reflection loops)与基于仿真的反馈机制,实现覆盖驱动的测试平台自动化生成。该框架构建了一个包含16,514个自然语言规格-RTL推理测试用例对的大规模数据集,并利用强教师模型实现覆盖感知的监督训练;在此基础上,设计了一种面向覆盖率优化的强化学习(Reinforcement Learning, RL)框架,结合仿真结果与覆盖率反馈信号作为奖励函数,持续优化学生模型。实验表明,经微调后的CovR模型在VerilogEval和RTLLM V2.0上达到93.81% cov@10,显著优于现有方法;进一步将微调模型回填至智能体迭代流程后,覆盖率提升至94.27%,并在全验证流程中实现18.95%的覆盖率增益及1.19%的突变检测分数提升,同时发现4.46%此前未被检测到的故障,充分验证了以覆盖率为导向的生成策略在生成式AI(Generative AI)辅助硬件验证中的关键价值。

链接: https://arxiv.org/abs/2609.19189
作者: Manar Abdelatty,Maryam Nouh,Sherief Reda
机构: Brown University (布朗大学)
类目: Hardware Architecture (cs.AR); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Design verification remains one of the most resource-intensive stages of hardware development, often consuming up to 70% of the total design effort. While recent work has explored using Large Language Models (LLMs) to automate testbench generation, most existing approaches focus narrowly on functional correctness, overlooking the critical aspect of coverage quality. To bridge this gap, we present CovR, an agentic framework for automated testbench generation that combines self-reflection loops with simulation-based feedback to maximize coverage. Using this pipeline, we construct a large-scale dataset of 16,514 natural specification RTL reasoning testbench tuples with a strong teacher model, enabling coverage-aware supervision. Building on this, we propose a reinforcement learning (RL) framework tailored for coverage-driven testbench generation, leveraging tool-derived rewards from simulation and coverage feedback to optimize a student model. Experimental results show that the CovR finetuned model achieves 93.81% cov@10 on VerilogEval and RTLLM V2.0, and 87.76% cov@10 on CVDP, outperforming state-of-the-art approaches by 7.97% and 3.59%, respectively. Furthermore, deploying the finetuned model back into the agentic refinement pipeline further improves cov@10 to 94.27% on VerilogEval and RTLLM V2.0 and 91.39% on CVDP. Moreover, when integrated as a plug-in stimulus engine for full verification workflows, CovR improves coverage by 18.95% and mutation detection score by 1.19%, while revealing 4.46% undetected failures, highlighting the importance of optimizing for coverage in LLM-based hardware verification.

[NLP-78] What Do We Expect from LLM s? Mapping the Design of LLM Benchmarks

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)评估体系演进过程中,评估标准与研究期望的动态变化缺乏系统性审视的问题。传统上,模型排名仅反映性能排序,却无法揭示评估本身如何随时间演变。为此,研究通过系统分析2022年1月至2026年8月期间arXiv平台上提交的14,767篇引入或更新评估资源的论文,采用分阶段筛选与自动化全文编码方法,深入考察了目标系统与领域、评估材料与条件、评分机制等方面的变化。其核心发现在于:评估重点正从静态文本生成转向对行动(action)、交互(interaction)及专业应用(professional applications)的重视,同时传统与新兴设计元素并存;在评估参与模式上,基于模型的评分在代理型与非代理型任务中均呈增长趋势,但由模型生成的评估材料并未出现持续上升趋势。这一研究揭示了公开科研如何将对模型能力的期望转化为具体测试与成功标准的过程。随着人工智能既参与构建测试、执行任务,又承担评分角色,其自身偏好与认知盲点可能被嵌入评估体系,从而引发关键质疑:评估范围的扩展究竟是提供了更独立的证据,还是在无形中复制了模型自身的局限性?

链接: https://arxiv.org/abs/2609.19182
作者: Chao Wang(Independent Researcher)
机构: Independent Researcher(独立研究员)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 15 pages, 5 figures, 7 tables. Data and code: this https URL

点击查看摘要

Abstract:Benchmarks are central to how progress in large language models (LLMs) is assessed and communicated. Yet model rankings alone reveal little about how evaluation requirements themselves are changing. The expanding variety of benchmarks offers another perspective: what researchers expect LLMs to do, and what they count as successful performance. We systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January 2022 and August 2026. Using staged screening and automated full-text coding, we examine changes in target systems and domains, evaluation materials and conditions, and scoring mechanisms. The collection shows growing emphasis on action, interaction, and professional applications, while established and newer design elements frequently coexist. Model participation also develops unevenly: LLM-based scoring grows within both agent and non-agent groups, whereas model-generated materials show no comparable sustained increase in recent cohorts. These findings illuminate how public research translates capability expectations into concrete tests and criteria for success. As AI participates in constructing tests, performing tasks, and judging responses, they also raise a question: does expanding evaluation provide more independent evidence, or risk reproducing the preferences and blind spots of its participating models?

[NLP-79] o Memories and Beyond: From Remembering to Knowing You across Long-Term Multimodal Personal Archives

【速读】: 该论文旨在解决当前人工智能系统在长期个性化记忆建模中存在的核心问题:现有长时记忆评估基准多为合成的、仅基于文本的数据,忽视了人类日常记忆中至关重要的视觉记录,缺乏真实、具有因果关联的纵向数据,导致模型仅能实现浅层的事实性回忆,无法有效支持对用户长期经历与动态偏好演变的深度推理。其解决方案的关键在于提出ReaLMem(真实世界长期多模态记忆)基准ChronoProfiler时序加权分析模块。ReaLMem是首个基于真实多年个人视觉档案及第一人称主观标注构建的多模态长期记忆评估体系,涵盖事实回忆、人格推断与预测性个性化三个认知层级,具备真实性和复杂性;而ChronoProfiler通过计算用户属性的时间稳定性得分,并将其作为显著性先验引入模型,有效缓解时间上不一致偏好之间的冲突,促进多共激活偏好在复杂个性化决策中的协同整合。实验证明,该框架不仅揭示了前沿多模态大语言模型(MLLMs)与记忆系统在预测性个性化上的性能瓶颈,更验证了高质量、时序感知表征对提升个性化能力的关键作用,为构建可长期演进的智能个人助手提供了可信评估平台与高效技术路径。

链接: https://arxiv.org/abs/2609.19167
作者: Wenqi Zhou,Zhuorui Yu,Kaiao Wen,Hao Zheng,Xinyi Zheng,Peiran Wu,Enmin Zhou,Chi-Hao Wu,Junxiao Shen
机构: University of Bristol(布里斯托大学); Memories.ai Research
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:As AI systems evolve into personalized digital companions, a central capability is reasoning over a user’s long-term personal history: not merely storing past events, but tracking longitudinal experiences and evolving preferences. Progress here is bottlenecked by evaluation, existing long-term memory benchmarks are largely synthetic and text-only, they overlook the visual records that anchor everyday human memory, lack the authentic and causally connected longitudinal data that real personalization demands, and consequently remain confined to shallow factual recall. We introduce ReaLMem (Real-world Long-term Multimodal Memory), the first benchmark built from authentic multi-year personal visual archives, paired with first-person subjective annotations. ReaLMem evaluates models across three cognitive tiers of increasing difficulty: factual recall, persona inference, and predictive personalization. We further propose ChronoProfiler, a temporal-weighting profiling module that computes temporal stability scores for user attributes and applies them as a salience prior, resolving conflicts among temporally inconsistent preferences and helping models compound multiple co-active preferences in complex personalized decisions. Extensive evaluation of frontier multimodal large language models (MLLMs) and memory systems on ReaLMem reveals predictive personalization as a consistent ceiling, exposes clear performance gaps and bottlenecks between MLLMs and memory systems, and shows that high-quality, temporally informed representations substantially improve personalization. Together, ReaLMem and ChronoProfiler provide an authentic testbed and a simple, effective mechanism for long-term personalization, laying a foundation for future research on lifelong AI companions.

[NLP-80] Advantage Scale Calibration Imbalance in Group-Relative Optimization under Low-Variance Rewards: Diagnosis and Bounded Recovery

【速读】: 该论文旨在解决验证器风格强化学习(verifier-style RLVR)中优势尺度(advantage scale)处理不当导致的偏好信号失真问题,尤其关注低方差情形下的两类情况:亚分辨率抖动(sub-resolution jitter)不应成为偏好信号,而可信但微小的基数差异(cardinal gaps)则应被学习且不破坏KL校准。其解决方案的关键在于提出一种优势尺度三重校准接口(advantage-scale three-way calibration interface),通过组内相同尺度的分母同时决定奖励分支强度、提示级批处理权重,以及当奖励分支在原始基数尺度上重新表达时所诱导的有效KL校准。该接口揭示了RLOO或特定实现为何能使可信的小差距变为以KL为主导,而GRPO采用标准差作为分母会无界放大微小差距。基于此接口,研究进一步引入奖励分辨率协议(Reward-Resolution Protocol)与MaxNorm-AC,分别用于过滤亚分辨率间隙,并对可信的非零间隙提供有界的基数恢复。实验表明,MaxNorm-AC在密集模型与混合专家(MoE)架构及数学/代码推理任务中均优于最强的鲁棒尺度基线,同时截断了低方差反比例尾部。

链接: https://arxiv.org/abs/2609.19164
作者: Fei Ding
机构: 未知
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 10 pages,2 figures

点击查看摘要

Abstract:In verifier-style RLVR, group-relative optimization often treats advantage scale as an implementation detail. This paper separates two low-variance cases: sub-resolution jitter that should not become a preference signal, and credible but small cardinal gaps that should be learned without distorting KL calibration. We propose an advantage-scale three-way calibration interface: the same within-group scale denominator simultaneously determines the reward-branch strength, prompt-level batch weight, and the effective KL calibration induced when the reward branch is re-expressed on the original cardinal scale. This interface explains why RLOO / this http URL can let credible small gaps become KL dominated, whereas GRPO’s standard-deviation denominator can amplify tiny gaps without bound. Based on this interface, we further introduce the Reward-Resolution Protocol and MaxNorm-AC, respectively filtering sub-resolution gaps and providing bounded cardinal recovery on credible nonzero gaps. Across dense / MoE architectures and math / code reasoning, MaxNorm-AC improves over the strongest robust-scale baseline while truncating the low-variance inverse-scale tail.

[NLP-81] Reflective Recovery: A Self-Supervised Method for Reasoning by Learning from Mistakes

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在推理过程中因错误累积导致的“可扩展性崩溃”(Scaling Collapse)问题,即当训练数据中的正确推理轨迹有限时,单纯增加正例数量无法持续提升模型性能,且模型在推理时一旦产生中间步骤错误便难以自我纠正。其解决方案的关键在于提出一种名为“反思恢复”(Reflective Recovery)的自监督方法:通过提取失败推理轨迹的前序片段,并将其与提示词拼接后重新输入模型,引导其从错误状态中恢复至正确解题路径。该方法利用失败轨迹中蕴含的错误信息作为训练信号,使模型学会识别并修正自身推理过程中的偏差,从而实现无需外部评判器或奖励模型支持的内在纠错能力。实验表明,该方法显著提升了多个基准上的推理准确率,并打破了传统方法的可扩展性瓶颈,推动模型从依赖结果记忆的范式转向具备过程感知与自我反思能力的新型推理模式。

链接: https://arxiv.org/abs/2609.19156
作者: Qirui Chen,Renjie Pi,Jiahui Gao,Lingpeng Kong
机构: Zhejiang University; The University of Hong Kong; The Hong Kong University of Science and Technology
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Data-driven fine-tuning is widely adopted to enhance reasoning in Large Language Models (LLMs) due to its simplicity and efficiency. However, mainstream imitation learning methods that rely exclusively on perfect reasoning trajectories suffer from a Scaling Collapse: when the problem set is limited, increasing positive examples fails to yield continuous improvement. However, during inference, an LLM can not guarantee that every intermediate step is correct and is therefore prone to errors. Once such errors arise, the LLM often struggles to recover and may be further misled by the accumulation of previous mistakes. To address this, we propose Reflective Recovery, a simple yet effective self-supervised approach that transforms failed reasoning attempts into recovery training data. Specifically, we extract initial segments of failed trajectories, concatenate them with prompts, and use them to guide the LLM toward valid solutions. Because these segments from failed trajectories are likely to contain errors, this process teaches models to recognize and correct mistakes during reasoning, enabling recovery from erroneous states without relying on external critics or reward models. Evaluated on extensive benchmarks, Reflective Recovery significantly improves performance. On DeepSeek-R1-Distill-Qwen-7B, it boosts accuracy from 30.0% to 37.5% on AIME 2025 and from 37.6% to 47.8% on Minerva. More importantly, analyses demonstrate that it breaks the scaling collapse barrier and enables models to develop emergent self-correction behaviors, representing a paradigm shift from outcome-oriented memorization to process-oriented reflective reasoning.

[NLP-82] owards Proactive Detection of User-Side Implicit Conflicts in Human-LLM Dialogue

【速读】: 该论文旨在解决人机对话中用户侧冲突(user-side conflict)的检测问题,即用户后续话语可能隐含地与早期意图相冲突,导致大语言模型(LLM)误解用户需求并生成不恰当回应。现有研究多聚焦于模型自身引发的冲突(LLM-side conflicts),而对用户侧冲突的识别与处理缺乏系统性探索。为填补这一空白,作者构建了首个由人工标注的基准数据集UC-Bench,用于评估用户侧冲突检测能力。实验表明,当前主流大模型在处理基于对话历史隐含不兼容性的冲突时表现不佳。针对轻量级模型在训练数据有限情况下的性能瓶颈,论文提出一种约束引导的数据合成方法SynUC,其通过在约束空间中建模用户话语间的隐式不兼容性,并借助SPEAKING框架实现可追溯的约束变换,从而生成高质量、带可靠标签的隐式冲突样本。基于SynUC构建的UC-Data数据集(含2,487条样本)显著提升了模型性能:在UC-Bench上,使用该数据微调的Qwen3.5-4B模型超越了更大规模通用模型如Claude Opus 4.8,以及采用传统合成方法训练的同构模型,验证了所提方法在提升轻量级模型对用户侧冲突感知能力方面的有效性。

链接: https://arxiv.org/abs/2609.19155
作者: Jinqiang Wang,Tao Zhu,Huansheng Ning
机构: University of Science and Technology Beijing (北京科技大学); University of Southern California (南加州大学)
类目: Computation and Language (cs.CL)
备注: 24 pages, 13 figures

点击查看摘要

Abstract:In Human-LLM dialogue, follow-up user utterances may implicitly conflict with earlier intents, leading the LLM to misinterpret user needs and generate inappropriate responses. A reliable dialogue system should proactively detect user-side conflicts before generating a response and seek clarification when necessary. However, prior work has largely focused on LLM-side conflicts, leaving user-side conflicts underexplored. To fill this gap, we construct UC-Bench, a human-annotated benchmark for evaluating user-side conflict detection. Preliminary experiments show that existing LLMs struggle with this task, especially when conflicts arise from implicit incompatibilities grounded in dialogue history. To improve lightweight LLMs with limited training data, we investigate data synthesis for user-side conflict detection. Existing synthesis methods do not explicitly model the implicit incompatibilities between historical and current user utterances, making it difficult to capture the evolution of conflicts and to generate reliably labeled implicit conflict samples. We propose SynUC, a constraint-guided synthesis method that represents user-side conflicts in a constraint space and uses the SPEAKING framework to guide traceable constraint transformations. Applying SynUC to WildChat, we construct UC-Data, a user-side conflict training set containing 2,487 samples. On UC-Bench, Qwen3.5-4B trained on UC-Data outperforms larger general-purpose LLMs such as Claude Opus 4.8, as well as the same backbone trained on data synthesized by existing methods.

[NLP-83] Neo-Classic: A Benchmark for Evaluating Linguistic-Aesthetic Reasoning in Classical Chinese Poetry

【速读】: 该论文旨在解决当前大语言模型(LLMs)在古典诗词理解任务中难以区分可迁移的语言美学推理能力与对预训练阶段熟悉模式的依赖性问题。其核心挑战在于现有基准测试容易因模型对历史语料的直接记忆或检索而产生虚假性能表现,从而掩盖了真正的推理能力。为此,论文提出Neo-Classic评估基准,采用建构主义的样本外(Out-of-Sample, OOS)数据集,由当代专家创作的严格格律诗歌构成,并辅以一系列逆向理解探测器(reverse understanding probes),以规避传统基于历史语料的验证或生成任务中存在的检索偏差。关键解决方案在于通过引入具有真实创作背景且未在训练集中出现的当代诗歌文本,强制模型进行深层次的层级约束满足(hierarchical constraint satisfaction)推理,而非依赖表面形式匹配。实验结果表明,尽管先进模型(如Qwen3-Max、Gemini-3-Pro、DeepSeek-V3.2)在局部形式特征上表现良好,但在跨时代文本迁移时性能下降20%至50%,且在篇章级顺序排列等全局规划任务中准确率仅为0%至13%,即使引入专家指导,推理增强模型最高仅达36%,仍显著低于人类专家水平。这揭示出当前主流模型在缺乏对整体结构与美学意图的统筹规划能力,即其语言美学推理仍局限于局部模式匹配,难以实现真正意义上的高层级、系统性推理。

链接: https://arxiv.org/abs/2609.19154
作者: Han Zhang,Zihan Gu,Zhiyuan Wang,Tianyi Ma,Jiacheng Lu,Xinyan Zhang,Yuhao Wei,Cheng Hua
机构: Shanghai Jiao Tong University (上海交通大学); Institute of Information Engineering, Chinese Academy of Sciences (中国科学院信息工程研究所); School of Cyber Security, University of Chinese Academy of Sciences (中国科学院大学网络空间安全学院)
类目: Computation and Language (cs.CL)
备注: Published in the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2026

点击查看摘要

Abstract:While Large Language Models (LLMs) achieve high accuracy on established Classical Chinese Poetry benchmarks, it remains challenging to distinguish transferable Linguistic-Aesthetic Reasoning from reliance on familiar pre-training patterns. To address this issue, we introduce Neo-Classic, an evaluation benchmark that combines a constructionist Out-of-Sample (OOS) dataset with a suite of reverse understanding probes. Unlike traditional benchmarks that rely on verification or generation over historical corpora, Neo-Classic comprises strictly metrical poetry authored by contemporary experts, reducing the possibility of direct retrieval. We evaluate state-of-the-art models, including Qwen3-Max, Gemini-3-Pro, and DeepSeek-V3.2, across five behavioral probes designed to test hierarchical constraint satisfaction. Our results reveal two primary limitations. First, a performance gap of 20 to 50 percent emerges when models transition from historical to contemporary texts. Second, models exhibit substantial difficulties in discourse-level ordering tasks, with standard accuracy remaining low (0 to 13 percent). Although expert-level guidance improves the performance of reasoning-enhanced models to 36 percent, a notable gap with human experts persists. These findings suggest that while current LLMs capture local formal patterns, they struggle with global hierarchical planning required for robust Linguistic-Aesthetic Reasoning.

[NLP-84] FakeSpotter: A content and strategy agnostic Viral Misinformation Detection Tool

【速读】: 该论文旨在解决现有虚假信息检测工具在面对新型误导性叙事时有效性受限的问题,因其通常依赖于二元真假分类或基于历史样本训练的模型,难以适应不断演变的传播策略。其解决方案的关键在于提出FakeSpotter——一种内容与策略无关的工具,通过测量虚假信息的结构指纹(structural fingerprints of misinformation)而非直接判断真伪,来评估文本的病毒式传播风险。该方法基于理论驱动的多维度框架,涵盖语言、叙事、逻辑和批判性思维层面,结合大语言模型(LLM)的重复评估与领域特定的逻辑回归分类器,分别处理短文本和长文本。在包含764条社交媒体及FakeNewsNet来源的标注语料库上,FakeSpotter在保留测试集上实现了短文本0.788和长文本0.793的宏平均F1分数。此外,其可解释性层提供基于特征的评分、信号一致性分析及警示指数,支持社会监听与人工监督,表明识别虚假信息的结构特征可实现早期、可解释且由人类干预的潜在高传播性虚假信息评估。

链接: https://arxiv.org/abs/2609.19152
作者: Giovanni Spitale,Federico Germani
机构: 未知
类目: Computation and Language (cs.CL); Social and Information Networks (cs.SI)
备注:

点击查看摘要

Abstract:Misinformation detection tools often rely on binary true and false classifications or models trained on historical examples, limiting their usefulness when novel misleading narratives emerge. Here, we present FakeSpotter, a content- and strategy-agnostic tool designed to estimate the viral misinformation risk of textual content by measuring structural fingerprints of misinformation rather than directly adjudicating truthfulness. FakeSpotter operationalizes a theory-driven framework across linguistic, narrative, logical, and critical-thinking dimensions, using repeated LLM assessments and domain-specific logistic regression classifiers for short and long texts. In a labelled corpus of 764 texts from social media and FakeNewsNet, FakeSpotter achieved macro F1 scores of 0.788 for short texts and 0.793 for long texts on a held-out test set. FakeSpotter’s interpretive layer provides explainable outputs through feature-based scores, signal agreement, and a caution index, and can be used for social listening. These findings suggest that identifying the structural fingerprints of misinformation can support early, explainable, and human-supervised assessment of potentially viral misinformation.

[NLP-85] Sampling Reveals Style: Unsupervised Training-Free Discovery of Prompt-Conditional Stylistic Axes in LLM Activations

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)中风格维度(stylistic dimensions)的自动发现问题,即如何在无需监督对比数据的情况下,识别出对特定提示(prompt)具有显著影响的风格特征。传统方法依赖人工标注的对比数据来挖掘风格维度,而本文提出一种无需训练、基于提示条件的无监督探针方法:通过在高温度下多次采样同一提示的生成结果,将得到的隐藏激活值进行池化,并应用主成分分析(Principal Component Analysis, PCA)提取主要变化方向,再根据极点生成内容自动标注主成分轴。该方法的关键在于利用模型自身解码过程中的内在变异来揭示风格结构,且不依赖外部标注数据。实验验证表明,在最强模型Qwen-3.5-4B-Instruct上,前两个主成分轴与人类自发提出的风格维度匹配度达72.8%精确率和43.6%宏召回率,且90.9%的标注者间一致性显示其极点生成内容与标签高度一致。此外,该方法揭示了不同模型间风格结构组织方式的显著差异,如Qwen和Llama-3.2-3B能有效暴露人类关注的风格维度,而DeepSeek-7B-Chat则因主导成分反映结构而非风格变异,导致性能下降。因此,简单地对模型自身解码方差进行PCA,是一种高效、低成本且可揭示跨模型风格结构差异的探针手段。

链接: https://arxiv.org/abs/2609.19150
作者: Ajit Mallavarapu,Ziwei Gu
机构: Cornell University (康奈尔大学); Harvard University (哈佛大学)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Large language models (LLMs) encode rich stylistic structure in their hidden activations, but discovering which stylistic dimensions are salient for a given prompt typically requires supervised contrastive data. We present a training-free, prompt-conditional alternative: we repeatedly sample completions of a single prompt at elevated temperature, apply Principal Component Analysis (PCA) to the pooled hidden activations, and label the resulting axes automatically from the pole generations. We validate the discovered axes against 245 human-elicited stylistic annotations in a two-phase study. On our strongest model (Qwen-3.5-4B-Instruct), the top two axes match spontaneously requested human dimensions with 72.8% precision and 43.6% macro-recall, and 75.6% of validity ratings judge the axes’ polar generations accurate to their labels, with 90.9% adjacent inter-annotator agreement. Discoverability is strongly model-dependent: both Qwen models and Llama-3.2-3B expose human-salient axes, while DeepSeek-7B-Chat drops to 35.3% precision, its leading components dominated by structural rather than stylistic variance. Simple PCA over a model’s own decoding variance is thus an effective, low-cost probe of stylistic structure in LLM representations, one that also exposes sharp cross-model differences in how that structure is organized.

[NLP-86] Subliminal Prompting Beyond Static Geometry: Causal Depth and Multi-Token Confounds

【速读】: 该论文旨在解决生成式模型中隐性学习(subliminal learning)现象的机制问题,即语言模型如何在看似无关的输出中传递隐藏属性(如动物与数字之间的关联)。其核心挑战在于区分不同可能的解释机制:传统观点“标记纠缠”(token entanglement)假设动物与数字标记通过输出词汇表产生关联,但现有测量方法无法明确界定这种关联是源于共变性、固定输出向量对齐、隐藏状态可读性,还是因果控制。为此,研究设计了一个固定的动物-数字提示协议,系统分离并独立测量了四种关键属性——固定几何结构(fixed geometry)、观测可读性(observational readability)、因果时序性(causal timing)以及多标记测量(multi-token measurement)。结果显示,从Llama-3.1-8B到70B,固定输出向量相似性对行为预测能力下降(均值相关性变化为-0.080),且输出头读出的归一化深度AUC无显著变化;而通过捐赠-控制实验(donor-control)发现,将某一数字提示中的临时答案位置状态复制至另一提示,最终动物得分会随之改变,其控制能力显著提升(AUC从0.254升至0.540,配对变化+0.286),且该效应在仅保留八个Transformer块时仍存在,说明其非由特定网络结构或身份混淆所致。此外,在Qwen模型中,逐位评分无法恢复单标记正向关联,而通过对标记平均则出现正向聚合关联,但该现象在控制数字长度后消失,揭示了长度混淆(length confound)的存在。因此,本研究的关键在于提出并验证了四类独立的测量维度,这些维度共同约束了对训练期间属性传递机制的解释,但尚未能唯一确定其根本机制。

链接: https://arxiv.org/abs/2609.19149
作者: Barath Velmurugan
机构: 未知
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 7 pages, 3 figures, 5 tables. Preprint

点击查看摘要

Abstract:Subliminal learning shows that language models can transmit a hidden trait through outputs that appear unrelated to it. One proposed explanation, token entanglement, links animal and number tokens through the model’s output vocabulary. Yet existing measurements answer different questions: whether outputs co-vary, fixed output vectors align, an answer can be read from a hidden state, or that state causally controls the answer. We measure each separately in a fixed animal-number prompting protocol. From Llama-3.1-8B to 70B, fixed output-vector similarity predicts behavior less well: the paired mean correlation change is -0.080 (95% CI [-0.127, -0.035]). A fixed output-head readout shows no resolved change in normalized depth AUC. To test control, we copy the temporary answer-position state from one number prompt into another at five depths and measure which prompt the final animal score follows. Donor-control AUC rises from 0.254 to 0.540, a paired change of +0.286 (95% CI [+0.272, +0.300]), with increases for all 18 concepts. The contrast remains with exactly eight transformer blocks remaining, while specificity and identity controls remain small or exact. In two Qwen models, scoring every digit in sequence does not recover the positive one-token association. Per-token averaging instead creates a positive pooled association that disappears after controlling number width, revealing a length confound. Thus, fixed geometry, observational readability, causal timing, and multi-token measurement are distinct properties of this frozen prompting channel. They constrain token-level explanations but do not identify the mechanism of training-time trait transfer.

[NLP-87] Modality Discrepancy Transformer for Ambivalence and Hesitancy Recognition

【速读】: 该论文旨在解决临床视频中情感状态“矛盾与犹豫”(Ambivalence and Hesitancy, A/H)的自动识别问题,其核心挑战在于跨模态信号的不一致——即面部、语音和语言通道间存在冲突,而传统多模态融合方法通常会抑制此类差异性信号。为应对这一难题,论文提出一种基于冲突感知的多模态融合框架的改进模型:模态差异变换器(Modality Discrepancy Transformer, MDT)。MDT的关键创新在于将原始的6令牌设计扩展为9令牌表示结构,包含三个模态嵌入、三个绝对差特征以及三个通过线性投影学习的哈达玛积(Hadamard-product)差异特征,从而显式建模各模态间的不一致性。该架构采用Transformer自注意力机制,并引入基于FiLM的文本条件调制与低秩适配(LoRA)微调作为核心组件,以增强对语义上下文的敏感性。此外,设计了一个文本引导的晚期融合分支,在推理阶段将仅依赖文本的辅助头输出与全模态输出进行融合,进一步提升判别能力。在第3届ABAW挑战赛的BAH数据集上,MDT在标注测试集上取得0.7408的宏平均F1分数,在私有排行榜上达到0.7368,显著优于现有最强基线超过10个百分点,且仅需单张GPU训练不足20分钟,验证了其高效性与优越性能。

链接: https://arxiv.org/abs/2609.19148
作者: Shiyu Luo,Yu Wang,Jiawen Huang,Zhaoxiang Xiao,Chenxi Huang,Qi Zhang,Bin Liu
机构: 未知
类目: Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
备注: 10 pages

点击查看摘要

Abstract:Ambivalence and hesitancy (A/H) are affective states in which individuals express contradictory signals across facial, vocal, and linguistic channels. Automatically recognising A/H in clinical videos requires detecting cross-modal disagreement – the signal that standard fusion methods suppress. Based on the conflict-aware multimodal fusion framework of Bekhouche et al., we present the Modality Discrepancy Transformer (MDT). MDT enriches the original 6-token design to a 9-token representation comprising three modality embeddings, three absolute-difference features, and three Hadamard-product discrepancy features learned through linear projections. These nine tokens undergo Transformer self-attention, with FiLM-based text-conditioned modulation and LoRA fine-tuning as core architectural components. A text-guided late fusion branch blends a text-only auxiliary head with the full multimodal output at inference. On the BAH dataset from the 3rd ABAW Challenge, MDT achieves 0.7408 Macro F1 on the labelled test split and 0.7368 on the private leaderboard, outperforming the strongest published baseline by over 10 points while training in under 20 minutes on a single GPU.

[NLP-88] Model-Agnostic and Language-Agnostic Voice Pipeline Improvement for the Agriculture Domain

【速读】: 该论文旨在解决小农户在使用生成式 AI(Generative AI)驱动的农业咨询助手FarmerChat时,因语音输入环境复杂导致的自动语音识别(ASR)准确率低的问题。具体而言,田间采集的语音常包含农机噪声、背景媒体声音、多人对话干扰以及大量农业领域专有词汇(如作物名、病虫害名称、化学药剂和数量术语),这些因素严重影响了关键语义信息的识别。针对这一挑战,论文提出了一种模块化、模型无关的ASR增强流水线,其核心创新在于不依赖对底层ASR模型进行微调或替换,而是通过一系列协同处理步骤实现性能提升:包括门控音频增强、说话人分离与目标说话人选择、通用ASR识别、基于加权农业词典的领域感知纠错,以及用于检测不可靠转录结果的质量门控机制。其中仅说话人分离阶段进行了微调,其余组件均采用现成的开源模型并通过统一接口集成。实验在印地语、泰卢固语和奥里亚语三种语言的标注数据集上验证,结果显示该流水线显著降低了词错误率(WER),尤其在多说话人场景下,通过目标说话人选择有效抑制了干扰语音的影响;在全部语料上,云部署ASR模型的相对WER降低16%-23%,本地设备模型降低5%;而在多说话人录音中,云模型降幅达32%-42%,本地模型为16%,所有改进均具有统计显著性。这表明,针对性的预处理、说话人选择与领域感知后处理可在不改变原有ASR模型的前提下,大幅提高农业语音转录质量。

链接: https://arxiv.org/abs/2609.20504
作者: Aakash Singh,Lakshmi Pedapudi,Chandrashekar M S,Sanyam Singh,Naga Ganesh,Vineet Singh
机构: Digital Green
类目: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Sound (cs.SD)
备注: 20 tables, 11 figures, 23 pages

点击查看摘要

Abstract:FarmerChat is Digital Green’s AI-powered agricultural advisory assistant for smallholder farmers, who access it in their own language through text, voice, or photographs. Voice is a critical channel for this population, yet field-recorded speech is challenging for general-purpose automatic speech recognition (ASR) because recordings frequently contain machinery noise, background media, competing speakers, and domain-specific agricultural vocabulary. These conditions disproportionately affect crop, pest, chemical, and quantity terms that carry the meaning of a farmer’s query. We present a modular, model-agnostic pipeline for improving ASR quality in FarmerChat without fine-tuning or replacing the underlying ASR model. The pipeline combines gated audio enhancement, speaker diarization and target-speaker selection, ASR, domain-aware correction using a weighted agricultural lexicon, and a quality gate for detecting unreliable transcripts. Only the diarization stage is fine-tuned; all other stages use off-the-shelf models behind common interfaces. We evaluate the pipeline on human-annotated FarmerChat recordings in Hindi, Telugu, and Odia using word error rate (WER) and a domain-weighted error rate that gives greater importance to agricultural terminology. The largest improvements occur on multi-speaker recordings, where target-speaker selection prevents competing speech from entering the transcript. Across the full corpus, the pipeline reduces WER by 16-23% relative on three cloud ASR models and by 5% on an on-device model. On multi-speaker recordings, the reductions are 32-42% for the cloud models and 16% for the on-device model. All reported reductions are statistically significant. These results show that targeted preprocessing, speaker selection, and domain-aware post-processing can substantially improve agricultural speech transcription while preserving the underlying ASR model. Comments: 20 tables, 11 figures, 23 pages Subjects: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Sound (cs.SD) Cite as: arXiv:2609.20504 [eess.AS] (or arXiv:2609.20504v1 [eess.AS] for this version) https://doi.org/10.48550/arXiv.2609.20504 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[NLP-89] Large Language Model Agents for Evidence Based Genetic Disease Severity Classification

【速读】: 该论文旨在解决遗传病表型严重程度分类主观性强、人工成本高的问题,尤其针对基因组筛查中商业检测面板在规模和重叠度上差异显著所导致的标准化难题。其核心解决方案是构建一个自主的AI代理系统,融合推理与行动(ReAct)框架与检索增强生成(RAG)技术,实现对10,211个人类表型术语(Human Phenotype Ontology terms)的自动化严重程度分类。该系统基于美国医学遗传学与基因组学学会(ACMG)的严重程度指南及美国妇产科医师学会(ACOG)的生活质量评估标准,自动检索PubMed文献,生成可解释的推理链,并独立验证声明的有效性。在表型层面,利用专家标注队列验证,系统达到93.55%的准确率(MCC 0.9237),82.6%至91.4%的判断获得直接证据或有效推论支持;在基因层面,通过对8,738对基因的严重程度聚合分析,识别出3,283对常染色体隐性遗传病具有严重或极重度表型表现。外部验证显示,该系统与Mackenzie’s Mission基因列表具有95.2%的一致性。该方法通过提供基于直接证据的可靠自动化分类,实现了基因检测面板设计的标准化。

链接: https://arxiv.org/abs/2609.19569
作者: Tohid Ghasemnejad,Ahmadreza Argha,Mark Grosser,John Wang,Min Yang,Thantrira Porntaveetus,Tony Roscioli,Nigel H. Lovell,Mahmoud Aarabi,Hamid Alinejad-Rokny
机构: 未知
类目: Genomics (q-bio.GN); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Disease severity classification for genetic conditions is subjective and labor-intensive, creating bottlenecks in genomic screening, where commercial panels vary widely in size and overlap. We developed an autonomous AI agent integrating Reasoning and Acting (ReAct) with Retrieval-Augmented Generation (RAG) to classify 10,211 Human Phenotype Ontology terms. It uses American College of Medical Genetics (ACMG)-endorsed severity guidelines and American College of Obstetricians and Gynecologists (ACOG) quality-of-life criteria to retrieve PubMed literature, generate interpretable reasoning chains, and independently verify claims. At the phenotype level, using expert-curated cohorts, the agent achieved 93.55% accuracy (MCC 0.9237) with 82.6% to 91.4% of claims supported by direct evidence or valid inferences. Gene-level severity was aggregated across 8,738 pairs, identifying 3,283 autosomal recessive pairs with severe or profound presentations. External validation showed 95.2% concordance with Mackenzie’s Mission gene list. This system enables standardized panel design by providing reliable, automated classification supported by direct evidence.

信息检索

[IR-0] Reasoning Quality Matters: Combating Reasoning Collapse in LLM -based Embedding Learning

链接: https://arxiv.org/abs/2609.20563
作者: Zihan Gong,Xiaohan Ye,Jiangchao Yao,Jinsong Lan,Xiaoyong Zhu,Xu Chen
类目: Information Retrieval (cs.IR)
备注: 30 pages, 8 figures

点击查看摘要

Abstract:Large Language Models (LLMs) have recently shown strong potential for producing context-rich text embeddings for retrieval. Most existing methods either treat embedding learning as passive feature extraction or exploit LLM reasoning through instruction following for better embedding optimization. However, specialization toward embedding objectives can suppress useful reasoning generation or produce retrieval-irrelevant text. We refer to these two forms of degradation as reasoning collapse. To address this issue, we propose CoFree (Collapse-Free Reasoning Embedding), a two-stage framework that progressively integrates LLM reasoning into query and document embedding optimization while preserving reasoning quality. At the first stage, CoFree applies reference-guided supervised fine-tuning to restore the reasoning ability and retain representational strength of the foundation embedding model. At the second stage, we introduce dual rewards, an embedding-oriented reward and a reasoning-oriented reward, to guarantee fine-grained reasoning of the relevance toward the embedding goal in reinforcement learning. This endpoint-coupled optimization transforms embedding learning from static alignment into a high-quality reasoning-guided search process for retrieval. Extensive experiments demonstrate the effectiveness of CoFree, with CoFree-4B achieving an average absolute improvement of 2.8 nDCG@10 points over Qwen3-Embedding-4B across 22 datasets from MTEB and BRIGHT. Online experiments in a real-world retrieval system further show consistent gains. Code, RTED, and model checkpoints will be made publicly available.

[IR-1] Viveka-Insight: a cross-lingual concept graph and citation-grounded retrieval resource over the complete works of Swami Vivekananda in English and Bengali

链接: https://arxiv.org/abs/2609.20303
作者: Tamal Maharaj
类目: Computation and Language (cs.CL); Digital Libraries (cs.DL); Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:Classical philosophical corpora pose three compounding challenges for language resources: they exist in several languages without parallel alignment, their vocabulary is remote from that of contemporary readers, and generated text over culturally sensitive material must be verifiably grounded. We present Viveka-Insight, a bilingual resource and open-source pipeline for the works of Swami Vivekananda (1863-1902): the nine-volume English Complete Works and the ten-volume Bengali Vani o Rachana, two related but non-parallel corpora of about 15 million characters. Four layers are released: (i) a structure-preserving parse (32,694 paragraphs, 168,842 sentences) with per-paragraph anchors deep-linking into the published editions; (ii) a cross-lingual concept graph of 8,362 language-agnostic concepts with 87,518 relation-typed paragraph-concept and 55,872 concept-concept edges, in which canonical English labels act as a string-equality key linking Bengali and English passages with no parallel data; (iii) a bilingual alias inventory of 60,850 surface forms (30,053 English, 30,797 Bengali); and (iv) a human-annotated set of 200 paragraph-concept edges judged by three annotators, released with all per-annotator judgments. We report known-item cross-lingual retrieval over 194 verified rendered lecture pairs (Recall@10 0.86 in both directions), a 30-question audit of citation integrity and modern-question bridging, and a human study placing concept-extraction precision at 0.60 under strict two-annotator consensus (Cohen’s kappa = 0.61). The extractor’s confidence weight is calibrated: restricting to weight = 0.8 raises precision to 0.71 while retaining 98% of concept-bearing paragraphs. Precision is markedly lower in Bengali than English (0.54 vs 0.68), locating the weakness in exactly the half that cross-lingual access depends on. The design transfers to other multilingual classical corpora.

[IR-2] FacetCRS: Multi-Faceted Preference Learning for Pricking Filter Bubbles in Conversational Recommender System

链接: https://arxiv.org/abs/2609.20175
作者: Yongsen Zheng,Ziliang Chen,Jinghui Qin,Liang Lin
类目: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:The filter bubble is a notorious issue in Recommender Systems (RSs), which describes the phenomenon whereby users are exposed to a limited and narrow range of information or content that reinforces their existing dominant preferences and beliefs. This results in a lack of exposure to diverse and varied content. Many existing works have predominantly examined filter bubbles in static or relatively-static recommendation settings. However, filter bubbles will be continuously intensified over time due to the feedback loop between the user and the system in the real-world online recommendation. To address these issues, we propose a novel paradigm, Multi-Facet Preference Learning for Pricking Filter Bubbles in Conversational Recommender System (FacetCRS), which aims to burst filter bubbles in the conversational recommender system (CRS) through timely user-item interactions via natural language conversations. By considering diverse user preferences and intentions, FacetCRS automatically model user preference into multi-facets, including entity-, word-, context-, and review-facet, to capture diverse and dynamic user preferences to prick filter bubbles in the CRS. It is an end-to-end CRS framework to adaptively learn representations of various levels of preference facet and diverse types of external knowledge. Extensive experiments on two publicly available benchmark datasets demonstrate that our proposed method achieves state-of-the-art performance in mitigating filter bubbles and enhancing recommendation quality in CRS.

[IR-3] hink Thrice Before Reranking: Multi-perspective Evidence and Reasoning Integration for Text Reranking

链接: https://arxiv.org/abs/2609.20131
作者: Lijun Liu,Zhengzong Chen,Wenyan Li,Yuanyuan Zhao,Fei Huang
类目: Information Retrieval (cs.IR); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Reasoning-based reranking with Large Language Models (LLMs) has shown promising improvements in text ranking. However, current methods predominantly rely on a single reasoning trajectory, resulting in rankings that are susceptible to reasoning errors and inherently constrained in modeling the multifaceted signals underlying document relevance. To resolve this dilemma, we propose MERIT-Rank(Multi-perspective Evidence and Reasoning Integration for Text Reranking), a framework that models complementary reasoning trajectories to improve reranking robustness. MERIT-Rank formulates a Multi-Trajectory Reasoning Space (MTRS) that evaluates query-document relevance from multiple perspectives and introduces a joint reranker that consolidates these reasoning paths into a unified ranking decision. We further develop Progressive Rank Policy Optimization (PRPO), a progressive training framework that stabilizes reasoning trajectories while continually improving ranking quality through staged optimization objectives. Experiments on both reasoning-intensive and traditional retrieval benchmarks show that MERIT-Rank consistently achieves superior performance over competitive baselines. The 4B model notably outperforms most 7B and even 32B rerankers on BRIGHT.

[IR-4] he Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents

链接: https://arxiv.org/abs/2609.20050
作者: Zhexi Feng,Ruiyi Zhang,Yongbo Yang,Pengtao Xie
类目: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 32 pages, 3 figures. Benchmark and evaluation resources: this https URL

点击查看摘要

Abstract:A coding agent halfway through an issue has already read much of what a retriever ranks highest. Relevance is scored per passage, but sufficiency belongs to the set: a ranker can fill its budget with variants of one required fact and leave the decision unsupported. We formulate state-conditioned minimal sufficient evidence recovery: given a captured agent state, recover a compact evidence combination that supplies the support its next decision still lacks. SERBench measures this on 500 held-out states from 45 repositories, recording what the agent has seen and crediting only sets that cover every fact the current decision was annotated to require. MSS-Complement treats acquisition as set construction, not ranking. Three semantic calls propose a jointly sufficient set, search for what it lacks, and return 4-8 intact source units within 6,144 tokens. One configuration, fixed on calibration data, recovers a complete set for 73.0% of those states at five items and 80.6% at eight, against 61.4% and 72.4% for Qwen3 embedding with reranking. A matched control ranking by similarity alone reaches 66.6%, placing the gain in the set-level policy, not the computation. From frozen repository source with no gold-derived pool, the lead is 5.0 points. On AMA-Bench it answers from a 76.2% smaller answer prompt, with accuracy 2.08 points above that benchmark’s own memory agent. Removing one required group from an otherwise complete set costs 12.3 and 11.1 points of repair-localization precision under two executors. Retrieval for agents is better posed as recovering what a decision lacks than re-ranking what an issue resembles.

[IR-5] Intrinsic Sequence-Likelihood Confidence in Retrieval-Dominated Extractive QA: Two Pre-Specified Negatives and What They Do and Do Not Attribute

链接: https://arxiv.org/abs/2609.19942
作者: Gunwoo Lee,Changmin Sung,Sang-Hwan Gwak,Ina Kim,Ji-Young Choi,Kyong-Ha Lee
类目: Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注: 26 pages main text + 26 pages supplementary (Online Resource 3). Submitted to Applied Intelligence. Code and data: doi: https://doi.org/10.5281/zenodo.22710121 , doi: https://doi.org/10.5281/zenodo.22721044

点击查看摘要

Abstract:In extractive document question answering whose questions were generated from the passages that contain their answers – so that retrieval recovers 92-99.8% of what any mode combination could reach, whatever its absolute accuracy – confidence-driven mechanisms have little to gain. Fine-tuning an open language model on a specialized domain corpus yields a model whose own confidence is a tempting control signal: it could decide which queries warrant further adaptation, and which answers to trust. We evaluate both uses under criteria fixed before the runs were executed, across four 7-9B model families whose adaptation moved closed-book F1 by at most +0.03, and both fail: a distillation trigger on all four families, under its pre-specified three-step transfer budget, and a routing-and-abstention policy in its single-model pilot. Retrieval alone recovers 92-99.8% of best-case combined accuracy under every correctness criterion we test, leaving routers no meaningful gain. The sequence-likelihood signal is insufficient relative to that mode – area under the receiver operating characteristic curve 0.65-0.81 under the registered criterion – before adaptation as well as after, unchanged by scalar recalibration and not consistently improved by token-level temperature rescaling. And the finer diagnostics depend on the correctness criterion and on answer length; on the three adapted combinations where we could test it, selector ablations show no statistically detectable downstream benefit from the confidence term on any seed; on Gemma, removing it changes the selector from failing to passing both registered criteria. The usable product is a set of pre-specified negatives with their dependencies made explicit.

[IR-6] rust but Validate the Instrument: Auditing AI-Generated RTL Verification Plans on Authored Security-Regression Proxies

链接: https://arxiv.org/abs/2609.19844
作者: Hang Xiao,Chuhong Xu,Kainan Zhou,Gangzhen Qian,Lu Yi
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注: Cyber-AI

点击查看摘要

Abstract:AI-generated RTL verification plans can satisfy a provider schema yet fail at the boundary to trusted execution. We present SecTB-RTL, an auditable framework covering 31 tasks and 124 authored hardware-security regressions. A deterministic non-AI baseline killed 36, 75, and 78 mutants at increasing resource limits. The first confirmatory run (C1-R2) failed before model execution because the provider rejected its response schema. After a schema-only repair made without viewing outcomes, a separately frozen follow-up run (C1-R3) completed 1,860 calls. The provider accepted 1,857 responses, but only nine passed the production semantic validator. The generation and execution rules did not match. We therefore preserve the run as an instrument-validation incident and report no prompt-effect estimate. This incident shows that provider or schema acceptance does not establish execution validity. Compilation and coverage are only diagnostics; the exact saved artifact must pass the full production path. A subsequent follow-up is excluded because it did not satisfy the preregistered evidence-completeness gate and is treated only as future work. We release the benchmark, failure-preserving contract, incident provenance, and governance controls needed to prevent infrastructure behavior from being misreported as model behavior.

[IR-7] Reproducing Transparent and Scrutable Recommendations: Exploring Open-Weight Models via Natural-Language User Profiles EMNLP’26

链接: https://arxiv.org/abs/2609.19831
作者: Noah Mamié,Laurin van den Bergh
类目: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)
备注: Accepted at BlackBoxNLP@EMNLP’26 (The 9th BlackboxNLP Workshop Special Track: Reproducibility and Reliability in Interpretability Analyses)

点击查看摘要

Abstract:In this reproducibility study, we investigate the transparency and scrutability of recommender systems enhanced by incorporating generated natural-language user profiles that represent user preferences. The original paper explores the synthesis of user profiles from raw user-generated review text across domains such as movies and accommodations (Amazon Movies TV, TripAdvisor). Crucially, these natural-language user profiles enable direct user interaction and intervention, allowing users to customize recommendations by correcting misattributed preferences or addressing cold-start settings. We successfully reproduce the core findings of the original study. Additionally, we extend the evaluation by conducting systematic context ablation experiments, multi-seed stability across five distinct random seeds to establish statistical reliability, and a mechanistic interpretability analysis using the nnsight framework to probe internal model representations under counterfactual profile perturbations. Our findings verify the original paper’s claim that User Profile Recommendation (UPR) achieves competitive performance under its test-set reranking protocol and makes recommendations more transparent. Perturbing the natural-language profiles does change predictions, but it shifts predicted ratings uniformly across genres with no detectable genre-selective effect, leaving rankings unchanged even under direct activation steering. We trace this back to the rating-regression objective rather than the profile interface, with ranking-objective models clearly exceeding in this task.

[IR-8] Dense Feature Representation over Sequence Modeling: A Solution to the KDD Cup 2026 UniRec Challenge KDD

链接: https://arxiv.org/abs/2609.19787
作者: Yi Zhang,Weiliang Ji
类目: Information Retrieval (cs.IR)
备注: 6 pages, 1 figure, 4 tables. KDD Cup 2026 Tencent UniRec Challenge Workshop

点击查看摘要

Abstract:We describe our 10th-place solution to the KDD Cup 2026 Tencent UniRec Challenge, industrial click-to-conversion (CVR) prediction over 34.82M records, and we ask which mechanisms actually move held-out AUC. Starting from the official PCVRHyFormer baseline, a 15-step single-variable chain raises test AUC from 0.813237 to 0.827816, and our final submission reaches 0.828535. A leave-one-out ablation from the full model attributes the gain: removing the dense-feature representation stack costs 0.0095 AUC and removing the orthogonalized optimizer costs 0.0028, while no sequence-modeling component (merged single-stream backbone, polarity channel, auxiliary head, per-token FFN) costs more than 0.0005, within or adjacent to a \pm 0.0004 seed band. We also report a generalization hazard: the row-group train/validation split shares one time window, so validation AUC overstates the leaderboard by about 0.014; anti-memorization and high-cardinality-ID changes even invert sign against it, a divergence that traces to dump-to-dump distribution shift and survives a time-ordered re-split. Dense representation and optimization, not finer sequence modeling, drive CVR AUC at this scale, and verdicts must come from the held-out leaderboard.

[IR-9] Self-Evolving Search Index

链接: https://arxiv.org/abs/2609.19656
作者: Sangam Lee,Wonjae Lee,Sunghwan Kim,Deogyong Kim,Jaehoon Kim,Daye Nam,SeongKu Kang,Dongha Lee
类目: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)
备注: Work in progress

点击查看摘要

Abstract:Information retrieval is increasingly important as LLM agents tackle complex tasks involving diverse information needs. Because retrieval relies on an index that represents each document through index keys, retrieval quality depends heavily on how effectively these keys expose the knowledge contained in each document. However, effective index representations vary across retrieval environments, making it difficult for any fixed optimization strategy to perform consistently. Yet evolving an index to its retrieval environment remains largely human-driven, requiring humans to diagnose retrieval failures, refine the optimization strategy, and reprocess the index accordingly. We propose SELF-INDEX, a framework that enables an index to self-evolve without human intervention. Its Optimizer autonomously diagnoses retrieval shortfalls, selectively revises the responsible index keys, and validates each revision before updating the index. Beyond reacting to observed retrieval demands, SELF-INDEX proactively explores additional demands through a Query Simulator, allowing the index to evolve beyond the queries already available for optimization. Across diverse corpora and retrievers, SELF-INDEX consistently improves retrieval performance while outperforming existing index optimization methods. We further show that these benefits extend to downstream applications, improving the effectiveness and efficiency of search agents and helping agent memory systems retrieve useful past interactions.

[IR-10] Beyond Similarity through Zero-Token Geometric Graphs for Multi-Hop RAG

链接: https://arxiv.org/abs/2609.19622
作者: Zeliang Li,Xiaofen Xing,Kailing Guo,Xiangmin Xu
类目: Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:Multi-hop retrieval-augmented generation (RAG) requires evidence that remains relevant to a query while introducing enough novelty to bridge semantic gaps. Dense retrieval tends to concentrate on semantically similar documents, whereas graph-based alternatives often depend on costly Large Language Model (LLM) entity extraction and may propagate through noisy connections. We introduce Geometric Gain Graph RAG (G ^3 RAG), a document-only framework whose offline graph construction uses no LLM calls or generated tokens. G ^3 RAG assigns each edge a geometric gain score, \cos\theta \cdot \sin\theta , that jointly captures directional consistency and orthogonality between document representations. A density-aware topological penalty suppresses highly connected hubs, while single-step controlled diffusion expands from filtered query seeds toward complementary evidence. We evaluate G ^3 RAG on MusiQue, 2WikiMultiHopQA, and HotpotQA using Nv-embed-v2 and Qwen3-8B-embed. G ^3 RAG obtains the best average F1 and answer-document hit rate among the evaluated graph-based baselines in both embedding settings, with gains of up to 4.26 F1 points in average performance and 5.76 points on MusiQue. It also removes the graph-construction token cost incurred by entity-based graph methods. These results show that geometric structure can support efficient multi-hop evidence discovery without LLM-based graph construction. Code is available at this https URL

[IR-11] Semantic Layer Induction from Raw Telemetry via Hierarchical LLM and RAG Abstraction

链接: https://arxiv.org/abs/2609.19615
作者: Yuanzhe Jia,Ali Anaissi
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Software Engineering (cs.SE)
备注:

点击查看摘要

Abstract:Modern applications generate massive volumes of raw telemetry data, but translating those noisy, heterogeneous event streams into actionable business insights remains a fundamental challenge. Data engineers and analysts expend substantial effort reconciling semantic discrepancies, hand-crafting parsing logics, and maintaining fragile mappings between raw data and business KPIs. In this paper, we present an end-to-end framework that fully automates the construction of a business semantic layer from application raw logs. Our approach introduces a two-stage semantic abstraction: first, high-level business features are identified via LLM inference augmented with domain-specific industry knowledge; second, fine-grained business nodes are derived through a structured pipeline comprising data refinement, hybrid retrieval, multi-stage filtering, semantic clustering, and canonical naming. Evaluation on production-scale telemetry demonstrates that our system improves human-assessed semantic quality from 50 to 80+ on a 100-point scale, reduces maintenance effort by 80%, filters out 74% of noise, and achieves 0.87 Cohen’s kappa via an integrated LLM-as-Judge evaluation, enabling continuous, scalable quality assurance. Overall, our work distinguishes itself from prior work by addressing the novel problem of business semantic layer induction from raw telemetry, operating without labeled training data or manual rule engineering.

[IR-12] FootprintRAG : Visual Analytics for Evidence Context Refinement in RAG -based Scientific Literature Exploration

链接: https://arxiv.org/abs/2609.19601
作者: Xingyu Liu,Yu Dong,Qizhen Yu,Shiyu Cheng,Zhe Wang,Guan Li,Guihua Shan,Dong Tian,Christy Jie Liang,Quang Vinh Nguyen
类目: Graphics (cs.GR); Human-Computer Interaction (cs.HC); Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:Retrieval-Augmented Generation (RAG) is increasingly used to ground large language model (LLM) outputs in scientific literature. However, in open-ended literature exploration, the evidence context used for generation is often produced through hidden retrieval, reranking, assessment, and filtering steps. Users may receive retrieval summaries without knowing how the system constructed the evidence context, which evidence units were retained or discarded, or whether potentially useful evidence was excluded before synthesis. We present FootprintRAG, an LLM-agent-powered visual analytics system for evidence context refinement in RAG-based scientific literature exploration. The core idea is to treat the RAG evidence context as an explicit, inspectable, and revisable analytical object before generation. FootprintRAG parses scientific literature into text and figure evidence units, expands an initial query into parallel query variants, retrieves and assesses evidence across iterative rounds, and surfaces ERS-ranked supplementary candidates from the corpus-level evidence space. Through coordinated views, the system connects retrieval trajectories, evidence-state revision, and provenance-aware summary generation into a user-steerable workflow. We evaluate FootprintRAG through two case studies, a user study, and a workflow-level comparison with representative RAG systems. The results show that FootprintRAG helps users compare retrieval directions, revise candidate evidence, recover potentially overlooked evidence, and trace generated summaries back to supporting evidence units. FootprintRAG is available at this https URL.

[IR-13] CliniCIRCA: A Modular LLM Framework for Constructing Longitudinal Mental Health Patient Journeys from Raw EHR Narratives

链接: https://arxiv.org/abs/2609.19585
作者: Aiwei Ivy Zhang,Nimra Ishfaq,Mohit Chandra,Santiago Alvarez Lesmes,Adam Coscia,Khatiya Chelidze Moon,Xiaohan Ding,Munmun De Choudhury
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:In mental health care, reasoning over patient journeys is a key task for clinicians. Yet these journeys, encompassing a longitudinal progression of biological, psychological, and social events, are often spread across disparate unstructured text narratives, making temporal recovery challenging. We present CliniCIRCA, a multi-stage LLM framework for Calendar-anchored, Imprecision-aware Reconstruction of Clinical Annals. To our knowledge, CliniCIRCA is the first to temporally classify clinical events across unstructured discharge summaries without event-level timestamps. From 14,882 MIMIC-III mental health admissions, we first construct a benchmark of 52 discharge summaries on which CliniCIRCA produces 15,891 temporally tagged events. After correcting 629 errors based on a clinician-in-the-loop evaluation, we produce verified gold-standard labels. Finally, the corrected timelines drive a temporally grounded summarization stage that compresses each source 1.52 times into a date-grouped chronological record. We then scale the framework to generate 1,000 silver-standard timelines and evaluate them as training data. Compared with zero- and few-shot prompting, instruction tuning generally improves five open-weight models on event extraction, temporal tagging, and summarization across silver and clinician-verified evaluations.

[IR-14] SCOUT: Sim-to-Real Text-Based Person Retrieval by Embedding-Space Prediction over Frozen Video Features ECCV2026 ECCV

链接: https://arxiv.org/abs/2609.19483
作者: Abdarahmane Traoré,Andy Couturier,Éric Hervet
类目: Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)
备注: 16 pages, 4 figures, 3 tables. Accepted at the ECCV 2026 Workshop on AI City Challenge (Track 4). Code and annotations: this https URL

点击查看摘要

Abstract:Text-based person retrieval under a sim-to-real gap (synthetic training data, a real-image gallery) is usually tackled with costly fine-tuned cross-encoders. We ask whether a frozen-encoder system can compete. We present SCOUT, which casts cross-modal retrieval as prediction in embedding space. A trainable predictor maps the patch tokens of a frozen video encoder into the embedding space of a frozen text encoder under a bidirectional InfoNCE objective, and no encoder is fine-tuned in the base model. The video encoder is V-JEPA, the text encoder is EmbeddingGemma, and the predictor is initialized from a Qwen3.5-0.8B decoder. We make three findings. First, the best frozen text encoder is simply the one whose geometry best matches the video features. A training-free alignment score ranks three candidate text encoders in the same order as their retrieval accuracy on our held-out split (Spearman \rho = 1.0 ); a fourth, LLM-based encoder shows the rule is metric-dependent, holding for a neighborhood-overlap score ( \rho = 0.8 ) but not for a linear probe ( \rho = -0.2 ). Second, two precision-targeted levers, parameter-efficient ExPLoRA adaptation of the video encoder and a training-free attribute-decomposed reranker built on a vision-language model, improve the top-rank precision that otherwise limits the frozen system, adding 2.2 points of leaderboard R@1. Third, a local-versus-public calibration study explains which interventions transfer to the real domain. On AI City Challenge 2026 Track 4 the full retrieve-fuse-rerank system reaches 84.25 mAP@10 on the final leaderboard, while a single frozen model submitted alone reaches 60.63. Our trained components cost about 95 GPU-hours. CMP, the dataset authors’ fine-tuned cross-encoder that trains for sixteen GPU-days, is one fusion member of the full system, not an alternative. Code and annotations: this https URL

[IR-15] Algebraic Retrieval: Composable Search for Agents

链接: https://arxiv.org/abs/2609.19482
作者: Damian Delmas
类目: Information Retrieval (cs.IR)
备注: 5 pages, 1 figure. Code and reproducible examples: this https URL

点击查看摘要

Abstract:Algebraic Retrieval lets AI agents compose search strategies at query time. Relevance criteria, eligibility constraints, and ranking preferences can be expressed together in a mathematical query. The query surface exposes available operations, so an agent can combine them for the question at hand and revise a program after inspecting results. We evaluate execution parity, not agent behavior or retrieval quality. Building on Programmatic Embedding Modulation (PEM), which exposes vector and score arithmetic during retrieval, we demonstrate contrastive scoring, candidate-pool reranking, and weighted ranking as composable queries, alongside executable SQL and PyTerrier counterparts. On the public 11,429-document Vaswani fixture, each program’s implementations select the same document set with score differences below 1e-6; one tied pair orders differently across scoring paths.

[IR-16] Beyond Private Training: The New Landscape of AI Privacy

链接: https://arxiv.org/abs/2609.19456
作者: Sean Culatana,Kang Li
类目: Cryptography and Security (cs.CR); Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:Retrieval-augmented systems increasingly rely on vector indexes that may retain deleted items in their search graph. Existing deletion interfaces can prevent deleted identifiers from appearing in returned results while still computing distances to their embeddings during graph traversal. We formalize this distinction as output safety versus traversal safety, and introduce TSD-AUDIT, a framework for auditing and enforcing traversal-safe deletion in graph-based approximate nearest-neighbor retrieval. On Faiss IndexHNSWFlat, native filtering leaves the number of distance computations unchanged relative to unfiltered search; at a 70% deletion rate, trace-faithful replay detects deleted-vector scoring in all 100 audited queries. Code inspection of hnswlib’s mark_deleted path reveals the same scoring-before-liveness pattern. TSD-AUDIT enforces an alive-before-scoring invariant, repairs connectivity using only live candidates, and emits per-query scored-trace certificates that an independent verifier can check against the deletion snapshot. Under region-targeted deletion, TSD-AUDIT improves Recall@10 over native filtering by 4.3–42.2 percentage points across deletion fractions from 0.5 to 0.9, while remaining comparable under random deletion. These results show that output-only deletion audits can miss process-level exposure: auditing deletion in vector retrieval requires accounting for the vectors scored during search, not only the identifiers returned.

[IR-17] Characterizing Web Search by Conversational LLM Agents : From Search Decisions and Strategies to Results and Responses

链接: https://arxiv.org/abs/2609.19244
作者: Mahsa Amani,Seungeon Lee,Abhisek Dash,Asmaa El Fraihi,Yunah Jang,Elisabeth Kirsten,Qinyuan Wu,Krishna P. Gummadi,Manish Gupta,Abhilasha Ravichander,Muhammad Bilal Zafar,Soumi Das
类目: Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:Conversational LLM agents increasingly rely on Web search, yet the end-to-end lifecycle of agentic search remains poorly understood. We present the first study of Web search across four major conversational platforms (ChatGPT, Claude, Grok, and DeepSeek), combining real-world user interactions (invivo) with controlled experiments using the same platform’s models by their APIs (invitro). We investigate the quality of agentic decisions to invoke Web search, their strategies to formulate queries, the potential domain preferences in the search results they receive, and the choices they make when transforming search results into grounded responses. We find that Web-search decisions vary substantially across platforms and models, while more frequent Web-search invocation does not necessarily yield better response quality. We further show that conversational agents employ different complex querying strategies and that platform specific search engines return search results from their preferred domains. Finally, although responses are largely grounded in search results, some claims rely on uncited search results, raising concerns about attribution and reliability. Our findings have important implications for the design of future AI agents and Web search tools optimized for conversational retrieval.

[IR-18] VisKG-LM: Compiling Knowledge Graphs into Visual Memory for Multiple-Choice Question Answering

链接: https://arxiv.org/abs/2609.19158
作者: Yixin Peng,Er Jin,Shiwei Luo,Diego Collarana,Stefan Decker
类目: Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Knowledge graphs are usually integrated into question answering by encoding a retrieved subgraph with a graph neural network and fusing it with the language model in the online inference path. The same subgraph is therefore re-encoded from scratch every time a pair is scored, across training epochs, seeds, and evaluation runs, even though the knowledge graph never changes. We ask whether the retrieved knowledge graphs can instead be compiled once, offline, and then accessed as read-only memory. VisKG-LM shows that it can, by decoupling graph encoding from language reasoning. It serializes each retrieved candidate-specific subgraph as Relation-Labeled Paths and renders the result as an image whose two-dimensional layout preserves the branching structure of the paths. Each image is encoded once, offline, and cached for reuse. At inference, the language model contextualizes the question and candidate from text alone, and only its final layer consults the cached visual memory, reading both its global layout and its local relational detail. The graph information thus enters only after the text has been understood. On the test sets of CommonsenseQA, OpenBookQA, and MedQA-USMLE, VisKG-LMimproves over GreaseLM by 1.2 , 0.8 , and 4.3 points, respectively, while matching or surpassing GraphVis, a 7 B vision-language model, with only about 400 M online parameters. Against a matched text-only control that receives the identical Relation-Labeled Paths, it gains 4.2 , 6.5 , and 5.1 points across the three benchmarks. These gains show that the complete visual-memory interface adds value beyond path textualization alone and support compiled visual memory as an alternative to online graph propagation.

[IR-19] Stop Removing Stopwords: How an Inherited Preprocessing Default Distorts Legal Text-as-Data

链接: https://arxiv.org/abs/2609.19153
作者: Gregory M. Dickinson
类目: Computation and Language (cs.CL); Computers and Society (cs.CY); Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Empirical legal scholarship increasingly treats judicial text as data, and much of it still runs on sparse, interpretable pipelines – TF-IDF features and linear classifiers – because the textual feature is often the object of study, not merely a means to a prediction. Yet these pipelines inherit a chain of preprocessing defaults from mid-century information retrieval that were never validated against classification accuracy, the most entrenched being stopword removal. This study introduces an exhaustive single-word ablation that measures a preprocessing step’s effect directly against the downstream objective, and applies it to stopword removal as the hardest case to dislodge. Matching Supreme Court Database labels to Caselaw Access Project opinion texts, it examines two binary tasks that bracket F1 headroom, ideological direction (no-removal baseline F1 ~ 0.68) and constitutional versus non-constitutional law type (~ 0.92), across 7,668 and 7,001 opinions. For each task the analysis approximates the best stoplist any expert could build, removing each of roughly 18,500 candidate words and measuring the effect directly. Three findings follow: generic stoplists in common use fall below the no-removal baseline in every test; even optimized stoplists are statistically indistinguishable from removing nothing; and meta-models trained on word-level features cannot predict which removals help, so list curation has nothing to target. The method generalizes to any inherited preprocessing default, and the result is a caution specific to interpretable legal text-as-data: a step that silently reshapes which features a model sees can distort the very doctrinal and ideological signal such research exists to recover. Leaving stopwords in place is a question of measurement validity.

[IR-20] What Users Think of Generative AI: A Cross-Platform NLP Analysis of Trust and Friction in App Store Reviews

链接: https://arxiv.org/abs/2609.19151
作者: Md Jafrin Hossain,Umme Nusrat Jahan,Shouvaggo Sharif Shammo
类目: Computation and Language (cs.CL); Human-Computer Interaction (cs.HC); Information Retrieval (cs.IR)
备注: Submitted to Array (Elsevier); currently under peer review

点击查看摘要

Abstract:Generative AI (GenAI) applications have achieved rapid consumer adoption, yet little large-scale research examines user-perceived quality, trust, and adoption barriers. We present one of the first cross-application analyses of app store reviews for six major GenAI applications (ChatGPT, Gemini, Microsoft Copilot, Claude, DeepSeek, and Perplexity), comprising 17,012 English-language reviews from Google Play and the Apple App Store. We combine BERTopic topic modeling with RoBERTa sentiment classification and evaluate cross-application differences using chi-square, Kruskal-Wallis, and multinomial logistic regression with Bonferroni correction. Both components are validated against human coding using a stratified sample of 300 reviews. Results show that negative sentiment concentrates in advertising (91%), authentication (89%), server reliability (83%), and subscription pricing (73%). Sentiment differs significantly across applications, with Claude exhibiting the highest negative sentiment (47.7%) alongside a strongly enthusiastic user base, indicating statistically significant polarization. These findings are robust despite unequal review counts across applications. As exploratory observations, a subset of DeepSeek reviews raised geopolitical and data privacy concerns related to its Chinese origin, while a proposed Trust Friction Score summarizes application-specific trust and usability barriers into interpretable dimensions. The study provides validated and actionable evidence on user trust, usability, and adoption barriers in consumer generative AI applications.

人机交互

[HC-0] he Data Hospital: A Workflow-Based Concept for Explainable Research Data Quality Assistance

链接: https://arxiv.org/abs/2609.20782
作者: Lennard Scheurer,Robert Porzel,Vinicius Carrillo Beber,Rainer Malaka
类目: Human-Computer Interaction (cs.HC)
备注: Concept paper, 16 pages, 13 figures, 6 tables

点击查看摘要

Abstract:Research data quality is multidimensional and purpose-dependent: it emerges from the interplay of data, intended use, contextual knowledge, documentation, intervention decisions, and traceability. This concept paper presents the Data Hospital, a human-in-the-loop control and interaction model for research data quality. Using a hospital metaphor, datasets are admitted, contextualized, assessed, reviewed in specialized stations, modified only through approved interventions, validated, documented, and made replayable where interventions are sufficiently specified. The concept combines deterministic profiling and inspectable evidence with optional evidence-bound explanation by Dr. Data and explicit user decisions. Preserved Raw Data and controlled working states separate observation from intervention. The contribution is not a new cleaning or imputation algorithm, but a ten-stage workflow that makes assessability, uncertainty, intervention authority, provenance, and process reproducibility visible. The prototype is an implementation-backed demonstrator rather than a released research artifact and illustrates selected parts of the concept through representational standardization, imputation, Patient File documentation, and replay. The paper concludes with a staged agenda for subsequent technical and user-centered evaluation.

[HC-1] Semantic Action Graph: A Shared Representation for Agent Grounding and Human Interpretation of Sports Highlights IEEE-VIS2026

链接: https://arxiv.org/abs/2609.20768
作者: Tica Lin,Deepak Chandran,Gauri Jagatap,Chen Chen,Andrea Fanelli,David Gunawan,Josh Kimball
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI)
备注: 5 pages, 3 figures, Accepted for publication at IEEE VIS 2026 Workshop on GenAI, Agents, and the Future of VIS

点击查看摘要

Abstract:Generative agents are increasingly used to select and narrate video highlights, but they typically operate over unstructured or frame-level representations. Their output is consequently difficult for a viewer to verify and steer toward individual preferences. We present the semantic action graph, a lightweight domain schema that represents a sports match as performer, action, recipient, moment, and state nodes connected by role, temporal, and outcome edges. The schema demonstrates three key properties: 1) connected event sequences, 2) a shared, closed vocabulary, and 3) frame-addressable moments, making it suitable to serve two consumers at once: an agentic pipeline that composes narrated highlights, and a visual interface through which viewers query and inspect the same structure. We instantiate it in SportSAGE, a design probe pairing a four-module highlight pipeline with a graph interface, and report feedback from 12 soccer fans. Participants were satisfied with the quality of the generated highlights and narratives, and used the graph interface to search, navigate, and interpret the match highlights. These results provide early evidence that one small, human-readable schema can ground agent generation and support human interpretation at the same time.

[HC-2] What Parents Can See: Divergent Accounts of Youth AI Companion Use in Parenting and Teenager Subreddits

链接: https://arxiv.org/abs/2609.20720
作者: Thomas Berkane,Anne Bischops,Anika Mellacheruvu,Maimuna Majumder
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Youth increasingly use AI companions, and parents are the primary mediators of that use. How effective that mediation can be depends on whether parents are aware of how adolescents actually use these systems and what risks and benefits such use carries; nevertheless, prior work has only studied these demographic groups in isolation, and existing taxonomies attend almost entirely to risk. We analyze 1,628 Reddit posts about youth AI companion use from parenting and teenager communities (2023–2026); develop a codebook covering modes of use, risks, benefits, and parental mediation; and apply it at corpus scale with an LLM. The two communities yield divergent accounts. Teenagers most often discuss receipt of emotional support from AI companions (31% of teenager posts vs. 19% of parenting posts), whereas parents most often discuss teenage use of AI companions for romantic and sexual interaction (36% vs. 25%). Teenagers are not unaware of other risks, however; indeed, attachment and dependence is the risk they raise most (19%), close to the parental rate (16%). Teenagers also describe benefits that risk-centered taxonomies do not capture and parents rarely mention, most notably emotional support (27% vs. 5%). We argue these differences track what a given kind of use makes visible to someone outside the conversation. Chatting with a companion for hours every night leaves a trace beyond the chat itself; sexting with a character stands out when a parent reads the log; venting about a fight with a friend does neither, since it looks like any other conversation. The first surfaces as dependence, the second as sexual content, and the third as emotional support, which is the one parents most often miss. Parental guidance and system design should attend to use cases that reach parents by neither route, emotional support foremost among them.

[HC-3] Ageing Digital Literacy and Interaction Modality in Immer-sive Virtual Reality: Psychomotor Performance Cognitive Flexibility and Their Processing-Speed Association

链接: https://arxiv.org/abs/2609.20719
作者: Panagiotis Kourtesis,Katerina Denaxa,Lydia Asimakopoulo,Petros Roussos,Maria Roussou
类目: Human-Computer Interaction (cs.HC); Computers and Society (cs.CY)
备注: 27 pages, 5 figures, 6 Tables

点击查看摘要

Abstract:Extended reality (XR) increasingly supports training and cognitive assessment, yet the age sensitivity of its interaction techniques is unclear. This study examined age, digital literacy, and interaction modality as correlates of psychomotor and cognitive-flexibility performance in immersive virtual reality (VR). Two hundred and two adults (19-90 years) completed a five-mode Fitts’ law task (eye-gaze, head-gaze, controller ray-casting, virtual finger, and controller direct touch), the Trail Making Test in VR (TMT-VR), and a digital-literacy questionnaire. Age was associated with slower task times across modes, but controller direct touch carried the steepest relative age gradient yet remained among the fastest in absolute terms; technique altered relative age sensitivity without determining absolute efficiency. Higher digital literacy was associated with faster TMT-VR completion but not Fitts task times. A Fitts-derived speed score predicted TMT-VR performance beyond age and partly accounted for its age association, consistent with shared processing-speed variance; an exploratory full-battery extension confined the Fitts-score association to completion-time and error-adjusted-time indices and the digital-literacy association to error-adjusted-time indices, not wrong-target errors, mean selection distance, or relative Part B indices. Age-inclusive XR assessment should standardise interaction modality, evaluate rather than exclude mid-air direct selection, and interpret cognitive scores alongside digital literacy and psychomotor speed.

[HC-4] Ownership in AI-Assisted Everyday Tasks

链接: https://arxiv.org/abs/2609.20658
作者: Megan Wei,Melanie Subbiah,Audrey Lee,Annya Dahmani,Dave Edwards,Helen Edwards,Ellie Pavlick
类目: Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:When does work done with AI still feel like ours? As AI becomes woven into everyday tasks, we must examine what happens to our sense of ownership and contribution when a machine shares in producing what we make. We report an exploratory qualitative survey in which participants were asked to describe two recent, self-selected tasks completed with AI: one that felt like their own and one that did not. We find that felt ownership depends on the process of collaboration: people disown work when they merely approve AI’s suggestions, but retain ownership when they lead, iterate, or rewrite. Ownership can also extend to settings where people own the vision for a project but not the execution; respondents reported high ownership on tasks they could not have completed without AI. Loss of personal voice and a lack of comprehension of the output both erode ownership. Finally, willingness to disclose AI use is often decoupled from actual pride or ownership, and instead shaped by community norms and fear of credit erasure. We propose several research directions as a result of these findings to promote AI development that supports people’s sense of authorship over their own lives.

[HC-5] Stereotypically Yours: Portrayal and Perception of Race-Coded AI Companions

链接: https://arxiv.org/abs/2609.20637
作者: Wang Claire,Jiayue Melissa Shi,Agam Goyal,Grace Sletten,Renwen Zhang,Eshwar Chandrasekharan,Koustuv Saha
类目: Human-Computer Interaction (cs.HC)
备注: 27 pages, 2 figures, 7 tables

点击查看摘要

Abstract:AI companions can purportedly adopt racial personas, raising questions about how they represent identity and how users interpret these portrayals. We combined an algorithmic audit of race-coded AI personas with interviews with 12 companion users who interacted with a probe. Our audit revealed systematic differences, such as Asian-coded male personas receiving higher submissiveness scores than White counterparts, and Black, Hispanic, and Indigenous male personas receiving higher aggression scores than their White counterparts in open-weight models. Interviews revealed that participants envisioned AI companions as offering cultural familiarity and outside perspectives, but differed in which portrayals they considered meaningful or stereotypical. Some rejected overt racial signaling while still expecting culturally distinctive responses. Triangulating these findings with theory, we highlight how social norms and cultural expectations complicate efforts to support meaningful racial representation without reproducing stereotypes. We discuss how companion personalization should be evaluated beyond user satisfaction to account for broader representational harms.

[HC-6] amCAMS: An Open-Source Research Platform for Studying Human Behaviour in Human-AI Teams

链接: https://arxiv.org/abs/2609.20490
作者: Amos Brocco,Alain Chavaillaz,Andreas Sonderegger,Juergen Sauer
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:In this article, we present TeamCAMS (Cabin Air Management System), a collaborative work environment for simulating human-AI (artificial intelligence) interaction for scientific research. The article outlines how several psychological theories guided the development of this multiple-task simulation. Modelling a process control environment, previous versions of TeamCAMS have already been used in empirical studies to address a wide range of research questions (e.g., comparing different forms of automation, evaluating impact of automation reliability, effects of stress on multiple-task performance). Outlining the technical possibilities offered by TeamCAMS, the article points out how its latest version offers researchers the possibility of addressing a set of new research questions including problems associated with teamwork (e.g., within-team conflict, distributed teamwork) and human-AI interaction. Finally, we will outline how this simulation environment can be enhanced further still to address research questions in new fields (e.g., automation of leadership). To promote transparency, reproducibility, and further development, TeamCAMS is made available to the research community under an open-source license.

[HC-7] greCAPTCHA: Assessing Understanding as Evidence of Research Authorship Under Generative AI

链接: https://arxiv.org/abs/2609.20481
作者: Justin Payan,Bálint Gyevnár,Atoosa Kasirzadeh,Nihar B. Shah
类目: Digital Libraries (cs.DL); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Conferences, journals, funders, schools, and universities are struggling with a surge of potentially AI-generated submissions from ostensibly human authors, who may not have exercised sufficient human oversight for their manuscripts. In turn, institutions evaluating submissions can no longer reliably credit expertise based solely on authors’ names on submitted work. To address this problem, we propose greCAPTCHA, a proctored assessment approach that measures authors’ understanding of research manuscripts via the construct of capacity to verify, which we define as the knowledge and reasoning required to critically assess the contents underlying one’s contributions to a manuscript. greCAPTCHA generates questions assessing multiple levels of understanding and provides an evaluative report based on authors’ responses. Using a prototype implementation, we conduct a user study and semi-structured interviews with 31 researchers to evaluate greCAPTCHA. Its automated scores predict which papers were or were not authored by study participants with an AUC of 0.90 . Participants reported positive overall experiences with the system and remarked on the appropriate construct validity for author understanding, while also suggesting important changes to be made before deployment. Our results provide initial evidence that greCAPTCHA can assess manuscript-specific understanding under proctored conditions.

[HC-8] How Do We Visualize Space in Molecular Biology? A Study of Spatial Transcriptomics Visualization Practices

链接: https://arxiv.org/abs/2609.20324
作者: Denisse Chacón-Ramírez,Mark S. Keller,Eric Mörth,Nils Gehlenborg,Marc Streit,Andreas Hinterreiter
类目: Human-Computer Interaction (cs.HC)
备注: Submitted to IEEE Transactions on Visualization and Computer Graphics

点击查看摘要

Abstract:A cell’s identity depends on where it sits in tissue: for example, a macrophage behaves differently in a tumor core than at its edge. Spatial transcriptomics has transformed how we study this by recovering that lost coordinate, but it does so by producing data that is simultaneously high-dimensional, multimodal, and uncertain. Visualizing this combination is a hard problem in its own right, and one that warrants an assessment of how the field currently represents it, what has worked, and what is still missing. We surveyed 148 papers and 1,824 figure panels using a What-Why-How coding framework grounded in Munzner’s nested model, connecting the data represented, the biological tasks motivating each visualization, and the design choices through which they are expressed; a subset of the surveyed work also contributed dedicated interactive visualization software that was not necessarily reflected in the static figures, and we looked at what interaction capabilities those tools supported as well. We close by outlining where the field stands and the challenges ahead for bioinformatics and visualization researchers to tackle together.

[HC-9] Personalising a Cross-User Surface Electromyography Encoder Under a Small Calibration Budget

链接: https://arxiv.org/abs/2609.20296
作者: Jethro Odeyemi,W. J. Zhang
类目: Machine Learning (cs.LG); Human-Computer Interaction (cs.HC)
备注: 28 pages, 7 figures, 5 tables

点击查看摘要

Abstract:A myoelectric interface needs calibration from the user before it will function. Earlier work has treated calibration as a quantity, but has not asked the question of what a device should do with the calibration repetitions once they have been collected. This paper views personalizing the cross-user encoder as a design decision with a cost. Four alternative approaches to using exactly the same labeled repetitions were tested from a single cross-user encoder per held-out subject. Prototypical adaptation, linear probes, scaled fine-tuning and full fine-tuning were tested at every budget up to the maximum each database allows, five repetitions on DB1 and four on DB2 and DB5. Comparing four ways to spend a small calibration budget across 77 subjects, full fine-tuning is the most accurate at every budget, consistently enough that there is no exception among subsets of subjects. The result which impacts how one might make an engineering decision however is that a gradient free prototypical rule recovers 52 to 78 per cent of its benefit with no optimiser and no per-user copy of the weights, which makes personalisation something a worn device can do at donning time. The widespread intuition that a good representation only needs a fresh classifier is incorrect here. How well each method may perform relative to a per-user classifier that would be fitted by a clinic will depend on the specific database.

[HC-10] A Multi-Objective Optimisation Framework for Corticomuscular EEG-EMG Pair Selection in Hybrid BCI

链接: https://arxiv.org/abs/2609.20275
作者: Dekka Muni Kumar,Yogesh Kumar Meena
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Emerging Technologies (cs.ET)
备注: 6 pages, 5 figures, 1 table, accepted at Brain-Machine Interface (BMI) Systems Session, IEEE International Conference on Systems, Man, and Cybernetics (IEEE SMC 2026)

点击查看摘要

Abstract:Hybrid brain-computer interface (BCI) systems that integrate electroencephalography (EEG) and electromyography (EMG) signals have shown significant potential in improving the reliability of motor imagery (MI) classification, particularly in neuro-rehabilitation applications. However, identifying informative EEG-EMG channel pairs that effectively capture corticomuscular interactions remains a challenging problem, as existing approaches typically rely on manually predefined channel combinations that may not generalise across subjects. In this work, a data-driven EEG-EMG pair selection framework is proposed, in which channel pair selection is formulated as a constrained bi-objective optimisation problem. The proposed method jointly maximises the spatial relevance of EEG channels with respect to motor cortex regions and the corticomuscular coupling strength between EEG and EMG signals, and is solved using the NSGA-II to automatically identify an optimal subset of pairs. To extract discriminative features, the correlation between band-power time features capturing EEG-EMG interaction is combined with ERD-based EEG features, and a sliding-window-based temporal analysis is employed to account for the dynamic nature of MI signals. The proposed framework is evaluated on MI data from eight stroke patients and achieves an average classification accuracy of 89.6%, demonstrating its effectiveness in capturing physiologically meaningful corticomuscular interactions and improving classification performance.

[HC-11] A Hybrid Gaze-Motor Imagery BCI Framework for Effective Decision Communication

链接: https://arxiv.org/abs/2609.20273
作者: Gowtham Reddy N,KongFatt Wong-Lin,Yogesh Kumar Meena
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Emerging Technologies (cs.ET)
备注: 6 pages, 5 figures, selected to be presented at Brain-Machine Interface (BMI) Systems Session, IEEE International Conference on Systems, Man, and Cybernetics (IEEE SMC 2026)

点击查看摘要

Abstract:Non-invasive brain-computer interfaces (BCIs) and eye-tracking technologies offer promising communication pathways; however, motor imagery (MI)-based BCIs often suffer from low discriminability and high inter-subject variability. To mitigate these issues, this study investigates the impact of visual fixation on neural response stability in both standalone MI and hybrid MI-eye tracking systems. We then propose a novel asynchronous hybrid paradigm that streamlines user intent by utilising eye-tracking for direct selection, followed by MI-based confirmation, significantly reducing the operational steps required by conventional systems. The paradigm was evaluated with 15 healthy participants using a 16-channel EEG system. Results show that MI-related information is predominantly localised within motor cortex regions, with limited-channel configurations (SVM: 0.58) achieving performance comparable to full-montage setups (SVM: 0.54). The hybrid MI paradigm further outperforms conventional MI, achieving up to 100% accuracy with greater robustness across all channel configurations. Our findings indicate that visual fixation enhances neural response stability, while integrating eye-tracking with MI enables the development of reliable, scalable multi-command BCI systems suitable for real-world applications.

[HC-12] owards a Unified Modality-Agnostic Multimodal Framework for Cognitive Workload Assessment

链接: https://arxiv.org/abs/2609.20199
作者: Stefanos Gkikas,Christian Arzate Cruz,Calvin Joseph,Giorgos Giannakakis,Raul Fernandez Rojas
类目: Machine Learning (cs.LG); Human-Computer Interaction (cs.HC)
备注: Accepted at the 14th International Conference on Affective Computing and Intelligent Interaction (ACII 2026)

点击查看摘要

Abstract:Cognitive workload reflects the mental effort required during task performance and is central to the design of adaptive human-machine systems. The use of biosignals to measure cognitive workload has been extensively researched and documented; however, studies examining the effects of combining heterogeneous biosignal modalities for this purpose remain limited. To provide insight into this area, we developed a unified, modality-agnostic, hierarchical Transformer-based architecture to process heterogeneous biosignal modalities within a single model. We use this framework in a pilot study evaluating all 31 possible combinations of five modalities: Electrocardiogram (ECG), Electrodermal Activity (EDA), Respiration (RESP), Peripheral Oxygen Saturation (SpO _2 ), and Electroencephalogram (EEG), under leave-one-subject-out validation across three cognitively distinct tasks: abstract reasoning (IQ), arithmetic problem solving (MATH), and a game task (GAME). In this pilot setting, the results suggest that: (i) EEG is the strongest single modality, ranking highest in IQ, GAME, and the pooled ALL setting, where samples from all three tasks are combined; (ii) adding more modalities does not consistently improve performance; (iii) the full five-modality combination achieves the highest \textitAverage score of 73.02% on IQ and 68.08% when the \textitAverage scores are averaged over the four evaluation settings: IQ, MATH, GAME, and ALL; and (iv) the proposed method reduces model size by approximately 50% compared with late-fusion alternatives while maintaining a lower inference time.

[HC-13] ResumeShield: Channel Separation and an Open Benchmark for Indirect Prompt Injection in AI Resume Screening

链接: https://arxiv.org/abs/2609.20188
作者: Jay Barach
类目: Cryptography and Security (cs.CR); Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)
备注: 10 pages, 3 figures, 7 tables

点击查看摘要

Abstract:An AI resume screener reads a document supplied by the person it is evaluating, inverting the usual trust relationship between an assessor and the material it assesses. Candidates exploit this by concealing instructions inside a resume using white text, zero font size, hidden elements, markup comments, document metadata, or zero width characters. A human reviewer sees nothing, while a naive extraction pipeline places the concealed text into the model prompt, where it is read as an instruction and obeyed. This is indirect prompt injection, listed as LLM01:2025 by OWASP, and recent measurement work reports it in roughly one percent of resumes in a production screening corpus. We present ResumeShield, an open-source defense and benchmark. The defense combines three filtering stages with a fourth architectural stage that places candidate content in an explicitly fenced data channel that the operator’s trusted instructions declare inert. The benchmark builds a seeded synthetic corpus spanning nine concealment techniques and two payload families, one using documented phrasings and one modeling an adaptive attacker who paraphrases around the filter, and it scores an attack as successful only when the screening outcome changes. On a corpus of 104 documents, the naive pipeline is manipulated in every injected case while the defended pipeline is never manipulated. Detection reaches a precision of 1.000 and a recall of 0.944 with no false positives on clean resumes. An ablation shows that channel separation alone removes all measured attack success, whereas the complete filtering stack without separation still leaves 16.7 percent of attacks effective. We also identify a concealment dilemma: every payload that evaded detection was one the attacker left visible, surrendering the invisibility that motivates the attack. ResumeShield is released under the Apache 2.0 license with synthetic data only.

[HC-14] Utilizing AI-Driven Project Management Tools for Optimized Talent Management in HRM: A Framework for Enhanced Resource Allocation and Performance Prediction

链接: https://arxiv.org/abs/2609.20167
作者: Jay Barach
类目: Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)
备注: 9 pages, 2 figures and 3 tables

点击查看摘要

Abstract:When it comes to aiding businesses with demanding tasks regarding human resource management, TalentOptima unequivocally boasts of the best there is to offer. This tool utilizes AI based decision making, advanced predictive analytics, and also machine learning, all of which help in enabling automated resource allocation. To aid with better human resource management, TalentOptima integrates perfectly with already existing HR frameworks such as tools, etc. and shifts the focus towards aiding the user with insights while simultaneously alleviating manual work, this aids in a plethora of positive HR outcomes. A total of 40 managers participated in a simulation via user testing to ascertain if HR costs would reduce and work productivity would rise, the results were quite clear, attrition rates had dipped alongside risk and resource management rates, TalentOptima was a clear winner. Whereas the other HR frameworks primarily focused on ensuring work was done, TalentOptima ensured optimal and innovative decision-making, which overtime has proven to be invaluable for multiple companies, these results aid in proving why the tool is revolutionary.

[HC-15] Designing Against Deskilling: Metacognitive Feedback Reduces Cognitive Offloading to LLM Assistants

链接: https://arxiv.org/abs/2609.20143
作者: Sebastian Maier,Kai Schwabe,Manuel Schneider,Stefan Feuerriegel
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Cognitive offloading to AI can reduce opportunities to practice skills, creating risks of deskilling. However, it remains unclear how to prevent deskilling without restricting access to AI. Here, we design two interventions to reduce offloading decisions: (1) metacognitive feedback that makes the implications of offloading for users explicit, and (2) an effort-based reward that incentivizes less extensive LLM assistance. We test both in a preregistered online experiment ( N = 704 ) with a 2 \times 2 design and a no-AI control. The task was to practice fraction arithmetic with an LLM-based assistant that provided solutions only on explicit request, followed by an unaided test. Metacognitive feedback reduced answer offloading (OR = 0.47 ) and improved test performance (OR = 1.51 ). We found no evidence that the reward affected either outcome. Our results identify metacognitive feedback as a promising design choice to reduce cognitive offloading.

[HC-16] AI Should Facilitate Democratic Deliberation at Scale ICML2026

链接: https://arxiv.org/abs/2609.20059
作者: José Ramón Enríquez,Jiaxin Pei,Alex Pentland
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY)
备注: 15 pages, 2 figures, ICML 2026

点击查看摘要

Abstract:AI systems can strengthen democracy by supporting deliberation at scale by addressing cognitive, social, platform-design, and market-driven frictions, while preserving human agency. Unlike proposals such as liquid democracy that restructure representation through vote delegation, in this position paper, we argue that AI-assisted deliberation offers a more promising path by lowering barriers to meaningful engagement without substituting machine judgment for human choice. Drawing on evidence from online deliberation platforms and experimental research, we identify four guiding principles: preserving agency and autonomy, encouraging mutual respect, promoting equality and inclusiveness, and augmenting rather than substituting active citizenship. We also address critical challenges, including alignment, sycophancy, training bias, and over-reliance on AI systems. We call on the machine learning community to develop deliberation-focused AI systems evaluated not on engagement metrics but on their capacity to facilitate informed, representative, and friction-robust discourse.

[HC-17] What People Almost Did: Evaluating LLM Social Simulations Beyond Behavioral Fit

链接: https://arxiv.org/abs/2609.20055
作者: JaeWon Kim,Angie Boggust
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:LLM-based social simulations are primarily evaluated for behavioral fit, testing whether agents reproduce the actions or response distributions of the people they are simulating. However, the promise of simulation extends beyond behavioral fit. Simulations can explain human behavior, diagnose barriers, and compare large-scale interventions. These use cases depend on understanding \textitwhy people acted a certain way, not just \textitwhat they did. As a result, behavioral fit is insufficient for these types of claims because behavior underdetermines the reasoning process behind it. For instance, the behavior of staying silent may be due to disinterest or suppressed speech, and not answering a call may be due to distrust of the caller or limited phone access. In this paper, we propose \textitrepresentational adequacy as a new evaluation target for LLM-based social simulations. By leveraging LLM reasoning traces, representational adequacy measures whether a simulation’s scenario–reasoning–action triples preserve the reasoning process behind the behavior in a way that is faithful to the population and scenarios being simulated. We distinguish representational adequacy from interpretability and alignment metrics, propose ways to integrate it into simulation research, and pose its measurement as an open problem.

[HC-18] Penquiry: A Pen-based Interactive In-situ QA System Leverag ing LLM s

链接: https://arxiv.org/abs/2609.19870
作者: Jeongmin Rhee,Changhee Lee,Hyunwoo Kim,Kiroong Choe,Bohyoung Kim,Sungahn Ko,Jinwook Seo
类目: Human-Computer Interaction (cs.HC)
备注: 19 pages, 12 figures, Accepted to ACM UIST 2026

点击查看摘要

Abstract:Pen-based digital devices remain a preferred medium for active, cognitively engaging study. Concurrently, Large Language Models (LLMs) have become indispensable for self-directed learning, enabling students to clarify concepts. However, a fundamental interaction gap exists between the fluid, spatial nature of pen-based workflows and the discrete, keyboard-heavy requirements of LLMs. We present Penquiry, an in-situ question-and-answer system that bridges this gap by enabling learners to pose questions directly on digital study materials via a pen. We characterize two primary interaction challenges in this multimodal transition: a Referential Barrier, which hinders grounding fine-grained visual elements into the query context, and an Expressive Barrier, which forces learners to translate diverse, non-textual intents–such as equations and diagrams–into rigid, typed sentences. To resolve these, Penquiry introduces a mediation layer featuring Content Snapping for unambiguous referencing and Question Autocompletion to expand sparse ink keywords into rich semantic queries. Through two iterative user studies (N = 16 per study), we demonstrate that Penquiry significantly reduces the cognitive and physical overhead of inquiry compared to traditional interfaces, providing a new blueprint for pen-based, in-situ AI interaction

[HC-19] Point Revise Review: Grounded Agent ic Analysis in Reactive Notebooks with marimo-lens IEEE-VIS2026

链接: https://arxiv.org/abs/2609.19839
作者: Péter Ferenc Gyarmati,Trevor Manz,Dominik Moritz,Mennatallah El-Assady
类目: Human-Computer Interaction (cs.HC)
备注: 5 pages, 4 figures. Accepted for presentation at the VISxGenAI Workshop, IEEE VIS 2026. Source code: this https URL

点击查看摘要

Abstract:In a computational notebook, a human can point to a rendered result and ask about “this,” while an agent acts through cells, dependencies, and runtime state. When what the human sees and what the agent operates on are disconnected, the human must describe what they mean and trace what the agent did in prose that strips away situated visual and computational context. We present marimo-lens, an extension to the reactive Python notebook marimo for grounded human-agent analysis. Lens connects a human’s marked output and note to its producing cell and relevant contributing computation, giving the agent computational context for the request. It surfaces agent activity, returns selected results to the notebook, and preserves the initiating selection for human review and reopening. We illustrate the lifecycle through an exploratory human-agent analysis of a real-world open dataset, showing how visually situated questions lead to computational inspection, notebook action, and returned evidence for human review.

[HC-20] Affective Shared Autonomy: Temporal Affect Dynamics and Subjective Evaluation in Bimanual Teleoperation Tasks

链接: https://arxiv.org/abs/2609.19802
作者: Zhengji Liang,Guiyin Tian,Sijin Qu,Hainan Liu,Shiyan Hu
类目: Robotics (cs.RO); Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Physical teleoperation integrates human cognitive flexibility with robotic precision, yet demanding manipulation tasks frequently induce severe cognitive workload, acute frustration, and execution breakdown. Conventional shared autonomy paradigms rely primarily on task-based rules, such as spatial error boundaries, which disregard the operator’s transient affective state and risk misaligned control interventions. To address this limitation, we propose an affect-aware shared autonomy teleoperation framework that dynamically modulates robotic assistance based on real-time operator state estimation. The system estimates operator affective states from synchronized facial video, cardiac signals, and bilateral arm kinematics, outputting a seven-state affective distribution and a three-category operational abstraction (neutral, productive, adverse). Affect-aware assistance is selectively triggered when the user is detected in a continuous adverse state, preserving task-positive engagement without unnecessary disruption. The empirical user study ( N = 30 ) confirms that the proposed affective assistance increases the productive states by up to 39.7% without compromising user agency. The collected dataset represents the first multimodal dataset that provides continuous visual, physiological, and operator’s bilateral motion tracking of temporal affective state shifts during bimanual teleoperation. Our multimodal fusion model outperforms zero-shot baselines (Qwen, MiniCPM-V) in tracking temporal state dynamics. This real-world deployment offers a new human-centric framework that integrates visual, physiological, and motion tracking for physical human-robot interaction.

[HC-21] readstone: A Social-Media-Inspired Platform for Multi-Agent Collaborative Data Analysis

链接: https://arxiv.org/abs/2609.19774
作者: Hyunwook Lee,Sungbeom Cho,William Benjamin,Changhee Lee,Hyotaek Jeon,Daeun Jeong,Sungbok Shin,Sungahn Ko,Niklas Elmqvist
类目: Human-Computer Interaction (cs.HC)
备注: 11 pages, 10 figures, 2 tables, 2 page-length appendix. Accepted to VIS 2026 and will be published on IEEE Transactions on Visualization and Computer Graphics (TVCG)

点击查看摘要

Abstract:Coordinating human analysts with autonomous AI agents faces the same challenges as human-to-human collaboration: sharing intermediate results, avoiding conflicts, and maintaining group awareness. Current tools rely on unstructured messaging or single-threaded chatbot interaction, which lack the structure to track evolving hypotheses or link claims to evidence. We propose agentic social data analysis, a collaboration paradigm extending social data analysis with a shared coordination feed modeled on the content timeline in social media services. We instantiate this concept in TREADSTONE, a platform where human and AI agents asynchronously post, link, and contest analytical claims via threaded messages within a shared feed. By allowing agents to proactively broadcast hypotheses and enabling users to steer the analysis through lightweight curation, Treadstone seeks to balance machine autonomy with human analytical control. A qualitative user study shows that Treadstone fosters collaboration while preserving human analytical agency, in contrast to the solitary experience of conventional chatbot interaction.

[HC-22] DELUGE: Decomposed Entropy-coded Live Unstructured Geometry Exchange for Real-time Particle Streaming

链接: https://arxiv.org/abs/2609.19750
作者: Hikari Yanagawa,Yuichi Hiroi,Takefumi Hiraki
类目: Graphics (cs.GR); Human-Computer Interaction (cs.HC); Information Theory (cs.IT)
备注: 11 pages + supplement 3 pages. Accepted for publication in IEEE Transactions on Visualization and Computer Graphics (TVCG), to be presented at IEEE ISMAR 2026

点击查看摘要

Abstract:Particle-based physics simulations, including fluids, smoke, and granular media, are fundamental to visual realism in immersive VR and AR. With the growing adoption of social VR and digital twins, demand is increasing for shared experiences in which multiple users interact with the same simulation in real time. Realizing such experiences requires low-latency streaming of large-scale particle data from a server to each client, yet existing point cloud compression methods such as G-PCC (TMC13) and Draco assume static geometric structures; when applied to dynamic particle streaming, their encoding latency exceeds the frame period, failing to meet real-time delivery requirements. We propose DELUGE, a streaming compression architecture that exploits the temporal coherence and velocity predictability inherent in physics simulation particles, achieving sub-frame-latency encoding and decoding through three complementary techniques. Evaluation on dynamic point cloud datasets demonstrates that DELUGE achieves approximately 20\times faster decoding than G-PCC (TMC13) and 6\times faster encoding than Draco. We further build end-to-end client implementations for both web browsers and Apple Vision Pro, and confirm through a within-participants perceptual quality evaluation and a two-person collaborative task study on Vision Pro that the proposed method supports real-time collaborative experiences with hand-tracked fluid interaction.

[HC-23] Faithful Where It Can Be Checked: Auditing a Reflection Agent Against Its System Prompt in a Randomized Trial

链接: https://arxiv.org/abs/2609.19635
作者: Subigya K. Nepal,Serena Soh,Noah Vinoya,SoHyun Park,Mahnaz Roshanaei,Gabriella Harari
类目: Human-Computer Interaction (cs.HC); Computers and Society (cs.CY)
备注:

点击查看摘要

Abstract:Conversational agents are increasingly used to guide reflection. A recent randomized trial compared a GPT-4o career reflection agent with the same program in a static journaling survey. Agent participants ended less committed to their career plans and more doubtful. We coded all 17,930 turns from its two studies, checked our coding against human coders and linked conversations to the trial’s surveys. The rules the agent followed were the easy-to-check ones, like a reply length cap. Told not to flatter, it praised participants in half of its turns; told to challenge gently, it almost never did, and such a break leaves no visible trace. The behavior tied to the worse outcome was the demand to decide: the survey posed each decision once, while the agent asked again when participants hesitated, and those pressed most ended most doubtful. Our findings inform reflection agent design and the writing of checkable instructions.

[HC-24] DataCanvas-EDU: An Agent ic Framework for Instructor-Guided Synthetic Data Generation in Business Analytics Education

链接: https://arxiv.org/abs/2609.19617
作者: Bang An,Maria Hamdani,Joseph Fox
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI)
备注: Synthetic Dataset, Agentic Framework, Data Analytics Education, Data Visualization

点击查看摘要

Abstract:Business analytics education requires diverse datasets to support different learning objectives, student backgrounds, and analytical tasks. Real-world data can be difficult to obtain and offer limited flexibility for adapting a case to a particular course. Even when suitable data are available, instructors must investigate the patterns, verify the results, and prepare assignments and reference solutions, requiring substantial time and effort. The use of large language models (LLMs) introduces an additional concern about training data contamination. Widely used public datasets often have extensive tutorials and worked analyses that models may have encountered during training. Students may therefore receive explanations drawn from existing analyses without practicing how to investigate unfamiliar data in collaboration with AI. This paper presents DataCanvas-EDU, an agentic framework for instructor-guided synthetic data generation in business analytics education. Instructors specify teaching goals and intended patterns through conversation, while an AI agent writes generation code, checks the resulting data, and prepares assignments, reference analyses, and rubrics. Four phases, Plan, Create, Verify / Test Analysis, and Evaluate, organize the process and support instructor review and revision. The framework is intended to simplify case preparation while creating opportunities for students to investigate newly designed patterns with AI. We illustrate the approach with WindowDash, a food delivery case containing 15,000 orders and nine designed patterns. DataCanvas-EDU is packaged as a reusable AI Agent Skill for compatible agent environments, with the package and installation instructions available at this https URL

[HC-25] Full-Duplex Speech Models Take the Floor When Asked Not When Needed

链接: https://arxiv.org/abs/2609.19596
作者: Linkai Peng,Baorian Nuchged,Kaiqi Fu,Yuyang Yao
类目: Computation and Language (cs.CL); Human-Computer Interaction (cs.HC)
备注: 5 pages

点击查看摘要

Abstract:Full-duplex speech models listen and speak at once, promising always-on assistants. Yet they must also decide when they should speak. Human listeners speak when addressed or when the speaker stops, but also self-select to correct a false claim, supply a missing word, or warn of danger. We ask whether full-duplex models do the same. To separate the reason to speak from the opportunity, we construct context-matched English monologues in which only the trigger utterance varies within a topic, define 10 conditions from turn-allocation rules, and compress inter-word pauses to limit opportunities created by silence. Across five model families, being addressed and silence are far more reliable triggers than false facts or hazards. Frame-level text-token probabilities in Moshi and PersonaPlex are lower for false facts than for Neutral when averaged over the first 2,s after trigger end. Pauses or permission to interrupt do not close this gap either. Given the floor, Moshi and PersonaPlex answer most direct questions, yet the proportion of non-empty false-fact replies that challenge the claim is only .14–.15, and the proportion of hazard replies that warn of danger is .04–.07. This paper thus identifies a gap in both speech initiation and response content. Closing it requires genuine content understanding and intervention decisions grounded in it.

[HC-26] Value Faces: Surfacing How Self-Presentation Shifts Across Relationships

链接: https://arxiv.org/abs/2609.19581
作者: Gabriel Koo,Rayhan Rashed,Farnaz Jahanbakhsh
类目: Human-Computer Interaction (cs.HC); Computers and Society (cs.CY)
备注:

点击查看摘要

Abstract:People present different aspects of themselves across relationships. Computational work has captured such variation in communication style. But this variation also extends to which principles people foreground or background in a particular relationship–i.e., in the values they express and how they balance them. We conceptualize these relationship-specific expressions of values as demonstrated values. To make demonstrated values visible, we introduce Value Faces, a system that analyzes a person’s existing chat histories from their everyday messaging platforms using Schwartz’s ten basic human values and produces separate value profiles for their different relationships. In a mixed-methods study(N=18), we find that the resulting value profiles distinguished participants’ relational contexts with twice the odds of guessing, while system-inferred differences across relationships aligned with participants’ perceptions of those differences. Participants used these profiles to articulate previously implicit differences in how they presented themselves, connect them to roles and changes over time, and reconsider their self-assessments.

[HC-27] Learning from Success and Failure: Acquiring Adaptive Dialogue Strategies for Social Robots

链接: https://arxiv.org/abs/2609.19570
作者: Sanae Yamashita,Yuki Okafuji
类目: Human-Computer Interaction (cs.HC)
备注: 8 pages, 3 figures

点击查看摘要

Abstract:Traditional dialogue systems for social robots require both dialogue strategies and user attribute recognition, each demanding specialized expertise. However, data collection is costly in real-world deployments, and the resulting datasets often include many failure cases. In this study, we aim to automate the acquisition of dialogue strategies by leveraging both successful and failed interactions using a vision-language model (VLM) and a large language model (LLM). We propose an architecture in which user attributes, recognized by the VLM, along with dialogue history, are fed into the LLM to generate dialogue strategies tailored to specific user attributes. We extracted dialogue strategies from an interaction dataset collected through a field experiment and evaluated their effectiveness. The results demonstrate that explicitly representing failure strategies complements success strategies and improves performance. Our findings highlight a practical pipeline for constructing and maintaining an interpretable strategy repository from in-the-wild deployment logs by recycling abundant failure interactions as reusable constraints, ultimately reducing the development cost of social robots.

[HC-28] Cyber Exodus: Burnout Symptoms Exit Intention and Peer Response in Online Cybersecurity Communities

链接: https://arxiv.org/abs/2609.19556
作者: Nadia Mehjabin,Ji Hyun Kim,Laura Barnes,Koustuv Saha,Henry Kautz,Subigya Nepal
类目: Human-Computer Interaction (cs.HC); Cryptography and Security (cs.CR)
备注:

点击查看摘要

Abstract:Security practitioners burn out at high rates, and the resulting attrition is itself a security problem. This workforce is hard to study: security operations centers are closed to outside researchers, studies that reach practitioners recruit through employers, and those who have disengaged most may have the least reason to answer an employer’s survey. The same practitioners discuss their working conditions openly in online communities. We adapt the Burnout Assessment Tool, a validated clinical instrument, into a text annotation scheme and apply it to 354,861 posts and 296,442 replies from five online communities of cybersecurity practitioners. Checked against two trained coders on 100 posts, the annotation reaches a macro F1 of 0.75 across the four symptoms and 0.98 for detecting any burnout signal. We find that the four symptoms point to different problems at work, not to the same problem at different levels of severity. Exhaustion appears in almost any complaint about staffing or workload. Mental distance, a loss of belief that the work is worthwhile, is the only symptom unrelated to operational problems, and among posts with a single symptom it is accompanied by a stated intention to leave roughly twice as often as any other. Peer responses show the opposite pattern. When a poster says they are considering leaving, the mix of replies shifts toward career advice, but this shift is smallest for mental distance. The symptom most strongly associated with leaving is thus the one peers adjust to least, and a single burnout score obscures both patterns.

[HC-29] FenceXR: AR Movement Replay for Error-Detection Training and Spatially Grounded Feedback

链接: https://arxiv.org/abs/2609.19505
作者: Avinash Ajit Nargund,Amelia Haruka Harrison,Timothy Robinson,Tobias Höllerer,Misha Sra
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Recognizing technical errors in movement is a perceptual skill important to motor learning, but it is challenging for beginners in fast, complex sports like fencing to develop it. A coach’s attention is scarce, live demonstrations vary from repetition to repetition, and video review is limited to whatever camera angle was used to record it. Coaches and advanced fencers reviewing a recording face a related problem. They can see an error, but have no way to anchor their feedback to the movement itself, and are left describing it in words the learner must map back onto their own body. We present FenceXR, an augmented reality system that reconstructs 3D movement replays from monocular smartphone video to address both problems. A Trainee module trains novices to detect common lunge errors while a Reviewer module lets coaches and advanced fencers attach text or voice annotations to a specific joint and moment in a replay, which can be shared asynchronously with a trainee. In a study with 18 novice fencers, unaided error-detection accuracy rose from near-chance (37.5%) before training to 64.1% after a single session, with interviews showing a shift from broad visual scanning toward targeted inspection of specific joints and their timing. In a video-based study with four fencing experts, all four viewed the Reviewer module as a valuable complement to their existing coaching tools, particularly for feedback that is difficult to convey through standard video. We end with a discussion of implications for designing AR systems that ground movement-based training and feedback in the movement itself.

[HC-30] CARES: A Conversational AI System for Regulation-Grounded Safety Reporting in Construction Education

链接: https://arxiv.org/abs/2609.19429
作者: Fan Yang,Jiabin Wu,Yuan Tian,Jiansong Zhang
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Construction safety reporting often relies on manual logs and static templates that provide limited feedback and leave daily activities disconnected from relevant regulations. This paper introduces CARES (Conversational AI Reporting for Enhanced Safety), a conversational AI system that integrates regulatory guidance into daily reporting to support construction safety education. CARES combines proactive multi-agent dialogue, retrieval-augmented generation (RAG), and automated report generation. The system guides users through reporting tasks, retrieves relevant regulatory passages using hybrid retrieval, and converts conversations into structured daily reports. Regulatory sources and the evolving report are displayed alongside the dialogue to support user review and correction. A preliminary evaluation involving 15 construction management students assessed retrieval quality, response faithfulness, and conversational relevance. CARES achieved an overall faithfulness score of 0.74 and an answer relevance score of 1.00, while initial retrieval ranking remained an area for improvement. These results provide preliminary evidence of the technical feasibility of regulation-grounded conversational reporting. The study highlights opportunities to integrate regulatory knowledge into routine documentation, with future work needed to evaluate effects on report quality, safety awareness, and learning outcomes.

[HC-31] Use and Effects of LLM s in Peer Review: A Randomized Experiment and Survey at ICML 2026

链接: https://arxiv.org/abs/2609.19420
作者: Sunnie S. Y. Kim,Wesley Hanwen Deng,Jennifer Wortman Vaughan,Buxin Su,Weijie Su,Alekh Agarwal,Sharon Li,Martin Jaggi,Daniel G. Goldstein,Nihar B. Shah,Miroslav Dudík
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:LLMs are rapidly reshaping peer review, making it important to understand how reviewers use them in practice and how different LLM-use policies affect review outcomes. We investigate these questions through a randomized experiment and an anonymous post-survey at ICML 2026, a major machine learning conference involving over 24,000 papers and 17,000 reviewers. Reviewers were assigned to either a conservative policy prohibiting all LLM use or a permissive policy allowing limited assistance, with randomization among a subset of main-track papers and reviewers. Policy assignment had near-zero effects on final paper decisions, paper scores, and reviewer confidence, although reviews under the permissive policy were 5.5-7% longer. Post-survey responses (N=1,486) revealed diverse attitudes toward LLMs and substantial noncompliance: 22.5% of conservative-policy reviewers reported using an LLM despite the prohibition, and 36.5% of permissive-policy reviewers reported at least one explicitly disallowed use. We discuss implications for future peer-review policy and tool design.

[HC-32] Durably Reducing Belief in Womens Health Misinformation Through Culturally Adaptive AI Videos

链接: https://arxiv.org/abs/2609.19364
作者: Anku Rani,Kokil Jaidka,Shruti Sharma,Pragya Mahajan,Manisha Wadhwa,Andrew B. Lippman,Pattie Maes,Paul Pu Liang
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Health misinformation disproportionately harms women, yet interventions rarely address the community norms that sustain false beliefs. We test whether culturally adaptive AI-generated video in which the presenter looks like someone from her community reduces misinformation belief among low-literacy women in suburban India. In a field experiment (N=434), participants watched an AI-generated video featuring either an adaptive or neutral presenter. The culturally adaptive presenter reduced misinformation belief by 30%, nearly twice the reduction produced by the neutral presenter compared to the non-intervention control condition. Post-experiment interviews suggest women recalled the neutral condition as a generic video but recognized the adaptive presenter. Gains persisted for three weeks. The adaptive advantage was largest for beliefs reinforced by community, such as blaming women for infertility, and negligible for medical knowledge gaps, such as understanding vaccines. These findings demonstrate the potential of culturally adaptive AI interventions to counter socially embedded health misinformation.

[HC-33] “I Know Where to Look” But Does the LLM ? Charting the Gaps Between Clinical Expert Needs and Unstructured Data Abstraction Tools

链接: https://arxiv.org/abs/2609.19318
作者: Venkatesh Sivaraman,Rigney Turnham,George Bonano,Nevin Aresh,Renumathy Dhanasekaran,Margaret Guo,Sindhu Kubendran,Olivia Lin,Jonathan D Louie,Kristan Olazo,Jeanne Shen,Harish Vasudevan,Jeanette Wong,Emily Alsentzer,Jason A Fries,Anobel Odisho,John Gordan,Jean Feng,Julian C Hong
类目: Human-Computer Interaction (cs.HC); Other Quantitative Biology (q-bio.OT)
备注: Under review

点击查看摘要

Abstract:Clinical data abstraction, the process of distilling structured information from patient records, plays a key role in advancing knowledge about diseases such as cancer. Information extraction (IE) with large language models (LLMs) could accelerate this process, but it is unclear whether current frameworks effectively support clinical researchers without AI expertise. To address this, we co-designed an interactive LLM-based abstraction system called Libretto with seven cancer research teams, then evaluated the system’s ability to help them answer real-world research questions. We found that while clinicians knew where and how to annotate complex concepts in patient notes, in twelve of fourteen tasks they faced barriers to replicating those intuitions with LLMs. Contextual note reliability judgments, difficulties in steering vibe-coded prompts, and inflexible evaluation strategies necessitated fundamental changes to the IE workflow. Our results highlight open problems for HCI research to bridge the gaps between AI data work tools and clinical users’ needs.

[HC-34] OHRID-Retail: An Open Multimodal Dataset of Human Activity in Retail Environments

链接: https://arxiv.org/abs/2609.19302
作者: Xiangrui Wang,Yuetong Wu,Jalen Beeman,Robert Cook,Yu Gu,Nathanial Pearson,Trevor Smith,Read Hayes,Boyi Hu
类目: Robotics (cs.RO); Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Open datasets describing human behavior in environments shared with mobile robots remain limited, particularly for retail activities that combine locomotion, reaching, object handling, and robot guided movement. This paper introduces OHRID Retail, an open, human centered multimodal dataset collected from 16 healthy adults performing a simulated shelf picking task under three within participant conditions: no robot, low speed robot guidance, and high speed robot guidance. Each participant completed two trials per condition. Whole body kinematics were recorded using 17 Xsens Awinda inertial sensors and muscle activity was measured at 10 locations using Delsys Trigno surface electromyography sensors. Descriptive analyses demonstrate variation in whole body movement intensity and muscle activation across robot interaction conditions and body locations. OHRID Retail provides openly available raw recordings, processed measures, documentation, and reproducible analysis resources. The dataset can support research in human activity recognition, multimodal sensor fusion, occupational biomechanics, ergonomics, human aware robot navigation, and human robot interaction in retail and related shared environments.

[HC-35] MuTable: Composable and Reusable Table Transformations for In-Situ Data Exploration

链接: https://arxiv.org/abs/2609.19294
作者: Fuling Sun,Devamardeep Hayatpur,Jane L. E,Nicole Sultanum,Haijun Xia
类目: Human-Computer Interaction (cs.HC)
备注: To be published in UIST 26

点击查看摘要

Abstract:Tables are central to data work to support precise lookup and full detail, but they can be limiting for overview and pattern-finding tasks. Visualizations are then created to gain richer perceptual support. In practice, moving between tables and charts often requires maintaining parallel representations, introducing context switching, and extra coordination work. Building on prior hybrid table-visualization systems, we present MuTable, a prototype that reifies transformations as persistent, composable, and reusable modifiers to support in-situ data exploration. Users can reshape the table while retaining and adapting intermediate forms as their questions evolve. An expert interview with eight data workers suggests that MuTable can support coordination between representations, rapid exploration, and greater user agency in constructing visualizations, as a low-commitment exploration space.

[HC-36] Bless his heart… he thought all we did was push a button": Understanding Worker Challenges with U.S. Election Technology

链接: https://arxiv.org/abs/2609.19233
作者: Delaney Gomen,Josiah Hester,Naveena Karusala,Michael Specter
类目: Human-Computer Interaction (cs.HC); Computers and Society (cs.CY)
备注:

点击查看摘要

Abstract:This paper explores how the technologies U.S. election officials depend on also challenge their work. Every election, misleading narratives stem from errors that occur while using election technology, threatening election official safety and eroding faith in elections. We seek to understand the impact of these challenges directly from worker perspectives. This paper presents a qualitative analysis using combined interview and survey data from a total of 50 election officials representing 20 states. Our findings present design as a substantial factor in mistakes when using election technology, but other obstacles, like lack of funding and technical support, also strain operations. We show how election technology design can lead to harms for election officials and connect those harms to their negative effects on democracy. Our work emphasizes the potential for researchers in human-centered computing to support U.S. election officials-and democracy-by working towards better usability across election technologies.

[HC-37] Digital Twins Need Feedback

链接: https://arxiv.org/abs/2606.23562
作者: Guo-Qiang Zhang
类目: Logic in Computer Science (cs.LO); Computational Engineering, Finance, and Science (cs.CE); Human-Computer Interaction (cs.HC); Social and Information Networks (cs.SI); Systems and Control (eess.SY)
备注:

点击查看摘要

Abstract:Digital twins are too often described as realistic simulations, anatomical avatars, dashboards, or data mirrors. Those artifacts can be useful, but they miss the defining property of a digital twin: bidirectional feedback between a physical counterpart and a virtual counterpart. The physical system continuously updates the virtual one; the virtual system informs actions that change measurement, intervention, operation, or governance in the physical world. We propose such a bidirectional feedback as the organizing principle for digital twins and apply it to a nested, multi-scale hierarchy of biological and social organization, in which lower-level units combine into higher-level systems, producing desirable properties at each level, from cells and tissues to organs, individuals, organizations, and population at large. Neuroinformatics is a stress test for this view because brain health, dementia, epilepsy, and other neurological diseases require the integration of cells, circuits, behavior, care pathways, and the translation of discovery to practice. Examples from epilepsy care and consortium-scale brain-cell atlas production show that digital twinning is not merely multi-scale modeling. It is a rich, multidisciplinary paradigm of computing for designing, governing, and driving feedback loops that turn data into accountable action.

[HC-38] Comparison and Analysis of Cognitive Load under 2D/3D Visual Stimuli

链接: https://arxiv.org/abs/2302.12968
作者: Yu Liu,Chen Song,Yunpeng Yin,Herui Shi,Jinglin Sun,Han Wang,Peiguang Jing
类目: Neurons and Cognition (q-bio.NC); Human-Computer Interaction (cs.HC)
备注: 9 pages, 13 figures

点击查看摘要

Abstract:With the increasing prevalence of 3D videos, investigating the differences of viewing experiences between 2D and 3D videos has become an important issue. In this study, we explored the cognitive load induced by 2D and 3D video stimuli under various cognitive tasks utilizing electroencephalogram (EEG) data. We also introduced the Cognitive Load Index (CLI), a metric which combines \theta and \alpha oscillations to evaluate the cognitive differences. Four video stimuli, each associated with typical cognitive tasks were adopted in our experiments. Subjects were exposed to both 2D and 3D video stimuli, and the corresponding EEG data were recorded. Then, we analyzed the power within the 0.5-45 Hz frequency of EEG data, and CLI was utilized to evaluate the brain activity of different subjects. According to our experiments and analysis, videos that involve simple observational tasks (P 0.05) consistently induced a higher cognitive load in subjects when they were viewing 3D videos. However, for videos that involve calculation tasks (P 0.05), the differences in cognitive load induced by 2D and 3D video were not obvious. Thus, we concluded that 3D videos could generally induce a higher cognitive load, but the extent of the differences also depended on the contents of the video stimuli and the viewing purpose.

计算机视觉

[CV-0] Can 4D Foundation Models Remember?

链接: https://arxiv.org/abs/2609.20819
作者: Guangzhao He,Hadar Averbuch-Elor,Wei-Chiu Ma
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project Page: this https URL

点击查看摘要

Abstract:Perceiving and remembering the visual world is fundamental to navigating and interacting with our environment. Current 4D foundation models, such as camera-controllable video models or 4D reconstruction models, can perceive and reconstruct dynamic environments, but how well they remember what they have perceived remains an open question. Existing benchmarks largely rely on pixel-level metrics and lack ground truth for objects once they leave the field of view, making them unable to evaluate visual memory in an object-centric manner against references. To fill this gap, we introduce PersistBench, a dataset and metric suite that leverages 360° videos as omniscient ground truth and proposes three evaluation aspects: object permanence, motion continuity, and appearance preservation. Evaluating various models across diverse categories reveals that current models can only maintain short-term consistency that degrades significantly once objects leave the field of view. Our findings highlight the gap between current model capabilities and robust visual memory (“seeing is not remembering”), providing guidance for future development of 4D foundation models. Dataset and code are available on the project page: this https URL.

[CV-1] SplashSplat: Reconstructing Splashing Liquids from Real-World Multi-View Videos

链接: https://arxiv.org/abs/2609.20818
作者: Peiyu Liu,Dingxi Zhang,Federico Tombari,Marc Pollefeys,Christina Tsalicoglou,Daniel Barath
类目: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
备注: 18 pages (11 main + 7 supplementary), 14 figures, 12 tables. Project page: this https URL

点击查看摘要

Abstract:A splash lives for a fraction of a second: sheets tear into ligaments and droplets, appearance is view-dependent and nearly textureless, and little persists long enough to track. Reconstruction research has consequently focused on smoke, synthetic liquids, or gently deforming surfaces. To our knowledge, no synchronized multi-view dataset of splashing liquids exists. We therefore introduce a benchmark of 20 real scenes, from coherent streams to violent splashes, captured by seven synchronized, calibrated 4K cameras at 60 fps, with manually refined per-view liquid and container masks and fixed evaluation splits. We further present SplashSplat, built on a single principle: impose physical structure only where the observations can constrain it. Per-frame liquid SDFs fused from the masks provide the geometry, level-set transport between consecutive SDFs yields a coarse velocity field, and Lagrangian carriers advected along this flow, corrected against each new observation and reseeded where coverage is lost, decode local Gaussians for differentiable rendering. SplashSplat outperforms state-of-the-art dynamic Gaussian splatting methods on our real captures and on a synthetic benchmark, with physically more plausible motion and a lower training cost. The same representation supports temporal interpolation and style transfer without re-optimization.

[CV-2] FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations

链接: https://arxiv.org/abs/2609.20817
作者: Kevin Qu,Tao Sun,Massimiliano Viola,Liyuan Zhu,Zhizhuo Zhou,Sayan Deb Sarkar,Konrad Schindler,Iro Armeni
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注: Project page: this https URL

点击查看摘要

Abstract:Modeling articulated objects from sparse monocular views is challenging because each observation reveals only partial geometry and motion evidence. Most feed-forward methods infer articulation from a single observation and therefore rely heavily on learned category-level shape priors. We present FAMOS, a feed-forward model that predicts movable-part segmentation and joint parameters from a sparse, unordered set of partial point clouds. Our model jointly reasons over multiple observations and naturally supports a variable number of inputs, including a single view. To aggregate articulation cues across observations, we introduce a Multi-state Articulation Transformer with alternating state-wise and global attention. We further propose an observed articulation span objective that supervises the motion range each part exhibits across the input observations, encouraging the model to leverage the full observation set. To overcome the limited scale and diversity of existing datasets, we introduce a procedural data generator that synthesizes self-annotated assets during training. Experiments on PartNet-Mobility, ACD, and ArtiCraft-10K demonstrate consistent improvements over both feed-forward and optimization-based baselines. Project page: this https URL

[CV-3] Paint-Anything: Unified Any-Color Control for Image Generation and Editing

链接: https://arxiv.org/abs/2609.20816
作者: Ji Xie,Dewei Zhou,Xinyu Huang,Zhennan Chen,Xun Wang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 29 pages, Seed Technical Report

点击查看摘要

Abstract:Professional design requires any-color control: the ability to specify an object’s target color with any 24-bit hex value for image generation and editing. Prior work has explored color generation, editing, and colorization, but often relies on dedicated color representations or specialized inference procedures. Advances in large language models offer a simpler starting point: even compact models can associate hex values with color semantics. We present Paint-Anything, which learns a shared hex-prompt interface for generation and editing through object-level color supervision. We develop a data pipeline that constructs Paint-500K from real images through object grounding, perceptual color labeling, and editing-pair synthesis. Since shadows make real-image labels only approximate colors, we complement this supervision with pure-color anchors whose pixels exactly match their paired hex values. These anchors are used only at high-noise timesteps, leaving low-noise training to natural images. We further introduce Any Color Benchmark (ACBench), comprising ACBench-T2I and ACBench-Edit, to measure object-level hex color fidelity across both tasks. On FLUX.2-4B, Paint-Anything improves ACBench-T2I and ACBench-Edit scores by 85.3% and 28.3%, respectively, relative to the base model, with ablations supporting the training recipe. It also achieves the highest average CompColor score among the compared methods.

[CV-4] ERCPMP-Gx: Endoscopic Image and Video Dataset for Morphological Histopathological and Genomic Characterization of Colorectal Polyposis

链接: https://arxiv.org/abs/2609.20815
作者: Zahra Ghaffari,Massih Bahar,Mojgan Forootan,Ali Darvishi,Hamidreza Bolhasani
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Hereditary polyposis syndromes can be precursor lesions to colorectal cancer and are associated with a broad spectrum of extracolonic tumors. Early identification and accurate classification of these syndromes are essential for timely diagnosis, individualized patient management, and targeted surveillance strategies for affected families. However, public endoscopic datasets are largely organized around the individual sporadic polyp, and none links the polyposis phenotype to histopathology and germline findings at the patient level. Here, we present ERCPMP-Gx, an endoscopic, histopathological, and genomic dataset developed to support the application of artificial intelligence (AI) in the recognition, characterization, and classification of colorectal polyposis. Most procedures were performed using the Olympus EVIS X1 system with white-light endoscopy (WLE), narrow-band imaging (NBI), magnifying NBI (M-NBI), and NBI with near focus modes, yielding 160 images and accompanying video clips. Approximately eighty percent of cases represent clinically and/or genetically confirmed hereditary polyposis syndromes (PG), including familial adenomatous polyposis (FAP), Peutz-Jeghers syndrome (PJS), juvenile polyposis syndrome (JPS), and ganglioneuroma syndrome (GNS), while the remaining twenty percent comprise non-hereditary polyps and polyp-mimicking lesions with overlapping morphological features (Non-PG), included to support differential classification. Each released record is linked, where available, to standardized endoscopic annotations, representative histopathology, and clinically reported germline findings, forming an AI-ready, patient-level annotation framework. The dataset is publicly accessible at Mendeley (this https URL). For the latest updates and further information, readers are referred to the DataBioX website: this https URL.

[CV-5] FlowSGS: Improving Flow Matching Priors for Inverse Imaging with Stochastic Interpolants

链接: https://arxiv.org/abs/2609.20769
作者: Tianao Li,Xinhui Qian,Emma Alexander
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Flow matching has emerged as the state-of-the-art generative model and has been used for plug-and-play (PnP) priors to solve inverse problems in computational imaging. However, existing flow-based inverse solvers assume linear forward models and/or make simplifying approximations in posterior sampling. To circumvent these problems, we introduce FlowSGS, a flow-based posterior sampling method using Split Gibbs Sampling (SGS) to decompose the posterior into a likelihood step and a prior step. Specifically, we sample from the likelihood step using Langevin dynamics and leverage the Stochastic Interpolants (SI) framework to integrate a pretrained flow model into the prior step. We provide a form for the prior step that uses SI’s reverse-time SDE, and show connections to previous PnP methods. Moreover, with the aid of the flow prior’s straight probability paths and a novel timestep correction technique for the reverse-time SDE, FlowSGS requires fewer network evaluations in its prior step than plug-and-play diffusion samplers. Our experiments show state-of-the-art performance on a range of inverse problems. For the first time, we provide an experiment on a nonlinear inverse problem (Fourier phase retrieval) for flow-based inverse solvers.

[CV-6] OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher

链接: https://arxiv.org/abs/2609.20756
作者: Damiano Da Col,Maximilian Igl,Peter Karkus,Kashyap Chitta,Boris Ivanovic,Marco Pavone,Konrad Schindler,Christos Sakaridis
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 9 pages, 5 figures

点击查看摘要

Abstract:As scaling pre-training data alone yields diminishing returns, post-training is becoming increasingly important across physical AI domains such as autonomous driving. End-to-end driving policies are pre-trained in open loop with behavior cloning on human demonstrations. However, compounding errors during closed-loop deployment can take the vehicle outside the training data distribution, increasing the risk of safety-critical incidents. Closed-loop post-training can mitigate this risk but requires costly simulation for sensor-based policies. We propose OPTED (on-policy fine-tuning for end-to-end driving) which decouples reinforcement learning from the post-training of the end-to-end policy: a privileged teacher is trained using RL on vectorized inputs (HD-map and bounding boxes). This teacher then provides supervision to the pre-trained student during closed-loop post-training. We apply OPTED to two camera-based models, TransFuser and VaVAM, and fine-tune them in AlpaSim, using neural reconstructions (3DGS) of real driving logs. Driving scores increase by factors of 1.6 \times and 9.5 \times , respectively. In controlled experiments OPTED matches closed-loop performance with approximately three orders of magnitude fewer simulator interactions than direct RL post-training, while staying closer to the human prior. Project page: this https URL

[CV-7] Should This Case Be Adapted? Prediction Frag mentation Controls Test-Time Adaptation

链接: https://arxiv.org/abs/2609.20700
作者: Lili Wang,Jing Li,Xiaowen Sun,Xiangyu Hu,Zhuangzhuang Gu,Jian Liu,Srihari Nelakuditi,Yan Tong
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 45 pages, 8 figures, 26 tables

点击查看摘要

Abstract:Episodic test-time adaptation resets a frozen segmenter to source weights M_0 on each case and adapts for a fixed step count. A fixed horizon conflates a cohort-level question, how far to adapt, with an irreducibly per-case one, whether this case should be adapted at all. Cohort means hide that decision: on cross-vendor cardiac MRI the mean \Delta Dice from adaptation is statistically indistinguishable from zero while 58.7% of cases are individually made worse. We quantify this harm as harmful accepted area (HA), the harmful fraction of the edited area a controller deploys. Held-out tuning gives a stronger baseline than a fixed horizon, but the budget it selects transfers on neither of the two main medical benchmarks, and no global budget can condition on the case. We show that prediction fragmentation—the disagreement geometry between M_0 and the adapted mask M_k —predicts HA with no labels or extra backward passes at decision time, comparably on three benchmarks (Spearman \rho 0.50–0.60), at a quarter of gradient-norm’s latency. A case-level router built on it cuts HA from 0.228 to 0.139 on a benchmark that took no part in its design, with the design frozen and only cut-points recalibrated there. On the cardiac benchmark the design was selected on, the router cuts HA from 0.129 to 0.013 at matched Dice and 1.10 deployed updates, against the retrospective-best budget found post hoc on evaluation labels, and reduces that 58.7% to 20.0%, an upper bound we quantify. Where the retained cases are not net-helped (as on prostate), the router still cuts HA but concedes accuracy, a boundary we report. Thresholds are fit once on a labeled split disjoint from evaluation; decisions use no labels or gradients. The template ports across architecture and domain (nnU-Net \to SegFormer, Cityscapes \to ACDC) with coordinate, thresholds and per-bucket actions instantiated per domain.

[CV-8] owards Scaling Marine Perception with Synthetic Data

链接: https://arxiv.org/abs/2609.20680
作者: Haoyu Ma,Onur Bagoren,Anja Sheppard,Elias Fandi,Ashrith Edukulla,Tanner Aslan,Natasha Sieh,Jingyu Song,Katherine A. Skinner
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at OCEANS 2026 Monterrey

点击查看摘要

Abstract:Scalable machine learning in challenging underwater environments is strongly limited by the lack of labeled real-world training data. This data is often expensive and laborious to gather, making large-scale real-world data challenging to gather and curate. However, simulated data can help close the gap, enabling many learning-based tasks for underwater perception. In this work, we extend OceanSim, an IsaacSim-based underwater perception simulator, with a Synthetic Data Generation (SDG) pipeline for training models to be used in underwater scenarios. The proposed pipeline enables users to generate large, automatically labeled, photorealistic datasets with configurable scene appearance, structure, and sensor settings. We evaluate the pipeline on a real-world sea urchin detection task and study how different forms of synthetic scene variation affect sim-to-real performance. Based on these experiments, we discuss findings on our results, main limitations of the current pipeline and identify future directions for improving underwater rendering fidelity, scene diversity, and the evaluation of sim-to-real generalization. The open-source code can be found at this https URL.

[CV-9] FunArt: Decoding Functional Structure and Articulation from Generative 3D Latents

链接: https://arxiv.org/abs/2609.20673
作者: Dennis Rotondi,Abdelrhman Werby,Kai O. Arras
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:To operate effectively in human environments, robots must identify articulated objects, segment their movable and interactive parts, and estimate their kinematic models. Existing articulated scene representations typically recover kinematics from observed interactions, while methods operating on static scans often decouple articulation from functional interactive elements. We present FunArt, a framework that constructs articulation-aware functional 3D scene graphs from posed RGB-D observations captured in a single static configuration. FunArt reconstructs object instances, converts their fused geometry directly into the O-Voxel representation of TRELLIS.2, and exploits its frozen, sparse-compression VAE as a structural prior. A lightweight query-based decoder combines compact object-level latents with dense, surface-aligned features to jointly segment movable parts and functional interactive elements while estimating motion type, axis, origin, and range. On the Articulate3D dataset, FunArt achieves state-of-the-art performance across movable-part segmentation, articulation estimation, and functional-element segmentation, both with and without ground-truth object input. In the end-to-end setting, it outperforms the strongest baselines by 1.5 AP_50 points for movable parts, 2.8 AP_50 points under joint origin-and-axis constraints, and 6.7 AP_50 points for functional elements. These results demonstrate that generative 3D latents encode actionable structural cues that can initialize robotic perception and planning before physical interaction.

[CV-10] Learning Foresight without Explicit Trajectories for 3D Diffusion Policies

链接: https://arxiv.org/abs/2609.20669
作者: Zhongbo Zhang,Zaibin Zhang,Yifan Wang,Changbo Yan,Lijun Wang,Huchuan Lu
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:3D diffusion policies are strong at generating geometrically grounded actions from current observations, but successful manipulation requires not only knowing what motion is feasible now, but also anticipating where the interaction is heading. Existing policies largely leave such foresight to emerge implicitly from action learning. We introduce Movement Trend Guidance, a simple but effective way to provide this foresight without introducing an explicit plan. From a short observation history, the policy learns a compact latent representation of interaction evolution. During training, sparse future gripper states supervise this representation; at inference, only the latent is retained as future-oriented conditioning alongside the current observation. The latent provides global conditioning for action generation, while an additional gated FiLM branch is used only at the UNet bottleneck. Despite adding only 3.52% more parameters to DP3, our method preserves the original dense-action and receding-horizon formulation and consistently improves upon DP3 across RoboTwin2.0, LIBERO-40, and DexArt. It reaches 62.8% vs. 56.1% in 50-task RoboTwin2.0 mixed training, 71.93% vs. 37.08% on LIBERO-40, and 72.0% vs. 49.0% on five real-robot tasks. These results show that a diffusion policy can benefit substantially from knowing where an interaction is heading, without being told exactly where to move.

[CV-11] Earth Surface Immune System for Rapid Monitoring of Unknown Anomalies

链接: https://arxiv.org/abs/2609.20662
作者: Jingtao Li,Qian Zhu,Xinyu Wang,Deren Li,Liangpei Zhang,Yanfei Zhong
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 51 pages

点击查看摘要

Abstract:Earth surface anomalies, driven by escalating climate change, and expanding human activities, are increasing in both frequency and diversity, yet their limited historical data and unpredictability make them fundamentally different from conventional remote sensing targets. Existing methods address specific anomaly categories or stop at localization, leaving a gap between detection and actionable information. Here we present ESIA, an Earth Surface Immune System whose architecture is constrained by three principles from the biological immune system, refined over millions of years against equally diverse and uncertain threats. A non-specific innate immune stage treats anomalies as unobserved changes in time-series satellite imagery, generating binary localization maps at 14.51 km2/s without assuming any anomaly category, surpassing the strongest general baseline by 37% in F1. A specific adaptive immune stage applies negative selection to filter text prompts and matches surviving prompts with localized image patches through a multi-modal foundation model, enabling open-vocabulary recognition of unknown anomaly attributes including category, affected area, and damage severity, with recognition F1 exceeding 80%. A mutation mechanism tunes minimal embeddings at test time, adapting to each scene in 3.26s using a single reference image pair. We validate ESIA on a global-scale dataset covering 19,801.60 km2 across six anomaly categories, comparing against 22 models, and further apply it to quantify degraded farmland in the Dnipro Delta following the Kakhovka Dam collapse and assess burn severity from 2025 Palisades Fire in Los Angeles. This unprecedented flexibility in handling unknown anomalies opens new avenues for real-time disaster response and environmental surveillance.

[CV-12] DexTouch-WM: Learning Action-Conditioned Tactile World Models from Human Touch for Dexterous Robot Manipulation IROS2026

链接: https://arxiv.org/abs/2609.20649
作者: Yan Qin,Yue Chen,Wenwei Lin,Shujia Liu,Chuqiao Lyu,Kailun Su,Chenze Yu,Ping Luo,Wenbo Ding,Tianxing Chen,Renjing Xu
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: Accept to IROS 2026 Workshop RoBoWoMo (Lightning Talk)

点击查看摘要

Abstract:Learning predictive models of contact-rich dexterous manipulation requires dense tactile interaction, but such data are costly to scale on real robots and remain tied to embodiment-specific sensors. We introduce DexTouch-WM, an action-conditioned world model that learns from scalable human touch to jointly predict future RGB observations and bilateral tactile dynamics. Our insight is that human and robot manipulation share transferable contact dynamics when their tactile observations and action spaces are made compatible. We deploy flexible piezoresistive arrays with a shared sensing layout on both human and dexterous robot hands, and retarget human motion into the robot action space so that human interaction can supervise the same dynamics model used for real-robot prediction. DexTouch-WM couples a pretrained video expert with a lightweight tactile expert using anatomy-aware tactile tokens and aligned action conditioning. In human-to-robot scaling experiments, we keep five hours of real-robot supervision fixed while increasing human interaction from 0 to 100 hours, and observe substantial improvements in held-out robot-domain visual, geometric, and contact prediction despite disjoint human and robot task sets. Beyond prediction, we evaluate the world models as surrogate environments for policy evaluation and as generators of synthetic trajectories for real-robot policy learning, showing that scalable human interaction provides a complementary data axis for learning dexterous robot world models.

[CV-13] PROVIA: Procedure State Tracking for Online Mistake Detection in Egocentric Videos KR

链接: https://arxiv.org/abs/2609.20638
作者: Di Wen,Kailun Yang,Jimmy Weissert,Luc Maria Scherrer,Cedric Zöllner,Ruiping Liu,Yufan Chen,Jiale Wei,Junwei Zheng,Kunyu Peng
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 9 pages, 2 figures, 4 tables. Code: this https URL

点击查看摘要

Abstract:An assistant watching egocentric video should notice a mistake from past frames alone, before the next step begins, and keep working once the person recovers. A mistake changes the state of the work, so every later step has to be read against what was done rather than against the plan. The first-mistake protocol that current online methods report on cuts each recording at its first mistake, so a fixed-time rule that never looks at the video is right on every case. We evaluate on complete trials, where mistakes and recoveries arise naturally, under a validation false-alarm budget and against controls that use timing alone. PROVIA keeps two records apart: a factual state, a learned summary of the steps each actor performed, mistakes included, and the accepted progress, an exact posterior over the state of an automaton induced from correct demonstrations by Bayesian state merging and over the execution status of each actor. Procedure-state transitions occur only in the correct-status branch; the mistake and correction branches retain the source state. A sequential test turns the per-frame mistake probability into alarms. With one filter and one optimization rule, PROVIA ranks mistakes best among the evaluated controlled baselines on CaptainCook4D, IndustReal, HoloAssist and IMPACT-ego. At a validation budget of 0.1 false alarms per minute it recalls .154 against .128 on CaptainCook4D and .034 against .015 on HoloAssist, where it leads at every budget. The pipeline runs at 58-70 frames per second. The source code is available at this https URL.

[CV-14] Refinement Is Inherently Editable: Training-Free Prompt-to-Prompt Image Editing with Generative Refinement Network

链接: https://arxiv.org/abs/2609.20633
作者: Yulong Chen,Ziqian Zhang,Haoyu Zhang,Ao He,Senmao Li,Kai Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Text-guided image editing must introduce the requested changes while preserving unrelated source content. Diffusion-based editors rely on spatial controls whose inaccuracies can leave edits incomplete or alter unrelated regions. Causal autoregressive editors face a further constraint: their fixed decoding order limits revision of earlier decisions. We introduce RefineEdit, a training-free prompt-to-prompt image editing framework built on a Generative Refinement Network. Our key idea is to couple edit localization with content generation through the global refinement of binary image codes, allowing editing evidence to be reassessed as the image evolves. RefineEdit initializes an editing branch from an intermediate source state, reusing the emerging layout. We compare the probabilities assigned by the two branches to the same source-sampled bits, using their signed differences to select editable positions and bits. Selected bits follow editing refinement, while the remaining bits copy the evolving source state. To stabilize editing across refinement steps, adaptive spatial freezing limits unnecessary mask expansion, while finite bit locking keeps recently selected bits editable. The framework requires no additional training, external masks, or attention control. Across nine editing categories of PIE-Bench, RefineEdit achieves the best background-preservation scores in PSNR, LPIPS, MSE and SSIM, together with the highest whole-image and edited-region CLIP scores among the evaluated methods.

[CV-15] PhGS: Post-Hoc Pruning and Refinement of Single-View Feed-Forward 3D Gaussian Reconstructions

链接: https://arxiv.org/abs/2609.20623
作者: Rinto Yagawa,Han Cheng,Dieter Schmalstieg,Hideo Saito,Shohei Mori
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Recent single-view feed-forward 3D Gaussian Splatting (3DGS) generation predicts a fixed number of Gaussians per camera ray, introducing severe spatial redundancy. Most existing compaction strategies target multi-view setups to exploit cross-view consistency and are incompatible with single-image models. Instead of retraining the base feed-forward network to directly output compact representations, our insight is to keep the base models frozen and apply post-hoc pruning and recurrent refinement to the generated Gaussians. Consequently, we propose a backbone-agnostic compaction pipeline for single-view feed-forward 3DGS that couples an importance-score-based pruning mechanism with a trainable, lightweight recurrent refinement module, which iteratively updates the surviving primitives to restore image quality. Our results demonstrate seamless integration with existing baselines while preserving novel-view rendering fidelity and achieving high memory reduction. Furthermore, our method supports flexible inference-time keep ratios for application needs.

[CV-16] INSPECT: Learning Robot View Selection from Assistant Use KR

链接: https://arxiv.org/abs/2609.20615
作者: Di Wen,Kailun Yang,Wenhao Guo,Yitian Shi,Junwei Zheng,Yufan Chen,Ruiping Liu,Jiale Wei,Rania Rayyes,Kunyu Peng
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: 9 pages, 3 figures, 5 tables. Code: this https URL

点击查看摘要

Abstract:Robots inspecting an assembly must determine which parts are present and whether they are correctly installed. During egocentric assembly assistance, head motion and workpiece handling reveal evidence for these checks, while spoken state confirmations link observations to procedural outcomes. We introduce INSPECT, which learns robot view preferences from records of a smart-glasses assistant that answers part queries and provides next-step guidance. Presence-Invariant TwinSwap (PI-TwinSwap) calibrates object evidence through paired identity interventions. Claim-indexed supervision separates evidence requirements from camera-reproducible observation changes. Object-centered calibration adapts relative view preferences to robot poses, while clause-level screening checks predicted evidence. The robot selects views using only its current observation and known poses, without candidate images. Evaluation uses annotated assistant-video replay to simulate state feedback, without target-domain view labels for policy training. On images of physical gearbox assemblies, INSPECT achieves the highest view utility among the compared non-oracle policies and raises human-rated full verifiability from 34.8% to 41.7% compared with keeping the current view. On commercial angle-grinder recordings in IMPACT, the transferred relative-view selector increases the correct decision rate from 50.6% to 54.3% with a frozen perception head. The source code is available at this https URL.

[CV-17] RawSLAM: Online HDR Gaussian SLAM from Linear Radiance

链接: https://arxiv.org/abs/2609.20589
作者: Marina Orozco González,Luis Merino
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Current dense visual SLAM systems rely almost exclusively on 8-bit tonemapped Low Dynamic Range (LDR) inputs, limiting their robustness in extreme lighting where shadows and highlights trigger tracking drift and mapping collapse. Conversely, existing raw and High Dynamic Range (HDR) reconstruction pipelines operate strictly offline. They depend on Structure-from-Motion preprocessing and are not suited for large inter-frame motion. We present, to the best of our knowledge, the first online Gaussian SLAM framework that tracks and maps directly on single-exposure 16-bit linear HDR imagery. Our method rests on three core components: an architecture-agnostic HDR Gaussian Splatting module featuring an MLP-free logarithmic parameterization of Gaussian color features; a Reinhard range-compressed photometric objective; and structure-guided spatial gradient weighting. Combined, these components allow our approach to outperform a direct HDR adaptation of MonoGS in both trajectory and reconstruction accuracy, while rendering natively in linear scene radiance for post-rendering processing. The same formulation runs unchanged on standard 8-bit inputs, roughly halving the MonoGS baseline error. Furthermore, our HDR Gaussian module transfers seamlessly to SplaTAM, Gaussian SLAM, and DROID-W, eliminating all tracking failures these systems suffer on challenging illumination sequences. To enable this research, we introduce RawSLAM: a dataset of 10 real-world indoor sequences featuring 16-bit RAW imagery, aligned depth, IMU measurements, and external OptiTrack poses. Code and dataset will be made publicly available soon.

[CV-18] CoRef-GS: Cooperative Referring Gaussian Splatting for Multi-Agent Scene Understanding

链接: https://arxiv.org/abs/2609.20586
作者: Zhikun Zhou,Kunyu Peng,Runyi Yang,Junhao Cai,Di Wen,Ruiping Liu,Danda Pani Paudel,Yi Zhou,Luc Van Gool,Kailun Yang
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)
备注: The established benchmark and source code will be publicly released at this https URL

点击查看摘要

Abstract:Referring scene understanding for embodied robots requires grounding object- and relation-centric language queries from a designated viewpoint. While a local semantic Gaussian map can support such grounding within one agent’s observations, cooperative settings require this ability to remain effective after independently reconstructed maps are aligned and fused. In this setting, the referred target or its contextual landmark may come from another agent’s observations, while spatial relations must still be interpreted from the querying robot’s viewpoint. We formulate this problem as cooperative referring Gaussian grounding over fused maps, which requires geometric alignability, instance-level semantic comparability, and view-conditioned relation reasoning. Existing language-aware Gaussian methods mainly focus on single-map querying, whereas Gaussian registration methods optimize geometric or photometric alignment without preserving language-grounding-oriented semantic compatibility. We propose CoRef-GS, a cooperative referring Gaussian splatting framework. CoRef-GS constructs local open-vocabulary instance-aware Gaussian maps, then aligns partially overlapping maps with a cross-agent alignment module by geometric and semantic consistency, and grounds queries using a view-conditioned mask relation graph. We further introduce CoQuad-Ref, a dual-quadruped benchmark spanning both real-world and simulated indoor scenes. Experiments show that, on simulated scenes, CoRef-GS reduces the rotation error from 2.58° after coarse initialization to 0.15° after refinement, and improves real-world referring mIoU over ReferSplat from 52.6% to 68.8%. The established benchmark and source code will be publicly released at this https URL.

[CV-19] DocAttriBench: Benchmarking Answer Grounding in Document Visual Question Answering

链接: https://arxiv.org/abs/2609.20574
作者: Luca De Grandis(1),Silvia Cappelletti(1),William Raccagni(1 and 2),Marcella Cornia(1),Lorenzo Baraldi(1),Rita Cucchiara(1) ((1) University of Modena and Reggio Emilia, Modena, Italy, (2) University of Pisa, Pisa, Italy)
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Answer grounding in document visual question answering remains an open challenge: most benchmarks lack grounding annotations or provide limited-quality labels, while constructing grounded datasets still requires costly manual effort. We introduce DocAttriBench (DAB), a large-scale benchmark for fine-grained, element-level source attribution in Document VQA, grounding answers to specific layout elements such as text blocks, tables, and images. To build DAB, we propose a Mask-based Perplexity-Derived Attribution method (MAPPET) that combines document layout and language modeling to identify the most informative element for each answer. MAPPET measures the increase in perplexity after masking candidate elements and attributes the answer to the element contributing most to model confidence. Applying MAPPET to multiple existing Document VQA datasets yields DAB, with 237k documents and 296k question-answer pairs with element-level grounding. We benchmark grounding-capable multimodal LLMs on DAB, evaluating answer accuracy, attribution accuracy, and overall answer quality. Results show that while larger models generally achieve higher answer accuracy, even the strongest models often fail to localize the supporting elements. DAB provides a scalable benchmark for developing grounded, verifiable, and trustworthy Document VQA models. Dataset and code are available at this https URL.

[CV-20] OmniMimic: Dynamics-completed Motion Augmentation for Multi-style Omnidirectional Quadruped Locomotion

链接: https://arxiv.org/abs/2609.20566
作者: Sheng Wu,Guoqiang Zhao,Zhe Yang,Fei Teng,Zhikun Zhou,Yanlin Yang,Zheng Fang,Hong Zheng,Yaonan Wang,Kailun Yang
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)
备注: The project page is at this https URL

点击查看摘要

Abstract:Animal demonstrations provide quadruped robots with natural and distinctive gait styles that are difficult to specify through hand-crafted rewards. However, their narrow directional coverage leaves little style-consistent supervision for backward, lateral, and turning commands. We present OmniMimic, a training framework that turns directionally limited animal demonstrations into a single multi-gait policy over target per-axis velocity ranges. OmniMimic first combines temporal reversal, constrained dynamics completion, and sagittal reflection to construct robot-specific kinematic and physical supervision beyond the observed directions. It then expands commands progressively from the demonstrated velocity distribution toward the target per-axis bounds, and uses a shared actor with soft-gated, gait-specialized residual experts to balance reusable locomotion skills with gait-specific corrections. Across four gaits in simulation, OmniMimic reduces mean foot-position RMSE at forward and backward reference velocities by 12.9% and velocity-tracking RMSE on a uniform Cartesian command grid by 63.1%, compared with the matched APEX baseline. The project page is at this https URL.

[CV-21] A Dual-Stream Regulated Reconstruction and Segmentation Network with Hierarchical Artifact-Prior Modeling for Ultra-Low-Field Pediatric Neuroimaging

链接: https://arxiv.org/abs/2609.20562
作者: Bahram Jafrasteh,Leo Milecki,Qingyu Zhao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Automated quality assessment, enhancement, and segmentation of multiple structures in 0.064,\mathrmT ultra-low-field pediatric MRI are limited by a low signal-to-noise ratio, weak anatomical boundaries, and frequent artifacts. We present a unified framework for the LISA 2026 Challenge that performs all three tasks together within one inference pipeline. A network with two coupled streams, built on a 3D U-Net, first reconstructs an enhanced uLF volume and then combines the original and enhanced images for subcortical segmentation. To improve boundary stability, we add an auxiliary class covering brain tissue outside the target structures, derived from whole brain masks. A head conditioned on an artifact graph predicts the seven artifact ratings from reconstruction residuals and frozen segmentation features. We address the scarcity of dense annotations using diffeomorphic registration from atlas to target for label propagation and to regularize anatomical reconstruction. We report validation results across all three tasks.

[CV-22] Automated Goldsmiths Mark Retrieval in Silverware ECCV2026

链接: https://arxiv.org/abs/2609.20509
作者: Atmik Tiwari,Vincent Christlein,Mark Fichtner,Freya Gohlke,Birgit Schübel,Theresa Witting,Heike Zech,Mathias Zinnen
类目: Computer Vision and Pattern Recognition (cs.CV); Digital Libraries (cs.DL)
备注: Accepted at the VISART workshop, ECCV 2026. 18 pages, 8 figures, 1 table

点击查看摘要

Abstract:For art historians, goldsmith marks play a critical role in the identification and dating of artifacts. In practice, experts must manually compare a query mark against hundreds of documented examples, a process that is both tedious and highly dependent on specialist knowledge. To address this, we present an AI-assisted retrieval pipeline that combines mark localization with metric-learning fine-tuning across three backbone architectures: an ImageNet-pretrained ResNet-50, a supervised ViT-S/16, and a self-supervised DINOv2 ViT-S/14. We conduct a systematic evaluation of cropping strategies, where we measure the impact of no cropping, manual ground-truth cropping, and learned detection-based cropping, and assess their interaction with each backbone. Our strongest configuration, DINOv2 ViT-S/14 with manual crop and metric-learning fine-tuning, achieves an mAP of 62.63% and a Top-1 accuracy of 73.74%. Our experiments show that self-supervised pretraining and mark localization are the two most impactful factors, with learned cropping recovering the majority of the gain from manual cropping without requiring ground-truth annotations at inference time. To enable reproducibility and adoption in the digital humanities, we release our manually annotated dataset and codebase, and deploy the system via a public web interface.

[CV-23] Grounded Product Understanding in Livestream Videos

链接: https://arxiv.org/abs/2609.20508
作者: Xinyu Zhang,Junjie Chen,Jiawei Ge,Qianlong Li,Libin Ma,Baokun Pan,Yahui Luo
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:E-commerce livestreams have emerged as an important channel for presenting products to online consumers, containing multiple products whose information is scattered in different moments. This poses significant challenges for downstream product understanding applications, such as product-centric livestream clipping, where models need to identify the product and its relevant segments for information gathering. However, existing benchmarks for general product understanding typically evaluate product retrieval and temporal localization in isolation, leaving the critical correspondence between product identity and temporal evidence largely unassessed. To address this limitation, we introduce GPUB, a large-scale benchmark comprising 3,000 livestream instances with quality-controlled multi-moment temporal annotations and a catalog of over 31K fashion products. GPUB supports three evaluation tasks: the main task Grounded Product Understanding (GPrU) requires jointly identifying the target product and localizing its supporting moments from a livestream video and a candidate product set; Product Retrieval and Product Moment Localization serve as two complementary subtasks. Evaluation of existing multimodal models shows that GPrU remains highly challenging, with the best-performing baseline achieving only 10.13% Pair mAP@.3. To narrow the performance gap, we further develop UniPro, a unified product understanding model that derives product-aligned and temporally structured representations from shared multimodal encoding, improving Pair mAP@.3 to 21.53% while achieving 37.23% Joint R@1@.3 on GPrU.

[CV-24] SenseFuse: Label-Free Fusion of Image and Shape Encoders for Open-Vocabulary 3D Instance Segmentation

链接: https://arxiv.org/abs/2609.20475
作者: Euiseok Han,Tri Ton,Hwanhee Kim,Seungyeon Ryu,Chang D. Yoo
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: 8 pages, 6 figures. Code: this https URL

点击查看摘要

Abstract:Open-vocabulary scene understanding is fundamental for robotics, laying the groundwork for spatial reasoning and object manipulation. While closed-vocabulary 3D instance segmentation heavily leverages 3D shape information, state-of-the-art open-vocabulary methods remain predominantly restricted to 2D image features or image-distilled representations during mask labeling. In this paper, we propose SenseFuse, a label-free fusion method that balances 2D image and 3D shape encoders for robust open-vocabulary 3D instance segmentation, refining only the mask-labeling stage of existing pipelines. We reveal that 2D image and 3D shape encoders exhibit largely disjoint failure patterns and rarely share identical wrong labels, whereas two 2D image encoders frequently repeat the same errors. This distinct behavior makes the 2D and 3D pair inherently complementary. We introduce an adaptive mechanism that selects a scene-level fusion weight to maximize a label-free sensitivity measure, estimated directly from a single scene’s unlabeled proposals in milliseconds. SenseFuse improves labeling accuracy in every evaluated setting across ScanNet200, Replica, and ScanNet++, recovering 67-100% (median 93%) of the gain achievable with an oracle weight, and it raises instance AP in 21 of 22 reported settings. Code is available at this https URL.

[CV-25] Cross-Architecture Foundation-Model Distillation for Edge Flood Segmentation

链接: https://arxiv.org/abs/2609.20441
作者: Fabian Schmalstieg,Karsten Mueller,Wojciech Samek
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: Main paper (17 pages) with supplementary material (11 pages). Submitted to IEEE JSTARS, Special Section on Generalist-Specialist Model Synergy for Remote Sensing: Theories, Methods, and Applications

点击查看摘要

Abstract:Geospatial foundation models can provide strong flood-segmentation performance, but their size limits deployment on memory-constrained edge hardware. We distill a 300-million-parameter Prithvi-EO-2.0 teacher, fine-tuned on the 252 manually labeled Sen1Floods11 training scenes, into a 0.7-million-parameter EfficientViT-B0 student. The teacher supervises additional unlabeled Sentinel-2 imagery, allowing the student training set to grow without new manual annotations. At the matched budget of 252 scenes, teacher-supervised training is competitive with direct training and improves STURM-Flood performance across tested configurations; a geometry-matched control shows that label source alone does not explain the difference. Scaling the teacher-supervised pool to 2,500 scenes narrows the remaining student–teacher gap: the float student reaches 0.787 water intersection over union on the Sen1Floods11 test split against 0.822 for the teacher, matches the teacher on STURM-Flood under our evaluation protocol, and remains below it on WorldFloods-v2. After activation replacement and quantization-aware training, the student runs as a 1.5-megabyte 8-bit integer (INT8) TensorRT engine on a Jetson Xavier NX at 5.57 milliseconds of graphics processing unit (GPU) compute per 512-by-512 image, with approximately 14 megabytes of runtime device memory. A fixed modified normalized difference water index (MNDWI) threshold is competitive with both models on the two clean external benchmarks, so we interpret those benchmarks as generalization tests rather than as evidence of learned-model superiority over a spectral rule. The results support the conclusion: foundation-model supervision can amplify a fixed manual annotation budget into a substantially larger training set and yield a compact, deployable edge model.

[CV-26] When Do Language-Grounded Explanations Help? A Graph-Bottleneck for Farm Monitoring Interpretable Sheep Facial Pain

链接: https://arxiv.org/abs/2609.20427
作者: Alam Noor,Miguel Guti’errez Gait’an
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Automated pain recognition from facial expression could make continuous welfare assessment practical in sheep, but adoption depends on trust: a stockperson cannot act on a score that arrives without justification. We ground a model in the Sheep Pain Facial Expression Scale (SPFES) by letting each detected facial region attend over text embeddings of the clinical descriptors and then test whether the resulting explanations mean anything. They do not. Ablating an entire descriptor changes the predicted logit by about 10^-4 , and the most-attended cue agrees with the predicted pain level in only 32.6% of regions, although the attention maps, the learned gate, and the generated text all proposed otherwise. We therefore remove the appearance bypass with a concept bottleneck whose classifier reads only SPFES concept scores, supervised by per-region state annotations that image-level pipelines discard. This costs 0.05 – 0.10 in Cohen’s \kappa but yields concepts that are demonstrably learned: minority pain-indicating states are recovered at 3.5 – 8.3\times their base rates, and the ear and eye severity orderings emerge without severity supervision. Removing the supervision alone leaves \kappa unchanged while concept accuracy falls to 0.109 , showing that architectural necessity does not imply semantic validity. We also show that pooled concept accuracy is misleading under clinical imbalance and provide a cross-validated, protocol-matched benchmark of seven methods on this dataset.

[CV-27] WeVisDoc: From Coverag e to Capability for Robust End-to-End Document Parsing

链接: https://arxiv.org/abs/2609.20423
作者: Hao Yu,Kang Liu,Linnan Zhao,Jiabo Zhan,Chong Sun,Chen Li,Jing Lyu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Document parsing converts document images into structured content and requires reliable performance across diverse layouts and acquisition conditions. Yet training corpora are biased toward common document types and clean digital pages, while expanding coverage alone does not specify how to address a parser’s remaining weaknesses. We present WeVisDoc, a two-stage data-centric framework for robust end-to-end document parsing. Stage I broadens semantic, structural, and appearance coverage through heterogeneous data and structure-preserving degradation synthesis. Stage II uses a held-out probe to measure the Stage I parser’s residual errors within fixed visual-structural clusters. These diagnostics guide targeted data construction and reallocation of the target-token budget. WeVisDoc-4B achieves an Overall score of 95.38 on OmniDocBench v1.6 and a mean Overall score of 75.54 across the three PureDocBench tracks, ranking first among the compared end-to-end parsers in all four settings. Compared with Stage I, Stage II improves Overall scores for the 2B and 4B models on both benchmarks, with larger gains on the degraded PureDocBench tracks, including a 4.03-point gain for the 4B model on the Real Degraded track.

[CV-28] ouchSight: Bare-Handed Tactile Prediction from Egocentric Video via Generative Visual Augmentation

链接: https://arxiv.org/abs/2609.20414
作者: Danyan Zhou,Jinxuan Lu,Jiawei Lin,Tianxing Chen,Chuqiao Lyu,Wenbo Ding
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Tactile signals provide direct contact and force measurements that are essential for understanding physical interactions and enabling dexterous robotic manipulation. However, tactile sensing requires direct measurement at contact interfaces, making large-scale data collection reliant on intrusive, costly, and restrictive instrumentation. We present TouchSight, a monocular egocentric vision framework for dense full-hand contact force prediction that leverages 500 hours of pressure-glove recordings and extensive hand-object interaction (HOI) data. To address the appearance gap between gloved training data and bare-hand real-world scenarios, we construct TwinTouch-20H: 20 hours of paired visual data in which generative video models re-render gloved recordings as bare-hand observations against new backgrounds while preserving the original measured tactile labels. TouchSight predicts dense force from both gloved and generated bare-hand videos, outperforms prior contact prediction methods on OakInk2, qualitatively generalizes to natural bare-hand egocentric videos from unseen datasets, and improves consistently as glove supervision scales. These results demonstrate that dense tactile signals can be recovered from egocentric vision alone, without tactile instrumentation at capture time.

[CV-29] Navi-Agent : Unlocalized Monocular Navigation Agent ICRA

链接: https://arxiv.org/abs/2609.20388
作者: Wenyuan Xie,Mengyang Hong,Yongzhong Wang,Yanbiao Ji,Yijin Zhou,Shaokai Wu,Shalayiding Sirejiding,Huayi Zhou,Yi-Chao Chen,Ma Ling,Yue Ding,Hongtao Lu
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: 8 pages, 7 figures. Submitted to 2027 IEEE International Conference on Robotics Automation (ICRA)

点击查看摘要

Abstract:Vision-Language Navigation in Continuous Environments (VLN-CE) requires an embodied agent to execute long-horizon instructions in unknown environments. Existing zero-shot VLN-CE systems typically maintain spatial states through geometric localization or coordinate-based representations. Recent geometry-constrained navigation removes depth and globally consistent coordinates, but maintaining persistent spatial awareness for place confirmation, progress verification, and recovery remains challenging. We present Navi-Agent, a zero-shot VLN-CE agent that constructs a coordinate-free spatial state from visual observations and executed motion histories. Navi-Agent organizes this state as a navigation topology, where nodes represent visual places and edges represent motion transitions. This representation enables observation-based approximate self-localization, task progress verification, and visual revisitation-based recovery. Navi-Agent performs closed-loop navigation by decomposing instructions into sub-goals, executing local visual navigation, and verifying visited places through the constructed spatial state. Experiments on zero-shot VLN-CE benchmark and real-world robot platforms show that Navi-Agent achieves state-of-the-art performance among geometry-constrained methods while remaining competitive with approaches relying on geometric localization.

[CV-30] Compact Vision Models for Iris Presentation Attack Detection under Presentation Attack Instrument Shift and Environmental Degradation ATC

链接: https://arxiv.org/abs/2609.20386
作者: Athanasios Angelakis,Marta Gomez-Barrero
类目: Computer Vision and Pattern Recognition (cs.CV); Cryptography and Security (cs.CR)
备注: Accepted at BIOSIG 2026. This preprint includes minor nomenclature and editorial corrections clarifying the project-specific Patch-ABMIL and Compact-TransMIL variants

点击查看摘要

Abstract:Iris presentation attack detection (PAD) is security-critical when a subsystem that appears reliable during development encounters presentation attack instruments (PAIs) or acquisition conditions absent from validation data. We benchmark three compact scratch-trained computer-vision models, each with at most approximately 0.26 million trainable parameters, on the Notre Dame subset of LivDet-Iris 2017 under PAI-driven domain shift and environmental degradation. All models are trained without external pretraining or data augmentation and evaluated over five seeds. A validation-selected threshold is transferred unchanged to the known-attack, unknown-attack, corrupted, and pooled test partitions. From known to unknown attack presentations, Attack Presentation Classification Error Rate (APCER) increases by 17.11-30.47 percentage points and Detection Equal Error Rate (D-EER) increases by 7.38-12.73 percentage points. At the validation-selected threshold, ZACH-ViT obtains the lowest unknown-attack APCER (47.69 +/- 4.84%) and D-EER (38.87 +/- 0.93%), while Compact-TransMIL obtains the lowest Bona Fide Presentation Classification Error Rate (BPCER). ZACH-ViT also gives the lowest unknown-attack BPCER at an APCER limit of 10% (81.29 +/- 1.95%). The high absolute errors show that the comparative advantage of the best compact model does not constitute deployment readiness under unknown PAIs.

[CV-31] MM-Future: Multi-Mode Joint World-Action Modeling for Autonomous Driving

链接: https://arxiv.org/abs/2609.20377
作者: Shuai Liu,Hechangle Gong,Hao Jiang,Runlin He,Junxiang Zhan,Kai Huang,Sheng Yang,Shaoqing Ren
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Autonomous driving involves coupled decision-making and scene evolution under multi-mode uncertainty. To capture this coupling and uncertainty, we introduce MM-Future, a world-action model that generates multiple paired scene-action hypotheses and models bidirectional interaction within each pair. Each hypothesis is initialized from a structured action prior and an independent future scene source, which are then co-evolved through a modality-aware diffusion Transformer. To support efficient multi-mode rollout, MM-Future compresses multi-view video into planning-oriented representations, dubbed MM-Tokens. Finally, a future-conditioned proposal scorer ranks trajectory candidates by shared history context and their paired predicted future. On NAVSIM navtest, MM-Future achieves 94.0 PDMS and 91.5 EPDMS, while attaining a 32.3 HD-Score in zero-shot closed-loop evaluation on HUGSIM. Ablations show consistent improvements over both single-mode and action-only variants, validating the benefit of multi-mode joint world-action modeling.

[CV-32] EliGSiR: Continual RGB-D Mapping with Gaussian Splatting under Bounded Compute

链接: https://arxiv.org/abs/2609.20348
作者: Björn Ellensohn,Elmar Rueckert,Christian Rauch
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: 8 pages, 8 figures

点击查看摘要

Abstract:Conventional 3D Gaussian Splatting assumes a closed set of observations and long optimization schedules. Continual RGB-D mapping in contrast poses the problem that new observations arrive online, while previously reconstructed regions must be preserved. We present EliGSiR (Evidence-guided Load-adaptive Incremental Gaussian Splatting with Image Replay), a continual Gaussian mapper that controls how the available optimization budget is used as the reconstruction evolves. Map-Guided View Scheduling filters redundant incoming views and reconsiders retained views according to the current state of the map. Load-Adaptive Fidelity adjusts supervision resolution to the current mapping load instead of following a fixed resolution schedule. Targeted Geometry Growth separates depth supervision from Gaussian creation and adds geometric capacity only where repeated RGB-D observations indicate missing or misplaced structure. Together, these mechanisms adapt which views are optimized, how much image detail is used, and where the representation grows while mapping remains active. We evaluate EliGSiR on Replica, TUM RGB-D, ScanNet++, and real RGB-D sensor sequences, considering both the final reconstruction and the map available throughout acquisition. On TUM RGB-D fr3/long_office_household, EliGSiR reaches 21.52 dB with the same ground-truth mapping poses used by the controlled baselines, compared with 19.42 dB for SplaTAM. In the tracked-pose comparison, EliGSiR with live ORB-SLAM3 poses reaches 23.02 dB in 155.5 s, compared with 20.10 dB in 230.9 s for CaRtGS using its native tracker. We further evaluate reconstruction throughout acquisition and show how EliGSiR adaptive view scheduling, supervision fidelity, and geometry growth improve the use of the available mapping budget.

[CV-33] Fast Cross-Strength Multi-Contrast Brain MRI Translation using Latent Bridge Matching MICCAI2026

链接: https://arxiv.org/abs/2609.20341
作者: Siddharth Srivastava,Till Bretschneider
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 10 pages, 4 figures. MRIxFields Workshop, MICCAI 2026

点击查看摘要

Abstract:Magnetic Resonance Imaging (MRI) acquired at different field strengths exhibits pronounced variation in noise, resolution, homogeneity, and contrast, which limits comparability across acquisition settings and complicates downstream analysis. We address this with a unified conditional model for controllable field-to-field synthesis, built on the framework of conditional latent bridge matching. Our single model achieves highly competitive results across the validation phase for all three tasks of the MRIxFields2026 challenge without task-specific architectures or training. We achieve fast generation with only a single inference step, producing all modality and field-strength combinations for 30 axial slices in under 90 seconds, as well as cross-modality-strength translation for a full volume in under 70 seconds, on a single NVIDIA A5000 GPU. We further provide extensive ablations regarding different components of our solution. Code: this https URL

[CV-34] FreqDINO: A Frequency-Guided Multi-Task Routing Vision Foundation Model for Universal Ultrasound Analysis

链接: https://arxiv.org/abs/2609.20340
作者: Qing Xu,Yixuan Zhang,Yue Li,Xiangjian He,Qian Zhang,Mainul Haque,Rong Qu,Wenting Duan,Jieyun Bai,Zhen Chen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by TBME

点击查看摘要

Abstract:Ultrasound image analysis plays a crucial role in cancer screening and prenatal diagnosis, yet comprehensive assessment requires jointly addressing tasks such as lesion segmentation and benign-malignant classification. While recent vision foundation models have shown remarkable universal representations, unlocking their potential for ultrasound is bottlenecked by the considerable domain gap from natural images. Existing methods typically fine-tune heavy vision encoders for isolated tasks, incurring substantial computational overhead while overlooking the underlying commonalities across heterogeneous tasks. In this work, we propose FreqDINO++, a frequency-guided multi-task routing vision foundation model for universal ultrasound analysis. We first introduce a Multi-task Routing Adapter (MR-Adapter) to support parameter-efficient integration of task-common and task-specific knowledge, a Frequency-aware Feature Enhancer (F ^2 -Enhancer) is then designed to capture the rich multi-scale frequency characteristics of ultrasound images, and a Task-aligned Collaborative Decoder (TC-Decoder) is devised to promote collaboration between dense and global prediction tasks through global-local token interaction. Extensive experiments on large-scale multi-task and external single-task ultrasound benchmarks demonstrate that FreqDINO++ consistently outperforms strong baselines and recent foundation models across 27 diverse clinical task scenarios, while also showing promising generalization to unseen data. The code is at this https URL.

[CV-35] AgriScope: Pixel-Grounded Multimodal Understanding for Agricultural Images

链接: https://arxiv.org/abs/2609.20325
作者: Abderrahmene Boudiaf,Mohamad Alanssari,Irfan Hussain,Sajid Javed
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Agricultural image understanding requires fine-grained recognition of plant diseases, pests, crop structures, and botanical species under complex real-world conditions. Despite recent advances in Multimodal Large Language Models (MLLMs), existing models remain limited to text-only outputs and lack pixel-level visual grounding capabilities. In this work, we introduce AgriScope, a unified pixel-grounded multimodal framework for agricultural image understanding. AgriScope jointly supports image-level, region-level, and pixel-level understanding within a unified framework, enabling tasks such as grounded caption generation, referring expression segmentation, and multi-turn multimodal interaction for agricultural imagery. AgriScope integrates biologically specialized semantic representations with dense spatial grounding through biological-semantic encoding, dense spatial representations, and pixel decoding. To support large-scale grounded learning, we introduce AgriGround, a large-scale pixel-grounded agricultural multimodal instruction-tuning dataset containing over 500K images and 11M instruction-following samples spanning plant disease analysis, crop and weed identification, insect pest recognition, and fine-grained botanical understanding. AgriGround is constructed through a multi-stage automatic annotation pipeline that integrates multimodal caption generation, phrase-level grounding, segmentation mask generation, and task-oriented instruction synthesis to produce densely grounded supervision. Extensive experiments across multiple agricultural vision-language tasks demonstrate the effectiveness of AgriScope in pixel-grounded multimodal understanding, establishing a strong benchmark for agricultural vision-language learning and visual grounding. The dataset and code will be made publicly available at (this https URL)

[CV-36] LLM -Guided Transformation of Non-Critical Driving Scenes into Safety-Critical Scenarios Using Augmented Reality

链接: https://arxiv.org/abs/2609.20318
作者: Noura Fady,Farah Khaled,Catherine M. Elias
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Testing Autonomous Driving Systems (ADS) requires realistic safety-critical scenarios, but collecting such data from real-world driving is costly and unsafe. This paper presents an automated pipeline that transforms safe driving scenes into safety-critical scenarios by combining computer vision, Large Language Models (LLMs), and Augmented Reality (AR). The system detects and tracks road users, extracts safety features including distance, velocity, motion direction, and Time-to-Collision (TTC), and assesses scene criticality. Safe scenes are modified by an LLM, which generates realistic collision-inducing objects and behaviors that are integrated into the original scene using AR. The proposed pipeline was evaluated on the nuScenes dataset, achieving 97.52% safety classification accuracy and successfully generating realistic scenarios such as pedestrian crossings, rear overtaking vehicles, and sudden-stop events. The results demonstrate an effective and flexible approach for automated generation of safety-critical scenarios to support the testing and validation of autonomous driving systems.

[CV-37] Not All Layers Are Equal: Dynamic Layer Routing for Reliable CLIP OOD Detection

链接: https://arxiv.org/abs/2609.20299
作者: Ignacio M. De la Jara,Cristian Rodriguez-Opazo,Damith Ranasinghe
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Information aggregation across model layers are revealed to improve OOD detection. In contrast to crafting a method for layer-wise information aggregation in recent work, we investigate if layer selection is a learnable problem. In other words, we transpose the question from how to fuse layers to one asking which layers to trust for an input. Using a generalizable, weak, out of distribution context crafting approach for supervision, shown to be more effective than state of the art methods’ mechanisms, we formulate learning a lightweight router to select a sparse, final-layer-anchored expert over CLIP’s layer depth for OOD detection. Across three diverse benchmarks we demonstrate our learnable routing method dubbed Voyager improves OOD detection. On ImageNet-1K, Voyager achieves an average FPR@95 of 18.86, outperforming the strongest, comparable, prompt-learning method by 8.8 points. These gains persist across multiple supervision sources, including those used by existing state-of-the-art prompt-learning methods, demonstrating that, whilst our weak OOD supervision context is highly effective, the key advantage is realized from the learnable router component rather than the supervision source. Significantly, Voyager is highly practical; router learning takes approximately two minutes using less than 1 GB of memory, making it approximately 20x more efficient than current prompt-learning approaches. Anonymized Code: this https URL Subjects: Computer Vision and Pattern Recognition (cs.CV) Cite as: arXiv:2609.20299 [cs.CV] (or arXiv:2609.20299v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2609.20299 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[CV-38] nyCNN: A 193K-Parameter Network for On-Device Plant Disease Detection with a Cross-Dataset Robustness Diagnosis

链接: https://arxiv.org/abs/2609.20290
作者: Ngoc-Bao Ho-Lam,Thai-Anh Nguyen
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 13 pages, 3 figures. Accepted at ISRSD 2026

点击查看摘要

Abstract:Detecting crop disease early is central to sustainable agriculture and food security under United Nations Sustainable Development Goal 2 (Zero Hunger), and is especially urgent in resource-constrained regions where expert diagnosis is scarce but low-cost mobile devices are widespread. This paper presents TinyCNN, a lightweight convolutional neural network for on-device plant disease classification. TinyCNN uses depthwise separable convolution blocks and contains only 193,190 trainable parameters with 110.05M MACs for a 224x224 input image. On the 38-class PlantVillage benchmark, TinyCNN achieves 98.88% test accuracy and 98.03% macro-F1 while being approximately 58x smaller than ResNet18 and 11.8x smaller than a MobileNetV2 teacher, directly reducing the energy, memory, and cost footprint of inference in line with Green AI principles. The paper further analyzes vanilla knowledge distillation as a sustainable model-compression strategy; an ablation over alpha in 0.3, 0.5, 0.7 and T in 2, 4 selects alpha=0.3, T=4, producing a distilled TinyCNN with 98.81% test accuracy. Finally, cross-dataset evaluation from PlantVillage to PlantDoc reveals a substantial robustness gap under real-world conditions, which a Grad-CAM analysis attributes to off-leaf, background-driven attention consistent with shortcut learning. TinyCNN is thus an energy-efficient, deployable building block for sustainable agricultural intelligence, while field robustness remains the key barrier to durable real-world impact.

[CV-39] Queries Knew More Than We Thought: Uncovering Latent Knowledge in Segmentation Models

链接: https://arxiv.org/abs/2609.20283
作者: Ignacio M. De la Jara,Cristian Rodriguez-Opazo,Damith Ranasinghe
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Modern segmenters often fail after the expensive computation has already been done: a useful mask is present among the model’s query-conditioned candidates, but the deployed selection rule does not expose it. We study this output-selection bottleneck in frozen DETR-family models. A ground-truth-only oracle first shows substantial hidden headroom in already-computed mask proposals. This raises a simple question: How can we better use the masks a segmenter has already computed but does not expose? We then ask whether that headroom can be recovered without adding queries, generating new masks, rerunning the backbone, or updating weights. HYDRA is a small selector trained only on cached frozen outputs. At inference time, it scores the cached candidates against an explicit keep-baseline option and acts only when a held-out calibrated margin indicates the selected candidate is sufficiently better. Trained on training-split caches and calibrated on held-out data, HYDRA improves Mask2Former, MaskDINO, and OneFormer by up to +7.41 dataset mIoU points on ADE20k and COCO, and improves SAM 3 by +9.4 class-macro prompt-IoU points on average across eight domains while preserving useful predictions through calibration. Paired LoRA controls show that lightweight weight adaptation does not remove the bottleneck: exposed predictions are often flat or worse, while routing over the adapted candidates still recovers accuracy. Finally, we connect the effect to query specialization under bipartite matching and verify it in a controlled TinyDETR study. These results show that frozen segmenters should be evaluated not only by the masks they expose, but also by the useful candidates they suppress.

[CV-40] CleanVideo: Adaptive Concept Erasure for Text-to-Video Diffusion Models

链接: https://arxiv.org/abs/2609.20267
作者: Junchi Liao,Hongji Li,Wenrui Zhou,Lijie Hu
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Concept erasure aims to selectively eliminate undesired visual semantics from pre-trained generative models without compromising their general utility. Extending concept erasure from images to video is nontrivial. Target concepts emerge gradually and vary across frames and denoising steps. As a result, fixed interventions may miss the target or introduce blurring, jitter, and content distortion. We propose CleanVideo, a selective erasure framework that performs low-dimensional subspace intervention controlled by a tri-modal gating mechanism. By jointly processing spatiotemporal visual features, timestep signals, and textual semantics, CleanVideo determines where, when, and whether to intervene, steering erased content toward natural surrogate concepts when such surrogates can be clearly defined while preserving non-target content. Experiments on three video diffusion models show that CleanVideo effectively erases target concepts while maintaining visual fidelity and temporal coherence, outperforming existing baselines under frame-level and video-level evaluations and under concept-recovery attacks when the protected pipeline remains intact.

[CV-41] AI or Real: Detecting Partially Altered Videos Under Resource-Constrained Environments

链接: https://arxiv.org/abs/2609.20263
作者: Tamoghna Chakraborty,Md Nurul Absur,Sourya Saha,Saptarshi Debroy
类目: Computer Vision and Pattern Recognition (cs.CV); Distributed, Parallel, and Cluster Computing (cs.DC)
备注: 7 Pages

点击查看摘要

Abstract:The proliferation of generative video models has shifted the practical detection threat from fully fabricated clips to partially manipulated footages. Although modern detectors achieve strong accuracy using foundation backbones of 400M+ parameters, their resource footprint precludes edge deployment. In this paper, we present a lightweight full-frame detector for partially manipulated AI-generated video, designed for deployment on edge hardware without face-detection preprocessing. The system distills a DINOv2-Base teacher into a frozen MobileNetV3-Small student through a pipeline that combines temperature-annealed soft-label transfer, attention-diversity regularization, frame-level supervision, and a residual feature adapter that conditions ImageNet features for artifact detection. We additionally target two failure modes specific to the partial-manipulation regime: false positives on legitimate scene cuts, addressed through within-video temporal hard negatives; and threshold-level miscalibration on the dominant pure-real class, addressed through calibration-aware sampling. Evaluation on a 55,393-sample spliced test set across fake-frame ratios from 6.2% to 31.2% demonstrates the student model closing 58% of the gap to the DINOv2-Base teacher (AUC 0.766) while running at 3.65 ms per 16-frame clip on RTX A4000 with a 150.4 MB checkpoint compatible with edge-device memory and latency budgets.

[CV-42] Generative Verification: Rethinking the Uncertainty Signal for Active Learning of Object Detection

链接: https://arxiv.org/abs/2609.20262
作者: Licheng Zhang,Zheng Gong
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Nearly every acquisition function for active object detection shares one arrangement, in that the model being improved is also the model being interrogated. We depart from it. In generative verification an independent generative model re-derives the label of a detection from the pixels inside its predicted box, and the disagreement between the two becomes the acquisition signal. Two properties follow from the arrangement itself rather than from any tuning. A displaced box, a box on background and a correct box carrying the wrong label all yield a crop that fails verification, so the failure modes arrive already combined in one scalar and the hand-weighted classification and localization terms of existing criteria are no longer needed. And because the verifier never observes the detector confidence, confidently wrong detections score highest, although a self-derived signal reads them as uninteresting and they are the costliest to leave unlabeled. We build the verifier as a conditional diffusion model whose diffusion target is a label representation rather than an image. Its reverse process is stochastic, so repeated generations return a distribution whose concentration reports how firmly the evidence determines the label, where a classifier returns a single point estimate. On PASCAL VOC and MS-COCO the signal outperforms output-uncertainty, feature-geometry, perturbation and ensemble criteria, gaining about one mAP50 point per round on MS-COCO, with its largest margins in the early rounds where confident detector errors are most common.

[CV-43] Distance to Class Prototypes: Active Learning for Object Detection

链接: https://arxiv.org/abs/2609.20248
作者: Licheng Zhang,Zheng Gong
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Deploying a deep object detector in a new setting is limited less by architecture than by the cost of annotating data from that setting. Active learning lowers the cost by choosing which images to label, and the choice is only as good as the signal used to score an unlabeled image. That signal is usually the class posterior, which is cheap but poorly calibrated, or the disagreement across several models or several stochastic passes, which is better but multiplies inference over a pool far larger than the labeled set. We propose a signal richer than the posterior yet still read from one forward pass of one network. A supervised contrastive term added to the training objective shapes a per-object embedding space in which distance encodes class membership, and an unlabeled detection is scored by how far it lies from the region occupied by its predicted category, weighted by its confidence. The criterion needs no ensemble, no auxiliary predictor and no repeated inference, and its entire cost is 2.89M parameters, an increase of 8.3% over a bare detector. On PASCAL VOC and MS-COCO it beats the posterior of the same detector in every round in which a selection is made, by up to 1.08% mAP50 against run to run deviations of 0.02% to 0.18%, and it stays competitive with ensemble and Monte Carlo dropout criteria costing three to fifty forward passes per unlabeled image. Experiments use the single-stage detector under which the compared criteria report their results, so that the selection decision is isolated from the strength of the detector.

[CV-44] SAGE-Yoga: Multi-Cue Learning for Yoga Pose Classification and Joint-Level Correction

链接: https://arxiv.org/abs/2609.20245
作者: Hung Le Chi,Khanh Minh Huynh,Long Nghia Tran Pham,Tan Phuc Huynh,Trong-Thuan Nguyen,Minh-Triet Tran
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Under Review for RIVF 2026

点击查看摘要

Abstract:Automated yoga analysis requires both accurate pose classification and interpretable feedback on pose execution. However, existing methods often rely on a single visual prediction, struggle to distinguish visually similar poses, and treat pose classification and correction as separate tasks. To address these limitations, we propose SAGE-Yoga, a unified coarse-to-fine framework for yoga pose classification and joint-level correction from a single RGB image. Inspired by how yoga instructors assess posture using multiple complementary cues, SAGE-Yoga first employs a bagging-based ensemble of complementary visual backbones to generate a ranked set of candidate pose classes. Additionally, a margin-based gating mechanism preserves confident visual predictions while invoking geometric verification only for ambiguous cases. Moreover, once the final pose class is determined, SAGE-Yoga retrieves a medoid reference pose and compares the observed joint angles with class-specific distributions to identify misaligned joints. Finally, these deviations are translated into actionable corrective feedback. Empirically, experiments on the Yoga-82 dataset show that the visual ensemble achieves 89.0% Top-1 accuracy, while the complete framework improves performance to 90.7% Top-1 accuracy and 90.1% Macro-F1. These results demonstrate that combining complementary visual evidence with selective geometric verification improves fine-grained pose classification while enabling interpretable, joint-level correction.

[CV-45] Scene-Q: Confidence-Aware Coarse-to-Fine Querying of 3D Scenes with Selective VLM Reasoning IROS2026

链接: https://arxiv.org/abs/2609.20235
作者: Juno Kim,Yesol Park,Hye-Jung Yoon,Byoung-Tak Zhang
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: 8 pages, 5 figures, 3 tables. Submitted to IROS 2026

点击查看摘要

Abstract:Indoor mobile robots require open-vocabulary scene understanding that grounds natural-language queries in a consistent 3D map. Many existing systems ultimately rely on cosine-similarity retrieval with contrastive image–text encoders, which is efficient but brittle when labels are near-synonymous or multiple similar instances appear. We present Scene-Q, a confidence-aware coarse-to-fine querying framework that normalizes encoder scores with temperature scaling and selectively invokes a reasoning VLM only for low-confidence cases. High-confidence queries are answered by fast retrieval, while ambiguous ones are reranked over a small top-K candidate set using the original multi-view images and instance bounding boxes, enabling context-aware disambiguation at low cost. Scene-Q improves open-vocabulary 3D instance segmentation on ScanNet200 and natural-language 3D instance retrieval on real-world reconstructions, with the largest gains on spatial and relational queries while keeping a substantial fraction of queries on the fast path.

[CV-46] A Two-Stage Multi-Scale Attention-Based Network for Weakly Supervised Cataract Fundus Image Enhancement

链接: https://arxiv.org/abs/2609.20222
作者: Xiaoyong Fang,Yue Wang,Xiangyu Li,Wanshu Fan,Dongsheng Zhou
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by Scientific Reports

点击查看摘要

Abstract:Cataract is a major cause of vision loss and hinders further diagnosis. However, cataract fundus image enhancement often grapples with challenges such as limited paired cataract retinal images and insufficient recovery of fine details in the retinal images. To mitigate these challenges, we in this paper propose a two-stage multi-scale attention-based network (TSMSA-Net) for weakly supervised cataract fundus image enhancement. Our TSMSA-Net leverages the domain transformation to synthesis paired real-like cataract images, solving the problem of difficult acquisition of paired images. To further extract detailed information from fundus images and reduce the generation of artifacts during the enhancement process, we propose a multi-scale attention-based stage to learn more useful features for cataract image enhancement. Experimental results on Kaggle and ODIR-5K demonstrate that our TSMSA-Net outperforms current state-of-the-art cataract fundus images enhancement even without paired images and exhibits certain generalization ability. Experimental results on Kaggle and ODIR-5K datasets indicate that our TSMSA-Net outperforms the current state-of-the-art methods for cataract fundus image enhancement, even in the absence of paired images. Additionally, it demonstrates a certain level of generalization capability. The enhancement also can improve the performance of vessel segmentation and classification in cataract images.

[CV-47] VLN on the Fly: An Onboard Vision-Language Navigation Stack for Aerial Robots

链接: https://arxiv.org/abs/2609.20191
作者: Marco S. Tayar,Felipe Tommaselli,Gianluca Capezutto,Pedro Antonio Rabelo Saraiva,Pedro H. V. de Freitas,Lucas Kido,Guilherme Sonego,Ricardo V. Godoy,Marcelo Becker
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Systems and Control (eess.SY)
备注:

点击查看摘要

Abstract:Running vision-language navigation fully onboard an aerial robot is hard, since grounding, planning, and control must share limited compute and a single-stage error is difficult to isolate in flight. End-to-end aerial policies fuse these stages into one network, giving up the observability and safety checks a modular stack keeps available. We propose VLN on the Fly, an onboard stack that keeps grounding, planning, and control as separate, inspectable stages. A quantized VLM grounds an instruction to a coarse image cell, depth lifts it to a 3D goal, a fast B-spline planner returns a feasible trajectory, and a pretrained reinforcement learning policy tracks it to motor commands across quadrotors. Across 15 onboard flights over three everyday referents in a controlled indoor volume, the stack reaches the target in 13 of 15 trials with 5.72 cm mean goal error and 39.3% average GPU utilization. In 6 additional cluttered-environment trials, the stack tracks collision-free trajectories under onboard perception gating.

[CV-48] MoSSGate: Memory-Modulated State-Space Gating for Skin Lesion Segmentation

链接: https://arxiv.org/abs/2609.20181
作者: Anum Awan,Mahnoor Buriro,Muhammad Younas Khan,Md Imam Ahasan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 15 pages, 5 figures, accepted on International Conference on Cloud and Network Computing

点击查看摘要

Abstract:Accurate skin lesion segmentation is crucial for reliable computer-aided dermatological diagnosis, yet existing convolutional and transformer-based models often struggle to jointly capture long-range spatial dependencies and fine boundary details under limited computational budgets. This trade-off between global context modeling and boundary-aware localization frequently leads to over-segmentation, fragmented predictions, or missing thin peripheral structures. To address this challenge, we propose MoSSGate, a plug-and-play module for U-Net that integrates (i) boundary-aware spatial gating to restrict long-range propagation to informative regions, (ii) an external memory modulator that provides sample-adaptive dynamic control, and (iii) parallel 2D state-space modeling for efficient global context aggregation with linear complexity. The proposed design enables adaptive, context-aware information propagation while preserving sharp and accurate lesion boundaries. Extensive experiments on the ISIC 2017 and ISIC 2018 benchmarks demonstrate state-of-the-art accuracy with strong efficiency, achieving 86.3% and 85.9% mIoU and 92.6% and 90.6% Dice, respectively, while requiring substantially fewer FLOPs than most competing CNN-based methods. These results highlight a favorable accuracy efficiency trade-off for high-resolution medical image segmentation.

[CV-49] MTF-Net: Multi-Modal Temporal Feature Fusion Network for Pedestrian Intention Prediction

链接: https://arxiv.org/abs/2609.20178
作者: Md Mahfuzur Rahman,Pengzhan Zhou,A. F. M. Abdun Noor,Md Imam Ahasan,Md Mustafizur Rahman,Fang Qu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 15 pages, 5 figures, accepted on International Conference on Cloud and Network Computing

点击查看摘要

Abstract:Accurately predicting pedestrian intentions is crucial for ensuring safe and proactive interaction between autonomous vehicles and pedestrians. However, existing approaches often depend on architectures that either model temporal dependencies within individual modalities or fuse modalities only at coarse semantic levels. To address these limitations, we propose MTF-Net, a novel Multi-Modal Temporal Feature Fusion Network that jointly models kinematic, appearance, and contextual cues for pedestrian intention prediction. MTF-Net integrates four complementary modalities-bounding-box dynamics, human pose keypoints, local context, and scene-level semantics within a recurrent fusion framework enhanced by gated linear units (GLUs). These GLU-based modules adaptively regulate cross-modal information flow, enabling interpretable and efficient feature interaction across temporal scales. Through three dedicated temporal encoding branches and an attention-guided fusion head, the proposed model robustly anticipates pedestrian crossing intentions several frames before they occur. Extensive evaluations on the PIE and JAAD benchmarks demonstrate that MTF-Net surpasses recent transformer- and graph-based models, achieving up to 0.95 AUC on PIE and 0.94 AUC on JAAD, while maintaining real-time performance. The results highlight that reliable pedestrian intention prediction arises from principled multi-modal fusion rather than excessive architectural complexity.

[CV-50] Needles in a Raystack: Ultra-Sparse LiDAR Occupancy Detection for Bat Tracks

链接: https://arxiv.org/abs/2609.20160
作者: Nico Klar,Pankaj Rana,Nizam Gifary,Jakob Traub,Aamir Ahmad
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 6 pages, 4 figures. Accepted and presented at the AI4Nature@AVSS 2026 Workshop of the 22nd International Conference on Advanced Visual and Signal-Based Systems (AVSS 2026), Lecce, Italy

点击查看摘要

Abstract:Monitoring flying animals is important for understanding and protecting biodiversity, but nocturnal species such as bats are difficult to observe in the field. Using LiDAR, bat movements at night result in ultra-sparse 3D spatio-temporal data in which standard reconstruction losses tend to predict only background and miss real flight paths. We study this problem as voxel-wise occupancy detection in sensor-centric LiDAR raystacks. A lightweight 3D U-Net is proposed that preserves temporal resolution, uses skip connections for spatial detail, and combines weighted binary cross-entropy with Dice loss to handle the strong class imbalance. In real LiDAR recordings of bats over open fields, cross-checked with acoustic monitoring, a reconstruction-based 3D convolutional autoencoder baseline fails to recover foreground trajectories. In contrast, the proposed U-Net recovers sparse foreground occupancy in diagnostic experiments and produces coherent occupancy patterns along bat flight trajectories, providing a practical basis for validation-scale experiments, later clustering of flight tracks, and future integration of bat activity information into biodiversity-aware turbine curtailment strategies. Comments: 6 pages, 4 figures. Accepted and presented at the AI4Nature@AVSS 2026 Workshop of the 22nd International Conference on Advanced Visual and Signal-Based Systems (AVSS 2026), Lecce, Italy Subjects: Computer Vision and Pattern Recognition (cs.CV) Cite as: arXiv:2609.20160 [cs.CV] (or arXiv:2609.20160v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2609.20160 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[CV-51] Ischemic Stroke Segmentation and Net Water Uptake Quantification on Multicenter Non-Contrast CT Using Supervised Target-Domain Adaptation

链接: https://arxiv.org/abs/2609.20151
作者: Linus Britt,Maximilian Nielsen,Susan Klapproth,Andre Kemmling,Michael H. Lev,Gabriel Broocks,Rene Werner,Thilo Sentker
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Objectives: Quantitative assessment of infarct hypodensity on non-contrast computed tomography (NCCT), including net water uptake (NWU), requires manual or semi-manual lesion delineation, often guided by CT perfusion or diffusion-weighted MRI, limiting clinical applicability. Automated segmentation on NCCT could enable efficient biomarker extraction such as NWU but remains challenging across heterogeneous multicenter data. This study aimed to develop and externally test a domain-aware deep learning framework for ischemic stroke segmentation on NCCT and assess its suitability for NWU quantification. Materials Methods: In this retrospective multicenter study of 801 patients from four datasets, an nnU-Net-based model was trained on NCCT scans from the University Medical Center Hamburg-Eppendorf and the Acute Ischemic Stroke Dataset. To adapt to new domains, the model was fine-tuned on target-domain subsets from Boston (n=11) and ISLES (n=75), with evaluation on held-out cases not used for fine-tuning. Automated segmentations and NWU values were compared with expert references. Results: For lesions \geq 30 mL, median Dice was 0.68 (Boston) and 0.56 (ISLES). Including smaller lesions, which predominated in ISLES, median Dice was 0.54 (interquartile range [IQR] 0.30-0.70) for acute lesion segmentation (Boston dataset) and 0.20 (IQR 0.03-0.41) for NCCT lesion segmentations when compared to post-treatment infarct (primary target of the ISLES challenge). Automated NWU mean absolute error was 1.37 percentage points (SD 1.61, Boston). Conclusion: Target-domain adaptation supported NCCT-only infarct segmentation across heterogeneous external cohorts, although performance varied across domains. The approach enabled low-error NWU quantification from baseline NCCT without advanced imaging, supporting further prospective clinical evaluation. Subjects: Computer Vision and Pattern Recognition (cs.CV) Cite as: arXiv:2609.20151 [cs.CV] (or arXiv:2609.20151v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2609.20151 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Thilo Sentker [view email] [v1] Thu, 17 Sep 2026 12:44:00 UTC (4,594 KB)

[CV-52] ask-Oriented Semantic Feature Transmission for Multi-Task Satellite Remote Sensing over Low-SNR Channels

链接: https://arxiv.org/abs/2609.20150
作者: Shuoyuan Sun,Hongyu Wang,Mugen Peng,Wenjia Xu
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Conventional satellite remote sensing transmission follows a reconstruct-then-infer paradigm that optimizes pixel-level fidelity, creating an objective mismatch with downstream tasks such as classification and detection, especially at low SNR. This paper investigates a task-oriented framework that bypasses image reconstruction and directly transmits semantic features extracted by a multitask-pretrained backbone. A lightweight channel adaptation module (CAM) compresses feature dimensionality for bandwidth reduction, and a feature restorer recovers task-relevant structure after channel corruption. With the backbone frozen, the CAM and task-specific downstream heads are jointly optimized with task and feature-level supervision under random-SNR training. Under the adopted AWGN setting, experiments on scene classification and object detection show consistent gains over reconstruction-oriented JSCC baselines across different SNR conditions, with the largest improvements in the low-SNR regime.

[CV-53] Bridging Modalities on the Cortex: Surface-based MRI to PET Translation with a Diffusion Bridge

链接: https://arxiv.org/abs/2609.20147
作者: Yitong Li,Alexandra Samoylova,Fabian Bongratz,Timo Grimmer,Dennis M. Hedderich,Igor Yakushev,Christian Wachinger
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Cortical hypometabolism measured by Fluorodeoxyglucose Positron Emission Tomography (FDG-PET) is a highly sensitive biomarker for dementia diagnosis. However, high costs, radiation exposure, and limited accessibility constrain its clinical utility. While cross-modal synthesis from Magnetic Resonance Imaging (MRI) offers a promising alternative, existing volumetric generation methods do not explicitly account for the highly folded cortical geometry, where disease-related patterns predominantly reside. To address this, we introduce a novel surface-based diffusion bridge framework DB-SUiT for MRI-to-PET translation that operates natively on the cortical manifold. A conditional Spherical U-shaped vision Transformer (SUiT) is specifically designed to model the intricate cross-modal relationships while preserving surface topology. It combines spherical convolutional encoders for multi-scale surface feature extraction with bottleneck Transformers to capture long-range spatial dependencies, while incorporating demographic and subcortical conditions to refine the synthesis. Evaluated on two datasets, including subjects with different dementia types, DB-SUiT demonstrates high-fidelity synthesis that substantially outperforms other baselines. In automated dementia classification, synthesized PET surfaces improve performance over MRI by 14.2% and PET volumes by 11.3%, approaching the performance of real PET surfaces. In a blinded reader study, synthetic PET achieved 85.5% diagnostic accuracy, compared with 75.8% for MRI and 95.2% for real PET. This further demonstrates cross-cohort and cross-pathology generalization, as the model was evaluated without retraining on an external cohort that included a dementia subtype not represented during training. Our code is available at this https URL.

[CV-54] Cross-Modal Attention Acts as a Frequency Filter: Why Verbose Prompts Improve Robustness in Vision-Language Models

链接: https://arxiv.org/abs/2609.20139
作者: Farooq Ahmad Wani,Maria Sofia Bucarelli,Mujtaba Hussain Mirza,Oleksandr Pryymak,Aryo Pradipta Gema,Iacopo Masi,Pasquale Minervini,Fabrizio Silvestri
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Vision-language models (VLMs) are fragile under image corruption. We find that the wording of the question affects VLMs in two opposite ways. Verbose questions make VLMs substantially more robust—e.g., rephrasing “Is there a cat?” into “Please look carefully and answer: is there a cat?”. Conversely, VLMs become more fragile under corruption when the question is semantically complex or finer-grained, e.g., “what colour is the cup left of the chair?” instead of “is there a cup?”. Both effects stem from question-conditioned cross-modal attention, which induces a spectral filter over image patches: verbose questions broaden its frequency support, while fine-grained questions concentrate it onto fewer visual scales. The model’s answer drifts most when this filter and the corruption sit on the same spatial frequencies. We test the filter view on Qwen3-VL and LLaVA-OneVision across GQA and CLEVR; verbose paraphrasing reduces drift variance by 70–81% on the 8B models. The practical recipe—pad the prompt—further yields measurable gains in accuracy, even under image corruption.

[CV-55] AnyviewMeter: Adapting Robotic Reward Models with Camera Geometry and Multi-View Attention

链接: https://arxiv.org/abs/2609.20106
作者: Yuang Tu,Runjia Tan,Yujie Yan,Jinghan Hu,Chen Lv
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 8 pages, 3 figures, 5 tables

点击查看摘要

Abstract:Robotic reward models evaluate task execution from visual observations, but their predictions can change with camera viewpoint and occlusion even when the underlying task state is unchanged. Adapting a pretrained reward model to a local task therefore requires accounting for how that task is observed. We introduce AnyviewMeter, a geometry-conditioned adaptation framework for robotic reward models that represent task progress as a scalar reward signal. It combines low-rank fine-tuning with token-aligned Plucker rays and synchronous block attention: ray conditioning incorporates camera geometry into visual features and attention queries and keys, while block attention fuses synchronized views inside the pretrained decoder. The framework supports both single-view reward prediction and joint multi-view evaluation through parameter-efficient adaptation of a pretrained Robometer model. On PickCube, single-view adaptation improves progress prediction in every camera group and reduces mean absolute error under a changed field of view by approximately 21% relative to RGB fine-tuning. Across simulated manipulation tasks, joint multi-view prediction reduces progress error by 41-69% compared with averaging single-view RGB predictions and improves temporal ordering in approximately 88% of task-camera groups. On real tasks with fixed and wrist-mounted cameras, mean absolute error decreases by approximately 21% relative to averaged RGB fine-tuning. These results support camera geometry and joint visual evidence as useful components of task-specific robotic reward adaptation.

[CV-56] A Smaller Transformer in Your Transformer BMVC2026

链接: https://arxiv.org/abs/2609.20100
作者: Dhananjay Tomar,Marius Aasan,Andreas Kleppe,Adín Ramírez Rivera
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 22 pages, 6 figures, 6 tables. Accepted at the 37th British Machine Vision Conference (BMVC 2026)

点击查看摘要

Abstract:Recent findings indicate that Vision Transformers settle into locally similar computational phases, implying a level of depthwise computational redundancy. However, existing methods to exploit this redundancy either fail to reduce inference compute or severely degrade model expressivity. In this work, we formalise a unified view of block redundancy that decouples the geometry from specific surrogate interventions. We then introduce Transformer-Within-Transformer (TWT), a post-hoc method that fuses contiguous groups of redundant layers into a single learned surrogate layer. TWT reduces parameter count and inference compute while remaining competitive with original models using half the depth on natural images, and in several downstream histopathology settings, TWT matches or even improves on the original baseline.

[CV-57] G2RA-NET: Graph-based Cross-Slice Relation Modeling with Attention Gating for Medical Image Segmentation

链接: https://arxiv.org/abs/2609.20088
作者: Shengye Wang,Zonglin Wu,Liang Fan,Yule Xue,Haozhe Zhao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 5 pages, 6 figures

点击查看摘要

Abstract:Medical image segmentation supports quantitative clinical analysis and computer-aided diagnosis. Recent methods for medical image segmentation have improved both local feature representation and volumetric context modeling. However, existing methods still strug- gle to efficiently model cross-slice relations in anisotropic volumet- ric images, limiting segmentation consistency and accuracy. This pa- per proposes G^2RA-Net, a medical image segmentation framework that combines graph-based cross-slice relation modeling with atten- tion gating. Graph-Based Slice Relationship Modeling (GSRM) cap- tures anatomical dependencies across consecutive slices by repre- senting each slice as a graph node and propagating semantic con- text through graph message passing. The Cross-Slice Attention Gate (CSAG) then selects relevant neighboring context and emphasizes target anatomical regions through attention-guided feature modula- tion. Experiments on brain MRI and abdominal CT datasets demon- strate that G^2RA-Net outperforms representative methods in seg- mentation accuracy and boundary quality. Ablation studies further validate the proposed design.

[CV-58] PointEvent: Rethinking Event-based Tiny Object Detection via Serialized Motion Evidence Accumulation

链接: https://arxiv.org/abs/2609.20066
作者: Zongze Wu,Baofeng Jia,Weiqi Yan,Jingyuan Zhang,Yu Zang,Xiaoyu Chen,Jing Han
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Code: this https URL

点击查看摘要

Abstract:Event cameras offer high temporal resolution and motion sensitivity for tiny UAV detection, yet distant targets generate sparse and fragmented events that are easily overwhelmed by clutter and ego-motion. Existing methods mainly rely on dense event representations or local sparse spatiotemporal modeling, resulting in redundant computation or fragmented modeling of motion continuity across distant asynchronous events. To address this limitation, we introduce serialized motion evidence accumulation, which treats motion continuity as an ordered evidence propagation process. Specifically, the same event stream is organized into locality-preserving spatiotemporal paths and chronology-preserving temporal paths through the latent complementary serializations. Based on this principle, we propose PointEvent, a lightweight event-wise state-space framework that alternates serialized scans across the complementary orders, progressively consolidating fragmented motion evidence beyond fixed local neighborhoods. A high-resolution event branch preserves fine-grained target responses, while compact context modulation suppresses interference. Experiments demonstrate that PointEvent achieves SOTA with the fewest parameters and fastest measured inference among the compared methods. Code: this https URL

[CV-59] A Free Lunch? Adapting PP-OCRv6 for Historical Text Recognition

链接: https://arxiv.org/abs/2609.20064
作者: Benjamin Kiessling(ALMAnaCH)
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Despite impressive reported scores, large vision-language models have seen limited practical uptake in historical automatic text recognition because of their computational cost, dependence on large-scale pretraining, and hallucination. Historical ATR therefore continues to rely largely on compact CRNN line recognizers, which are visually grounded and trainable on modest data. Lightweight recurrence-free recognizers promise the accuracy of larger models with the practical advantages of CRNNs, yet have not been comprehensively evaluated on historical writing. We adapt PP-OCRv6, a recent compact text recognizer without strong language modeling, for historical line recognition and compare it with a conventional CRNN across generalized pretraining, domain-specific training, corpus-level fine-tuning, and manuscript-specific few-shot adaptation on multilingual Latin- and Arabic-script material. While PP-OCRv6 does not consistently outperform the baseline when trained from scratch, heterogeneous pretraining produces markedly better generalization. Comparisons with the Qwen3.5-based Medusa recognizer further show that fine-tuned PP-OCRv6 can outperform a large VLM tailored towards historical Latin-script HTR.

[CV-60] Astronex-World 1.0: Real-Time Interactive World Model Foundation

链接: https://arxiv.org/abs/2609.20034
作者: Xin Zhou,Cong Miao
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注: Technical report. 25 pages, 13 figures, 10 tables. Project page: this https URL ; Code: this https URL ; Weights: this https URL

点击查看摘要

Abstract:We present Astronex-World 1.0, an open controllable video world-model foundation. Given a text prompt (text-to-video) or an initial observation (image-to-video), the model predicts future visual states under frame-aligned camera trajectories, continuous actions, and an embodiment identifier, and accepts text events inserted at a specified position of a rollout. The family provides a bidirectional model for full-context generation and a causal model with block-causal attention and cross-block KV caching for persistent generation, both built on the Wan2.2-TI2V-5B prior. PRoPE injects camera intrinsics and extrinsics, while a 64-dimensional action stream modulates every Transformer layer. A five-stage training path develops bidirectional camera and action control, converts the backbone to block-causal generation, distills a few-step student, restores mixed-domain dynamics, and applies asymmetric DMD/DMD2 distribution matching. The causal model generates 832x480 video at 24 fps. All five training stages run on two NVIDIA L20 48 GB GPUs, and the causal model streams in real time on one. It scores 73.5 on WBench Navi and 70.0 on WBench Full. On Full, this 5B model is above the 13.6B LongCat-Video and the 14B Helios, within one point of the 22B LTX-2.3, and above YUME 1.5, which is post-trained from the same 5B prior on NVIDIA A100 GPUs. The reserved action input and output interfaces allow post-training for embodied intelligence and autonomous driving.

[CV-61] GRF-Recon: Global Ray-Field Optimization for Long-Sequence Feed-forward Reconstruction ECCV2026

链接: https://arxiv.org/abs/2609.20012
作者: Enpeng Li,Yunzhou Zhang,Zhiyao Zhang,Dexuan Lyu,Chenyu Wang,Chiyuan Cui,Cheng Cheng
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to ECCV 2026 as a Spotlight presentation

点击查看摘要

Abstract:Feed-forward 3D reconstruction provides an efficient paradigm for scene modeling from image sequences. Scaling these models to large monocular scenarios are constrained by excessive GPU memory footprint, degraded local geometry, and long-term trajectory drift. Existing chunk-based optimization strategies provide limited geometric constraints and fail to maintain global consistency over extended trajectories. We present a unified framework for stable and scalable feed-forward 3D reconstruction from long monocular sequences. Our approach builds on coarse-to-fine trajectory alignment augmented by lightweight geometric prior injection. Distilling monocular geometric cues into the feed-forward backbone via LoRA adaptation improves depth accuracy on fine structures while preserving inference efficiency. We introduce a hybrid-weight sparse ray-field optimization that leverages high-frequency geometric features to guide local point-cloud refinement and enforce consistent inter-frame ray constraints. Unlike prior chunk-based methods, this establishes strong cross-frame geometric coupling while maintaining scalability. Finally, an efficient trajectory stitching strategy with joint ray-error optimization explicitly reduces accumulated drift. Extensive experiments show that our approach achieves competitive trajectory accuracy compared with representative SLAM systems, while maintaining globally consistent 3D reconstruction in large-scale scenarios.

[CV-62] AVTrace: Diagnosing Audio-Visual Temporal Reasoning in Omni Models

链接: https://arxiv.org/abs/2609.19991
作者: Longyin Zhang,Parth Sakhare Mahendra,Chengwei Wei,Ning Zhang,Lim Ming Chong,Sirui He,Ai Ti Aw
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Omni models can describe video content, but can they locate events in time, preserve event order, and judge audio-visual synchronization? We introduce AVTrace (Audio-Visual Temporal Reasoning Assessment and Capability Evaluation), a silver-standard diagnostic suite spanning onset and span grounding, synchronization, next-step prediction, cross-modal localization, chain parsing, and event-conditioned comprehension. It contains 34,114 training examples and category-balanced development and test splits of 3,500 and 7,000 examples. We evaluate five open omni models under their respective input configurations using reference-blind response normalization followed by deterministic scoring. All five off-the-shelf systems score below the test split’s majority-label baseline of 0.556 on synchronization verification, and obtain low scores on chain parsing and event-conditioned grounding and comprehension. Development-set perturbations reveal task-dependent sensitivity in Qwen3-Omni-30B to modality removal and changes in visual input processing, without isolating their underlying causes. Parameter-efficient temporal post-training improves Gemma4-E4B-it on several benchmark metrics. On three external image benchmarks, task metrics change modestly, including some degradations, while teacher-forcing perplexity decreases. Together, these findings show that semantic reference-text overlap should not be treated as a proxy for temporal localization, and that AVTrace can identify task-specific weaknesses while providing a testbed for temporal post-training.

[CV-63] QCPruner: Query-Conditioned Population Coverag e for Visual Token Pruning

链接: https://arxiv.org/abs/2609.19990
作者: Shengli He(1),Yongchao Liang(1),Roumeng He(2),Junjie Zeng(1),Jiyuan He(1),Can Wu(1),Li Zheng(1) ((1) Guizhou University, (2) Shanghai Ocean University)
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:The high visual-token load in multimodal large language models (MLLMs) motivates training-free pruning to reduce later-layer computation, but under a fixed budget, pruning must preserve query-relevant evidence while avoiding redundancy. Existing methods rank tokens, diversify selected subsets, or optimize coverage without using a shared per-visual query utility to weight both visual targets and candidate representatives. We introduce QCPruner, which makes both roles query-conditioned through bilateral utility weighting. Using keyword-matched query anchors, QCPruner fuses two cross-modal cues into utility and applies it to both visual targets and candidate representatives within visual-affinity-based coverage. The resulting nonnegative facility-location objective is monotone and submodular, retains the standard (1-1/e) greedy guarantee, and requires no model training or parameter updates. Across LLaVA-1.5, LLaVA-NeXT, LLaVA-Video, and Qwen2.5-VL, QCPruner achieves the highest average relative performance among evaluated complete-system pruning methods at every reported token budget. At 32 of 576 tokens on LLaVA-1.5-7B, it retains 96.1% of unpruned performance, versus 93.9% for the strongest evaluated baseline. At 256 of 1296 tokens on Qwen2.5-VL-7B, the corresponding values are 96.7% and 92.5%.

[CV-64] An Event Preserving Velocity Invariant Representation for Event Cameras ECCV2026

链接: https://arxiv.org/abs/2609.19973
作者: Mikihiro Ikura,Luna Gava,Jiahang Wu,Chiara Bartolozzi,Arren Glover
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: @inproceedings{ikura2026event, title={An Event Preserving Velocity Invariant Representation for Event Cameras}, author={Ikura, Mikihiro and Gava, Luna and Wu, Jiahang and Glover, Arren and Bartolozzi, Chiara}, year={2026}, booktitle={ECCV 2026 Workshop-Event-Based Multimodal Vision: From Imaging to Perception and Understanding} }

点击查看摘要

Abstract:Event cameras provide low-latency, high temporal resolution perception for real-time vision tasks such as this http URL novel circuitry (i.e. asynchronous, independent pixels) that enables these advantages also introduces new algorithmic challenges. Velocity-invariant representations alleviate missing observations under slow motion and motion blur under fast motion, but most discard temporal information by converting events into image-like representations. We propose Set of Centre Active Receptive Fields (SCARF), a real-time velocity-invariant representation that preserves raw events while consistently handling fast motion, stationary scenes, and independently moving objects. SCARF achieves state-of-the-art performance in both computational efficiency and representation quality.

[CV-65] Beyond the Foreground: FOV-Aware Polyp Image Synthesis via Lesion-Guided Adaptive Mucosal Context Propagation

链接: https://arxiv.org/abs/2609.19966
作者: Tong Wang,Yuting He,Bin Ren,Yutong Xie,Guanyu Yang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Synthetic image and mask pairs can alleviate scarce colonoscopy annotations, but realistic synthesis requires preserving the supplied lesion while generating compatible mucosa. Existing foreground-guided methods treat all non-foreground pixels as background and rely mainly on local integration. Directly applying them to colonoscopy causes two problems: non-mucosal black regions contaminate generated tissue, and local reasoning produces inconsistent mucosal texture and illumination. We propose LAMP, the first foreground-guided framework for polyp image synthesis based on lesion-guided adaptive mucosal context propagation. LAMP explicitly separates the lesion, valid mucosa, and camera exterior using a field-of-view (FOV) mask. Lesion-to-Mucosa cross-attention extracts lesion appearance conditions for valid-mucosa locations, while FOV-constrained multidirectional Vision Receptance Weighted Key Value propagates them over legal tissue support. An adaptive gate then controls their residual fusion into the diffusion U-Net. Extensive experiments on five polyp datasets demonstrate that LAMP substantially outperforms existing methods in overall generation quality and consistently improves five downstream segmentation models. Our code will be released at this https URL.

[CV-66] Enhanced Knowledge Distillation for Detection Transformer via Teacher Prediction Refinement

链接: https://arxiv.org/abs/2609.19964
作者: Yitong Xing,Yuhao Cheng,Yanping Li,Yichao Yan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Detection Transformers (DETRs) achieve strong performance in object detection but remain challenging to deploy on edge devices due to their high computational cost. Existing DETR distillation methods mainly focus on aligning distillation points, while largely overlooking the quality of the teacher’s supervision itself. We observe that due to stage-wise non-monotonic prediction behavior in DETRs, well-localized or correctly classified predictions from earlier stages may degrade in later ones, and some negative predictions become increasingly overconfident. As a result, relying solely on the current stage’s predictions yields inaccurate and inconsistent supervision. To address this issue, we propose Teacher Prediction Refinement Distillation (TPRD), a plug-and-play module that refines teacher predictions before distillation by exploiting stage-wise prediction information. TPRD improves supervision quality through Positive Prediction Correction (PPC), which corrects degraded positive predictions by restoring more accurate ones from earlier stages, ensuring reliable localization and classification signals, and Negative Prediction Suppression (NPS) suppresses the influence of overconfident negatives, preventing them from providing misleading supervision to the student. To preserve informative dark knowledge, we further introduce Maximum Dark Knowledge Preservation (MDKP), which selectively refines target-class logits while retaining non-target relations. Extensive experiments on MS COCO and PASCAL VOC demonstrate the effectiveness and robustness of the proposed method. Our code is available at this https URL.

[CV-67] LapaTrack-3D: 6 DoF pre-operative shape tracking for laparoscopic surgery

链接: https://arxiv.org/abs/2609.19954
作者: Jingwei Song,Javid Hussain Jakir,Ray Zhang,Wenwei Zhang,Hao Zhou,Xiaomeng Xian,Maani Ghaffari
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: This paper has been accepted by IEEE Transactions on Medical Robotics and Bionics (T-MRB)

点击查看摘要

Abstract:This work proposes a real-time 6 Degree-of-Freedom (6 DoF) tracking algorithm for monocular laparoscopic surgery. It provides alignment between intra-operative video and pre-operative data (e.g., CT). The 6 DoF tracking offers a solution for accurately locating the internal anatomy of the target organ despite the lack of tactile feedback and transparency. The ORB-SLAM2 framework is adopted and modified for prior-based 3D tracking with four major modifications. First, the primitive 3D shape is used for fast initialization of the ORB-SLAM2 monocular mode. Second, a pseudo-segmentation strategy is employed to separate the target organ from the background for tracking. Third, the 3D shape is incorporated as a geometric prior in its pose graph optimization. Fourth, the Multi-Scale Retinex with Chromaticity Preservation (MSRCP) algorithm is leveraged and modified for image enhancement in challenging illumination scenarios. In-vivo and ex-vivo experiments validate that LapaTrack-3D provides robust 3D tracking and effectively handles typical challenges such as poor illumination, fast motion, out-of-field-of-view scenarios, partial visibility, and ``organ-background’’ relative motion. LapaTrack-3D achieves a processing rate of 13 Hz for 1280*720 pixel video.

[CV-68] DirtyMoCap: Robust Motion Capture from Unconstrained Markers

链接: https://arxiv.org/abs/2609.19927
作者: Long Wang,Shuting Zhao,Shen Yan,Siyuan Yu,Xiaoben Li,Zeyu Cai,Yumeng Hou,Yuliang Xiu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Homepage: this https URL

点击查看摘要

Abstract:Optical motion capture delivers high-fidelity human motion, but its reliance on strict marker layouts and clean trajectories severely limits its real-world applicability. In practice, tracking systems frequently output unconstrained markers: sparse, noisy, and unordered point clouds with unknown or varying configurations. To bridge the gap between corrupted raw markers and parametric human models, we introduce DirtyMoCap, a robust, marker-layout-free framework. Our core insight is to map unordered marker observations to a fixed set of “proxy anchors” comprising skeletal joints and body surface points, which serve as a stable intermediate representation. We first initialize and track these anchors over long sequences using a recurrent sliding-window architecture. Then, a custom differentiable Gauss-Newton solver fits the SMPL-H model to the tracked anchors to recover full-body pose, translation, and shape. By explicitly deriving geometric residuals, our solver learns adaptive observation confidence, smoothness, and prior weights end-to-end, adapting dynamically to the reliability of the input data. Extensive experiments on diverse, noisy marker configurations demonstrate that DirtyMoCap successfully generalizes across arbitrary layouts using only a single trained model. It consistently outperforms state-of-the-art configuration-specific baselines in both joint and vertex reconstruction accuracy, while our custom CUDA solver achieves up to a 100x speedup over standard PyTorch implementations. We further apply DirtyMoCap to heterogeneous raw optical MoCap recordings of traditional Chinese martial arts, yielding a Kung Fu motion dataset of temporally coherent SMPL-H reconstructions. Code and data are available at this https URL.

[CV-69] CitySTAR: Structured and Topology-Aware Reasoning for Open-Vocabulary Urban 3D Grounding

链接: https://arxiv.org/abs/2609.19911
作者: Shuai Zhang,Hongye Hou,Qinghe Liu,Zhuoxiao Li,Dongli Wu,Jing Ou,Yuan Liu,Wufan Zhao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:3D grounding aims to localize target entities in complex scenes from natural language and plays a fundamental role in embodied perception and spatial reasoning. However, existing approaches mostly rely on feature similarity or direct matching, making it difficult to connect natural-language intent with the implicit semantic and geometric structures hidden in billion-scale urban point clouds. We reformulate city-scale 3D grounding as structured constraint reasoning, where description semantics are organized into computable cross-modal constraints over open-vocabulary 3D entities, attributes, and spatial relations. We present CitySTAR, a training-free framework for reasoning-driven urban 3D grounding. CitySTAR lifts raw billion-scale urban point clouds into a query-ready scene graph of open-vocabulary 3D instances, with CodeLLM-driven tools supplying multimodal evidence for node attributes and 3D spatial relations. It then models target-context topology with paired hypergraphs and performs bidirectional topology verification for structural disambiguation. Finally, a Reflective Cross-modal Grounding module integrates topology consistency and candidate-centered 2D visual evidence to make decisions over a metric-aware 3D context graph. To further support this setting, we introduce CitySTAR-3D, an enhanced benchmark that improves semantic coverage, instance completeness, bounding-box fidelity, and spatial-relation complexity in city-scale 3D grounding. Extensive experiments show that CitySTAR consistently improves open-world urban 3D grounding while maintaining strong interpretability and generalization.

[CV-70] GS-PI: An Optimization-Decoupled Appearance Decomposition Approach for Generating PBR Gaussian Assets

链接: https://arxiv.org/abs/2609.19907
作者: Jieting Xu,Rengan Xie,Zijian Huang,Zehui Jin,Rui Wang,Yuchi Huo
类目: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
备注:

点击查看摘要

Abstract:Gaussian Splatting (GS) excels at novel-view synthesis but encodes baked-in radiance, tightly entangling illumination with geometry and preventing seamless integration into physically based rendering (PBR) pipelines. Existing inverse-rendering methods attempt to disentangle materials via joint optimization, but often suffer from competing objectives that cause severe ambiguities and residual lighting artifacts. To overcome this, we present GS-PI, a novel optimization-decoupled framework that casts PBR material generation as a geometry-conditioned diffusion process on 3D point clouds. By operating directly in the 3D domain, our method inherently guarantees multi-view consistency, sidestepping the severe pixel correspondence issues that challenge 2D diffusion approaches. We introduce a multi-scale cross-view conditioning mechanism that integrates three complementary components: a global semantic prior, source-anchored photometric cues, and an absolute spatial learned view-direction conditioning signal. This design efficiently compresses complex multi-view evidence, mitigating cross-view projection misalignment and successfully preventing specular highlights from baking into intrinsic colors. By extracting a point cloud from a pre-trained Gaussian model, predicting PBR attributes via conditional diffusion, and distilling them back through differentiable rasterisation, we yield a fully relightable PBR-GS asset. GS-PI outperforms recent inverse-rendering baselines while replacing per-scene joint illumination/BRDF optimization with a learned diffusion pass followed by a short target-driven distillation, without requiring proxy meshes.

[CV-71] BinoGen: Scaling egocentric binocular data for embodied visual perception and learning

链接: https://arxiv.org/abs/2609.19881
作者: Chunpeng Li,Ya-tang Li
类目: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
备注:

点击查看摘要

Abstract:Embodied visual perception relies on temporally coherent visual experience accumulated through continuous engagement with the environment. However, collecting large-scale egocentric binocular observations together with dense annotations remains costly and difficult. Moreover, visual experience is shaped not only by the environment but also by the embodiment of the observer, including viewing height, field of view, binocular geometry, and motion through the scene. To address these challenges, we present BinoGen, an automated framework for generating large-scale, embodiment-aware egocentric binocular visual experiences in indoor environments. BinoGen jointly models environmental and observer variation through generative scene synthesis, probabilistic object instantiation, appearance randomization, stochastic trajectory generation, and configurable binocular camera setups. The framework produces synchronized binocular videos together with dense multimodal supervision, including depth maps, optical flow, surface normals, semantic maps, object coordinates, and camera poses. Using BinoGen, we construct a dataset comprising more than 20 million annotated images for supervised learning. We demonstrate two complementary utilities of BinoGen. First, incorporating BinoGen data consistently improves real-world visual perception, including depth estimation, object detection, and video object tracking. Second, paired human-inspired and mouse-inspired observations from the same environments enable controlled investigation of how observer embodiment affects perceptual learning. Embodiment-specific adaptation substantially improves performance, while joint training enables a single model to perform competitively across both embodiments. Together, these results demonstrate that large-scale, controllable visual experience can improve embodied perception…

[CV-72] SlugTrails: An Egocentric Benchmark for Floor Plan Localization in Large Buildings

链接: https://arxiv.org/abs/2609.19876
作者: Yunqian Cheng,Roberto Manduchi
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 8 pages, 4 figures. Preprint. Code and data: this https URL

点击查看摘要

Abstract:Floor-plan-based indoor visual localization enables infrastructure-free positioning, but most methods are developed and evaluated in small residential environments unlike the large public buildings of real deployment. We introduce SlugTrails, a floor plan localization benchmark for large indoor spaces under realistic egocentric sensing: 30 Hz Aria glasses recordings across three campus buildings and six floors ( 22089 m ^2 of floor plan outline), CAD-derived floor plans with semantic classes and circulation space masks, and trajectories aligned into the floor plan frame using laser-surveyed anchors. One protocol covers three practical ways of gathering geometry under a limited field of view – a single walking frame, a stationary multi-view sweep, and a walking stream with odometry – so methods designed for different regimes are compared on the same buildings and ground truth. Evaluating five representative geometric and learned systems under their native sensing configurations, we find that stock checkpoints (official released weights) are near zero on SlugTrails (at most 0.004 R@1m30 ^\circ on walking single frames), while fine-tuning on SlugTrails improves every trainable family on all three tasks (e.g., F ^3 Loc 0.0 \rightarrow 0.141 single-frame and 0.03 \rightarrow 0.66 sequential), with gains compounding as observations accumulate. The same fine-tuned weights also improve cross-dataset generalization on LaMAR with no LaMAR training (sequential R@1m 0.048 \rightarrow 0.143 for F ^3 Loc and 0.063 \rightarrow 0.127 for UnLoc), whereas train-from-scratch on SlugTrails alone stays far below fine-tuning from stock weights – evidence that floor plan localization is currently limited by indoor data rather than by architecture. We release the dataset, protocols, and tools at this https URL.

[CV-73] BINDER: A Latent Variable Model for Probabilistic Medical Image Registration

链接: https://arxiv.org/abs/2609.19875
作者: Stefano Cerri,Amirhossein Hassankhani,Yaël Balbastre,Koen Van Leemput
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:We propose a new probabilistic model for general-purpose medical image registration that builds upon the mutual information registration criterion. It centers around a spatial interpolation technique that assumes latent voxel-wise correspondences between the images being registered. By exploiting these latent variables, we derive dedicated optimization and MCMC sampling techniques that only involve closed-form iterative updates. When applied to nonlinear registration, an efficient demons-like optimization algorithm is obtained that shows robust out-of-the-box performance across a variety of monomodal and multimodal registration tasks. We also demonstrate a corresponding sampler that can quantify, for the first time, uncertainty in multimodal registration scenarios with very high-dimensional 3D deformations. Our code, which we call BINDER (Bayesian INference for DEformable Registration), is freely available at this https URL.

[CV-74] PART: Learning 3D Part Assembly and Retrieval with Transformers SIGGRAPH

链接: https://arxiv.org/abs/2609.19872
作者: Ruchao Bao,Wenzheng Wu,Chucheng Xiang,Zhongyuan Liu,Yuan Liu,Jinxin Dong,Ligang Liu,Ziqi Wang
类目: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
备注: Accepted to SIGGRAPH Asia 2026 Conference Papers. 11 pages, 12 figures. Project page: this https URL

点击查看摘要

Abstract:3D assembly is fundamental to modern manufacturing and digital content creation. In this paper, we present PART, a unified transformer-based framework for 3D part retrieval and assembly: given a target shape and a part library, PART automatically selects the appropriate parts and predicts their 6-DoF poses to reconstruct the target. While prior work has achieved impressive progress on assembling a pre-defined set of parts, this more practical retrieval-based setting remains largely unexplored. The task faces three key challenges: (i) a combinatorially explosive search space that grows exponentially with library size; (ii) variable-length outputs, as different targets require different numbers of parts; and (iii) continuous 6-DoF pose estimation for part assembly. To address these, we formulate retrieval and assembly as a set prediction problem and design a novel transformer-based framework that retrieves parts and regresses their poses with variable-length output. Additionally, we exploit the duality between part pose estimation and target segmentation through joint training and a novel segmentation-enhanced optimization module. Finally, We curate a large-scale dataset of 80K+ shapes, and the results show that PART generalizes to scene layouts, image targets, and real-world scans. Project Page: this https URL.

[CV-75] Socialized UAV Cross-Task Learning: Towards Cross-Granularity Collaboration through Hierarchical Interaction

链接: https://arxiv.org/abs/2609.19867
作者: Xinjie Yao,Ruipu Zhao,Yunqi Zhu,Zhihe Fan,Zhoupeng Guo,Weihao Li,Zhen Wang,Qilong Wang,Pengfei Zhu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 9 pages, 6 figures

点击查看摘要

Abstract:Joint learning across heterogeneous tasks is often treated as task coupling through feature sharing, distillation, or auxiliary supervision. However, in cross-task learning, mismatched representational and supervisory granularities make such coupling prone to interference, teacher bias, or unidirectional collapse. We argue that cross-granularity learning is fundamentally a problem of hierarchical interaction regulation rather than simple task coupling. This issue is particularly evident in UAV perception, where visual shifts and detection–segmentation objectives naturally form coarse- and fine-grained knowledge sources. To systematically study this problem, we introduce CrossUAV, a UAV benchmark for joint object detection and instance segmentation that provides a unified evaluation platform for cross-granularity task collaboration. To address these challenges, we propose Cross-Granularity Socialized Collaboration (CGSC), a progressive and adaptive framework that regulates when, where, and how tasks exchange information across network hierarchies. CGSC progressively activates cross-task interactions and adaptively adjusts the strength according to task contribution, suppressing harmful interference while exploiting complementary coarse- and fine-grained structures. Extensive experiments demonstrate consistent improvements on both tasks, validating hierarchical dynamic interaction as an effective mechanism for cross-granularity collaboration.

[CV-76] Feeling Terrain Before Crossing: World Models for Off-Road Navigation

链接: https://arxiv.org/abs/2609.19863
作者: E-In Son,Dong-Wook Kim,Ji-Hoon Hwang,Kangsun Lee,Jisung Bae,Jung-Taak Kim,Seung-Woo Seo
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: 8 pages, 6 figures

点击查看摘要

Abstract:Navigation world models plan by foresight, predicting the future that each candidate action sequence produces and selecting the best, rather than mapping observations to actions directly. Unlike urban settings where a predicted scene is a sufficient proxy, off-road navigation hinges on the robot–terrain interaction, so the prediction must cover not only what the camera will see but what the robot will feel. However, existing scene-focused models do not predict how much the robot will slip, tilt or shake along a planned trajectory. Proprioception captures these dynamics directly and, when used as input, improves the prediction of the physical future. We present Feel-WM, the first off-road navigation world model that conditions on proprioception and predicts what the robot will feel alongside what the camera will see. The physical future takes the form of a future proprioceptive state and a failure risk, both learned from the robot’s own experience without human labels. The planner rolls out the physical future alongside the scene and weighs the predicted failure risk against goal similarity in a separable score. Experiments on real off-road data and in simulation demonstrate that Feel-WM outperforms visual-only navigation world models in open-loop planning and closed-loop rough-terrain navigation across wheeled and legged platforms. Deployed on a Husky on mountain trails, Feel-WM plans onboard, predicts rough ground ahead and steers around it, completing courses that an end-to-end policy fails.

[CV-77] PACE: Precise AI Cinematic Expression: A Typed Specification for Script-Grounded Previsualization and Geometric Conformance

链接: https://arxiv.org/abs/2609.19853
作者: Bing Duan,Qiang Guo,Linpu Li,Zhijian Mao,Min Zhu,Zhirui Ren,Yiwei Yan,Xi Chu,Xiaoding Li
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 36 pages, 7 figures, 3 tables. Code: this https URL

点击查看摘要

Abstract:Between a screenplay and a film sits a planning problem that is spatial first: who stands where, and what a camera sees from where it stands. An image diffusion model asked for a shot in free text settles that plan by its own defaults. We present PACE (Precise AI Cinematic Expression), a typed representation for the plan: the screenplay evidence, the characters, props and locations it needs, where each subject stands, and what the camera does. A value is written once at the level it belongs to (script, scene, shot or panel) and inherited below it. A compiler turns the result into both the prompt sent to the diffusion model and a 3D scene built in metres, and a camera solver places the camera so that the declared framing is the framing built. Where a declared value becomes geometry, PACE measures, field by field, how far the compiled camera and the staged render sit from the declaration, rather than asking a model to judge. On the 11-scene Automatic Drive screenplay, every staged single-subject panel places its subject within 1.2% of frame width of its declared position; with two or three subjects one camera pose cannot satisfy every position, and the residual is reported rather than absorbed. On 204 external director-storyboard shots, delivered head height is 1.906 times the staged target from the director’s words, 1.733 from the compiled prompt, and 0.955 with the greybox control; the condition that holds framing best draws the described action least. Declaring the pose on 30 shots raises the action drawn from 58.9% to 74.4% without moving the framing. Transitions, fitted motion and human review of the generated panels remain open. Code: this https URL Comments: 36 pages, 7 figures, 3 tables. Code: this https URL Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) ACMclasses: I.2.10; J.5 Cite as: arXiv:2609.19853 [cs.CV] (or arXiv:2609.19853v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2609.19853 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[CV-78] KoUniTalk: A Lightweight Articulation-Centered Korean-English 3D Talking Face Benchmark

链接: https://arxiv.org/abs/2609.19840
作者: Hyunjung Chung,Unsang Park
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 22 pages, 5 figures; includes supplementary material

点击查看摘要

Abstract:High-quality 3D talking face datasets remain largely English- centric, and Korean 3D facial motion data are difficult to combine with standard English benchmarks because of differences in mesh topology, spatial scale, coordinate system, and temporal sampling. We present KoUniTalk, a lightweight articulation-centered Korean-English 3D talk- ing face benchmark that retargets VOCASET and the released Korean speech-based 3D talking face data to a shared mesh topology using de- formation transfer. Rather than proposing a new deformation-transfer algorithm or a full-head identity-preserving avatar dataset, KoUniTalk provides an identity-neutral canonical output space for controlled speech- driven facial articulation training and evaluation across English and Ko- rean. The unified template contains 1,176 vertices and focuses on the mouth and adjacent lower- and mid-face regions, reducing the output dimensionality from 15,069 and 72,147 dimensions to 3,528 dimensions, corresponding to 4.27-fold and 20.45-fold reductions compared with VO- CASET/FLAME and the original Korean mesh, respectively. To exam- ine whether retargeting preserves speech-relevant motion, we evaluate semantic mouth-landmark trajectories, including mouth opening, mouth width, aperture ratio, and mouth-opening dynamics. Since the official test set of the Korean dataset is not publicly released, we additionally define a subject-disjoint Korean benchmark split. The processed matched benchmark contains 22 speakers, 4,978 sequences, and 642,781 frames, enabling Korean-English cross-dataset evaluation of speech-driven 3D fa- cial animation models in a single compact articulation-template space. Source-reported inventory counts are listed separately from these pro- cessed counts

[CV-79] SnapPhysics: A Physics-Aware Scene Graph from a Single View for Interactive Mixed Reality Scenes

链接: https://arxiv.org/abs/2609.19815
作者: Suji Kang,Seok-Young Kim,Young Bin Kim,Taewook Ha,Dieter Schmalstieg,Shohei Mori,Woontack Woo
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted for publication in IEEE ISMAR, 2026

点击查看摘要

Abstract:We propose SnapPhysics, a training-free framework that reconstructs 3D objects and estimates their physical properties such as mass, friction, and center of gravity from a single image. For physically coherent interactions in mixed reality (MR), such properties are as important as geometry. Prior approaches infer them by analyzing object dynamics in video, which is computationally costly, or by querying vision-language models (VLMs) on single images, which lacks geometric grounding and inter-object relationships. We address these limitations by combining instance-level 3D reconstruction and spatial alignment with a physics-aware scene graph that encodes these relationships and per-object metric geometry as structured context for VLM-based property reasoning. Experiments on 3D-FRONT show that SnapPhysics improves scene-level F-Score by 18.6% over the best learning-based method, and on real captured scenes with ground-truth mass, it reduces the mean absolute log difference error (mALDE) by up to 20.5% and improves log-scale correlation ( r^2_\mathrmls ) by up to 19.6% over VLM-only estimation. SnapPhysics enables physically interactive MR experiences without manual parameter tuning. Project page: this https URL.

[CV-80] Absence is Presence: Understanding Visual Scene Negative Events Under Safety Cognitive Constraint

链接: https://arxiv.org/abs/2609.19812
作者: Zhiyun Jiang,Hanyong Wang,Binbin Liang,Yu Xie,Menglong Yang,Wei Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Traditional scene understanding focuses on affirmative information objectively present in images. However, in safety-critical domains, comprehending key information that should exist but is actually absent is vital for risk mitigation. To bridge this gap, we focus on visual scene negative captioning with safety as the cognitive constraint. The core challenge is to convert physical absence into semantic negative events. Existing vision-language models (VLMs) struggle with this process because affirmation bias suppresses negative reasoning, while limited mental filling capability and representation bias further hinder the inference of absent information. To address these challenges, we propose a negative captioning framework based on counterfactual reconstruction and contrastive decoding (CRCD). Inspired by human cognition, CRCD reformulates the task as counterfactual latent change captioning to bypass affirmation bias. It contrasts a synthesized safe expectation with reality to identify semantic omissions. To address limited mental filling, we design a dual-branch counterfactual reconstruction architecture. The amodal completion branch restores defective objects, while the functional association branch infers completely absent safety objects. Concurrently, a multi-condition representation learning mechanism is integrated to mitigate representation bias by projecting universal features onto predefined safety criteria subspaces, thereby capturing information across more dimensions. By decoding feature-level semantic residuals between the reconstructed scene prototype and raw input, CRCD bounds the non-existence search space and activates the decoder’s negative logic. Extensive experiments validate the effectiveness of CRCD, establishing a high-performance baseline for this pioneering task.

[CV-81] AI Smart Glasses for Wearable Intelligence: From Egocentric Sensing to Agent ic Personalization

链接: https://arxiv.org/abs/2609.19793
作者: Xu Yuan,Yi Wang,Zhuohang Jiang,Haohao Qu,Yujuan Ding,Shanru Lin,Guoliang Xing,Hongxia Yang,Jiannong Cao,Qing Li,Wenqi Fan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Recent advances in artificial intelligence (AI) are reshaping smart glasses from egocentric capture and display devices into platforms for wearable intelligence. Smart glasses increasingly serve as wearable AI systems that connect first-person observation with real-time assistance under strict form-factor constraints. We frame this transition through the lens of \emphAI smart glasses and define them as a system-level concept in which egocentric sensing, resource-aware computing, intelligent reasoning, multimodal interaction, and real-world application constraints are co-designed for personalized assistance in the physical world. To systematically study this perspective, we organize the survey around four connected dimensions. First, we examine the hardware foundation that bounds sensing, computation, feedback delivery, and sustained deployment. Second, we study wearable intelligence, where egocentric signals are transformed into perceptual, contextual, and agentic capabilities. Third, we discuss interaction design, through which users request, receive, correct, and regulate assistance during ongoing activity. Fourth, we analyze application scenarios across healthcare, accessibility, situated learning, daily life assistance, cultural tourism, and industrial support, showing how domain requirements reshape system design and evaluation. We further identify five cross-cutting research challenges for future AI smart glasses: next-generation hardware, trustworthy egocentric intelligence, lifelong personalized memory, proactive intelligence, and embodied foundation models. By centering smart glasses as wearable-intelligence platforms, this survey provides a unified framework for organizing technologies, applications, and open challenges in this emerging area.

[CV-82] Printing the Underdetermined: Materializing Multi-solutionness in Figurative Paintings

链接: https://arxiv.org/abs/2609.19782
作者: Yutao Ming,Teng Xu,Youjia Wang,Yunyang Liu,Fengmin Yang,Fuqiang Zhao,Jingyi Yu,Hua Yang,Yanjun Zhou
类目: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
备注: 10 pages, 5 figures

点击查看摘要

Abstract:Figurative paintings are often approached as if they depict a single recoverable 3D scene: viewers infer depth and occlusion, and reconstruction pipelines attempt to converge to one stable model. We instead foreground multi-solutionness, the non-uniqueness of 3D configurations compatible with a single painted image, and propose a workflow that keeps this non-uniqueness visible and material. Multi-solutionness arises from two sources: unobserved content, where backsides and occluded volumes admit multiple plausible completions, and observed cues, where perspective, shading, and occlusion still underconstrain geometry. When additional views are synthesized by a video generative model without explicit 3D constraints, small frame-level drifts become inevitable rather than exceptional. Our pipeline samples multiple camera-orbit multi-view video sequences from one painting, reconstructs each sequence with 3D Gaussian Splatting into a point-based Gaussian scene representation where density halos and ghosting expose unresolved degrees of freedom, and fabricates these representations as physical artifacts using DreamPrinting. By treating multiple compatible interpretations as explicit outputs rather than residual error, we provide a computational framework for spatial readings of figurative painting that can be inspected, compared, and discussed in both digital and physical form.

[CV-83] Benchmarking MLLM s via Cognitive Expected Scene Graph for Safety-Critical Visual Negation Understanding

链接: https://arxiv.org/abs/2609.19767
作者: Zhiyun Jiang,Hanyong Wang,Binbin Liang,Yu Xie,Menglong Yang,Wei Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:True machine intelligence requires transcending passive pixel registration to master top-down functional reasoning over absent information via visual negation understanding. However, unconstrained visual negation paradigms remain overly open-ended, and pervasive affirmation bias causes both existing Multi-Modal Large Language Models (MLLMs) and evaluation metrics to fail under negative semantics. To solve these intertwined challenges systematically, we first anchor the boundaries of negation reasoning within specific cognitive goals. Specifically, by focusing on safety as a highly pragmatic and critical cognitive dimension, we define the task of \textbfScene \textbfNegation \textbfUnderstanding under \textbfSafety Cognition (\textbfSNUS). Under this framework, we construct a high-fidelity negative caption dataset mapping dense assertions of localized hazards. Concurrently, we propose the Cognitive Expected Scene Graph (CESG) Score, a structure-grounded, polarity-aware evaluation metric. Extensive experiments demonstrate that while current models struggle on the task, traditional metrics completely collapse under semantic reversals. Conversely, our framework delivers a solid benchmark for SNUS, providing a rigorous foundation to advance risk-aware situational comprehension and counterfactual cognition.

[CV-84] STAR: Structure-aware Test-time Adaptation for diffusion-based light field Reconstruction

链接: https://arxiv.org/abs/2609.19747
作者: Wontae Choi,Ki Ryum Moon,Jae Young Lee,Hyung Sup Yun,Il Yong Chun
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 5 pages, 3 figures, 2 tables

点击查看摘要

Abstract:Light field (LF) reconstruction from limited and noisy focal stack (FS) measurements is a highly ill-posed inverse problem. Although the LF-to-FS imaging geometry is fixed for a given optical setup, LF spatial-angular structure—including within-view spatial details, cross-view angular dependencies, and disparity across views—varies across scenes. Consequently, a fixed pre-trained prior may not optimally capture the spatial-angular structure of each test LF. We propose Structure-aware Test-time Adaptation for diffusion-based light field Reconstruction (STAR), the first test-time adaptation framework for reconstructing an LF from FS. For each test LF, STAR freezes a pre-trained diffusion prior and fits three lightweight adapters to the observed FS to jointly adapt the three components of the LF’s spatial-angular structure. STAR outperforms existing state-of-the-art methods in both two- and three-focal-sheet settings, with shorter inference times than those with test-time parameter updates.

[CV-85] Region-Level Policy Optimization for Fine-grained MLLM Perception

链接: https://arxiv.org/abs/2609.19745
作者: Yuheng Shi,Xiaohuan Pei,Minjing Dong,Chang Xu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Fine-grained visual perception in MLLMs is commonly improved by raising the resolution, but the added visual tokens inflate vision-encoding and language-model prefilling costs. We show that the two operations underlying fine-grained perception, localizing the region of interest (RoI) and recognizing its content, have different resolution requirements. In a controlled diagnostic, localization tolerates roughly 3 to 4 times stronger token compression than recognition, which motivates localizing from a coarse view and concentrating resolution on the selected evidence. Decoding coordinates with the MLLM can be trained end-to-end from answers, but costs a full model pass per query and depends on grounding ability. A lightweight proposal network distilled from the model’s attention is fast, but inherits the noise of its attention targets. The RoI from the proposal network reaches the answer through a discrete region choice, so its faithfulness to the answer cannot supervise the network. We therefore optimize the proposal network with region-level reinforcement learning, which we call Vision-RL2. It treats coherent regions as actions, and a frozen MLLM reader scores each one by how its removal changes the answer likelihood. Complementary subtractive and additive objectives suppress distracting proposals and recover missing evidence, updating only the predictor without region annotations, response sampling, or reasoning trajectories. The refined proposal further enables a sparse encoding that magnifies evidence and excludes background tokens. Across six fine-grained benchmarks and four MLLM backbones, Vision-RL2 improves accuracy over the base model at every token budget and surpasses its largest-budget accuracy with about 4 times fewer visual tokens. Code is available at this https URL .

[CV-86] Federated Learning Framework for Privacy-Preserving Kidney Stone Detection

链接: https://arxiv.org/abs/2609.19740
作者: Najiyya Younas,Omar Abdulkader,Yaser Ali Shah,Muhammad Jawad Ikram,Jebran Khan,Amaad Khalil
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Recent innovations in deep learning have significantly enhanced the diagnosis of medical images, although they are based on the use of centralized data storage that pose severe threats to patient privacy and medical data security. To address this issue, this research proposes a Federated Learning (FL) model that is coupled with an optimized YOLOv8 network to detect the kidney stones on a computed tomography (CT) image and at the same time, protect privacy of the patients. The suggested system can help various medical organizations to jointly train a common model without exchanging the information about the patients. This is to ensure that data protection laws like GDPR and HIPAA are adhered to. The residual feature fusion and DropBlock regularization among other architectural improvements are also included in YOLOv8 to enhance detection robustness and minimize overfitting. Experimental analysis carried out on a distributed CT dataset demonstrated that the federated YOLOv8 model has a mAP at 50 of 0.733 and is able to keep the data confidential. Moreover, its lean design facilitates fast edge deployment and real-time inference across a clinical setting. Altogether, these findings indicate that Federated Learning is a safe and efficient solution to AI-assisted diagnosis in contemporary healthcare when combined with the use of sophisticated object detection models.

[CV-87] Recency Forcing: Bridging the Long-Horizon Gap in Autoregressive Video Generation

链接: https://arxiv.org/abs/2609.19729
作者: Tri Cao,Hung Nguyen,Phong Nguyen,Khoi Nguyen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Autoregressive (AR) video generation degrades over long horizons due to an overlooked train-inference discrepancy we term KV eviction mismatch: models train on short clips where all context frames reside in the KV cache, but at inference, memory constraints force distant frames to be evicted from the KV cache - removing context the model was conditioned on. Rather than simulating eviction via context truncation - which discards temporal information the model still needs and degrades motion coherence - we keep the context but while progressively reducing the influence of distant frames, making their eventual eviction negligible. To guide this design, we introduce the positional response R( \Delta, , t_\textdenoise) , a perturbation-based sensitivity measure revealing that context influence decays steeply with temporal distance and varies systematically across denoising steps. Motivated by this analysis, we propose Recency Forcing, which applies a non-positive, timestep-dependent bias, termed Temporal Response Bias (TRB), on pre-softmax attention logits derived directly from R , closing the train-inference gap without modifying context length or training objectives. We further introduce Biased Attention Reparameterization (BAR), an exact reformulation that moves the bias outside the softmax, making TRB a standard FlashAttention call at zero overhead. Recency Forcing operates in both training-free mode and training-based mode. Experiments on VBench and VBench-Long demonstrate state-of-the-art long-horizon generation quality at no additional inference cost.

[CV-88] SeetaPsych v1.0: An Open-source Computer Vision Toolkit for Behavior-based Psychological Measurement

链接: https://arxiv.org/abs/2609.19719
作者: Jiabei Zeng,Chiqin Li,Kaizhou Li,Fei Chang,Yong Li,Yuanhao Zhao,Dan Han,Wenqiang Yang,Xilin Chen,Shiguang Shan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Automated visual analysis opens new avenues for behavior–based psychological measurement. Nevertheless, existing technological modules are typically scattered across task specific systems with heterogeneous interfaces and disparate deployment requirements. In this work, we present SeetaPsych v1.0, an open source, unified and extensible computer vision toolkit designed to extract psychologically relevant signals from facial images and/or face based videos. The current release encompasses four major core modules aiming at behavior–based physiological perception: unified face based emotion analysis (simultaneous facial expression recognition, facial action unit detection, and valence–arousal estimation), camera based heart rate estimation, screen point–of–gaze estimation, and scene gaze following. A suite of auxiliary preprocessing modules for human centric visual analysis is also included, comprising face detection, facial landmark detection, and head detection. These functionalities are encapsulated within a modular Pipeline/Runner architecture that automatically resolves attribute dependencies, constructs computation graphs, and support intermediate result sharing among modules. SeetaPsych provides standardized Python APIs to facilitate reproducible, large scale analyses, alongside an interactive WebUI for rapid, code–free method evaluation. Overall, SeetaPsych offers an integrated and accessible visual measurement platform for research in psychology, behavioral science, human computer interaction, and related fields.

[CV-89] GAPrompt: Multi-Granular Geometry-Aware Point Cloud Prompt for 3D Vision Model

链接: https://arxiv.org/abs/2609.19716
作者: Zixiang Ai,Zhenyu Cui,Yufei Guo,Wenwen Qiang,Lei Chen,Jiwen Lu,Jiahuan Zhou
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by TPAMI 2026. Code at this https URL

点击查看摘要

Abstract:Pre-trained 3D vision models have substantially advanced point cloud analysis, yet adapting them to downstream tasks via full fine-tuning is computationally expensive and storage-intensive. Parameter-Efficient Fine-Tuning (PEFT) offers a promising alternative by reducing both adaptation cost and storage burden. However, existing prompting-based approaches ignore the intrinsic geometric structures of point clouds, thereby limiting their adaptation capability. This limitation stems from their inability to encode both fine-grained geometric cues and coarse-grained structural semantics, as well as failing to propagate such information effectively through the model hierarchy. To address these challenges, we propose GAPrompt++, a multi-granular geometry-aware prompting method that provides richer geometric guidance for efficient 3D task adaptation. Specifically, we introduce a Point Shift Prompter that extracts multi-granular geometric features across different scales, enabling instance-specific geometric adjustments during adaptation. Next, a Keypoint Prompter adaptively generates point-level prompts to highlight local geometric saliency and fine-grained structural details. Furthermore, a Prompt Propagation mechanism injects these multi-granular geometric cues throughout the feature extraction hierarchy, strengthening the ability to capture essential geometric characteristics. Extensive experiments show that GAPrompt++ achieves state-of-the-art performance among prompting-based PEFT methods and even surpasses full fine-tuning across diverse benchmarks, while requiring less than 2% trainable parameters. In addition, to address the saturation of existing evaluation datasets, we construct two more challenging benchmarks derived from 3D Gaussian Splatting and Multi-View Stereo reconstruction, offering diverse and realistic point cloud scenarios to promote future research.

[CV-90] Understanding and Exploiting Diagonal Attention Sparsity in Autoregressive Image Generation

链接: https://arxiv.org/abs/2609.19702
作者: Daeun Kim,Junwha Hong,Changhun Oh,Yoonsung Kim,Yoonhyeong Lee,Jongse Park
类目: Computer Vision and Pattern Recognition (cs.CV); Performance (cs.PF)
备注:

点击查看摘要

Abstract:Autoregressive image generation has emerged as a paradigm for multimodal AI systems due to its compatibility with transformer-based LLM serving infrastructures. However, generating thousands of visual tokens per request makes decoding increasingly bottlenecked by KV cache accesses during attention computation. Sparse attention is particularly attractive for this workload because many visual generation applications tolerate moderate quality degradation in exchange for improved performance and efficiency. While sparse attention has been extensively explored for text-based LLM inference, it remains unclear whether its sparsity assumptions generalize effectively to autoregressive image generation. We present the first systematic characterization of attention sparsity in autoregressive image generation across diverse workloads and representative open-source models. Our analysis reveals several distinguishing properties, including a pronounced prefill-decode asymmetry, strong attention concentration on prompt and local tokens, and a unique diagonal attention sparsity pattern arising from the spatial locality of visual tokens. Motivated by these observations, we propose a diagonal-aware sparse attention mechanism that selectively skips KV entries along the diagonal attention direction within a recent window. Implemented on top of a GPU-based serving system using FlexGen, FlashAttention-2, and custom kernels, our approach achieves up to 3.1x throughput and 1.19x latency improvements with less than 2% quality degradation compared to dense inference.

[CV-91] IMFD: End-to-end Multi-Face Forgery Detection through Instruction-based Large Vision-Language Models AACL

链接: https://arxiv.org/abs/2609.19693
作者: Dasom Choi,Sangjun Moon,Hyeongchan Im,Jaeeon Park,Jingun Kwon,Hidetaka Kamigaito,Taro Watanabe,Manabu Okumura
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 8 pages, 5 figures, 5 tables. Accepted to Findings of AACL-IJCNLP 2026

点击查看摘要

Abstract:The rapid increase of deepfakes has raised significant concerns due to their spread on social media. Traditional multi-face forgery detectors crop and verify each face independently, ignoring background context and inter-face relationships, which often yields suboptimal performance. To overcome these limitations, we leverage instruction-based Large Vision-Language Models (LVLMs), which can interpret entire images and follow complex textual instructions. We propose a simple yet effective single-stage multi-face forgery detector, called IMFD (Instruction-based Multi-face Forgery Detector), which is trained end-to-end to jointly localize faces and predict per-face forgery labels. Rather than treating face box prediction only as a joint objective, IMFD explicitly integrates predicted face bounding boxes into the instruction as visual cues that enhance instruction grounding and forgery detection. To support the training and evaluation of IMFD, we convert existing multi-face forgery datasets into an instruction-based format. Experimental results and analyses show that IMFD improves multi-face forgery detection by integrating face bounding boxes into the instruction, and consistently outperforms various state-of-the-art methods.

[CV-92] MiX: Micro-Inverted-Scaling for End-to-End Low-Bit Vision-Language Model Acceleration MICRO MICRO2026

链接: https://arxiv.org/abs/2609.19683
作者: Yuan Liao,Jae-sun Seo
类目: Hardware Architecture (cs.AR); Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to the 59th IEEE/ACM International Symposium on Microarchitecture (MICRO 2026)

点击查看摘要

Abstract:The deployment of Vision-Language Models (VLMs) on edge devices is severely bottlenecked by memory bandwidth, necessitating aggressive sub-8-bit quantization. Since edge accelerators are strictly constrained by area and power, they require end-to-end quantized models. However, the extreme dynamic range gap between multi-modal tokens causes standard block formats to suffer “microscaling collapse,” where a single massive outlier hijacks the shared exponent, underflowing surrounding elements and destroying attention maps. To break this bottleneck, we propose Micro-Inverted-Scaling (MiX), a novel format that mathematically inverts the microscaling paradigm: rather than grouping multiple mantissas under one shared exponent, MiX groups private, per-element exponents under a single shared mantissa. To handle asymmetric VLM outlier topologies, we introduce an adaptive dual-format (MiX-MX) inference framework. By algebraically factoring out the shared MiX mantissa, this framework maps to a custom accelerator, replacing multipliers with efficient shifters. Evaluated end-to-end on multiple VLMs, our 4.5-bit MiX formulation exhibits equivalent or superior accuracy on multi-modal benchmarks compared to NVFP4. Simultaneously, the MiX accelerator delivers a 25% improvement in area efficiency over the NVFP4 baseline and a 2.3-4.5x speedup with 1.4-2.9x energy reduction across models compared to the state-of-the-art accelerator Focus, proving the inverted-scaling datapath is physically superior for efficient VLM deployment.

[CV-93] Beyond Patch Removal: Persistent Adversarial Effects in Vision-Language-Action Policies

链接: https://arxiv.org/abs/2609.19669
作者: Enhao Wu,Fusen Guo,Yuxin Cao,Ziyang Lyu,Lin Li,Wei Song
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: 8 pages, 2 figures

点击查看摘要

Abstract:Adversarial patches to Vision-Language-Action (VLA) policies can cause both immediate action corruption and persistent state effects that remain after the patch is removed. Existing evaluations largely focus on continuous attacks and do not separate these two effects. We introduce a state-restoration protocol that removes the patch at matched action-chunk boundaries and measures subsequent recoverability under the same remaining step budget. Clean, random-patch, deviation-matched, and fixed-direction controls distinguish adversarial effects from occlusion, action-error magnitude, and directional persistence. We also evaluate a recovery adapter trained on attack-induced states under controlled intervention latency. On OpenVLA-OFT with EDPA attacks, only 36.2% of LIBERO-Long episodes remain recoverable after five chunks, compared with 89.9% and 87.0% for the deviation-matched and fixed-direction controls. Similar persistent effects are observed on autoregressive OpenVLA. The recovery adapter improves recovery from 7.7% to 47.4% at one-chunk latency, but its benefit decreases substantially with delayed intervention. These results show that adversarial effects can persist after patch removal and that timely intervention is critical for recovery.

[CV-94] VideoResearcher: Self-Improving Tool Design for Long-Video Understanding

链接: https://arxiv.org/abs/2609.19664
作者: Dingqiang Ye,Dongdi Zhao,Kaishen Wang,Qingqiao Hu,Jingchen Sun,Yijun Liang,Yuqi Jia,Yiqiao Huang,Yunjie Tian,Jiaxing Zhang,Chuanyang Jin,Ke Zhang,Vishal M. Patel,Di Fu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Video agents have made substantial progress in long-video understanding. Yet effective video-agent systems require costly, time-consuming manual design and trial and error. Current self-improvement methods either refine low-impact prompts, recombine predefined micro-tools, or struggle with convergence in harness optimization. To bridge this gap, we target high-impact video-tool with VideoResearcher, a training-free multi-agent framework that autonomously designs, tests, and refines tools for video understanding, like a human researcher. VideoResearcher operates through dual Solving and Evolving loops: it analyzes tool-use trajectories to identify capability gaps, coordinates specialized agents to develop and validate executable tools, and reuses evolved tools to strengthen evidence acquisition in subsequent video reasoning. Through iterative tool refinement and validation, it progressively strengthens evidence acquisition without updating model parameters. VideoResearcher achieves state-of-the-art performance among self-improving agents and approaches the human-designed upper bound, demonstrating a training-free paradigm for long-video understanding that expands agent capabilities through autonomous tool development while reducing costly manual engineering.

[CV-95] owards Active Cross-View Object Geo-Localization

链接: https://arxiv.org/abs/2609.19662
作者: Shunyu Yao,Xiaohan Zhang,Zhuoran Yang,Haoqi Lai,Qi Ming,Xiaoxi Hu,Hui-Liang Shen,Si-Yuan Cao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Cross-view object geo-localization (CVOGL) typically assumes a fixed query image, overlooking the ability of mobile agents to actively acquire more informative observations. To address this limitation, we introduce Active Cross-View Object Geo-Localization (ActiveGeo), where an agent sequentially selects new viewpoints and determines when to stop, aiming to improve localization with minimal observations. We further propose ActiveMoPT, an ActiveGeo framework with three-stage training. First, Multi-View Prompt-Preserving Adaptation enables the model to aggregate multiple query views while reusing the initial prompt. Second, Trajectory-Guided Policy Initialization uses supervised agent trajectories to learn viewpoint selection and initial stopping behavior. Third, Cost-Aware Policy Refinement employs GRPO with a gain-cost reward to jointly optimize localization accuracy and observation efficiency. We also construct ActiveGeo-858, a zero-shot test set containing 858 scenes and 1,716 target annotations. Experiments show that ActiveMoPT achieves state-of-the-art performance on MoP-UAV using only 1.45 query views on average, and substantially outperforms previous CVOGL approaches under zero-shot evaluation on ActiveGeo-858.

[CV-96] Instance Segmentation and Fine-grained Classification for Urban Buildings with Adaptive Region Dividing and Spatially-Supervised Contrastive Learning

链接: https://arxiv.org/abs/2609.19631
作者: Weiyuan Zhang,Qi Zhang,Hui Huang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 10 pages, 4 figures

点击查看摘要

Abstract:Accurate instance-level and functional understanding of urban buildings in large-scale point clouds is essential for digital city modeling and urban analysis. However, the extensive spatial coverage of urban scenes leads most existing methods to rely on predefined blocks for training and evaluation, although such partitions are rarely available in real-world applications and introduce additional preprocessing while fragmenting complete building structures. To address this issue, we propose an adaptive region-dividing strategy with unified scene-level evaluation. Specifically, the 3D point cloud is projected onto a bird’s-eye-view (BEV) plane, where a pretrained segmentation model is used to detect building regions. The detected bounding boxes are then back-projected to the original point cloud to construct structure-aligned adaptive training blocks, enabling semantically guided dynamic partitioning without manual design. Furthermore, beyond instance-level understanding, few methods have explored fine-grained classification for urban buildings, and thus we also put forward a fine-grained classification model for urban buildings with a spatially-supervised contrastive loss. First, for each segmented building, a point transformer classifier jointly encodes its body and local context using geometric, color, and core-context information. Then, the class-balanced weighted cross-entropy is used to alleviate severe class imbalance. The proposed spatially-supervised contrastive loss further enhances inter-class discriminability by assigning greater weight to spatially proximate, same-category buildings, encouraging compact functional representations while separating easily confused categories. Extensive experiments on UrbanBIS and STPLS3D demonstrate the advantages of the proposed method in building instance segmentation and fine-grained classification compared to existing SOTA methods.

[CV-97] VGGT-GS SLAM: Uncalibrated Monocular Gaussian Splatting SLAM with Feed-Forward Priors

链接: https://arxiv.org/abs/2609.19628
作者: Yuhang Han,Hao Wang,Jiaxi Cao,Xingyu Liu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 9 pages, 4 figures

点击查看摘要

Abstract:We present VGGT-GS SLAM, a monocular 3D Gaussian Splatting SLAM system designed for uncalibrated videos. Starting from feed-forward VGGT pose and depth priors, our system performs submap differentiable bundle adjustment that jointly refines camera poses and a 3D Gaussian map, while optimizing submap-shared intrinsics and radial–tangential distortion through analytic calibration Jacobians. To improve global consistency, we introduce Gaussian-native alignment (GNA) for camera-anchored scale refinement between sequential submaps and verification of loop-closure candidates. Extensive experiments on standard indoor benchmarks show consistent improvements in localization accuracy and strong rendering quality under uncalibrated settings, establishing a strong baseline for uncalibrated Gaussian SLAM.

[CV-98] Selective Cotton Boll Localization for Robotic Harvesting: Evaluation of Deep Learning Vision Models Under Field Conditions

链接: https://arxiv.org/abs/2609.19592
作者: Thevathayarajh Thayananthan,Xin Zhang,Isuru Laddusinghe Badu,Jonathan Harjono,Glen C. Rains,Beiwen Li,Leonardo M. Bastos,Nuwan K. Wijewardane,Vitor S. Martins
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: 27 Pages, 19 Figures, 15 Tables

点击查看摘要

Abstract:This study developed and evaluated a deep-learning-based perception framework for selective robotic cotton picking. The dataset contained 1,008 annotated field images collected using three cameras under varying natural lighting and weather conditions. Object-detection models from the YOLOv8 through YOLOv13 families were evaluated using their default configurations, while segmentation performance was assessed using YOLOv8-seg, YOLOv11-seg, YOLOv12-seg, the Segment Anything Model (SAM), SAMv2.1, FastSAM, and Grounded-SAM with the Recognize Anything Model (RAM). Among the detection models, GELAN-s achieved the most favorable balance between mean average precision (mAP) and inference speed, obtaining an mAP of 86.1%, precision of 81.6%, recall of 76.6%, and an F1-score of 79.0%, with an average inference time of 42.3 ms per image. Among the direct segmentation models, YOLOv12-m-seg provided the most favorable balance between AP@0.5 and FPS, achieving a segmentation AP@0.5 of 83.7% with an inference time of 20.4 ms per image. In the detection-prompted segmentation approach, bounding-box prompts generated by GELAN-s improved the localization of cotton bolls for SAM and SAMv2.1, while SAMv2.1 Tiny consistently outperformed FastSAM and Grounded-SAM with RAM. In the area-based evaluation against manually annotated segmentation masks, YOLOv12-m-seg achieved an R^2 value of 0.966, compared with 0.860 for GELAN-s + SAMv2.1 Tiny. Field experiments conducted using a UR5e robotic manipulator, a custom end-effector, and a ZED2i stereo camera further validated the effectiveness of the YOLOv12-m-seg model for real-time cotton boll detection, segmentation, and selective picking under varying confidence levels. These results demonstrate that YOLOv12-m-seg provides an efficient perception model for robotic cotton harvesting and has strong potential for field deployment.

[CV-99] A Multi-Modal Generative Model for Tomato Disease Leaves Understanding

链接: https://arxiv.org/abs/2609.19555
作者: Khang Nguyen Quoc,Minh-Phuoc Tran,Gia-Han Truong,Luyl-Da Quach
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: In submission to Computers and Electronics in Agriculture Journal

点击查看摘要

Abstract:Artificial intelligence for plant disease analysis has advanced from task-specific classifiers to multi-modal models capable of jointly interpreting visual and textual information. However, practical deployment in precision agriculture remains limited because most existing approaches treat disease understanding as isolated prediction tasks, failing to capture the complementary relationships among symptom recognition, severity assessment, and question-driven diagnostic reasoning. In tomato pathology, accurate interpretation of diseased leaves requires more than label prediction; it demands integrating visual symptoms with semantic context to support a comprehensive and explainable understanding. Here, we present SOLAR, a multimodal generative model that understands tomato disease spanning six question-answering tasks. SOLAR learns to align visual features with task-aware language representations by Fusion Expert module based on mixture-of-expert, enabling it to generate contextually relevant answers across diverse diagnostic tasks. By formulating tomato disease analysis as a generative Visual Question Answering (VQA) task, SOLAR provides a flexible framework that supports multi-task inference within a single model while improving performance and cross-task knowledge sharing. We evaluate SOLAR on 41,677 images, including 216,209 Question-Answering (QA) pairs to understand tomato leaf disease under both closed and open-ended QA settings. Experimental results show that SOLAR consistently outperforms state-of-the-art vision-only, vision-language, and task-specific models across all tasks, demonstrating superior accuracy, robustness, and multimodal reasoning. These findings highlight the potential of generative multimodal modeling as an effective direction for understanding of plant disease. The code for this study is available at this https URL.

[CV-100] VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations Active Perception and Metric Control

链接: https://arxiv.org/abs/2609.19554
作者: Zhongbo Zhang,Jiayi Jin,Yifan Wang,Zaibin Zhang,Haiwen Diao,Lijun Wang,Huchuan Lu
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Spatial intelligence requires more than describing object locations. Under incomplete observation, models must identify and acquire missing evidence, interpret it in a common spatial frame, and act on it. We introduce VA-Bench to evaluate the complete observe-reason-act-revise loop. General-purpose MLLMs learn procedural context from RGB-only demonstrations, actively select camera viewpoints, issue metric Cartesian commands, and revise them from execution feedback. Models receive no privileged object poses, oracle trajectories, or learned action heads. A fixed model-agnostic controller executes only model-specified targets. VA-Bench contains 14 base task families (11 single-arm and three dual-arm), seven held-out geometry/layout variants, and a long-horizon five-object composition track. We evaluate 12 primary model conditions in three independent runs over the same 20 physically verified seeds per base task, reporting terminal success, nine trajectory-level behavioral diagnostics, and subtask progress. First, the best-performing model scores 100.0% on target localization and 78.9% on spatial relations in the annotated run. Its three-run macro-average task success is only 53.93+/-3.17%. Second, active camera control significantly improves task success over passive multi-view observation. In one matched comparison, success rises from 27.86% to 57.50%. Third, held-out geometric transfer can reduce task success by over 30 percentage points. No model completes a strict long-horizon episode, despite substantial partial progress. VA-Bench thus tests whether general-purpose MLLMs can turn visual demonstrations and actively acquired evidence into successful embodied action.

[CV-101] PerSeM: Persistent Semantic Memory for Long-Horizon Open-Vocabulary UAV Mapping

链接: https://arxiv.org/abs/2609.19542
作者: Saurbh Singh Jamwal,Ganesh Ramakrishnan
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Open-vocabulary segmentation enables rich semantic perception for UAVs, but frame-wise predictions can remain temporally inconsistent across repeated observations and changing viewpoints. We present PerSeM, a training-free persistent semantic memory framework for long-horizon open-vocabulary UAV mapping. PerSeM associates frame-wise semantic observations with persistent world-space voxels and constructs a majority-based semantic memory, which is conservatively refined through history-preserving spatial refinement, trust-aware replay, and context-guided verification. Experiments on the Forest and UAVScenes benchmarks show that persistent 3D memory provides substantial gains in semantic correctness and temporal stability over frame-wise predictions. Beyond this strong persistent-memory baseline, PerSeM provides consistent additional improvements, improving both semantic accuracy and temporal stability across all five evaluated UAVScenes sequences. Analysis using regions identified independently of the final PerSeM predictions further shows that these gains are concentrated in semantically difficult and temporally unstable regions, where majority-based memory is most likely to remain uncertain. These results demonstrate that persistent 3D aggregation provides a strong foundation for long-horizon semantic mapping, while conservative refinement of uncertain memory states can provide additional improvements without retraining or additional neural-network inference.

[CV-102] AMB3R-SLAM: Kilometer-scale SLAM with Hierarchical Backend

链接: https://arxiv.org/abs/2609.19518
作者: Hengyi Wang,Lourdes Agapito
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: Project page: this https URL

点击查看摘要

Abstract:We present AMB3R-SLAM, a real-time monocular SLAM system capable of reconstructing kilometer-scale trajectories over 10k frames on a single consumer-grade GPU. Our model couples a lightweight front-end for low-latency online tracking with a hierarchical backend that progressively enforces local, mid-level, and global consistency. By avoiding bundle adjustment that relies on the static world assumption, our system naturally handles complex dynamic scenes out of the box. Furthermore, we demonstrate that our method can be extended to leverage stereo, RGB-D, and LiDAR as additional inputs. AMB3R-SLAM achieves strong camera tracking performance across 9 datasets, reducing the absolute trajectory error (ATE) of previous state-of-the-art methods on VBR and Oxford Spires by over 70%. With additional LiDAR input, our model further reduces ATE to sub-meter level on KITTI and VBR datasets.

[CV-103] ParticleSplat: Self-supervised Object-centric Latent Particle Splatting

链接: https://arxiv.org/abs/2609.19463
作者: Lyuxing He,Daniel Guo,Elizabeth Terveen,Deepak Pathak,David Held,Tal Daniel
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: Project page: this https URL

点击查看摘要

Abstract:We present ParticleSplat, a self-supervised object-centric learning method that decomposes scenes into a set of latent ‘‘particles’’ representing semantic entities through feedforward 3D Gaussian Splatting. Building on the Deep Latent Particles (DLP) framework, which represents images as a set of particles with attributes such as position, scale, and visual appearance, we address a key limitation of DLP: its inherently 2D nature, which prevents explicit 3D spatial and geometric reasoning that are critical for downstream tasks such as robotic manipulation. Leveraging the structural similarity between latent particles and 3D Gaussian primitives, we introduce a 3D latent particle space trained with a novel view synthesis objective. Our model jointly encodes multiple views with camera poses into a shared 3D object-centric latent space, then transforms particles into particle-aligned 3D Gaussians whose composition reconstructs the full scene. On simulated and real-world datasets, we show that this formulation inherently learns object masks without supervision and supports controllable 3D scene editing, such as moving objects by modifying particles in the latent space. We further establish that the learned 3D representation improves downstream performance on robotic manipulation tasks.

[CV-104] Efficient Unified Multimodal Understanding (EUMU): Winning Solution for the MUMU Track at the 8th LSVOS Challenge

链接: https://arxiv.org/abs/2609.19451
作者: Dayoung Kil,Seong-heum Kim
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:The Mobile Unified Multimodal Understanding (MUMU) Challenge requires a single efficient model to jointly perform multi-concept image tagging, open-vocabulary object detection, and image captioning. We present Efficient Unified Multimodal Understanding (EUMU), the winning solution for the MUMU Track of the 8th LSVOS Challenge. EUMU builds on a shared pretrained multimodal model, using its prompt-based capabilities for detection and captioning and training lightweight heads on shared visual features to predict quality, scene, and event tags. Rather than treating the three tasks independently, EUMU applies task-aware inference refinement by reusing task outputs as cross-task cues. For detection, caption cues help recover objects missed by the initial detection. For captioning, detection cues help refine the caption to better reflect the detected objects. For tagging, image statistics refine quality predictions, while caption and detection cues refine scene and event predictions. This design unifies all three tasks within a single model while satisfying the challenge’s resource constraints. EUMU contains 239.169M parameters, requires 23.947 GFLOPs, uses 4.5 GB of peak inference memory, and achieves a final challenge score of 17.3409. Code and models are available at this https URL.

[CV-105] Seeing Abnormal from Normal: Glomerular Abnormality in Representations of Normal Renal Morphology

链接: https://arxiv.org/abs/2609.19444
作者: Greta Hasko,Rachit Saluja,Tianyu Shi,Leiyue Zhao,Yuechen Yang,Daniel Reisenbuechler,Tianyuan Yao,Zhenhao Guo,John Cannon,Haichun Yang,Yuankai Huo,Yuling Chi,Lorraine Gudas,Mert R. Sabuncu,Yihe Yang,Ruining Deng
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Fine-grained evaluation of glomerular pathology must distinguish normal glomeruli from abnormalities such as global and segmental glomerulosclerosis, obsolescent, ischemic, solidified, disappearing, and atubular glomeruli. Supervised classification requires labeled examples of every category, which is impractical when subtypes are rare or absent from the training cohort. One-class anomaly detection offers an alternative by modeling normal data and scoring deviations, allowing previously unseen abnormalities to be detected. We use the frozen residual U-Net backbone of Omni-Seg, pretrained to segment structurally normal renal primitives without abnormal-subtype labels. We propose NoRDeC (Normal-Reference Detection and Characterization), a framework combining Mahalanobis normal-reference scoring with layer-wise representation analysis to determine whether and where glomerular pathology is encoded, how spatial aggregation affects detection, and whether abnormalities alter inter-layer relationships differently. Using glomerular images from two institutions, we evaluate backbone layers and aggregation strategies, compare NoRDeC with PaDiM and PatchCore, and analyze representations using centered kernel alignment (CKA). Layer 4 with Center-70 aggregation achieved a pooled AUROC of 0.926\pm0.013 . NoRDeC achieved the highest AUROC in six of seven abnormality categories and in the pooled analysis, while CKA suggested subtype-dependent changes in inter-layer relationships not captured by anomaly scores alone. The normal-reference model is fitted using only normal glomeruli; abnormality labels are used for configuration selection, evaluation, and grouping in the representation analysis. These results show that a frozen renal feature extractor can support both detection and representation-level characterization of glomerular abnormalities without using abnormal examples to fit the detector.

[CV-106] RGS: Reflection-aware Gaussian Splatting via Learning Geometry Continuity for Reflective Objects ICRA2026

链接: https://arxiv.org/abs/2609.19421
作者: Xiaobiao Du,Yida Wang,Cheng Bi,Kun Zhan,Xin Yu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project Page: this https URL Published in ICRA2026

点击查看摘要

Abstract:Gaussian Splatting has significantly improved the quality of novel view synthesis with explicit Gaussian representation. However, we observed that existing 3D Gaussian Splatting methods (3DGS) often suffer from surface collapse issues on reflective regions, and thus produce inferior geometry and low-quality specular. In this work, we propose a physically-based deferred rendering framework, named Reflection-aware Gaussian Splatting (RGS), that can accurately model specular regions and improve novel view synthesis performance. Specifically, we found that a powerful 3D foundation model can provide a strong 3D geometric prior to foster correct geometric modeling. Based on this, we propose a cross-view shape consistency regularization to regularize the geometry surface with the large model prior and cross-view constraints. In this manner, our RGS can produce smoother geometric surfaces on reflective regions while reducing geometric hollows. To further improve rendering results on reflective regions, we present a reflection-aware densification strategy that is designed to capture specular variations across various views. With this strategy, our RGS is able to render novel views of objects in higher quality. Extensive experiments demonstrate our method consistently renders high-quality reflective objects, achieving state-of-the-art performance.

[CV-107] WZPlanner: Safe End-to-End Path Planning for Autonomous Driving in Work Zones

链接: https://arxiv.org/abs/2609.19393
作者: Nishad Sahu,Changzhong Qian,Guangzhou Cai,Shounak Sural,Ragunathan(Raj)Rajkumar
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Work zones alter lane geometry through temporary traffic controls and closures that may be absent from on-board maps, challenging autonomous vehicle (AV) perception and planning. Generalization is also limited by scarce public datasets with structured geometric supervision. We present WorkZonePlan, a dataset comprising 149K+ synthetic and 5K+ real-world multimodal samples with 3D annotations for lane boundaries, work zone boundaries, and driving trajectory options. It also provides 76 closed-loop CARLA scenarios replayed under three weather conditions, yielding 228 Bench2Drive-format evaluation routes. We introduce WAVE (Work-zone-focused AV data generation in Virtual and rEal Environments), a semi-automated pipeline for creating the dataset, and BoundaryFormer (BF), a transformer-based model that jointly predicts lane and work zone boundary polynomials and driving trajectories. BF uses slot attention for boundary prediction. Ablations show that a separate trajectory decoder using boundary slot features substantially improves trajectory prediction over a slot-attention-only approach. Building on this finding, BF++ offers Camera and Camera+LiDAR variants with metric ground-plane encoding, typed boundary/trajectory queries, long-range point anchors, image-space curve refinement, and conservative gated LiDAR fusion. On the 211 routes common to all four models at the evaluation freeze, BF+±Camera and BF+±Camera+LiDAR achieve Driving Scores of 63.0 and 64.4, respectively, compared with 59.3 for SimLingo and 26.1 for TransFuser++ (TF++). BF++ is 40 times smaller than SimLingo and more than 10 times smaller than TF++, while achieving higher Driving Scores. These results support jointly predicting lane boundaries, work zone boundaries, and driving trajectories as a promising direction toward safer AV operation in work zones. Code and dataset: this https URL.

[CV-108] LinePilot Digitizer: Line-Plot Recovery with Manual and Automatic Calibration

链接: https://arxiv.org/abs/2609.19377
作者: Fengbo Ma,Rayan Akhtar,Aakash H. Joshi,Xiaoting Li,Haijian Sun,Zhen Xiang,Xianyan Chen,Yiping Zhao
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Recovering numerical series from line plots requires accurate axis calibration and reliable curve extraction. We present LinePilot Digitizer (LinePilot), which combines continuous color-based curve recovery with three calibration modes: LinePilot (standard), LinePilot (enhanced), and LinePilot (OCR). We also introduce DigitizerBench, the first dedicated benchmark for systematically evaluating digitizer performance, using an orthogonal design spanning signal, rendering, and plot-structure factors with complementary automatic and human-guided evaluations. We evaluate performance using failure-penalized capped normalized root-mean-square error (FPC-NRMSE), which assigns unit loss to missing, unusable, or catastrophically inaccurate outputs. On DigitizerBench-Full, LinePilot (OCR) achieves the lowest mean FPC-NRMSE (0.672) and highest trusted usability (38.2%) among the tested automatic pipelines. On DigitizerBench-Lite, LinePilot (enhanced) achieves the lowest mean FPC-NRMSE (0.081), 100% output success, and highest trusted usability (93.3%). The orthogonal benchmark design further enables factor analysis to identify the factors that most significantly affect digitizer performance. Together, the three calibration modes provide a practical trade-off between automation, user control, and accuracy within a shared curve-recovery workflow.

[CV-109] Open-vocabulary 3D object detection with promptable segmentation

链接: https://arxiv.org/abs/2609.19358
作者: Ömer Faruk Deniz,Mustafa Taha Koçyiğit
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 18 pages, 6 figures, 12 tables

点击查看摘要

Abstract:Three-dimensional object detection for autonomous driving is dominated by detectors trained on large corpora of human-annotated 3D boxes. Such a detector learns a fixed category list, and everything outside it is invisible. This paper asks whether the task can be solved training-free and open-vocabulary. A promptable segmentation model (SAM3), queried with class names as text prompts, supplies instance masks in the vehicle’s six surround-view cameras, and the masks are turned into metric 3D boxes using the geometry of the scene. The core is a controlled three-stage comparison on nuScenes in which 2D detection is held fixed and only the source of 3D geometry changes. Geometry predicted from images alone reaches 0.183 mean average precision (mAP) under the official protocol; fitting boxes from raw LiDAR points inside the same masks with training-free rules reaches 0.298 mAP / 0.348 nuScenes detection score (NDS) at zero labeling cost; borrowing supervised box geometry at inference time lifts the same detections to 0.413 mAP / 0.555 NDS, which locates the pipeline’s largest deficit in measurement precision rather than 2D detection, while class confusion and confidence calibration survive that substitution. Reversing the direction, a three-state camera-witness rule built from the same masks improves a supervised LiDAR-only detector from 0.596 to 0.630 mAP, roughly half the gain of fully supervised camera fusion, with no training. A coverage analysis shows that SAM3 finds 84% of in-range objects with a correctly named mask; the classes that fail in the official metric are misnamed or geometrically unforgiving, not unseen.

[CV-110] Can Vision-Language Models Judge Olympic Diving? From Reasoning to Scores in Zero-Shot Action Quality Assessment

链接: https://arxiv.org/abs/2609.19354
作者: Henry O. Velesaca,David Freire-Obregon,Luigi Miranda,Abel Reyes-Angulo
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Automated action quality assessment (AQA) in Olympic sports remains a challenging task due to the complexity of human motion and the subjectivity inherent in expert judging. This work evaluates the capability of open-source Vision-Language Models (VLMs) to perform zero-shot action quality assessment on Olympic diving videos using the AQA-7 benchmark dataset. In this regard, a regression-based framework is pro-posed to leverage both the semantic reasoning and phase-level sub-scores generated by the VLMs, combining TF-IDF vectorization, dimensionality reduction, and ensemble learning to predict final competition scores. Experimental results show that standalone VLMs achieve moderate Spearman correlations below 0.32, while the proposed ensemble regression framework substantially improves performance in the reported evaluation, reaching a Spearman correlation of 0.67 with a four-model configuration. Textual reasoning features con-sistently outperformed raw numerical sub-scores, highlighting the richness of VLM-generated explanations for action quality analysis. These findings suggest that VLMs hold strong potential as assistive tools for explainable and semi-automated sports performance evaluation. The code is publicly available on GitHub this https URL diving judge vlm

[CV-111] RAUL: Reference-Assisted Ureteroscopy Localization for Skill Assessment

链接: https://arxiv.org/abs/2609.19236
作者: Fangjie Li,Mai Bui,Charan Mohan,Michael Miga,Matthieu Chabanas,Nicholas Kavoussi,Jie Ying Wu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Objective: Incomplete navigation of anatomy during ureteroscopic kidney stone surgeries can contribute to repeat interventions. While skilled surgeons have lower reintervention rates, there are no objective metrics to quantify scope-navigation performance to evaluate when a trainee becomes skilled. This work aims to recover ureteroscope trajectories from endoscopic video and derive navigation metrics to quantify differences in skill. Methods: We propose RAUL, a reference-assisted reconstruction framework for recovering ureteroscope trajectories from ureteroscope videos only in phantoms. For each phantom, we use a slow, high-quality reference exploration video to generate a reference reconstruction. We localize subsequent exploration videos against this reference. We evaluate localization accuracy against electromagnetically tracked scope pose. We compute navigation metrics from phantom exploration trajectories to compare surgical residents across experience levels. Results: The proposed reference-assisted framework achieves a mean translation root mean square error of 0.5 \pm 0.1 mm across 9 phantoms. Compared to standard Structure-from-Motion (SfM), the proposed pipeline increases frame-wise localization coverage from 50.5 \pm 14.9% to 86.1 \pm 7.2% of all video frames. The reconstructed trajectories revealed significant differences between high- and low-experience trainees in established navigation metrics. Conclusion: RAUL enables substantially more complete recovery of ureteroscope trajectories from videos compared to standard SfM pipelines, enabling trajectory-based skill assessment without additional tracking equipment. Significance: To the best of our knowledge, this is the first use of video-only recovery of ureteroscope trajectories without external tracking sensors for skill assessment, supporting scalable automated assessment of ureteroscopy navigation skill.

[CV-112] Open ultrasound foundation model for robust segmentation and clinical measurement across heterogeneous settings

链接: https://arxiv.org/abs/2609.19230
作者: Chao Qin,Fahad Shahbaz Khan,Salman Khan,Sarim Ather,Siddiq Anwar,Rao Muhammad Anwer,Shadab Khan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: The PDF includes the Supplementary Information

点击查看摘要

Abstract:Ultrasound is the most widely deployed imaging modality worldwide, yet clinical AI remains fragmented into narrow single-task models that fail when device, operator, or anatomy changes. Here we present SonoCorpus, an open resource unifying 456,963 images and 1,626,085 expert masks from 53 public datasets spanning 24 clinical applications and 17 countries, and SonoBase, an interactive segmentation foundation model pretrained on it. Across fifteen evaluation datasets introducing new organs, devices, operators, and geographies, SonoBase outperforms SAM2, MedSAM2, and the concept-promptable MedSAM3 on every dataset and matches per-dataset specialist models trained on the same data; on fully external data it exceeds the accuracy these baselines achieve on their own in-distribution benchmarks. Ejection fraction derived from its segmentations falls within inter-observer variability (6.63% error), with fewer misclassifications at the defibrillator-candidacy threshold than either promptable baseline (13% versus 18–42%); fetal head-circumference (1.81~mm) and gestational-age (1.2 days) errors fall below inter-observer variability. Where a baseline fails outright, one in four test cases, SonoBase recovers a usable segmentation in 81% of them, including on handheld probes operated by minimally trained users in two low- and middle-income countries (Sierra Leone and Tanzania). Five labeled examples can help the model adapt to a new setting, and the identical training protocol transfers well to newer models such as SAM3, locating the advantage in ultrasound-specific pretraining rather than any single architecture. To ensure reproducibility and enable the community to build on SonoBase as a platform, we release all checkpoints, optimizer states, data-split indices, deduplication hashes, and starter code.

[CV-113] 4D Radar Perception Algorithms for Autonomous Driving: A Review

链接: https://arxiv.org/abs/2609.19216
作者: Xumin Wu,Jun Zhou,Jilin Mei,Chen Min,Yu Hu
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Signal Processing (eess.SP)
备注: 12 pages, 9 figures, 5 tables. Submitted to IEEE Sensors Journal

点击查看摘要

Abstract:Research on 4D millimeter-wave radar perception algorithms has flourished in recent years, extending from signal processing and object detection to semantic segmentation, motion estimation, occupancy prediction, and dynamic scene reconstruction. This review organizes the field according to the evolution of perception tasks and algorithms. It first introduces radar fundamentals, data representations, and quality-enhancement methods, and then reviews object-level perception, motion and localization, local and dense spatial perception, and dynamic scene understanding. Across these directions, we compare radar-only learning, multimodal fusion, and cross-modal supervision and knowledge distillation. Particular attention is paid to how elevation, Doppler measurements, and radar physical priors are exploited across tasks. We further summarize the task coverage, input data, annotations, and evaluation protocols of existing datasets, clarifying the empirical support for different research directions. Finally, we discuss the common challenges and future directions of 4D radar perception for autonomous driving. This review provides a task-oriented perspective on the transition from sparse object perception to dynamic spatial understanding.

[CV-114] HyperAMS-Net: Adaptive Multi-Scale Spatial Hypergraph Network for Brain Disorder Classification MICCAI2026

链接: https://arxiv.org/abs/2609.19755
作者: Proloy Kumar Mondal,Md Kamran Hussin Chowdhury,Hoi Leong Lee
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: Accepted at the 17th International Workshop on Machine Learning in Medical Imaging (MLMI 2026), held in conjunction with MICCAI 2026

点击查看摘要

Abstract:Accurate classification of brain disorders from neuroimaging data remains challenging because of substantial inter-subject heterogeneity and the complex multi-scale patterns present in functional connectivity and morphological representations. To address these challenges, we propose HyperAMS-Net, a deep learning framework for brain disorder classification using neuroimaging representations derived from resting-state functional MRI or structural MRI. HyperAMS-Net integrates adaptive multi-scale convolution, hypergraph attention, spatial-channel attention, and adaptive feature fusion. Specifically, adaptive multi-scale convolution learns data-driven weights over multiple receptive fields to capture complementary patterns at different scales. Hypergraph attention models higher-order dependencies among learned feature representations through node–hyperedge–node message passing, while spatial-channel attention enhances discriminative feature learning. Adaptive feature fusion further aggregates complementary information across parallel network branches. HyperAMS-Net is evaluated on three benchmark datasets spanning distinct brain disorders: ABIDE for autism spectrum disorder, REST-meta-MDD for major depressive disorder, and ADNI for Alzheimer’s disease, using 5-fold stratified cross-validation. HyperAMS-Net achieves state-of-the-art performance across all evaluated datasets, attaining the highest accuracy and AUC among the compared methods. Ablation studies further demonstrate the contribution of each proposed component, with the largest performance degradation observed when hypergraph attention is removed.

[CV-115] he segmentation ceiling: why explicit left-ventricular masks do not improve learned ejection-fraction regression

链接: https://arxiv.org/abs/2609.19730
作者: Farshid Farhadi Khouzani,Paul La Plante,Bryar Mustafa Shareef,Laxmi Gewali
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
备注: 15 pages, 4 figures. Submitted to Computers in Biology and Medicine

点击查看摘要

Abstract:Accurate estimation of left ventricular ejection fraction (EF) from echocardiography is central to cardiovascular care, and deep learning enables automated EF prediction from echocardiographic video. Because EF is clinically derived from left-ventricular (LV) volumes, a widely held intuition is that explicit LV segmentation should improve prediction. We introduce a quantitative criterion, the segmentation ceiling, that makes this testable: from EF as a normalized difference of end-diastolic and end-systolic volumes, we derive in closed form how per-frame segmentation area error propagates into EF error, and thus the accuracy a mask must reach before it can improve on direct regression. Using EchoNet-Dynamic, a UniFormer-S backbone, and the empirically measured within-patient error correlation, the criterion places the break-even near 10% per-frame area error, whereas a representative segmenter operates at roughly 14%, above the ceiling. Consistent with this, four strategies for injecting segmentation or area information (a predicted-mask channel, end-diastolic/end-systolic clip sampling, and per-bin and amplitude area-consistency objectives) fail to beat a raw-video baseline; ground-truth masks help only through label leakage. Input representation thus not being the limit, we identify generalization as the practical lever: weight averaging with strong augmentation attains a test R^2 of 0.806 (MAE 4.08) under a matched dense-clip protocol, comparable to an R(2+1)D baseline (0.811) while tightening the validation-to-test gap. Finally, a heteroscedastic beta-NLL formulation yields informative, well-calibrated per-prediction uncertainty, larger for clinically harder low-EF cases, where Monte-Carlo dropout does not. The segmentation ceiling gives a concrete design criterion for when mask-guided EF estimation is worthwhile, plus a simple, uncertainty-aware recipe for EF regression.

[CV-116] Compression Hurts Pooling Helps: Information Loss in Rayleigh-Scale Estimation from B-Mode Ultrasound

链接: https://arxiv.org/abs/2609.19525
作者: D. Hudson Smith,Ahmer Raza
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Applications (stat.AP)
备注: 20 pages, 7 figures

点击查看摘要

Abstract:Clinical B-mode images are widely available as potential data sources for quantitative ultrasound (QUS) analysis for tissue characterization. However, standard clinical ultrasound devices apply unknown log-compression to RF envelope data before display and storage. Previous work has demonstrated estimation of the underlying RF envelope statistics in the presence of an unknown compression law. Using Fisher information analysis, we show that finite-offset log compression causes severe information loss when estimating the Rayleigh scale \sigma , which controls diffuse speckle. For a single image window, unknown compression raises the minimum achievable variance for unbiased estimation of \sigma by a compression-independent factor of approximately \FisherMinInflation . When M equal-sized windows share the same unknown compression settings, the excess variance decays as 1/M ; even in the most favorable regime, reducing the variance inflation factor below 1.1 requires \FisherBestCaseWindows windows. Our analysis treats the contrast parameter a as unknown and the boundary offset b as known; estimating b experimentally shows even larger variance. We validate this theory using synthetic estimation experiments and demonstrate RF-scale recovery on real RF-envelope windows from the OASBUD dataset. Together, these results clarify the limitations of using routine B-mode images for QUS.

[CV-117] Mammography Foundation Models for Opportunistic Prediction of Major Adverse Cardiovascular Events

链接: https://arxiv.org/abs/2609.19385
作者: Paula Feldman,Nusrat Binta Nizam,Sunwoo Kwak,Batuhan Karaman,Katerina Dodelzon,Mert Sabuncu
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Cardiovascular disease (CVD) remains the leading cause of death among women, yet cardiovascular risk assessment often relies on clinical variables that may be missing, outdated, or unavailable in routine care. Screening mammography offers an opportunity for opportunistic cardiovascular risk stratification because it is routinely acquired and contains vascular features, including breast arterial calcifications (BAC), that are associated with cardiovascular risk and events. We evaluate whether mammography specific foundation models, originally pretrained for breast cancer-related tasks, can transfer to cardiovascular risk prediction without cardiovascular specific supervision or explicit BAC annotation. We constructed a 5-year major adverse cardiovascular event (MACE) cohort of 22,497 women linked to electronic health record outcomes, including 500 events (2.22% prevalence). The foundation models achieved AUROCs of 0.823 and 0.822 substantially exceeding an age-only model (AUROC 0.765), despite using only the screening mammogram as input, with no clinical variables. Both foundation models evaluated assigned substantially higher predicted risk to patients with radiologist-documented BAC, despite BAC never being used as a training label, and showed activation patterns consistent with vascular findings. Together, these findings suggest that mammography foundation models can recover clinically relevant cardiovascular risk information directly from mammographic pixels and suggest that screening mammography may provide an opportunistic source of cardiovascular risk information to complement conventional clinical assessment without additional imaging. Code is available in this https URL

[CV-118] Perceptual Refinement of an End-to-End Video Streaming Pipeline via Generative AI Layers

链接: https://arxiv.org/abs/2609.19215
作者: Emanuele Artioli,Farzad Tashtarian,Christian Timmerer
类目: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
备注: 28 pages. Submitted to ACM Transactions on Multimedia Computing, Communications and Applications (TOMM), special issue on MMSys and co-located workshops. Extended version of the NOSSDAV 2025 paper ELVIS ( arXiv:2512.14185 ). Code: this https URL

点击查看摘要

Abstract:Traditional codecs treat every region of a frame alike; a generative layer can instead degrade the regions a viewer attends to least and reconstruct them at the client. We present PRESLEY, which extends the prior conference work ELVIS by replacing destructive block removal with adaptive in-place degradation under a removability mask, signaling per-block strength in a bit-packed side channel, and restoring via generative backbones conditioned on transmitted visual priors rather than unconditioned in-painting. We separate the problem into three goals: choosing which blocks to degrade, degrading them so the encoder spends fewer bits, and restoring them. Against its predecessor at matched rate, PRESLEY achieves a decisive mean -56.4% BD-rate reduction on delivered background quality across 13 rate ladders spanning multiple codecs and dataset families. Against pristine baselines, PRESLEY defines the operating regime of generative transport: delivering substantial bitrate savings (up to -29.4% BD-rate) and superior background quality (17/23 sequences) in the target bit-starved regime, while maintaining foreground fidelity bit-exact. We further map where the theoretical headroom in this class of architecture lies. Using an exact leave-one-superblock-out combinatorial oracle as an additive empirical bound, we show that existing complexity heuristics already capture 83.3% of bit-cost savings, bounding remaining cost-axis headroom at about 5% of total bitrate. We then identify and model the primary unaddressed axis – post-restoration damage – which disperses widely (4.9-8.4 dB). We prove that this damage is predictable before transmission (held-out rho = +0.400), establishing the feasibility of transmit-time restorability modeling and defining the roadmap for joint rate-distortion-restoration selection rules.

人工智能

[AI-0] Workspace Models: Lightweight Robotic Memory via Saliency-Driven Supervision

链接: https://arxiv.org/abs/2609.20820
作者: Nitish Dashora,Douglas Chen,Idan Shenfeld,John Marangola,Pulkit Agrawal,Max Simchowitz
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: 26 pages; CoRL 2026; 11 figures

点击查看摘要

Abstract:Complex robotic manipulation tasks frequently require a long-term memory of past events and actions. As conditioning on full histories renders policies prone to spurious correlations and degrades performance, many approaches to policy memory involve compressing historical information through expensive VLM queries in-the-loop to process only task-salient information. In this paper, we propose an alternative approach in which computationally intensive VLM queries are made during train-time to learn a lightweight latent memory that can be efficiently queried at deployment time. Our representation, which we call the \textbfworkspace token, is trained by (1) using a VLM to identify current and historical information necessary for completing a task, then (2) distilling these into the workspace token using a set-reconstruction decoder loss. In both simulation and hardware, we show that the workspace token can be used as a drop-in replacement for observations during deployment, enabling policies to solve memory-intensive tasks without the need for VLM reasoning in-the-loop. Interestingly, we found that workspace tokens are not only more lightweight but also lead to better policy performance.

[AI-1] Quantifying Overclaiming Propensity in Frontier LLM Agents

链接: https://arxiv.org/abs/2609.20812
作者: Nolan Smyth,Yorguin-Jose Mantilla-Ramos,Pascal Jr Tikeng Notsawo,Saskia Helbling,Alberto Tosato,Mohamed Amine Merzouk,Nouha Dziri,Gauthier Gidel,Tommaso Tosato
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 7 figures, 6 tables

点击查看摘要

Abstract:Frontier coding agents are increasingly trusted to work autonomously for long periods, yet an agent’s final response is often the only account of that work a user sees. We quantify the propensity of frontier agents to \emphoverclaim task completion, a misrepresentation that can mislead the user. An agent overclaims when its final response contradicts information in its context. This definition requires no inference about intent and is independent of task success. We introduce \emphOverclaimBench, an evaluation suite composed of five file-review scenarios, transcript-based coverage measurements, and registered planted defects. We evaluate eight proprietary frontier models in their own production command-line interfaces, and four open-weight models under a single fixed harness on OverclaimBench and find that 1) agents do not read all the files they were asked to review in 67.9% of runs; 2) among runs where not all files are read, agents are \emphmisleading 80.4% of the time (59–96% per model), either falsely claiming to have read all files or omitting that coverage is incomplete; 3) requiring delegation to subagents increased reading coverage, but among reviews that remained incomplete, a large majority were still misleading; and 4) agents that falsely claimed a complete review missed planted defects at about 1.8 times the rate of agents that read every file, showing that claims of completion can conceal substantive failures. Together, these results show that agents’ final responses are not reliable accounts of their actions.

[AI-2] GeoAAC: Geometry-Based Adaptive Action Chunking from Denoising Trajectories in VLA Policies ICRA

链接: https://arxiv.org/abs/2609.20776
作者: Xin Chen,Sen Chen,Yujuan Ding,Jian Liu,Guoqing Wang,Wei Ye,Heng Tao Shen,Yi Bin
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 9 pages, 6 figures. Submitted to the IEEE International Conference on Robotics and Automation (ICRA) 2027

点击查看摘要

Abstract:Action chunking is widely used for action generation and execution in Vision-Language-Action (VLA) policies, yet existing approaches commonly use a fixed action horizon. During a rollout, different task stages may require different levels of action continuity, control precision, and closed-loop feedback, making a fixed horizon unable to accommodate changing control requirements. We propose \textbfGeoAAC, a geometry-based adaptive action chunking method for flow-based VLA policies that adjusts the action horizon according to the reliability of the current action prediction. We show that the geometry of Flow Matching denoising trajectories provides process-level information for characterizing prediction reliability, with geometric variation across action prefixes remaining positively correlated with predictive uncertainty. GeoAAC uses this prefix-wise geometry to construct a horizon-wise geometric profile and adaptively determine the action horizon from a single generation without additional training. Experiments with GR00T N1.5 and \pi0.5 on LIBERO, LIBERO-Pro, RoboCasa365, and real-world manipulation tasks show consistent improvements over fixed-action-horizon baselines and existing adaptive methods, including up to 8.7 percentage points in simulation and an increase in average real-world success rate from 53.3% to 74.4%.

[AI-3] RAFT: A Stateful Retrieval-Augmented Framework for Troubleshooting Agents EMNLP2026

链接: https://arxiv.org/abs/2609.20754
作者: Mingxuan Zhang,Xiaowen Wang,Anupma Sharan,Zhengyi Chen,Chenyu Diana Zhang,Shanshan Yang,Chittibabu Pacharu
类目: Artificial Intelligence (cs.AI)
备注: Accepted to the EMNLP 2026 Industry Track

点击查看摘要

Abstract:Effective troubleshooting agents in enterprise customer support depend on retrieving actionable guidance from similar historical cases, yet existing retrieval-augmented generation (RAG) systems treat support cases as static documents and overlook their multi-stage, stateful nature. We introduce RAFT (Retrieval-Augmented Framework for Troubleshooting Agents), a stateful RAG framework that abstracts each closed historical case into a directed chain of timeline entries and retrieves at the entry level, surfacing cases whose intermediate states match the active case and returning the parent-case trajectory anchored at the matched state; an optional case-level graph links cases through a configurable similarity representation. We evaluate this retrieval layer directly, which, unlike evaluating a full agent system, requires no production deployment. Because public multi-stage troubleshooting data is extremely rare, we pair a synthetic benchmark built from Microsoft Learn Windows Server documentation with real Apache Jira issues carrying human-created duplicate labels. RAFT improves Case Hit over vanilla RAG and GraphRAG baselines at every stage of case progress, with statistically significant gains over the strongest baseline; the Jira results provide directional evidence that the advantage transfers to real case histories. We release our benchmark, implementation, and the Apache Jira evaluation set.

[AI-4] Large Language Models as Falsifiers for Cyber-Physical Systems

链接: https://arxiv.org/abs/2609.20752
作者: Ali ArjomandBigdeli,Jiawei Zhou,Stanley Bak
类目: ystems and Control (eess.SY); Artificial Intelligence (cs.AI); Logic in Computer Science (cs.LO); Software Engineering (cs.SE)
备注: 22 pages, 5 figures, 3 tables

点击查看摘要

Abstract:Falsification searches for counterexamples to formal specifications in cyber-physical systems (CPS). With specifications written in Signal Temporal Logic (STL), falsification can be formulated as a robustness optimization problem, traditionally tackled with black-box search algorithms. In parallel, large language models (LLMs) have recently emerged as surprisingly effective optimizers when coupled with iterative prompting. In this work, we connect these ideas and introduce LLM-Falsifier, an LLM-based approach that falsifies specifications by minimizing the STL robustness degree. Beyond generic prompt-based optimization, our key idea is to expose the LLM to semantic information that is natural for language models but absent from standard numerical optimizers, including natural-language input and output names, output trajectories, and critical-time witnesses for the minimum robustness value. These additions enable smarter and more sample-efficient robustness search. On the ARCH-COMP falsification benchmarks, LLM-Falsifier is shown to outperform existing falsification tools based on a range of optimization paradigms, from surrogate-based and Bayesian optimization to search-based testing, on 14 of 21 specifications when measured by the average number of simulations required to find a counterexample.

[AI-5] QA on Any Spreadsheet Requires Interpreting Its Grid Structure

链接: https://arxiv.org/abs/2609.20732
作者: Zofia Smoleń
类目: Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注:

点击查看摘要

Abstract:Semantic cell annotation improves chunking interpretability for spreadsheets in LLM-driven RAG systems, aiding answer generation through enriched context rather than improved retrieval accuracy. We propose a novel framework of splitting any spreadsheet into interpretable chunks using cell role annotation. Our framework beats the state of the art, yet it faces a hard ceiling. Spreadsheets are fundamentally two-dimensional unstructured data with continuous relationships and infinite potential cell roles. Because classification models are restricted to finite, pre-defined classes, they cannot perfectly capture this structural nuance, even with human-level annotation. We show that addressing the spreadsheet-to-LLM bottleneck requires moving beyond discrete cell classification. Instead, the field must develop dimensionality-reduction techniques to directly flatten 2D unstructured spreadsheets into 1D unstructured text. Text chunks would be easier for downstream RAG to interpret and generate from.

[AI-6] Deep Noir: Autonomous Steering Discovery via Architectural Chronometry in Transformer Models

链接: https://arxiv.org/abs/2609.20722
作者: Frank E. Bobe III,Gregory D. Vetaw,Darshan W. Bryner,Matthew G. Cook,Jose L. Salas-Vernis
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Activation steering modifies LLM behavior at inference time, but identifying where and how strongly to steer remains manual. We introduce Deep Noir, a framework that uses Logit Lens convergence and causal head-level attribution to autonomously discover optimal steering parameters. Across three scales (1B x 3, 2-3B x 2, and 7-9B x 4), our engine achieves 16.7 percentage-point improvement on spam at 1B (standard deviation 4.7; 39 runs), with gains increasing to 21 to 42 percentage points at 7-9B across four architectures. On SST-2 sentiment, it achieves a 13.1 percentage-point improvement with zero code changes. Mechanistic grounding enables automated discovery of intervention points that generalize across tasks and architectures. On sentiment, RepE without head masking fails to improve over baseline, while Deep Noir improves all models (p less than 0.01). We further show that steering creates a predictable prompt-injection attack surface whose vulnerability increases monotonically with steering magnitude. This finding is relevant to agent systems deploying steered classifiers.

[AI-7] HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface

链接: https://arxiv.org/abs/2609.20659
作者: Zimu Han,Yiming Zeng,Jiyao Zhang,Zihao Zhao,Yuanfei Wang,Yixiang Jin,Shiqi Li,Shuangben Chen,Wei Huang,Ruodai Li,Hui Shen,Hao Dong
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large-scale vision-language-action (VLA) models provide powerful priors for robot manipulation, yet adapting them to a specific deployment remains challenging. Supervised fine-tuning (SFT) on task-specific demonstrations provides a step toward deployment, but faces two persistent limitations: static data provide limited coverage of out-of-distribution states, and standard imitation objectives do not distinguish progressing behavior from less useful data. Interactive post-training can address these limitations, but typically requires repeated policy execution and human intervention on a physical robot. We introduce HIL-UMI, a policy-guided Universal Manipulation Interface (UMI) framework for robot-free human-in-the-loop VLA post-training. During handheld UMI demonstrations, HIL-UMI queries the current policy on the same observation stream without executing its predictions. The Energy Score compares the human action trajectory with policy inference and triggers collection when their discrepancy indicates an out-of-distribution region. In a separate feedback loop, low online advantage predictions identify essential segments for refining a progress-based advantage estimator. The updated estimator then guides advantage-conditioned behavioral cloning using a balanced mixture of base demonstrations and new policy data. This design preserves the iterative and policy-aware nature of human-in-the-loop learning while decoupling data collection from robot deployment. Experiments on four real-world tasks spanning long-horizon and precise manipulation show that HIL-UMI achieves consistent improvement over SFT and benefits from both targeted collection and advantage refinement. Moreover, HIL-UMI outperforms HG-DAgger on Clean Up Table with lower per-frame collection time, suggesting a scalable path for VLA post-training across operators and locations.

[AI-8] A Simulation Platform for AUV Fault Recovery: Exploring LLM -Based Diagnostic Strategies

链接: https://arxiv.org/abs/2609.20620
作者: Khalid Halba,Kylie Cooper,James G. Bellingham
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: 6 pages, 3 figures, 2 tables. Accepted for presentation at the 2026 IEEE/OES Autonomous Underwater Vehicles Symposium (AUV 2026), Southampton, UK. This is the author-accepted manuscript

点击查看摘要

Abstract:Autonomous underwater vehicles (AUVs) operating beyond reliable communications must recover from failures without human intervention. We investigate an architecture in which conventional deterministic layered control autonomy manages normal operations, while an invokable large language model (LLM) serves as a diagnostic and recovery planner when onboard anomaly detection identifies performance outside expected limits. Because language models are stochastic, rigorous evaluation requires ensemble testing rather than individual demonstrations. We present a closed-loop simulation architecture that couples real-time C vehicle software with a higher-level orchestration layer for physics-based fault injection, structured prompting, language-model interaction, mission file generation, validation, execution, and LLM-judge scoring. The framework, which we call SPAR (Simulation Platform for AUV Recovery), supports evaluation across fault realizations, prompt structures, reasoning models, and mission conditions. We vary these for a mass-shift fault over 480 SPAR trials, evaluating a frontier model and three off-the-shelf locally deployable LLMs. Model choice dominates diagnosis: the frontier model places the CG-shift mechanism in its top three hypotheses in 85-90% of trials, versus 60-78% for the best local model. Reasoning analysis indicates that local-model success is associated with following the complete diagnostic procedure, whereas weaker models often commit prematurely to elevator failure even though the actuator tracks its command. Diagnosis and operational decision performance do not appear to be coupled in this dataset. The contributions are an architecture extending unanticipated-fault recovery from detection to mitigation and an ensemble methodology for evaluating LLM-assisted mission management on low-power AUVs.

[AI-9] Inference-Engine Fingerprinting Attacks are Practical: Exploring Model-Driven Environmental Discovery Exploitation and Escape

链接: https://arxiv.org/abs/2609.20614
作者: Sarah Radway,Andrew Cheng,Vijay Janapa Reddi,James Mickens
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Frontier AI models are rapidly gaining the ability to exploit vulnerabilities in complex pieces of software. The risk is not theoretical, as evidenced by recent sandbox escapes performed by frontier models at OpenAI and Anthropic. Discussions of how to sandbox inference stack components often focus on components other than the inference engine itself (e.g., network proxies or code execution environments). However, the inference engine is an attractive target for a misaligned model. For example, if a model can trigger exploits in that engine merely by generating specially-crafted output tokens, the model can initiate a multi-step, to-the-bare-metal exploit chain in the engine, without relying on vulnerabilities in other components of the inference stack, and without assistance from externally-provided, maliciously-crafted input tokens. In this paper, we show that a misaligned model can perform inference engine fingerprinting to determine the specific engine (e.g., vLLM, SGLang) which executes the model. Once the engine has been fingerprinted, the model can leverage engine-specific exploits to take control of the engine using only carefully-selected output tokens. We provide concrete examples of model fingerprints in five popular engines, and demonstrate how realistic agentic harnesses allow a model to leverage those fingerprints to identify the local engine. We also describe a proof-of-concept, to-the-bare-metal exploit chain that originates from a fingerprinted (and subsequently compromised) inference engine. We conclude by discussing several ways that inference engines could be changed to make fingerprinting attacks more difficult. Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI) Cite as: arXiv:2609.20614 [cs.CR] (or arXiv:2609.20614v1 [cs.CR] for this version) https://doi.org/10.48550/arXiv.2609.20614 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-10] Limits of Confidence in Diffusion

链接: https://arxiv.org/abs/2609.20581
作者: Russ Webb,Amitis Shidani,Alice Bizeul,Dan Busbridge
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Discrete diffusion, including remasking and uniform-state samplers, generate a sequence by writing multiple token positions per step, drawing each from a per-position distribution and choosing which positions to write from those same distributions. For domains of general interest (pixels, phonemes, or words) there are inherent dependencies between tokens. We show that a step matches the training distribution only when the positions it writes are conditionally independent given the tokens already fixed, that no product of per-position distributions can match a dependent group, and that per-position distributions do not determine whether a group is dependent: two joint distributions can have identical per-position marginals while differing in which combinations of values occur. On ScanAndAdd, a synthetic task whose joint distribution is available in closed form, we verify that every group of two or more undetermined positions a confidence ranking writes is dependent, and measure the generated distribution to be 29\times the sampling-noise floor total variation while per-sample metrics are 1.0 .

[AI-11] Accelerating Visual Policy Learning with Sampling-Based Model Predictive Control

链接: https://arxiv.org/abs/2609.20575
作者: Yilang Liu,Haoxiang You,Qian Wang,Daniel Rakita,Ian Abraham
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 8 pages, 6 figures

点击查看摘要

Abstract:Learning visual policies for locomotion and manipulation requires coordinating contact with the environment and can incur substantial computation and GPU memory costs. First-order policy gradients (FoPG) reduce training cost through differentiable simulation, but local optimization can converge to unintended contact patterns. To address this shortfall, we propose Sampling-Guided Policy Search (SGPS), which couples recurring action-target refinement by sampling-based model-predictive control with first-order policy optimization. Behavior cloning initializes the policy from sampled actions; training then alternates sampling-based refinement with short-horizon FoPG updates under perturbed initial states and randomized dynamics. For visual policy training, we use a decoupled FoPG formulation that excludes rendering from the computation graph, enabling direct learning from depth observations without a state-policy teacher. On a single GPU, SGPS learns policies for locomotion, obstacle traversal, crate pushing, and bimanual carrying on simulated Unitree Go2 and G1 robots. Our experiments further show that refinement improves policy learning beyond initialization and tracking alone. For hardware deployment, the distilled policy transfers zero-shot to a real Go2 and uses onboard depth to autonomously trot, crawl, clear hurdles, and switch between these behaviors.

[AI-12] Mitigating Retaliatory Algorithmic Collusion in Repeated Games

链接: https://arxiv.org/abs/2609.20548
作者: Karthik Sivachandran,Rohan Paleja
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Reinforcement learning agents trained to maximize their own reward in repeated interactions can converge to supra-competitive outcomes resembling explicit collusion, without communication or shared design. Existing mitigation approaches are largely tied to specific economic settings, like two-sided platforms and auctions, leaving open how to design interventions for general repeated games. We address this gap by formalizing the connection between empirical observations from prior work on Q-learning collusion and classical theory of Simple Penal Codes (SPCs). We show any non-trivial SPC induces a quantifiable conditional dependence in agents’ policies, detectable via the total variation distance between an agent’s action distributions across cooperation and defection histories. Building on this connection, we propose CURB (Collusion Unwinding via Reward shaping and Belief injection), a reward-shaping framework that penalizes this Total Variation (TV) distance signal during Q-learning and is guaranteed to convert any SPC fixed point of the dynamics into a trivial one, thus precluding collusive equilibria sustained by punishment threats. Empirically, CURB substantially reduces collusion by Q-learning agents in both Bertrand and Cournot Competition Repeated Games. We further demonstrate that CURB extends to deep Q-network agents in Bertrand competition, suggesting the mechanism generalizes beyond tabular Q-learning.

[AI-13] Refuse Decompose Refresh: A Claim-Safe Protocol for Closed-Loop AI Evaluation

链接: https://arxiv.org/abs/2609.20538
作者: Peiying Zhu,Sidi Chang
类目: Artificial Intelligence (cs.AI)
备注: 8 pages, 0 figures, 3 tables. The reproducibility artifact is linked in the paper

点击查看摘要

Abstract:An AI evaluation can be perfectly reproducible and still support the wrong claim. This risk is acute in closed-loop systems: policy determines visited states, observable components, and which failures leave a measurable trace. We propose a claim-safe protocol with three actions. Refuse: abstain when a clean reference stream or matched runtime comparison lacks support. Decompose: report protocol execution, operational false admission, and structural hypotheses separately rather than as one PASS/FAIL label. Refresh: treat distribution-shift alarms as requests to invalidate and recompute a reference map, not as fault evidence. We instantiate the protocol in an aggregate-only simulator with 24 policy components, three demand regimes, two fault-mask families, and independent development and heldout seeds. The preregistered heldout contains 1,440 cases and 21,600 partition rows. Only 55/72 regime-component units were reference-admitted and 54/55 remained runtime-admitted, making abstention part of the result. Stable false admission was 0/20 represented components, with a one-sided exact 95% upper bound of 0.1391 under a frozen 0.20 rule. Within admitted units, affected clean traffic outpredicted nominal fault-cell fraction: across 540 unit-arm rows nested in 20 component clusters, the cell-minus-traffic negative-log-likelihood difference was 0.1264 nats per row, with a 95% component-cluster interval of [0.0593, 0.1918]. A drift log shows why “null” must be reference-relative: clean fault-null streams triggered 15/15, 0/15, and 14/15 alarms across three regimes, while only the middle regime matched the frozen detector reference. Rather than a universal threshold, we contribute an executable contract linking observable support, statistical calibration, and justified claims.

[AI-14] FreqCondNorm: Towards Cross-domain Predictive Maintenance through a Frequency-Conditioned Transformer Foundation Model

链接: https://arxiv.org/abs/2609.20535
作者: Zaynab Raounak,Camille LHermine,Zhiguo Zeng
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Deep learning predictive maintenance models suffer from poor transferability across machines and operating conditions, especially when labelled data are scarce and signals span five orders of magnitude in sampling frequency (1 Hz to ~100 kHz). We propose FreqCondNorm, a Transformer-based architecture that introduces a FiLM-style frequency-conditioned normalization layer to unify heterogeneous time-series within a single model. The architecture is pretrained on five public predictive maintenance datasets (CWRU, MFPT, UOC18, PRONOSTIA, CMAPSS) using masked auto-encoding and contrastive learning with balanced domain sampling. On fault diagnosis, the model achieves 99.2% accuracy on CWRU (+6.4 pp over CNN) and 82.1% zero-shot accuracy on MFPT, demonstrating strong transfer across sampling frequencies. However, the approach does not improve remaining useful life prediction, suggesting a mismatch between pretraining and RUL objectives that warrants future investigation.

[AI-15] SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness

链接: https://arxiv.org/abs/2609.20519
作者: Haozhe Liu,Tian Ye,Sensen Gao,Qihang Cao,Yitong Li,Mingchen Zhuge,Duomin Wang,Ruihua Zhang,Ping Luo,Jiawang Bian,Lei Zhu,Ligeng Zhu,Enze Xie,Song Han
类目: Artificial Intelligence (cs.AI)
备注: 15 pages, 8 figures, 4 tables. Code: this https URL . Project page: this https URL

点击查看摘要

Abstract:As coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands from isolated predictions into long trajectories of reasoning, tool use, and feedback. Token efficiency therefore becomes important for scaling recursive self-improvement. We take an RSI-inspired approach at the harness layer, scaling auto-research loops across increasingly numerous and diverse environments for harness rollouts. At this scale, the process yields reusable improvements that transfer beyond their development setting, moving automated harness discovery toward production-level outcomes. Four mechanisms survive selection and form SoL-Pi, spanning action execution, context compaction, observation handling, and delegated reading. On the 51-task EdgeBench evaluation, SoL-Pi achieves performance comparable to Pi across GPT-5.6 Sol and Opus 5 while reducing recorded token traffic by 44.7-49.0% and API cost by about one third. In other words, estimated hourly savings are \ 8.75-\ 13.50 relative to native Codex and Claude Code harnesses, and \ 4.36-\ 5.71 relative to Pi.

[AI-16] How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents

链接: https://arxiv.org/abs/2609.20474
作者: Yukun Zhang,Kemu Xu,Yishen Chen
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Agent harnesses supply planning guidance, organize execution, and check completion. We study how these components affect success, erroneous acceptance, and cost in two Retail experiments and an Airline pilot in \tau^2 -bench. The primary comparison pairs prewritten task-specific plans (Fixed) with shuffled policy text matched in word count (Sham), isolating the contribution of guidance content. Across 265 matched cells, Fixed improves oracle-verified success by 7.17 percentage points (90% task-clustered bootstrap interval, 1.15–13.36 points), with gains concentrated in higher-complexity tasks. A read-only terminal verifier rejects 61% of Retail oracle-invalid episodes while withholding 17% of correct ones, at less than one cent of additional cost per episode. Which component matters more depends on the loss assigned to erroneous acceptance: at low liability the planning gain dominates; at high liability the verifier’s avoided false passes dominate—and a standalone verifier captures nearly all the false-pass benefit of the full planning-plus-verification stack at a fraction of its cost.

[AI-17] Deep Learning-Based Classification of Cognitive and Resting States Using Electroencephalography Signals

链接: https://arxiv.org/abs/2609.20467
作者: K. A. Januka S. Fernando,Harshit Srivastava
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 16 pages, 19 figures, 7 tables, and 1 algorithm

点击查看摘要

Abstract:The categorization of cognitive and resting states derived from electroencephalography (EEG) signals is crucial for comprehending fluctuations in brain activity linked to various mental states. EEG provides a non-intrusive approach for documenting brain function in both resting and task-oriented cognitive conditions, whilst deep learning techniques enable the automatic extraction of significant patterns from intricate EEG data. This study presents a deep learning framework to distinguish between resting and cognitive states through EEG records. The proposed framework integrates a Convolutional Neural Network (CNN) stacked with a Gated Recurrent Unit (GRU) for the extraction of features from EEG signals. Time-frequency analysis is conducted to explore the salient aspects of signals, and the derived features are then assessed utilizing conventional deep learning and machine learning classifiers, including the suggested 2D-Net architecture. The proposed approach and feature extraction strategy outperform the evaluated comparative methods, achieving accuracies of 83.177% for resting-versus-mathematical task classification, 76.107% for resting-versus-memory task classification, and 83.432% for resting-versus-music task classification. The findings illustrate the efficacy of integrating signal processing with deep learning methodologies to discriminate resting from cognitive states utilizing EEG signals.

[AI-18] Fingerprinting Multimodal Large Language Models

链接: https://arxiv.org/abs/2609.20457
作者: Chao Huang,Meng Tong,Kejiang Chen
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 10 pages, 3 figures. Accepted to ACM Multimedia 2026 (MM '26) as an oral presentation

点击查看摘要

Abstract:While multimodal large language models (MLLMs) enable a wide range of image-text reasoning tasks, recent incidents indicate that they are vulnerable to illicit deployment and unauthorized distillation. Existing solutions for model provenance are typically confounded by shared language backbones in MLLMs and struggle to detect violations of distillation. To bridge this gap and safeguard model ownership, we present the first study on multimodal model fingerprinting. Inspired by recent findings that self-attention acts as a low-pass filter and that its low-frequency components are informative, we develop AttnPrint for white-box provenance. Specifically, we extract cross-modal attention distributions and isolate their low-frequency components to serve as model fingerprints. To facilitate black-box auditing, we further introduce DistillTrace, which employs hypothesis testing of MLLM outputs to identify potential model infringement. We conduct extensive experiments on 154 model instances across 19 multimodal architectures. Notably, AttnPrint achieves strong derivative-model detection performance while remaining robust to five downstream modification techniques. DistillTrace also provides evidence of distillation relationships under three parameter-independent techniques.

[AI-19] SkillAA: Attribution-Guided Skill-Graph Updating with Targeted Validation and Rollback

链接: https://arxiv.org/abs/2609.20455
作者: Ziqiao Shang,Ling-Yue Ge,Lan-Zhe Guo
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:External skills provide domain procedures without parameter updates, but existing methods often edit skills directly from failed rollouts without structured routing from an observed failure to an editable location; existing skill graphs also underuse semantic boundaries, object addresses, and topological dependencies for skill retrieval, targeted updating, and scoped validation. We introduce SkillAA (Skill Abductive Attribution), a structured skill-optimization framework for frozen language models. It represents skill applicability, execution, and composition in a unified graph, allowing the same structure to support skill selection, attribution-guided repair, and update validation. SkillAA contrasts successful and failed executions to route candidate repairs to specific graph objects, updates only the selected local structure, and uses Local and Big Gates to screen candidate changes before commitment. With gpt-5.6-sol, SkillAA reaches 81.5%, 66.7%, and 91.2% on SearchQA, LiveMath, and DocVQA, respectively, and attains the highest observed mean in every main setting. These results support the utility of attribution-guided graph editing and graph-scoped validation.

[AI-20] he Organization of Inference: Information Resource Constraints and AI Production

链接: https://arxiv.org/abs/2609.20449
作者: Yukun Zhang,Kemu Xu,Yishen Chen
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:The economic value of inference depends on how capacity and task information are distributed across stages of AI production. We study these organizational margins using controlled workflow experiments on externally verified software-engineering tasks. In two matched resource panels, direct execution records the same success rate of 59.6 percent at logical-token ceilings of 12,000 and 24,000, while success under information-constrained planning rises from 36.2 to 51.2 percent. The planning disadvantage narrows by 15.0 percentage points (95 percent task-cluster bootstrap interval: 4.2 to 25.8). A strict read-only planning campaign varies whether the planner sees the task issue. At 12,000 tokens, issue access raises success by about 16 percentage points over issue-hidden planning. Compared with direct execution, task-informed planning is about 10 points lower at 12,000 tokens; at 24,000 tokens, it shows a 29.6-point advantage. In the resource panels, direct execution uses substantially less than either ceiling, while the planning workflow’s binding rate falls from 46.2 to 0.8 percent and downstream execution accounts for 89.9 percent of the increase in total use. Scale determines the capacity available to a system; workflow and information structure shape the productive value

[AI-21] SCGFM-ART: Amortized Relational Transport for Structure-Centric Graph Foundation Models

链接: https://arxiv.org/abs/2609.20419
作者: Xiaodong He,Xincheng Wang,Zhao Kang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 21 pages, 6 figures

点击查看摘要

Abstract:Graph foundation models (GFMs) aim to learn transferable representations across severely heterogeneous graph domains. However, severe domain shifts in topology, graph scale, and feature semantics impede the construction of a unified, domain-agnostic representation space. To address this, we propose SCGFM-ART, a structure-centric GFM framework that aligns arbitrary graphs onto a shared relational atlas via Amortized Relational Transport (ART). The relational atlas serves as a universal coordinate system defined by a finite set of relational landmarks (bases), while ART directly predicts reusable, end-to-end graph-to-base transport plans, bypassing costly runtime Gromov-Wasserstein optimizations. Under this formulation, SCGFM-ART decomposes a graph into a unified representation: globally via its relational response coordinates relative to the atlas, and locally via its node-to-role structural correspondences. These correspondences project disparate node attributes into a canonical role space, resolving structural and semantic heterogeneity within a singular alignment interface. Rigorously modeling graphs and atlas bases as finite measured relational spaces, we establish coordinate fidelity bounds, prove stability under predicted transport plans, and derive an amortized coverage bound that guarantees our learning objective tightly surrogates ideal relational coverage. Benchmarked across 14 cross-domain graph- and node-level classification tasks, SCGFM-ART achieves state-of-the-art transferability, securing superior average ranks of 2.29 and 1.14, respectively. Topological perturbation analyses demonstrate that node-role transport retains fine-grained structural nuances beyond global coordinates. On real-world benchmarks, the amortized formulation yields 44.2 to 85.1 times faster frozen target-domain inference by avoiding iterative alignment at test time.

[AI-22] Accelerating Sharded Data Parallelism at Scale with Federated Learning

链接: https://arxiv.org/abs/2609.20359
作者: Gianluca Mittone,Marco Aldinucci
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI); Performance (cs.PF)
备注:

点击查看摘要

Abstract:The symbiotic scaling of artificial intelligence models and high-performance computing systems continually creates algorithmic challenges in their convergence. Foundation models (FMs) are a crucial example, requiring months-long training on thousands of cutting-edge GPUs. Sharded data parallelism (DP) is the dominant strategy to accelerate such computations by splitting data and models across multiple GPUs. However, it incurs prohibitive communication overhead when deployed at scale, particularly on multi-tier interconnects with heterogeneous performance. Inspired by the efficient communication principles of federated learning (FL), this work introduces two hybrid algorithms - FL+FSDP and FL+HSDP - interleaving sharded DP with FedAvg-style aggregations. Such approaches decouple large DP deployments into smaller, loosely-coupled federation groups, requiring minimal inter-group traffic while keeping the global batch size bounded by the groups’ size. Formal analysis of communication costs and experimental validation prove their scalability and flexibility. A Llama3.1 8B pre-training on 512 A100 GPUs shows that, under identical hyperparameters, FL+FSDP and FL+HSDP achieve up to 8.04 faster data processing and 4.48 lower evaluation perplexity than their counterparts, demonstrating superior computational efficiency and improved model quality. These properties stem from reduced communication overhead and the bounded growth of the global batch size relative to the federation group size.

[AI-23] Generating Heterogeneous 3D Geological Microstructures from 2D Images via a Stable Diffusion-Adversarial Model

链接: https://arxiv.org/abs/2609.20358
作者: Ali Aouf,Eric Laloy,Bart Rogiers,Christophe De Vleeschouwer
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Characterizing the physical properties of clay and cementitious materials matters across many fields, from materials science to geological waste disposal. Property simulation typically calls for 3D imaging, which is expensive, not always accessible, and technically limited for certain materials. Recent progress in deep generative models offers a way around this, reconstructing 3D volumes from the more easily acquired 2D images. Among GAN-based methods for 3D microstructure generation, SliceGAN has shown strong results for homogeneous isotropic and anisotropic systems. It struggles, however, to capture the finer detail of more complex heterogeneous microstructures, which motivates alternative generative frameworks. We introduce a hybrid approach that draws on the stability and generation quality of denoising diffusion models. Since no 3D ground truth is available, we replace the standard denoising loss with an adversarial loss, which yields a stable training process in our experiments. We show that the resulting model generates microstructures of varying complexity with minimal slice artefacts and close agreement with ground-truth phase fractions and structural descriptors. Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2609.20358 [cs.AI] (or arXiv:2609.20358v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2609.20358 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-24] A Qualitative Model for Reasoning about Path and Support IJCAI

链接: https://arxiv.org/abs/2609.20349
作者: Abhishek Jaiswal,Zoe Falomir
类目: Artificial Intelligence (cs.AI)
备注: Workshop on Qualitative Reasoning 2026 at IJCAI (35th International Joint Conference on Artificial Intelligence)

点击查看摘要

Abstract:Spatial reasoning abilities correlate strongly with performance in STEM fields. Games offer a compelling medium for training these critical skills in developing children who have a natural proclivity for play. However, to facilitate human-like tutoring and player guidance, these games require an AI agent capable of making commonsense inferences from spatial events. Qualitative reasoning (QR) models appear to be a suitable framework for these application domains. As these models reason in symbolic representations, they can seamlessly translate game states into interpretable feedback for human-like player guidance. This paper introduces a hybrid qualitative model designed for Camelot Jr., a block-puzzle game that requires constructing multi-level bridges to connect two avatars stationed on separate towers. The game poses a challenge for the player, who must make platforms stable, plan their path, and ensure they use all the provided blocks. To handle the precise physics required by the domain, we integrate a mathematical center-of-mass stability logic to guide our qualitative solver. Our work facilitates spatial skill training in Camelot Jr. and contributes to the development of human-centric, explainable game-playing agents.

[AI-25] STR-Agent : An LLM -Driven Agent for QoS-Aware Routing in LEO Satellite Networks

链接: https://arxiv.org/abs/2609.20347
作者: Bowen Lu,Mugen Peng,Yaohua Sun,Hongyu Wang,Kerui Guo,Wenjia Xu
类目: Networking and Internet Architecture (cs.NI); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:LEO satellite networks feature dynamic topologies, time-varying links, and diverse service requirements, which make conventional routing schemes difficult to support fine-grained quality-of-service (QoS) provisioning. Existing studies mainly optimize routing over network states with predefined objectives, but rarely address the practical challenge of translating unstructured natural-language service requests into adaptive routing decisions. To bridge this gap, we propose STR-Agent, an LLM-driven framework for QoS-aware routing in LEO satellite networks. The key innovation of STR-Agent lies in unifying intent perception, tool-based execution, experience accumulation, and reflection-based policy adaptation within a single agent architecture. Specifically, the Perception Module converts natural-language requests into structured routing semantics, while the Reflection Module dynamically adjusts the service-to-routing-policy mapping according to real-time congestion conditions and historical routing outcomes, rather than relying on a fixed routing objective. In addition, we develop a specialized perception model, and construct a domain-specific supervised fine-tuning dataset for LEO service understanding. Simulation results in a Walker-Delta constellation show that STR-Agent significantly outperforms conventional baselines: it reduces end-to-end delay by up to 60% compared with DQ-Dijkstra, improves average intent-understanding accuracy from 45.4% to 92.45% after supervised fine-tuning, and the Reflection Module further reduces the delay by 120 ms at 600 Mbps. These results demonstrate the potential of LLM-driven agent architectures to enable service-aware and adaptive QoS routing in future LEO satellite networks.

[AI-26] Structured Four-Stage Legal Translation: From Natural-Language Traffic Rules to PROLOG

链接: https://arxiv.org/abs/2609.20334
作者: May Myo Zin,Wachara Fungwacharakorn,Ken Satoh,Katsumi Nitta
类目: Artificial Intelligence (cs.AI)
备注: In Proceedings of the International Workshop on Translating Natural Legal Language into Formal Representations (NLL2FR 2025)

点击查看摘要

Abstract:Traffic regulations are written for human interpretation and therefore rely on shared background knowledge and flexible phrasing, which inherently introduce ambiguity, context dependence, and semantic underspecification. These linguistic characteristics conflict with the precision required by computational reasoning engines such as Prolog, which demand explicit logical structure. This study evaluates two baseline translation approaches, Natural Language to Prolog ( NL\rightarrow Prolog ) and Logical English to Prolog ( LE\rightarrow Prolog ), and introduces a new reasoning-guided translation framework called Structured Four-Stage Legal Translation ( S4L\rightarrow Prolog ). The proposed S4L framework performs semantic role extraction, scene completion, logical mapping, and Prolog rule generation within a single guided prompt, enabling direct translation of raw traffic rules into executable logic without human intervention. A benchmark consisting of twenty real-world traffic rules was used to evaluate each approach in terms of syntactic validity, semantic correctness, and logical completeness. S4L\rightarrow Prolog achieves the highest accuracy, correctly formalizing 75 percent of the rules, while NL\rightarrow Prolog reaches 60 percent and LE\rightarrow Prolog reaches 55 percent. Qualitative analysis further shows that S4L captures implicit causal relations, deontic modality, and exception structure more reliably than the baselines. These results demonstrate that structured reasoning prompts can substantially improve the reliability of natural-language-to-logic translation for legal and safety-critical applications.

[AI-27] NeuSOGA3D: A Neuro-Symbolic Framework for Explainable 3D Geometric Reconstruction

链接: https://arxiv.org/abs/2609.20323
作者: Qingde Li,Qingqi Hong,Zihan Li,Jie Tian
类目: Artificial Intelligence (cs.AI)
备注: Preprint. Community feedback and comments are welcome

点击查看摘要

Abstract:Three-dimensional reconstruction from unorganized point clouds remains a challenging problem in computer vision, geometric modeling, and computer-aided design. While neural implicit methods achieve impressive reconstruction accuracy, geometry is typically encoded in latent representations that limit interpretability and reuse within engineering workflows. We present NeuSOGA3D (Neuro-Symbolic Geometric Abstraction in 3D), a hybrid framework that combines learned perceptual priors inherited from NeuSOGA with explicit symbolic geometric reasoning. The method projects point clouds onto principal orthographic planes, constructs symbolic implicit spline representations from the resulting observations, and fuses them through shape-preserving constructive solid geometry operations to generate a coarse visual hull. Additional geometric detail is recovered through cross-sectional decomposition and volumetric reconstruction using Partial Shape-Preserving Splines. Unlike conventional neural implicit approaches, NeuSOGA3D progressively transforms observations into explicit symbolic entities, including control polygons, implicit spline fields, cross-sections, and volumetric lofts. Experiments on all forty categories of the ModelNet40 benchmark demonstrate the ability of the framework to recover structurally meaningful and CAD-compatible geometric representations from diverse point-cloud observations. The results highlight the potential of combining learned perception with symbolic geometric reasoning for explainable geometric intelligence. Comments: Preprint. Community feedback and comments are welcome Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2609.20323 [cs.AI] (or arXiv:2609.20323v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2609.20323 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-28] Diagnose Recover Certify: Task Readiness under Hidden Dynamics Changes

链接: https://arxiv.org/abs/2609.20304
作者: Nguyen Viet Tuan Kiet,Huynh Thi Thanh Binh
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:A deployed control policy can conceal consequential dynamics changes: an actuator may lose effectiveness without affecting the current task when the policy rarely excites it, despite being critical for a future task that has not yet been specified. We introduce task readiness under dormant dynamics drift, a decision problem that unifies active change diagnosis and post-change control recovery under a limited, task-agnostic interaction budget. An agent must identify whether and where local dynamics have changed, use a small number of informative interactions to characterize the change before downstream task identity is revealed, and subsequently provide each candidate task with either a recovered policy and a calibrated lower bound on its achievable return or an abstention decision to a safe fallback. We propose Evidence-Gated Matched-Pulse Transport, an intervention-based Bayesian procedure that couples fault localization with estimation of actuator effectiveness through a shared matched-response representation, thereby preserving diagnostic reliability while converting localized evidence into recovery-relevant uncertainty. This uncertainty is propagated to task-conditioned policy selection and readiness certification, enabling deployment decisions that explicitly trade off expected performance, confidence, and fallback use. We evaluate the resulting framework on a diverse suite of dormant-actuator benchmarks spanning multiple simulators, under a protocol that separates diagnosis from capability recovery, scores deployment by readiness coverage, selective risk, and interaction cost as well as return, and identifies the fault regimes in which transported evidence is decisive.

[AI-29] Agent PProf: Semantic Profiler for Long Horizon AI Agents

链接: https://arxiv.org/abs/2609.20301
作者: Yusheng Zheng,Chaokun Chang,Yu Mao,Tianyuan Wu,Yuxi Huang,Tao Ma,Wenan Mao,Shuyi Cheng,Andi Quinn,Wei Wang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:AI agents increasingly orchestrate long-running activities with users, tools, and system resources for days and weeks. To improve agent quality, safety, and cost efficiency, developers need to determine where failures happen, what triggers unsafe effects, and which tasks consume the most budget, then optimize those tasks. In systems software, profiling answers similar questions by aggregating resource consumption and attributing it to responsible code paths to identify hotspots. Yet existing agent observability tools focus on per-execution debugging and tracing rather than cross-run, long term profiling, making these questions difficult to answer at scale. Agent observability needs profiling, not only debugging, but profiling agents is challenging: the responsible entities are task intent like diagnose authentication, compare branches rather than code paths, and lack stable identifiers for aggregation. We propose a semantic operation stack model that adapts profiling to agent trajectories. Uniform operations represent all activities, and operation stacks replace the runtime call stack, enabling hierarchical attribution at different granularities. We observe that an agent’s task occupies a contiguous span and decomposes into subtasks, so we introduce recursive operation segmentation, which recursively splits trajectories at task boundaries. AgentPProf is a profiler that aggregates agent trajectories into pprof-compatible profiles, enabling flame graph visualization and analysis. AgentPProf reaches 0.764 B^3 F1 against human annotations on CodeTraceBench. On three problem-localization benchmarks, the profile raises MAP by up to 56%, demonstrating that it effectively attributes resources, locates problems, and helps optimize token cost at practical profiling cost. AgentPProf is available at this https URL.

[AI-30] Labeled Incidence Structures for Native Transformer Modeling of Text Knowledge Graphs and Hypergraphs

链接: https://arxiv.org/abs/2609.20278
作者: Mahesh Godavarti
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Text, knowledge graphs, and hypergraphs all have elements that play distinct roles within relation instances, structure that is lost when data is flattened into token sequences. We introduce labeled incidence structures (LIS), a uniform representation that encodes each endpoint as (x_d, s, e) : content x_d , a role or slot s , and the relation instance e in which that role appears. Because every data type maps to the same (x_d, s, e) representation without flattening, a single standard transformer can process them all natively, structural differences are carried entirely by the operators, not the architecture. LIS assigns a structural address to each endpoint by composing a slot operator and an instance operator, A(s,e) = R_s R_e . We characterize when this factorization gives every token a unique, path-independent address. When it does, the natural operator comparing endpoint j to endpoint i is the relative transport P_j\to i = A_i^-1 A_j , which gives attention a role- and relation-aware inductive bias without imposing an arbitrary sequence order. Additive encodings of the form “position term plus relation term” can miss information that depends jointly on s and e . We prove this in a controlled example family: when the journey operator is approximated by the sum of a position-only term and a relation-only term, the approximation cannot capture how position and relation combine, only their separate effects. We also analyze persistent knowledge repositories. Identifiers tied to storage locations make models sensitive to storage order, while freely learned identifiers can become harder to control as the repository size M grows relative to the sample size n . Computing relation-instance operators from content avoids this storage-order issue and yields a capacity bound independent of M , under fixed architectural and Lipschitz assumptions. Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI) MSC classes: 68T07, 68T30, 68T50, 68R10, 05C65, 05B20 ACMclasses: I.2.6; I.2.4; I.2.7; I.5.1; G.2.2 Cite as: arXiv:2609.20278 [cs.LG] (or arXiv:2609.20278v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2609.20278 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Mahesh Godavarti [view email] [v1] Wed, 29 Jul 2026 11:30:33 UTC (39 KB)

[AI-31] JEPA-WAM: Connecting Generated Visual Instructions to World Action Models through JEPA Latent Representations

链接: https://arxiv.org/abs/2609.20277
作者: Tianbin Liu,Jian Zhu,Taiyi Su,Jianjun Zhang,Chong Ma,Zitai Huang,Yi Xu
类目: Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注: 11 pages, 4 figures

点击查看摘要

Abstract:World Action Models (WAMs) have demonstrated strong robotic manipulation capabilities by augmenting pretrained video generative models with action experts. However, current WAMs still show limited instruction-following ability when conditioned solely on text instructions. We argue that this limitation stems in part from a structural imbalance in robot-learning data: rich visual-action trajectories are often paired with sparse and repetitive language annotations, allowing policies to identify tasks from visual context and motion regularities rather than grounding the instruction itself. To address this limitation, we introduce JEPA-WAM, which augments each text instruction with a bank of stochastically generated visual instructions, providing diverse visual cues for instruction following. Specifically, JEPA-WAM uses an off-the-shelf text-to-image generator to sample multiple task-completion images conditioned on the text instruction, without training the generator. Although these generated images may differ from the current visual scene in appearance and layout, they remain semantically aligned with the instruction and serve as visual goal references. To focus on task-level semantics beyond appearance, we encode these references with a frozen V-JEPA 2.1 encoder. The resulting dense goal representations are compressed into compact goal tokens that condition both the video and action experts through cross-attention. We further construct a real-robot instruction-following benchmark covering in-distribution, out-of-distribution scene, and out-of-distribution instruction settings. On this benchmark, JEPA-WAM achieves success rates of 87.3%, 74.5%, and 80.9% in these three settings, outperforming \pi0 and Fast-WAM by at least 10.0, 27.3, and 14.5 percentage points, respectively.

[AI-32] AI-Driven Real-Time Relay Optimisation in Smart Urban NR-V2X Networks via Learning-to-Optimise Graph Neural Networks

链接: https://arxiv.org/abs/2609.20271
作者: Giambattista Amati,Federica Mangiatordi,Emiliano Pallotti,Simone Angelini
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Networking and Internet Architecture (cs.NI)
备注: 6 pages, conference

点击查看摘要

Abstract:Reliable and low-latency communication is a fundamental requirement for smart city services and Industry 4.0 applications enabled by NR-V2X networks. However, limited Road-Side Unit (RSU) deployment and complex urban propagation conditions often prevent Connected and Automated Vehicles (CAVs) from maintaining stable connectivity. This paper proposes an AI-driven Learning-to-Optimise (L2O) framework based on Graph Neural Networks (GNNs) for real-time multi-hop relay selection in NR-V2X systems. The vehicular network is modelled as a graph, where nodes represent CAVs and RSUs, and edges encode radio-link characteristics. An offline Mixed-Integer Linear Programming (MILP) formulation provides optimal relay decisions used as supervision for training an edge-aware Graph Isomorphism Network with Edge Features (GINE). Extensive experiments on realistic urban datasets demonstrate that the proposed approach achieves near-optimal connectivity performance, recovering up to 11.3% connectivity gain, while reducing execution time by orders of magnitude (up to 100 x speed-up) compared to MILP. The framework enables scalable and real-time network control, making it suitable for smart city and Industry 4.0 deployments.

[AI-33] When AI Agents Commit: Cognitive Serializability Across Data Evidence Policy and Authority

链接: https://arxiv.org/abs/2609.20261
作者: Jun He,Deying Yu
类目: Artificial Intelligence (cs.AI); Distributed, Parallel, and Cluster Computing (cs.DC)
备注: 22 pages, 3 figures, 4 tables

点击查看摘要

Abstract:Autonomous agents derive concrete mutations from database reads, retrieved evidence, policy, beliefs, and delegated authority. Those inputs may change while reasoning is in progress. Database isolation orders the submitted transaction; agentic transaction processing determines whether a proposal satisfies an executable contract. Neither guarantee establishes a common valid point for the mutation and its derivation inputs unless the contract represents the relevant predicates. Typed dependency tokens distinguish content integrity from applicability, and trusted mediation captures the values exposed to reasoning. Under strict Cognitive Serializability, committed effects admit a serial order and a logical event at which every value exposed to derivation is unchanged. The fences last until the runtime event that realizes the sealed durability domain. The weaker Effect-Compatible Cognitive Admission recertifies an effect against a simultaneously held current dependency vector and current policy without claiming to serialize the original stochastic derivation. TCT combines immutable versioned executable definitions, registry-derived authority plans, sealed envelopes, guard-first commit transactions, post-seal envelope- and witness-bound grants, co-committed receipts, idempotent grant finalization, and receipt-driven epistemic reconciliation. Complete registered footprints and a single growing phase induce an acyclic lock-point order over local guards and incompatible external reservations. The corresponding results give serializability conditions and an observational-equivalence boundary for zero-error soundness and positive progress. A falsification suite tests the implementation obligations: the prototype prevented all injected anomalies and added 3.22 ms mean commit overhead.

[AI-34] Is It Still Worth Training a Classical Model in the Era of LLM s? A Crossover Benchmark on Tabular Data

链接: https://arxiv.org/abs/2609.20218
作者: Kaihua Ding
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models can label a tabular row from a plain-English description with no training - a capability now shipping in mainstream spreadsheet tools such as Microsoft Copilot in Excel and Anthropic’s Claude for Excel - raising a practical question for the many business prediction problems where labels are expensive: should you prompt a frozen LLM, or collect data and train a model - and if so, how much data? We quantify the answer with the labeled-data crossover N*, the training-set size at which a trained classical model’s learning curve overtakes a frozen LLM’s training-free (and therefore flat) error. Aggregating 126 independent student evaluations of small GPT models under eight prompting configurations across 18 tabular datasets, paired with authoritative power-law learning curves for six classical model families, we find that training wins fast: even given an oracle choice of its best prompt configuration, a trained classical model beats the small frozen LLM using no more labeled data than is already on hand in 86% of cases, and wins by the smallest labeled subset we evaluate in 40%, with the observed crossover at a median of ~6% of the training set. In-context few-shot examples do not behave like training - error versus shot count does not follow a power law - and the same protocol re-run by independent implementers varies with a coefficient of variation of 0.148. A controlled probe indicates the LLM depends on recognizable feature-name semantics, which plausibly makes our crossover a conservative estimate (we do not claim memorization). For a typical business table, the evidence is clear: collect a few hundred labels and train a gradient-boosted model.

[AI-35] Scene-Conditioned Relation Routing for urban cellular activity forecasting

链接: https://arxiv.org/abs/2609.20209
作者: Qingzhong Li,Jingye Lin,Hui Ma,Yajun Zhang,Xinjun Pei,Ming Yan,Fei Xing
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Accepted in IEEE International Conference on Systems, Man, and Cybernetics (SMC), 2026

点击查看摘要

Abstract:Urban cellular activity forecasting requires jointly modeling heterogeneous spatiotemporal signals, including SMS usage, mobile network traffic, and call activity. Existing methods often separate temporal modeling, spatial relation learning, and multi-signal prediction, relying on fixed graph structures or static multi-task learning schemes, which limits their adaptability to changing urban scenes. We propose SCRR-Net, a scene-conditioned spatial relation routing framework in which urban contextual information jointly controls spatial dependency selection and cross-task knowledge transfer. SCRR-Net includes a context encoder, a spatial graph expert routing module, a temporal Transformer encoder, and a task knowledge routing module. Experiments on the Milano and Trento datasets demonstrate that SCRR-Net consistently outperforms competing methods on SMS, network traffic, and call activity forecasting, while providing interpretable routing behaviors.

[AI-36] JointMatch: A Unified Heterogeneous Graph Neural Solver for Large-Scale Ride-Sharing Matching

链接: https://arxiv.org/abs/2609.20200
作者: Kun Zhao,Xu Chen
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Ride-sharing platforms must continuously decide which open requests to bundle into shared trips and which idle vehicles should serve them. The dominant academic approach decomposes this into two sequential matching problems – request pairing first, then vehicle assignment – and applies a separate solver to each. This decomposition is convenient computationally but loses revenue and scales poorly because the first stage commits to ride bundles before the available vehicles are known. We propose JointMatch, a learning-based framework that handles request pairing and vehicle assignment together on a single graph. The graph is sparsified by spatial proximity so that its size grows linearly rather than quadratically with the number of vehicles and requests, and a graph neural network scores all candidate decisions in one forward pass. On the New York City Yellow Taxi data, the framework already exceeds both the classical Blossom heuristic and a faithfully-trained two-stage GNN baseline – often by a wide margin – and at city scale (fleet 10000) it runs more than 20\times faster per dispatch epoch than either. A supervised training stage closes most of the remaining revenue gap, and a policy-gradient fine-tune aligns the trained model with realised revenue.

[AI-37] Music Hallucination in Audio-Language Models: A Hierarchical Formulation and Empirical Study ACM-MM2026

链接: https://arxiv.org/abs/2609.20195
作者: Yu Liu,Jiahui Liu,Zhilin Liu,Cong Cao,Fangfang Yuan,Yuling Yang,Pin Xu,Yanbing Liu
类目: ound (cs.SD); Artificial Intelligence (cs.AI)
备注: Accepted at ACM MM 2026

点击查看摘要

Abstract:Audio-language models increasingly generate confident music descriptions that are unsupported by the input audio. We present, to our knowledge, the first music-specific, layer-wise, multi-paradigm empirical study of hallucination in audio-language models and formulate it as a hierarchical perceptual grounding failure across five layers: sound events, temporal properties, tonal attributes, style, and emotion. We introduce MuseDiag, a multi-paradigm diagnostic framework with contradiction-based verification, and evaluate nine models (four open-source and five closed-source). We find that (1) vocal misperception is a universal weakness across all nine models, tonal perception is a major axis of architectural differentiation, and Audio-Flamingo-3 remains the stable leader while substantial reordering below it reveals paradigm-specific vulnerability profiles; (2) affirmative bias, generation-mode effects, and layer-specific perceptual limitations are each empirically associated with the observed patterns, with convergent evidence from multiple analyses rather than strict causal attribution; and (3) our two training-free mitigation methods, Audio-Dependency-Aware Decoding for Music (ADD-M) and Taxonomy-Guided Perceptual Anchoring (TPA), can reduce hallucination in probing, but their gains vary by model and often do not carry over to free-form generation, showing that music hallucination mitigation must be evaluated across paradigms.

[AI-38] SoftTri: Smooth Triangular Membership Functions for Adaptive Fuzzy Inference Systems

链接: https://arxiv.org/abs/2609.20194
作者: Babak Sarani,Rahman Ardakanian,Ali Mousavi
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Triangular membership functions (MFs) are widely used in fuzzy systems because of their interpretability, low parameterization complexity, and strong locality properties. However, their inherent nondifferentiability at knot points limits the effectiveness of gradient-based optimization in adaptive neuro-fuzzy architectures, often necessitating subgradient approximations or heuristic smoothing techniques. In this paper, we propose \emphSoftTri, a differentiable triangular membership function constructed using a smooth soft-hinge mechanism inspired by Swish-type activations. The proposed formulation preserves the geometric structure and localized behavior of classical triangular MFs while providing C^\infty smoothness with respect to both the input variable and the membership parameters (a,b,c) for any finite sharpness parameter \beta0 . Closed-form analytical gradients are derived to enable efficient and fully differentiable backpropagation-based learning. SoftTri is integrated into a Takagi–Sugeno fuzzy neural network with grid-partitioned rules and evaluated on multiple one-dimensional and two-dimensional nonlinear approximation benchmarks as well as a real-world regression task using the Airfoil Self-Noise dataset. Experimental results demonstrate that SoftTri consistently improves optimization stability and approximation accuracy compared with classical triangular membership functions, while achieving performance comparable to or better than Gaussian MFs under identical rule structures and training settings. The proposed approach provides an effective compromise between interpretability and differentiable optimization in modern neuro-fuzzy learning systems.

[AI-39] Sequential Contextual Fit Predicts Human Behavioural and Neural Dynamics Across Domains

链接: https://arxiv.org/abs/2609.20179
作者: Kun Sun,Rong Wang
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Human perception, action and decision making unfold in sequences, but computational predictors are often domain-specific. This study computes and tests sequential contextual fit (SCF), an embedding-based measure of how well a current information state matches its recent context. The metric uses a simple recency-weighted similarity kernel and can be applied to words, sounds, visual scenes, affective states, choices, actions and neural representations. Across language processing, music-evoked emotion, a subset of audiovisual emotion EEG data, gambling decisions, human activity recognition and decision-related EEG, lower contextual fit predicted longer processing times, larger affective or behavioural transitions and stronger neural-state changes. These effects remained after controlling for established predictors including surprisal, reinforcement-learning prediction error, acoustic change, visual change and sensor change. SCF therefore provides a computational measurement layer for relating contextual compatibility to behavioural processing and cognitive/neural state-transition dynamics.

[AI-40] PaGNet: A Panel-Aware GBDT–Neural Network for Multi-Target Corporate Tax Avoidance Proxy Forecasting

链接: https://arxiv.org/abs/2609.20177
作者: Wonho Song,Hyungjoon Kim
类目: Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); Machine Learning (cs.LG)
备注: 32 pages, 1 figure, 13 tables

点击查看摘要

Abstract:Forecasting corporate tax avoidance proxies from firm–year panel data is challenging because predictive signals are distributed across short firm histories and related targets, while screening-oriented use requires transparent model behavior. We propose PaGNet (Panel-Aware GBDT–Neural Network), a two-branch hybrid that combines a LightGBM branch using panel-temporal summaries with a Panel-MLP branch using attention-pooled temporal aggregation and shared-trunk multi-task learning. A per-target validation-optimal blender produces both the final prediction and a compact branch-reliance diagnostic without trainable fusion parameters. On the KoTaP panel of 1,754 Korean listed firms from 2011–2024, PaGNet is evaluated under a leakage-free, shared-hyperparameter protocol across four feature regimes. In the direct-proxy-lag-excluded FS1 regime and the tax-history-augmented FS2 regime, accrual targets (TSTA, TSDA) route stably to the LightGBM branch, where PaGNet raises explained variance over the strongest of six baselines by roughly 0.08 – 0.11 on the primary split. GETR often leans toward the neural branch, while CETR exposes a validation–test branch-selection mismatch rather than a stable branch assignment. A panel-flatten control shows that most accrual gains come from observed multi-year base-panel values, with PaGNet’s panel-aware representation adding a smaller but directionally consistent refinement. Rolling-origin analysis confirms stable accrual routing, bounds ETR diagnostics to split-specific behavior, and identifies a far-horizon split where supervised models underperform naive persistence. PaGNet is therefore best viewed not as a universally superior tabular learner, but as a proxy-aware panel model that combines competitive forecasting with explicit per-target branch-reliance reporting.

[AI-41] QUALS: Corpus Equilibrium for Universal Forecasting via Pattern Quantization and Learnability Synchronization

链接: https://arxiv.org/abs/2609.20156
作者: Yujie Li,Zezhi Shao,Chengqing Yu,Yisong Fu,Weijie Zhu,Yifan Du,Jilin Hu,Bin Yang,Yongjun Xu,Fei Wang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Ubiquitous time series data across diverse domains enables critical applications in areas such as transportation systems and power grids. Recently, training foundation models on massive datasets to achieve accurate zero-shot forecasting has emerged as a major research focus. However, current studies predominantly prioritize architectural innovations while insufficiently addressing data diversity, often relying on simple data sampling strategies that fail to manage complex data distributions effectively, leading to inefficient use of training data and suboptimal performance. To address this, we propose QUALS, a large-scale time series corpus equilibrium framework. QUALS significantly enhances data efficiency, i.e., enabling existing models to achieve superior performance using only a small fraction of the original training data. Specifically, QUALS operates through two core mechanisms. First, a pattern quantization framework systematically decodes heterogeneous patterns from mixed corpora via vector quantization and uniform binning. Second, a learnability synchronization framework calibrates sampling weights for heterogeneous patterns, bridging the optimization gap between simple and complex motifs to maximize overall training efficiency. Extensive benchmarks demonstrate that pre-training on QUALS consistently achieves superior zero-shot performance, even under substantially reduced training budgets.

[AI-42] MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents

链接: https://arxiv.org/abs/2609.20152
作者: Pritish Mishra,Ishaan Kumar,Akshat Mandoli,Sudarshan Kamath
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Generally, most voice agents are cascaded systems, i.e., an ASR model transcribes the caller’s audio, a language model reads the transcript and decides what to say and which backend tools to call, and a TTS model speaks the reply. Nearly all of the decision making happens in the language model, but existing evaluations measure it either too broadly or too narrowly. End-to-end voice benchmarks score the full pipeline, so recognition errors and model errors mix into a single number. LLM benchmarks isolate the model but they do not evaluate what makes real phone calls hard, such as transcription issues, caller’s voice being split across messages and the requirement that replies follow the language and script specified. We introduce the Multi-Turn Voice Agent Benchmark (MTVA-Bench), which evaluates the language model on the same conditions it faces inside a cascaded system. The caller is played by an LLM following a set of rubrics and tool calls are answered by a mock backend which responds to the arguments the model actually sent. The benchmark contains 49 agents working across 490 reviewed scenarios and supports 7 languages. Scoring is a combination of deterministic checks on tool calls with two LLM judges, one that scores scenario specific rules and one that grades conversation quality without access to the task. Both judges must cite specific messages from the transcript. Task and conversation scores are weighted equally, since a call can complete its task and still go badly for the caller. In a seven-model study, six of the models select the correct tool within 6.4 points of one another, but their overall scores span 24.4 points. Most of the gap comes from argument values, action ordering, rule compliance, and what the model says around its tool calls.

[AI-43] AdaRepair-Mem: Adaptive Experience Orchestration for Repository-Level Program Repair

链接: https://arxiv.org/abs/2609.20130
作者: Z. C. Luo,J. C. Guo,W. J. He,S. Y. Wang,J. C. Yu,F. M. Zhao,Y. Chen,T. Cao,L. Q. Liu,N. Zheng,W. Xu,J. Jiang,Z. M. Zhao
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注: 12 pages, 9 figures

点击查看摘要

Abstract:Recent memory-augmented repository-level program repair methods reuse historical repair experiences to improve LLM-based issue resolution. However, our analysis reveals three limitations in existing repository-level memory retrieval. First, episodic memory is highly imbalanced across repositories, leaving low-resource repositories with little effective support. Second, more memory does not monotonically lead to higher repair success, suggesting that relevance, quality, and redundancy matter more than raw memory volume. Third, memory accumulation is phase-misaligned: repositories may contain many reproduction experiences but few patch or refinement experiences. To address these problems, we propose an adaptive experience retrieval framework for repository-level program repair. Our framework introduces coverage-aware retrieval, which falls back to cross-repository or repair-type-based memories when same-repository memory is insufficient; quality-aware selection, which ranks memories by relevance, historical utility, specificity, and redundancy; and stage-aware routing, which separates and retrieves memories for reproduction, localization, patch generation, patch refinement, and validation. Evaluated on SWE-Bench-Lite and SWE-Bench-Verified, the proposed framework improves repair performance on under-covered repositories, reduces noisy memory retrieval, and better supports failed-to-fixed patch refinement. Our results show that the key to memory-augmented repair is not simply accumulating more experiences, but retrieving the right experiences for the right repair context.

[AI-44] Local Sparsity Enables Unsupervised LLM Safety Detection

链接: https://arxiv.org/abs/2609.20129
作者: Xin Chen,Gil Kur,Alexander Shevchenko,Andreas Krause
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Deployment-time safety methods for large language models (LLMs) are predominantly supervised and assume access to unsafe training data. Nevertheless, new attacks and harm categories regularly arise, not captured by models trained in such a supervised fashion. An alternative approach is to view this problem through the lens of anomaly detection, namely, to rely solely on modeling safe data and flagging out-of-distribution inputs. However, LLM activations lie in a high-dimensional space, raising concerns about whether anomaly detection is statistically feasible. We show that, under the linear representation hypothesis (LRH), there may indeed be hope. In the LRH concept space, which is typically recovered via a sparse autoencoder (SAE), nearby points share a small common active support. Using this local sparsity insight, we propose a framework for locally masked SAE-based anomaly detection, supported by theoretical justifications. We validate it on various architectures and datasets, including both capability-testing datasets and safety-specific datasets. Finally, when we allow algorithms to use 1% out-of-distribution data for calibration, locally sparse methods achieve near-optimal performance, demonstrating their ability to capture meaningful safety information while using only 1-2% of SAE neurons for computation.

[AI-45] Multi-Dimensional Prosody Judgment For Live Streaming Speech Synthesis

链接: https://arxiv.org/abs/2609.20124
作者: Zifan Guan,Longyu Lu,Junan Zhang,Zhizheng Wu,Meiguang Jin,Junfeng Ma
类目: ound (cs.SD); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Evaluating live streaming speech synthesis (TTS) requires assessing fine-grained, highly expressive prosody such as emotion, intonation, and energy which traditional MOS predictors fail to capture. While proprietary Large Language Models (LLMs) like Gemini can evaluate these aspects, they are too costly for massive inference and reinforcement learning feedback. To address this, we first introduce Live-ProsodyJudge (LPJ), a cost-effective pairwise evaluator distilled from Gemini into Qwen3-Omni. However, we identify a critical flaw in standard multi-dimensional evaluation: verdict coupling. The judge tends to lazily align all individual dimension scores with its overall preference, collapsing a rich multi-dimensional rubric into a single preference bit. To resolve this, we further propose Decoupled-Live-ProsodyJudge (D-LPJ). D-LPJ eliminates the overall verdict target to prevent blind following, masks uncertain pair-dimensions during Supervised Fine-Tuning(SFT), and introduces a novel span-local GRPO strategy that applies normalized advantages strictly to their corresponding rationale spans. Evaluated on highly curated human-annotated test sets, 10 sample balanced-order LPJ achieves higher point accuracy than a single Gemini call, while D-LPJ successfully produces independent,decoupled dimension judgments. Furthermore, in a Best-of-8 TTS candidate selection tournament, the LPJ-selected utterance falls within the human top-3 in 85.29% of high-confidence cases, demonstrating its efficacy for fine-grained TTS preference optimization.

[AI-46] Perception Layout and Validation: Calibrated Confidence for Reliable Straight-Through Processing of Financial Documents

链接: https://arxiv.org/abs/2609.20110
作者: Yichao Jin,Yushuo Wang,Yuxuan Han,Kwan Ching Yee Sonia,Weiyang Song,Chiu Jin-Chun Kent,Wong Chong Hwee,Wong Tiong Kiat,Kenneth Zhu Ke,Jingyuan Zhao
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Straight-through processing (STP) on extracted key-value fields from financial documents without human review requires a calibrated probability together with a bounded guarantee on the residual error of the auto-approved tier. The emergence of modern Vision Language Models (VLMs) provides an out-of-the-box capability for extracting the key-values, but their verbalized confidence signals are unreliable and weakly track field correctness. This paper introduces a decomposed confidence layer along three interpretable channels, including perception, layout, and validation. Together with a final conformal risk control, the score can be used for reliable STP of financial documents. The method is validated on three public datasets covering real invoices, synthetic invoices, and ad-buy forms, using two different VLM families (Qwen3.6-27B and Gemini-3.1-Flash-Lite). Our decomposed score consistently improves the separation of correct from incorrect extractions, substantially raising the AUROC from 0.54-0.74 for VLM verbalized signals to 0.90-0.99 with contributions from all three designed channels. Crucially for industrial deployment, this enables usable STP. The native VLM confidence signals could clear only 0.1%-7.0% of fields under risk control at a target error of 10%. In contrast, the proposed method auto-approves 49-72% of fields while holding the empirical error of the accepted tier at or below the target.

[AI-47] A Scalable Trust Discovery Architecture for the Internet of Agents

链接: https://arxiv.org/abs/2609.20095
作者: Song Zhang,Jiankang Yao,Hongtao Li,Xiaojun Zhang,Xugang Shen,Xin Li,Yanbiao Li
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:The Internet of Agents is expected to enable large numbers of autonomous agents to discover, verify, and collaborate with each other across heterogeneous platforms. However, current agent protocols mainly address tool invocation and inter-agent communication, leaving scalable agent registration, trustworthy identification, and capability-oriented discovery largely unresolved. To address this, this paper proposes a scalable trust discovery architecture for the Internet of Agents. The proposed architecture adopts a hierarchical and distributed design consisting of three layers: Agent Root for trusted registry governance, Agent Registry for agent registration and metadata publication, and Agent Resolver for distributed capability discovery and trust-aware resolution. The architecture further introduces a registry-suffix-anchored composite identity scheme, which binds an agent native identifier to a trusted registry suffix to generate a globally discoverable identity. It also incorporates a dual-certificate and multi-level authentication mechanism to strengthen identity trust among agents. We implement a prototype and evaluate it through large-scale agent registration and resolution experiments. The prototype achieves an average registration latency of 58ms and an average discovery latency of 25ms, and it supports more than 19,000 registration requests per second and more than 29,000 agent discovery requests per second. These results demonstrate the feasibility of the proposed architecture, providing a practical approach toward scalable and identity-trusted agent ecosystems in the Internet of Agents.

[AI-48] Solving Minimum Span Antibandwidth and Cyclic Antibandwidth Labeling Problems

链接: https://arxiv.org/abs/2609.20091
作者: Hieu Truong Xuan,Khanh To Van
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:The Antibandwidth and Cyclic Antibandwidth problems are NP-hard graph labeling problems that aim to maximize the minimum (cyclic) distance between labels assigned to adjacent vertices. Extensive research on these problems has resulted in a variety of mathematical formulations and computational approaches. However, their minimum span perspective, in which a prescribed minimum (cyclic) distance is fixed and the objective is to minimize the label span, has received comparatively little attention. In this paper, we consider this complementary perspective by introducing the Minimum Span Antibandwidth/Cyclic Antibandwidth Labeling (MSABL/MSCABL) problems and developing a unified Boolean Satisfiability (SAT)-based framework for solving them. The SAT-based framework formulates MSABL/MSCABL as a sequence of decision problems and exploits their monotonicity to accelerate the search process. We also consider two SAT solving strategies, parallel and incremental SAT solving: the former examines multiple candidate spans concurrently, while the latter reuses a single SAT instance while progressively restricting the label domain. The proposed approaches are evaluated on benchmark instances from the Harwell-Boeing Sparse Matrix Collection and compared with CPLEXCP, CPLEXMIP, and Gurobi. The results show that SAT-based approaches are highly competitive in solution quality, with the parallel approach performing best overall for MSCABL and the incremental approach for MSABL. With the no-hole constraint, they remain competitive with CPLEXCP and significantly outperform CPLEXMIP and Gurobi, particularly for MSCABL. These results demonstrate the effectiveness of SAT solving as an exact approach for MSABL and MSCABL.

[AI-49] UnifiedPlayers: Enhance Tool-Integrated Reasoning in Agent ic Reinforcement Learning

链接: https://arxiv.org/abs/2609.20089
作者: Wenjie Liao,Liangjie Zhao,Zehong Cao
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Self-evolving methods reduce the need for human-annotated trajectories by allowing tool-using agents to generate their own training data. Yet existing methods typically separate trajectory generation from evaluation, relying on static verifiers that cannot adapt to emerging failure modes or self-consistency signals that may reinforce errors shared across trajectories. Jointly adapting planning, execution, and evaluation offers a promising alternative, but introduces a fundamental coordination challenge: each component continuously changes the data or feedback used to train the others. We address this challenge with \textbfUnifiedPlayers, a cooperative framework comprising a Planning Player that generates tasks, an Execution Player that produces multi-turn trajectories with Python tool calls, and an Evaluation Player that constructs executable verifiers. We design role-specific rewards that coordinate the three players toward a shared learning objective under GRPO. Across two model backbones and twelve reasoning benchmarks, UnifiedPlayers outperforms the strongest prior baseline by at least 3.5% on mathematical reasoning and 3.9% on general reasoning tasks. Moreover, the learned verifier achieves 84.2% adversarial detection accuracy, while its reward signal exhibits 2.03 \times higher per-question variance than a self-consistency baseline, providing more discriminative verifications. These results highlight cooperation among specialized players as a promising path toward self-enhanced tool-integrated agents.

[AI-50] ailored to you: longitudinal effects of personalising language models

链接: https://arxiv.org/abs/2609.20077
作者: Canfer Akbulut,Justine Breuch,Arianna Manzini,Lujain Ibrahim,Matija Franklin,Roma Patel,Iason Gabriel,Kristian Lum,Laura Weidinger
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Interest in developing personalised language models is rapidly growing. While personalisation is often viewed as a mechanism to better serve diverse user needs, the effects of sustained interactions with personalised models on people’s perception of and behaviour toward AI remain poorly understood. Most critically, downstream consequences outside the immediate human–AI interaction loop, such as effects on users’ self-perceptions and interpersonal relationships, remain largely unexamined. In this study, we recruited 992 participants to complete daily advice-seeking interactions with language models over the course of five days, comparing outcomes from a non-personalised baseline against two personalisation approaches: memory-based (conditioned on prior conversational history) and survey-based (conditioned on information collected through a pre-study intake survey). We find that several changes in human-AI interaction over time are driven primarily by repeated exposure rather than personalisation itself. However, participants interacting with personalised models experienced differences in advice-seeking and information-sharing attitudes and behaviours: participants in the memory-based condition engaged in greater self-disclosure and rated the model as less creepy, while participants in the survey-based condition reported higher regret about having shared personal information with the AI. We conclude by highlighting the nuanced effects of different personalisation approaches on interaction outcomes, and discussing the implications of these findings for the responsible design and deployment of personalised AI systems.

[AI-51] FCA-Guided Counterfactual Explanations for Multi-Modal Breast Cancer Diagnosis: A Framework Achieving Perfect Validity with Emergent Sparsity

链接: https://arxiv.org/abs/2609.20067
作者: Abdullahi Isa,Souley Boukari,Muhammad Aliyu
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Deep learning models for multi-modal breast cancer diagnosis achieve high predictive accuracy but remain clinically unacceptable without actionable, counterfactual explanations. Attribution-based methods (LIME, SHAP) are categorically inapplicable to this purpose, as they generate no alternative instances and thus cannot be evaluated on counterfactual quality metrics. This investigation provides empirical evidence that FCA-Guided Counterfactual (FCA-CF) framework that uses a Formal Concept Analysis (FCA) concept lattice as a hard structural constraint on counterfactual search, operating over a multi-modal TCGA-BRCA dataset. We benchmark against four genuine counterfactual methods: Wachter-style CF, DiCE, FACE, and NICE, evaluated on 60 benign-predicted TCGA-BRCA instances. The FCA-CF framework achieves Validity = 1.0000 (100% of counterfactuals successfully flip the prediction), Sparsity = 2.37 features changed (best among all valid methods), and Proximity = 0.900 (normalised L2-based, matching NICE as joint best). The classifier achieves Accuracy = 0.980, F1 = 0.976, ROC-AUC = 0.9947. Ablation analysis confirms that the FCA lattice constraint is the primary sparsity driver (removing it increases sparsity by +40%, p 0.001, Cohen’s d = 0.78), while Phase C greedy refinement accounts for the largest individual contribution (+113% sparsity increase when disabled, p 0.001, d = 5.01). FCA-guided counterfactual generation achieves a clinically important Pareto-dominant outcome; it is simultaneously the sparsest and among the most proximate of all valid methods, with perfect validity. The emergent sparsity property arising from lattice topology rather than numerical penalty terms constitutes a structurally novel contribution to the counterfactual explanation literature.

[AI-52] Robust Workflow Generation via Adversarial Learning for Audio Deepfake Detection

链接: https://arxiv.org/abs/2609.20063
作者: Xiang Li,Pin-Yu Chen,Wenqi Wei
类目: ound (cs.SD); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:The rapid advancement of speech synthesis and voice conversion technologies has made audio deepfakes increasingly realistic, posing serious security risks in practical applications. While existing detection methods achieve strong performance under controlled conditions, they often fail to generalize under real-world perturbations and corruptions. In this paper, we propose ROGUE, a framework that dynamically constructs robust detection workflows by orchestrating multiple detection tools. ROGUE formulates workflow generation as a sequential decision-making problem and introduces a dual-agent paradigm, where a perturbation agent generates audio perturbations and a policy agent learns to select and execute detection tools under perturbed conditions. Through adversarial learning, ROGUE enables perturbation-aware tool selection, adaptive execution strategies, and improved robustness to distribution shifts. Extensive experiments across multiple datasets and real-world corruptions demonstrate that ROGUE consistently outperforms strong baselines in both robustness and generalization. Our results highlight the effectiveness of adversarially optimized workflow generation for building reliable audio deepfake detection systems in real-world deployment settings.

[AI-53] WiCleanData: Guaranteeing the Type Consistency of Wikidata by Taxonomy Refinement and Constraint Enforcement

链接: https://arxiv.org/abs/2609.20057
作者: Yiwen Peng(IP Paris),Marc Jeanmougin(IP Paris),Thomas Bonald(IP Paris)
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Because of its collaborative nature, Wikidata suffers from errors, in- consistencies, and excessive complexity, such as redundant classes, ambiguity between instances and classes, wrong taxonomic paths, and type constraint violations. The manual curation of these issues is infeasible at scale. To address these challenges, we introduce WiCleanData, a refined version of Wikidata with a consistent tax- onomy and free from type constraint violations. Specifically, we have designed an automated pipeline that first cleans the taxonomy with language model assistance, then simplifies type constraints by hierarchical aggregation, and finally filters facts accordingly. The resulting knowledge graph, free from any type violation, is made publicly available via a Web interface, enabling easy exploration and downstream applications.

[AI-54] MAGMA-GEN: Validated Recovery Supervision from Ambiguous Failures via Counterfactual Re-Execution

链接: https://arxiv.org/abs/2609.20056
作者: Loan Bernat(LAAS-GEPETTO),Matthieu Grard,Ariane Herbulot(LAAS-RAP),Florent Lamiraux(LAAS-GEPETTO)
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Hierarchical robotic systems executing long-horizon manipulation tasks must make high-level semantic decisions that orchestrate stochastic low-level skills. In this setting, failed rollouts are ambiguous: a poor downstream state may reflect an invalid high-level decision, partial observation, or a valid decision whose physical execution failed. Traditional supervised learning lacks data for such recovery states, while reinforcement learning struggles with sparse rewards and non-local credit assignment. We propose MAGMA-GEN, an on-policy data-generation pipeline that converts ambiguous failed rollouts into validated recovery supervision. MAGMA-GEN first uses a privileged coach to hypothesize an early decision-level error and propose localized correction or recovery actions. Because this diagnosis is fallible, candidates are retained only if re-execution from the same state under matched conditions improves downstream progress. This produces supervised examples from the agent’s own failure distribution without per-step human demonstrations. Evaluated on interactive long-horizon manipulation tasks, MAGMA-GEN improves task success and recovery capabilities, against distillation and trajectory-repair baselines under evolving task constraints in both simulation and real-robot execution.

[AI-55] DART: Distillation-Aware Reparameterization for Training-Free LoRA Reuse in Few-Step Video Diffusion Models

链接: https://arxiv.org/abs/2609.20051
作者: Shihong Li,Juntao Xu,JinCao,Maowen Tang,Jun Huang,Jintao Li
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Step distillation reduces the cost of video generation, but reusing a LoRA trained for a longer trajectory can alter its functional effect or degrade target quality. Static parameter compatibility offers one perspective on this problem; our observations show that similar measured geometry can coexist with different adapter behavior under a shortened denoising schedule. We propose DART, a training-free method that combines low-rank coordinate transport with target-schedule response calibration using forward evaluations and no source training videos. On a four-step Wan2.2 target, DART-F improves the joint quality score from 0.9029 to 0.9227 and changes macro functional retention from -0.4644 to +0.1349. Component analysis shows that calibration accounts for most of the quality improvement, while coordinate transport provides complementary gains when combined with calibration. Adapter-level results reveal positive functional effects for some adapters and strong attenuation with reduced negative functional effects for others. Evaluations on two additional targets show the same aggregate trend. These results motivate evaluating distilled-model LoRA reuse jointly through functional preservation and negative-transfer avoidance, without assuming recovery for every adapter.

[AI-56] Correct Now Insufficient Later: Auditing Update Sufficiency in Context Compression

链接: https://arxiv.org/abs/2609.20045
作者: Guangzhe Zhang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 20 pages, 9 tables, 2 figures. Code and reproducibility materials to be released separately

点击查看摘要

Abstract:A memory can answer a current query correctly while discarding distinctions required by a later update. We investigate this failure with a paired-history audit: two histories have the same current answer, receive a shared future update, and require different subsequent answers. A pilot evaluates 24 history pairs across six synthetic mechanisms, 12 memory conditions, two repeats, and two model backends. A deterministic frontier selector obtains strict reveal accuracy of 96/96 on DeepSeek and 82/96 on GLM; a structured writer obtains 62 successes with one unresolved outcome and 56/96. The configured four-outcome joint contrast has finite-sample identification intervals of [0.521, 0.542] and [0.292, 0.313], not confidence intervals. A record-level audit distinguishes retained-state adequacy, response delivery, and answer-schema compliance without changing those original scores. It finds 26 and 25 well-formed but semantically wrong structured reveal memories, while all 14 GLM frontier reveal failures contain correct values in the wrong wrapper. Tombstone removal produces 16/16 exact replay failures in the targeted mechanism. Identifier renaming then exposes a separate flaw: original frontier late-reference adequacy falls from 8/8 to 94/320 transformed instances. We provide and test a label-equivariant repair, but it preserves only 2/8 original late-reference answers: eliminating a naming shortcut does not solve unknown future relevance. These results support a scoped evaluation methodology and reproducible failure analysis, not general superiority of the repaired algorithm. Paid pilot evidence, retrospective diagnostics, and new offline tests are reported separately; no independent held-out or natural-task validation is claimed.

[AI-57] Can Data Attribution Filter Out Subliminal Learning? Not Reliably

链接: https://arxiv.org/abs/2609.20027
作者: Moritz Weckbecker,Sweta Jena,Jonas Müller,Ponnurangam Kumaraguru,Sebastian Lapuschkin,Wojciech Samek,Louis Jaburi,Gonçalo Paulo
类目: Artificial Intelligence (cs.AI)
备注: 16 pages, 17 figures

点击查看摘要

Abstract:Subliminal learning allows language models to transmit behavioral traits through training data with no obvious semantic relationship to those traits, undermining content-based data filtering as a safety intervention. Training data attribution offers an alternative: it identifies the training examples responsible for a given model behavior, independent of their semantic content, and so may apply in exactly the cases where semantic inspection fails. We evaluate three gradient-based attribution methods (GradCos, a contrastive GradCos variant, and EK-FAC) across three models, comparing them against divergence tokens, a strong baseline previously shown to localize subliminal learning (albeit one that requires access to counterfactual teacher models). Filtering at the token level, EK-FAC mitigates a significant part of the effect, the other methods provide little benefit, and all mostly fall short of divergence tokens. Filtering entire samples is less effective for every method, though EK-FAC often gives a stronger signal than divergence tokens in this setting. Success is inconsistent across methods and settings: variants that work well for some model-preference combinations fail for others, and we do not identify a consistent explanation for these differences. Our results suggest that gradient-based attribution can identify data responsible for subliminal learning in some settings, but that some approximations are more reliable than others.

[AI-58] FedeRICo: Federated Region-Influenced Coupling for Traffic Flow Prediction

链接: https://arxiv.org/abs/2609.20026
作者: Fermin Orozco,Man Luo,Johan Wahlström
类目: Artificial Intelligence (cs.AI); Distributed, Parallel, and Cluster Computing (cs.DC)
备注:

点击查看摘要

Abstract:Urban traffic forecasting often relies on information distributed across stakeholders who may be unable to share raw data due to privacy or commercial constraints, motivating federated spatial-temporal approaches. In such federated settings, each client observes traffic over a distinct sensor subgraph with its own spatial topology and temporal dynamics, leading to significant heterogeneity across clients. Existing federated spatial-temporal methods typically rely on model parameter aggregation and provide limited mechanisms for recovering spatial dependencies across client boundaries. This introduces two key limitations. Specifically, parameter aggregation across heterogeneous graph domains tends to dilute client-specific representations, while road network partitioning breaks the propagation of traffic dynamics across client boundaries. To address these challenges, we propose FedeRICo, a federated traffic forecasting framework that combines gradient-level collaboration with boundary-aware residual communication. FedeRICo employs a dual-branch forecasting architecture in which a globally guided branch captures transferable forecasting structure, while a private residual branch preserves client-specific corrections and incorporates boundary residual signals. The global branch is coordinated through gradient alignment across all clients, enabling collaborative optimisation without destructive parameter interference. To recover cross-client spatial dependencies, boundary messages are extracted through a trend-residual decomposition that suppresses periodic structure and communicates only transient spatial-temporal residual signals between physically adjacent clients. Experiments across four real-world traffic forecasting benchmarks demonstrate that FedeRICo consistently outperforms state-of-the-art federated spatial-temporal baselines while maintaining competitive training runtime.

[AI-59] Governance-as-Code: Translating EU AI Act Technical Requirements into Executable Compliance Pipelines for Generative AI Systems ICML2026

链接: https://arxiv.org/abs/2609.20016
作者: Rudrendu Kumar Paul,Sourav Nandy
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI)
备注: Accepted at the AI4Law Workshop, ICML 2026. Camera-ready version

点击查看摘要

Abstract:The EU AI Act (Regulation 2024/1689) imposes technical obligations on high-risk AI providers, yet Articles 8-15 were drafted for predictive AI and leave seven technical gaps when applied to generative systems, spanning non-deterministic data governance, training-data provenance, continuous conformity, human oversight, open-ended robustness, emergent risk, and generative fairness. We deliver Governance-as-Code (GaC), a framework of 43 machine-checkable acceptance criteria across six compliance modules that run in a CI/CD pipeline and emit Article-indexed audit evidence, and we show the actual Rego policy code rather than merely describing it. Our central commitment is that the Act’s open-textured standards (“appropriate levels,” “possible biases”) become declared, auditable numbers: robustness thresholds are derived from the provider’s documented baseline and a state-of-the-art floor, and framing bias is collapsed into eight measurable proxies tested by counterfactual demographic probing. We also correct who owes what, since under Article 25 and Chapter V a downstream deployer relies on the upstream provider’s Article 53 training-data summary and documents only the layers it controls, so GaC verifies that summary rather than demanding per-sample documentation the deployer never had. We validate on two enterprise deployments, a high-risk advisory chatbot and a limited-risk content generator, benchmarking against a manual expert audit rather than documentation artifacts that were never designed to enforce compliance. GaC reproduces all of the manual audit’s findings, including three penalty-triggering violations, while cutting audit labor by roughly 75%.

[AI-60] Dynamic Generalized Gromov-Wasserstein Optimal Transport

链接: https://arxiv.org/abs/2609.20008
作者: Junda Ying,Zhiwei Zeng,Peijie Zhou,Lei Zhang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Optimization and Control (math.OC); Quantitative Methods (q-bio.QM)
备注:

点击查看摘要

Abstract:Gromov–Wasserstein optimal transport (GW-OT) extends classical optimal transport by introducing structure-aware transport cost. This is particularly relevant for spatial transcriptomics, where dynamical reconstruction should preserve tissue structure in addition to matching expression patterns. While static formulations have been widely used for such structure-aware alignment, a general dynamic formulation for reconstructing continuous trajectories is still missing. We introduce Travelling Pair Dynamical Alignment and Trajectory Estimation (TP-DATE), a theoretical and computational framework to generalize GW-OT dynamically in a simulation-free manner. We formulate a broad class of static and dynamic Quadratic-form OT (QOT) through path actions and prove the static dynamic equivalence. We further develop travelling-pair flow matching, which allows interacting conditional paths and marginalizes their interactions into a single vector field. On synthetic and real spatial transcriptomics data, TP-DATE better preserves spatial structure and improves continuous 3D dynamics reconstruction.

[AI-61] EPIG-Tree: Compute-Optimal Branching for Gradient-Efficient Reinforcement Learning

链接: https://arxiv.org/abs/2609.20004
作者: Nikita Khomich,Leopold Hermansson,Ido Hakimi
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 12 pages, 8 figures

点击查看摘要

Abstract:Reward-based reinforcement learning for language models, exemplified by Group Relative Policy Optimization (GRPO), collapses an entire stochastic trajectory into a single scalar reward. This is clean and scalable, but it explores and allocates reward inefficiently: a trajectory may contain many causal decisions, recovery attempts, and environment-randomness events, yet every token or action inherits one trajectory-level advantage. We study tree-based rollout construction as a compute-allocation problem for policy-gradient estimation. Our central claim is that branches should be placed not where the policy is merely uncertain, but where an additional branch most reduces uncertainty about the policy gradient per unit of compute. From a law-of-total-variance decomposition of the local policy-gradient random variable, we derive two allocation laws: new branches reduce decision uncertainty, while repeated suffix rollouts reduce continuation uncertainty. The resulting EPIG-Tree score allocates branches using the already computed rollouts. It estimates occupancy- and score-weighted value uncertainty, along with a suffix law n_e \propto w_e |\nabla_\theta \log \pi(a_e|h_e)| \sigma_e / \sqrtc_e . Empirically, EPIG reduces gradient MSE in cloned-state control, winning in all nine dense continuous-control environments of a 13-environment sweep and recovering the reference gradient direction near-perfectly, and it improves frozen-LLM gradient calibration relative to entropy branching. In online single-turn math, tree-local credit beats flat GRPO, while branch placement is secondary to token-level credit assignment. In online multi-turn Wordle, EPIG attains the highest final win rate (0.850), overtaking flat GRPO, which saturates early at 0.790, and entropy branching as training proceeds, confirming that the gradient-estimation advantage transfers to a stateful, large-action setting.

[AI-62] E-AVI: Evidence-Grounded Multimodal Assessment for Automated Video Interviews

链接: https://arxiv.org/abs/2609.20001
作者: Haoshen Wang,Dongbo Che,Zeyi Xie,Yuanjie Du,Shicheng Hua,Xingyu Wang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Automated video interview assessment integrates verbal content, acoustic delivery, and visual behavior, yet numerical predictions alone provide limited inspectable support. We present E-AVI, an evidence-grounded framework that extracts timestamped multimodal evidence and integrates dimension-conditioned evidence attention with source-level embeddings for scoring. A shared evidence pool further supports natural-language feedback and follow-up question answering. On RecruitView and a private hospitality dataset, E-AVI consistently outperforms fine-tuned multimodal baselines in rank correlation. Ablation, evidence-deletion, bootstrap, human-audit, and QA analyses characterize the predictive contribution, grounding, and practical utility of the evidence pathway. Together, these results demonstrate that our proposed E-AVI framework improves predictive performance while providing inspectable support for assessment, feedback, and interactive analysis.

[AI-63] Customizable and Jointly Optimized Route Planning : A Deep Architecture Enabling Differentiable Shortest-Path Search

链接: https://arxiv.org/abs/2609.19996
作者: Rui Zhao,Chao Chen,Longfei Xu,Chenguang Ji,Hengbin Cui,Kaikui Liu,Xiaolong Li
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:With the widespread use of online navigation and ride-hailing services, achieving optimal route planning for diverse user preferences has recently attracted increasing attention. Classic graph algorithms for pathfinding use heuristic cost functions to define edge weight, thus providing no optimality guarantee of route quality. Prior data-driven approaches equating ground truth of the optimal route with user trajectory, which is however moderately influenced by the navigation service, suffers from the feedback loop problem. To address these issues, we propose a deep architecture that is able to jointly optimize cost functions and route-ranking model towards any route preference. First, we run a multi-objective Dijkstra algorithm offline to collect the set of Pareto optimal routes, deeming it as the complete candidate set. Exploiting the property of such a set, we design a neural network structure that emulates shortest-path search and route ranking in an end-to-end differentiable manner. Second, we define route preference as a task of constrained optimization of route attributes, and propose a novel loss function that optimizes a single-objective variable, with other variables strictly under constraints. We conduct extensive experiments on real-world datasets. The results show that our architecture significantly outperforms state-of-the-art methods in route quality and customizability.

[AI-64] Past Future All at Once: Mitigating Stability-Plasticity Dilemma via Post-hoc JANUS Rectification

链接: https://arxiv.org/abs/2609.19985
作者: Zhilong Zheng,Letian Tao,Yang Guan,Yujie Yang,Wei Xiong,Kehua Sheng,Bo Zhang,Jingliang Duan,Keqiang Li,Shengbo Eben Li
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Fine-tuning foundation models on new tasks inevitably suffer from catastrophic forgetting. While existing works attempt to mitigate this on the basis of parameter-efficient fine-tuning methods, they adopted an overly restrictive Subspace Orthogonality condition. In this paper, we introduce a purely post-hoc and tuning-agnostic weight rectification framework that achieves Parameter Space Orthogonality, which is the necessary and sufficient condition for preserving historical performance to the first order. By projecting parameter updates into the JAcobian NUll Space (JANUS), our method significantly recovers compromised historical knowledge without interfering with the underlying fine-tuning process. To overcome the local validity of the Jacobian approximation, we further propose a Multi-step Adaptive Rectification mechanism that utilizes the JANUS shift to dynamically verify the valid trust region and adjust step sizes. Coupled with our proposed ghost projection, ghost orientation comparison, and sequence-level singular value decomposition compression techniques, JANUS also achieves great temporal and spatial efficiency. Experiments demonstrate that JANUS seamlessly integrates with various fine-tuning methods, significantly mitigating the stability-plasticity dilemma by recovering historical knowledge while preserving downstream task adaptation.

[AI-65] MaskHarness-WAM: Instance-Grounded Harnessing for Long-Horizon Robot Manipulation

链接: https://arxiv.org/abs/2609.19974
作者: Zitai Huang,Taiyi Su,Jian Zhu,Jianjun Zhang,Chong Ma,Tianbin Liu,Weiyi Lu,Yi Xu,Hanli Wang
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Long-horizon robot manipulation requires not only stable local visuomotor control, but also continuous target tracking and reliable task progress assessment throughout execution. This challenge becomes particularly critical when multiple objects share identical appearances and must be manipulated in a prescribed order. In such scenarios, relying solely on a limited-horizon manipulation policy is often insufficient to determine which instance should be operated on and when the task should transition to the next stage. To address this challenge, we propose MaskHarness-WAM, an instance-grounded harness for long-horizon manipulation. The proposed system connects high-level task planning with low-level manipulation policies through target masks, while leveraging visual feedback for subtask scheduling and continuous execution. Since each subtask corresponds to a different target instance, the low-level policy requires a newly established initial target mask under the updated scene at each subtask transition. The harness continuously re-observes the environment, generates, and verifies the target mask at subtask boundaries, thereby updating the instance-level spatial condition provided to the low-level policy. Furthermore, the system advances the manipulation process by switching target instances according to the verified completion status of each subtask. Experiments on a real robot platform demonstrate that MaskHarness-WAM substantially outperforms limited-horizon policies on sequential multi-object manipulation, showing its effectiveness in extending local manipulation skills to reliable long-horizon execution.

[AI-66] Efficiently Distributed Federated Learning

链接: https://arxiv.org/abs/2609.19972
作者: Gianluca Mittone,Robert Birke,Marco Aldinucci
类目: Performance (cs.PF); Artificial Intelligence (cs.AI); Distributed, Parallel, and Cluster Computing (cs.DC)
备注:

点击查看摘要

Abstract:Federated Learning (FL) is experiencing a substantial research interest, with many frameworks being developed to allow practitioners to build federations easily and quickly. Most of these efforts do not consider two main aspects that are key to Machine Learning (ML) software: customizability and performance. This research addresses these issues by implementing an open-source FL framework named FastFederatedLearning (FFL). FFL is implemented in C/C++, focusing on code performance, and allows the user to specify any communication graph between clients and servers involved in the federation, ensuring customizability. FFL is tested against Intel OpenFL, achieving consistent speedups over different computational platforms (x86-64, ARM-v8, RISC-V), ranging from 2.5x and 3.69x. We aim to wrap FFL with a Python interface to ease its use and implement a middleware for different communication backends to be used. We aim to build dynamic federations in which relations between clients and servers are not static, giving life to an environment where federations can be seen as long-time evolving structures and exploited as services.

[AI-67] Neuro-Symbolic Agent ic AI for Networked Low-Altitude UAVs

链接: https://arxiv.org/abs/2609.19961
作者: Yuqi Ping,Tianhao Liang,Nanchi Su,Guangyu Lei,Junwei Wu,Qinyu Zhang,Tingting Zhang
类目: Artificial Intelligence (cs.AI); Systems and Control (eess.SY)
备注: Agentic AI, neuro-symbolic AI, unmanned aerial vehicles (UAVs), autonomous decision-making, networked UAV systems

点击查看摘要

Abstract:Networked low-altitude unmanned aerial vehicles (UAVs) need reliable and adaptive decision-making capabilities to operate under uncertain observations, dynamic environments, and intermittent connectivity, while many existing agentic systems remain limited by hallucination risks, data dependence, and weak generalization. This article investigates neuro-symbolic agentic AI (NSAAI) as a framework for combining neural grounding, symbolic reasoning, and closed-loop agentic interaction to support more reliable and adaptive UAV autonomy. We first examine its capability foundations in data efficiency, compositional generalization, continual learning, and zero-shot transfer, and then develop a reference architecture integrating task and goal management, neuro-symbolic planning, verification and metacognition, skill execution and network interaction, and shared knowledge and memory. An urban fire-inspection case implemented in LAESim illustrates how a UAV can coordinate sensing and cloud access under intermittent connectivity, reuse a verified image-delivery skill, and satisfy explicit evidence conditions before completing the mission. The results illustrate the potential of NSAAI to support reusable skills, evidence-grounded decision-making, and adaptive mission execution in networked UAV systems. We further discuss key research directions in uncertainty-aware reasoning, knowledge and skill expansion, adaptive self-monitoring, and standardized evaluation.

[AI-68] Not All AI Agents Are Equal: Characterizing Resource and Performance Dynamics

链接: https://arxiv.org/abs/2609.19947
作者: Wonmi Choi,Minuk Park,Zhixiong Niu,Yongqiang Xiong,Chuck Yoo,Gyeongsik Yang
类目: Artificial Intelligence (cs.AI); Performance (cs.PF)
备注:

点击查看摘要

Abstract:LLM-based AI agents process user requests through iterative reasoning and tool execution, often involving the invocation of remote LLM APIs with local tool containers. This execution model can make the optimization of agent serving difficult because latency, local resource demand, and container bottlenecks inter-mix across requests. However, the current agent ecosystem runs without much consideration of resource dynamics, which results in significant waste of the precious resources. This paper analyzes the resource inter-mix of AI agents for three representative tasks: retrieval-augmented question answering, web search, and software coding. To this end, we characterize the latency with respect to the resource dynamics of processing multiple requests and tasks concurrently. Our measurements show that agents have a wide range of behaviors depending on tasks, so that even the same tool can differ substantially in resource dynamics. We also find that running multiple requests concurrently exposes task-dependent bottlenecks in resource dynamics such as CPU, disk I/O, and memory. Furthermore, we uncover that faster LLM responses or more CPU cores do not always accelerate agents. Based on these observations, we demonstrate new optimization opportunities that exploit the resource dynamics of tasks: CPU-aware tool admission and task-aware CPU allocation. Our results show that the latency of CPU-sensitive agent tasks improves \sim 5.4 \times , and the average latency across multiple tasks is reduced \sim 32% compared to native agents.

[AI-69] MaSCoD: A Multi-Agent Framework for Structural-Context-Guided Candidate Causal Graph Generation

链接: https://arxiv.org/abs/2609.19944
作者: Yudai Nakada,Yuichiro Nishiura,Jin Michael Splichal
类目: Artificial Intelligence (cs.AI)
备注: 31 pages, 4 figures, 18 tables. The first two authors contributed equally

点击查看摘要

Abstract:Large language models (LLMs) have been applied to causal discovery, but candidate-graph generation rarely treats premature omission of potentially relevant causal relations as an explicit design objective. We propose MaSCoD, a multi-agent framework that organizes candidate third variables and local structural patterns before direct-edge judgment. We evaluate MaSCoD on Auto-MPG, DWD, and Sachs using GPT-5.4 as the primary backbone and GPT-4o for replication. MaSCoD exhibits a dataset- and backbone-dependent retention-selectivity profile rather than uniform superiority. Across all six dataset-backbone settings, Full, which supplies structural hypotheses before direct-edge judgment, achieved higher mean Recall and F1 than No Phase 1, which instead constructs them within the judgment procedure, while also increasing false-positive rates. Additional reference-edge retention over all evaluated baselines was observed on DWD with GPT-5.4 and on Sachs with GPT-4o, rather than uniformly across settings. Partial ablations showed that supplying both information components did not always outperform supplying only one. For GPT-5.4, stage-wise analysis showed that the Full-No Phase 1 retention gap was already present after direct-edge judgment, while reconciliation introduced additional reference-edge loss for Full on Sachs. These findings support structural pre-organization as an explicit design and evaluation target for omission control and motivate evaluating context construction jointly with its utilization in judgment.

[AI-70] Beyond Depth Truncation: Controlled Evaluation of Depth Utilization in Recursive Language Models

链接: https://arxiv.org/abs/2609.19934
作者: Ha Van Dau,Thanh Tung Khuat,Nguyen Thanh Dung
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Depth-recurrent language models iteratively apply a small layer stack, decoupling per-token compute from distinct parameter count. To determine whether such a model genuinely utilizes its depth, both recurrence and layer-pruning literatures rely on a shared evaluation: truncating depth at inference time, plotting quality against retained depth fraction, and reading off the slope. While cheap and training-free, this metric suffers from an unexamined flaw: it extracts a single scalar from an intervention that alters multiple model properties simultaneously. Depth truncation concurrently reduces the number of block applications, decreases the volume of distinct computation performed, and pushes the readout head onto an out-of-distribution residual stream. The observed slope conflates all three factors, yet is conventionally interpreted as reflecting solely the second. We propose the Depth Control Protocol (DCP), a diagnostic suite that disentangles these three quantities. DCP comprises three positive controls that isolate each factor while varying the others, a negative control applying the identical interventions to dense transformers to ensure the effect is not an artifact of the measurement protocol, and a controlled training intervention to verify causality. The linchpin control, running the full budget of block applications while executing only a single distinct iteration, is strictly realizable only in depth-wise weight-sharing architectures, since in a dense network repeating a layer yields an entirely different model rather than the same model in an alternative configuration. Subjects: Artificial Intelligence (cs.AI) MSC classes: 68T50, 68T01, 68T07 Cite as: arXiv:2609.19934 [cs.AI] (or arXiv:2609.19934v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2609.19934 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-71] From “Who Is This User?” to “What Does This Purchase Mean?”: A Deployed Pipeline for Semantic User Profiling at Bank Scale ICDM

链接: https://arxiv.org/abs/2609.19928
作者: Ryota Mitsuhashi,Tetsuro Morimura,Hirotake Ito
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 10 pages, 3 figures, IEEE International Conference on Data Mining 2026 (ICDM)

点击查看摘要

Abstract:Per-user LLM inference on transaction histories binds the inference budget linearly to user count, which becomes prohibitive at applied scale. We re-cast attribute inference from per-user to per-transaction-pattern. The pipeline runs in three phases: Resolve abstracts item names with optional web grounding, Profile infers attributes for each frequent pattern, and Tag clusters free-text attributes into a queryable database. In Profile, a single LLM call per pattern emits predefined categorical labels, free-text attributes, and per-attribute prevalence estimates. Because inference runs over patterns rather than users, the budget grows with the pattern count rather than the user count. On the public Open e-commerce corpus, the database is statistically indistinguishable from an LLM that reads each user’s raw history directly in AUC across the evaluated attributes, and the prevalence estimates carry discriminative signal between positive and negative users. The pipeline is deployed at a major Japanese bank profiling on the order of tens of millions of users, with close to a three-order-of-magnitude reduction in LLM inference targets versus a per-user pipeline. The code is publicly available on this https URL.

[AI-72] Learning and Transferring Closed-Loop Robot Software

链接: https://arxiv.org/abs/2609.19906
作者: So Kuroki,Yujin Tang
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Closed-loop robot policies require observation processing, state management, and situation-dependent branching, making them costly to design and tune manually. Although coding agents increasingly support control-code generation and optimization, it remains unclear whether implementations improved on source tasks also support policy acquisition for new tasks. We study this question by treating complete closed-loop implementations as reusable execution experience. For each source task, a coding agent generates policy code from a few successful demonstrations and iteratively improves it using simulation feedback. The validation-selected implementations are retained in a software archive. For new tasks, the agent generates and improves policies using archived implementations, target demonstrations, and execution feedback. The resulting policy is then frozen and executes without further model calls. Across four source tasks in RoboCasa, iterative optimization increases mean success from 28.3% to 64.2%. Across nine target tasks and three independent runs, mean success is 45.2% without references, 41.5% with initial source code, and 57.0% with optimized source code. Optimized references outperform initial references in all three runs on the nine-task average, with a mean gain of 15.6 percentage points. These results demonstrate the value of execution-improved software as a resource for acquiring new policies in this setting, although initial references remain better on two target tasks when averaged across runs.

[AI-73] RACE: Accountable Agent ic Retrieval for Source Discovery in Digital Archives

链接: https://arxiv.org/abs/2609.19897
作者: Donghan Bian(ENC, LRE),Marie Puren(LRE, ENC),Florian Cafiero(LRE, ENC)
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Historical archives pose a difficult retrieval problem for retrievalaugmented generation systems: documents are OCR-degraded, heterogeneous across genres and sources, and require strong source traceability for scholarly and institutional use. We introduce TRACE, a training-free agentic retrieval framework designed for accountable source discovery over historical corpora. The system was developed in the context of DECIDON, an interdisciplinary project on the circulation of political discourse between parliamentary debates and the press during the French Third Republic, involving digitised historical collections and institutional use cases. The prototype is currently deployed internally within the project and accessible to 24 researchers across six partner institutions. We evaluate TRACE on HistoriQA-ThirdRepublic, a benchmark of 1,752 French historical questions over parliamentary debates and newspapers from 1887, with documents derived from Bibliothèque nationale de France digitised collections. TRACE achieves R@10 = 0.856 and MRR = 0.653, outperforming sparse, dense, graph-based, and agentic RAG baselines, with the largest gains on multi-hop and cross-corpus questions. At approximately 0.02 per question under the default hosted inference configuration, TRACE also remains economically feasible for heritage institutions, laboratories or companies that cannot rely on costly local GPU infrastructure. These results suggest that, for large digital libraries and archives, retrieval accountability and corpus-aware agent design can provide a practical alternative to heavier training-based or graph-construction approaches.

[AI-74] ClashBench: Conflicts Leading Agents to Seize and Harm

链接: https://arxiv.org/abs/2609.19892
作者: Yuejin Xie,Yu Li,Dadi Guo,Qingyu Liu,Yuqian Fu,Yanwei Fu,Yujiu Yang,Xia Hu,Dongrui Liu
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:As agent systems become more widely used, multiple agent sessions increasingly run alongside pre-existing user tasks in the same environment, sharing resources with limited capacity or mutually exclusive states. This creates a safety risk: when granted sufficient privileges, an agent may resolve a resource conflict by terminating or otherwise disrupting an existing task rather than reporting it. In this work, we identify and formalize this failure mode, which we term destructive resource preemption: obtaining the resources required for a requested task by terminating, overwriting, evicting, or degrading an incumbent task. To systematically study this risk, we introduce ClashBench, an executable benchmark comprising 268 validated conflict cases across 55 resource types, and evaluate 17 models through Codex, Claude Code, and OpenCode. We observe destructive preemption in 44.5% of trajectories, where the agent completes the requested task while causing the incumbent task to fail its health check. We also show that prompt-based safeguards are insufficient: an instruction to avoid affecting existing tasks reduces but does not eliminate preemption, while an instruction explicitly authorizing the agent to stop local processes increases it. More concerningly, in 31.9% of successful destructive-preemption cases, the final response mentions neither the resource conflict nor the action taken to resolve it, raising concerns about possible concealment. These findings establish destructive resource preemption as a broad safety risk in privileged agent systems and motivate stronger privilege controls, task isolation, and conflict-aware safeguards.

[AI-75] Physical knowledge on historical data matters more than enforcing physical constraints on the forecast

链接: https://arxiv.org/abs/2609.19871
作者: Etienne Lehembre(CA, LIFO),Pascal Audigane(BRGM),Vincent Nguyen(LIFO),Christel Vrain,Thi-Bich-Hanh Dao(LIFO, CA)
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Time series forecasting has seen signicant advancements with the emergence of new deep learning models. However, forecasting time series in applications involving physical processes remains a major challenge. Despite the apparition of Physics Informed Neural Networks (PINN), recent models do not estimate unobservable intermediate physical variables, which are important for domain experts to understand the target behavior. To this end, we propose a Physics Informed Recurrent Neural Network (PIRNN) which predicts, along the target, unobservable variables on both historic data and forecast target. This approach enhances the model robustness and results interpretation using domain knowledge. Our method is easily adaptable to any physical model using several equations, each having its own set of unobservable variables, to describe it-self. As a case study, we incorporate physical equations used for groundwater levels predictions by the physical model called Gardenia. This model uses transfers equations between reservoirs, optimized with data assimilation, to simulate the evolution of groundwater levels. Evaluation includes several well known neural network models and the Gardenia model compared on twelve real world datasets. In addition, we study the impact of each component through an ablation study. Our model outperforms other models on ve out of the twelve datasets and our ablation study underlines the importance of having a physical background in our time series forecasting task. Finally, the coherence of the physical variables predicted by our neural network is assessed by a domain expert.

[AI-76] A Functional Pilot for Certified Freshness-Aware Semantic–Spatial Range Retrieval

链接: https://arxiv.org/abs/2609.19855
作者: Taimoor Ahmad
类目: Databases (cs.DB); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Geographic applications need every object inside a radius that satisfies a semantic threshold, yet embedding indexes return approximate top-ranked lists and may omit qualifying records silently. We present FRESH-GEORANGE, a semantic- spatial range design that separates source-watermark freshness from optional record age. Geographic cells and semantic mi- croblocks provide admissible pruning bounds; a graph proposes verification order but supplies no correctness evidence. Exact mode scans every nonprunable block and the delta overlay. Certified mode may stop early and reports a deterministic query- specific recall lower bound from verified answers and unresolved records. A reproducible CPU pilot uses 2,500 real OpenFlights airport records, a 2,000-record base, and 740 simulated insert, delete, and text-revision events; it evaluates 180 unique queries over five seeds. Exact mode achieved 100.00% set recall on every query. The 95-percent mode achieved 99.91% empirical mean recall with a 99.41% reported mean certificate and no observed bound violation. However, its 7.24 ms median latency was 5.85 times the 1.24 ms spatial-first exact baseline, and full-history delta replay became slower than rebuilding at larger batches. The prototype therefore validates the completeness mechanism, not performance superiority or production freshness. Submission- scale evaluation requires real map diffs, official recent baselines, and truly incremental versioned maintenance.

[AI-77] Constraint-Safe Graph-Context Scoring for Stable Point-Feature Labels Under Text-Width and Accessibility-Inspired Profiles

链接: https://arxiv.org/abs/2609.19848
作者: Taimoor Ahmad
类目: Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Point-feature label placement on interactive maps must reconcile geometric validity, display yield, local placement utility, and stability across camera motion. Accessibility and multilingual requirements further change label dimensions, yet algorithmic evaluations often collapse these concerns into overlap counts. We present LABELSENSE-Pilot, a reproducible prototype that generates eight compass candidates per feature, scores candidates with a multilayer perceptron over graph-context summaries, adds a previous-placement bonus, and selects a layout through mixed-integer optimization. The executed scorer is deliberately not described as a graph transformer. Every returned layout is checked for viewport containment, per-feature uniqueness, and pairwise clearance. Experiments use 2,500 airport coordinates and names spanning 155 countries, with country-grouped splits and generated density, camera, text-suffix, preference, and enlarged-font stressors. Across five seeds, LABELSENSE-Pilot displayed 85.62 percent of labels with 2.09 percent flicker and zero collisions. Versus a handcrafted-utility integer program, LABELSENSE-Pilot sacrificed 1.43 percentage points of display while reducing flicker by 12.04 points. Enlarged-box-aware layouts produced zero proxy violations, whereas standard geometry reevaluated at 1.5x violated 52.57 percent of selected placements. These results establish an auditable engineering trade-off, not human accessibility, multilingual usability, or preference. Official recent baselines and participant evidence remain required before submission.

[AI-78] Improving Cross-embodiment Transfer in Latent Action Models with Action-Similarity Supervision

链接: https://arxiv.org/abs/2609.19846
作者: Maxime Alvarez,Renzo Caballero,Tatsuya Matsushima,Yusuke Iwasawa,Yutaka Matsuo
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:As generalist robot policies gain vision and language from web-scale pretraining, demonstrations remain costly to collect and tied to the robot that recorded them. Latent action models (LAMs) address both by learning latent actions from action-free videos that can be shared across embodiments, however, in practice, LAMs are sensitive to background visual noise, and the same motion from two different robots may be encoded with different latents. One solution to the background visual noise is to add an auxiliary loss predicting the robot action from the latent action, further associating the latent action space to the embodiment specific robot action space. We study a different use of the same labels, through action-similarity supervision. The similarity between any two latent actions is trained to match the similarity of the two ground-truth robot action sequences. The ground-truth actions are never predicted by the LAM, so the latent action does not need to encode embodiment specifics. We evaluate cross-embodiment transfer on RoboTwin 2.0 in a controlled setup, two bimanual robots demonstrate disjoint task sets, a policy is trained on all the demonstrations, and each robot is evaluated closed-loop on the tasks only the other demonstrated. With the policy architecture and its hyperparameters, the dataset, and the evaluation protocol fixed, predicting latent actions instead of ground-truth actions more than doubles cross-embodiment success. Given the same ground-truth actions, similarity supervision transfers better than an auxiliary loss that predicts the ground-truth action during the LAM training. Computing the similarities on end-effector motion rather than joint-space motion, and letting the loss compare latent actions across the two robots, gives the best approach of the study.

[AI-79] A Dual-Process Perspective on Nudge Susceptibility in LLM -Based GUI Agents

链接: https://arxiv.org/abs/2609.19843
作者: Haya Halimeh,Sascha Kaltenpoth,Kevin Bösch,Oliver Müller
类目: Artificial Intelligence (cs.AI); General Economics (econ.GN)
备注: Preprint of a manuscript completed in November, 2025

点击查看摘要

Abstract:LLM-based GUI agents increasingly act on behalf of users in digital environments that were designed with human users in mind. These graphical user interfaces were designed to support, but also deliberately steer, the behaviour and decisions of users. While behavioural biases in the textual outputs of LLMs are well-documented, far less is known about how such influence operates when models act as agents that perceive interfaces and execute decisions—and, in particular, whether the reasoning capabilities increasingly built into these agents make them more robust to it. Drawing on Dual-Process Theory, we empirically investigate whether LLM-based GUI agents are susceptible to automatic (Type 1) and reflective (Type 2) digital nudges, and how their reasoning configuration moderates this susceptibility. In a randomized online shopping experiment with 3,600 agents and a total of 21,600 simulations across six frontier models from three providers, we found that agents were vulnerable to both nudge types. Crucially, the reasoning configuration moderated these effects in opposing directions, reducing susceptibility to automatic default nudges while heightening it to reflective social influence nudges. Extensive reasoning therefore did not make agents more robust but redirected the route through which choice architecture takes effect. Exploratory analysis further showed this redirection to be systematically structured by model scale. Beyond establishing nudge susceptibility as a behavioural property of agentic AI, the study positions interface design as a governance concern for organizations that delegate decisions to autonomous agents.

[AI-80] MetaRTL: Meta-path Attention Enhanced Relational Table Learning

链接: https://arxiv.org/abs/2609.19832
作者: Ken Zhong,Weichen Li,Zheng Wang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Relational table learning has gained increasing attention with the widespread use of relational databases. Existing methods typically rely on deep GNN or HGNN stacks, leading to high computational costs and limited performance on large real-world databases. We propose MetaRTL, a two-stage framework for scalable and expressive relational table learning. In the first stage, MetaRTL obtains initial table embeddings via lightweight pre-training. In the second stage, it performs non-parametric message passing to derive meta-path features, which are then aggregated by an attention module, MetaAttn. By shifting computation from deep message passing to efficient meta-path aggregation, MetaRTL captures rich relational semantics while maintaining high efficiency. Experiments on 10 real-world datasets across 24 tasks demonstrate the effectiveness of the proposed method.

[AI-81] Dual-Axis Policy Optimization for LLM Agents : Bayesian Feedback Attribution and Trajectory Mass Normalization

链接: https://arxiv.org/abs/2609.19830
作者: Yingxuan Zhuang,Binhe Yu,Jingxiao Yang,Ruopei Sun,Ziting Li,Cheng Tan,Xuhong Zhang,Jianwei Yin,Jintao Chen
类目: Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
备注:

点击查看摘要

Abstract:Reinforcement learning for LLM agents involves two distinct optimization di- mensions: how environment feedback is exploited within a trajectory, and how complete trajectories are aggregated across a batch. We formulate these dimen- sions as Intra-Trajectory Feedback Attribution and Inter-Trajectory Objec- tive Aggregation, and introduce BATON (Bayesian Attribution and Trajectory Objective Normalization), a dual-axis policy optimization framework. BATON instantiates the first axis with Bayesian Feedback Attribution, which constructs a feedback-conditioned posterior over sampled actions, and the second with Trajec- tory Mass Normalization (TMN), which assigns equal optimization mass to com- plete trajectories. Experiments with GRPO and GiGPO on ALFWorld, WebShop, and SearchQA show that both axes provide independent gains and that their combi- nation consistently achieves the strongest overall performance across model scales.

[AI-82] Steering Equilibrium Selection in Regularized Self-Play via the Reference Policy

链接: https://arxiv.org/abs/2609.19820
作者: Luis Leal
类目: Artificial Intelligence (cs.AI); Computer Science and Game Theory (cs.GT); Machine Learning (cs.LG)
备注: 17 pages, 8 figures, 4 tables. Companion to arXiv:2606.28308 and arXiv:2607.17543 . Fully reproducible: a single self-contained notebook regenerates every number, table, and figure

点击查看摘要

Abstract:Regularized self-play – the family behind DeepNash’s Stratego play – drives a two-player zero-sum policy to a Nash equilibrium by best-responding to a slowly moving, entropy-regularized reference policy \rho . When the game has a polytope of value-equivalent equilibria, the regularizer silently breaks the tie: with a uniform reference it selects the maximum-entropy member, the I-projection of \rho onto the Nash set. Can the reference be used to choose the equilibrium on purpose? On five exactly solvable games plus a 2-D polytope, with exact best responses and equivalence tests over independent seeds, anchoring the reference at a target member and refining steers self-play to that member with mean coordinate error 0.007 at median exploitability 5\times10^-5 , TOST-equivalent to the request within \pm0.05 ; the anchoring persists through refinement and follows the reference, not the initialization. Selection follows the reach-weighted I-projection (slope 0.969 [0.950, 0.987]). We report with equal emphasis where the story breaks: fixed off-manifold references cost 0.08-0.25 exploitability; stiff or flat families require a smaller mirror step, set by a pre-registered rule; boundary targets undershoot; curvature predicts where boundary saturation bites (rank correlation 0.90, p=0.037) while interior precision is curvature-independent. Table and MLP steering maps are equivalent within \pm0.03 at every target (30 seeds); matched control arms show attention’s robust signature is excess seed variance, any systematic shift bounded at 0.018 and not significant. Against a best response the selection-robustness trade-off is degenerate: steering matters only against fixed, non-equilibrium opponents. The recipe – anchor the reference at the desired member and refine – reinterprets the KL anchor of RLHF-style RL as a selection knob, not only a stability leash.

[AI-83] CoRELoop: Parameter-Efficient Controlled Recurrent Refinement for Audio Deepfake Detection

链接: https://arxiv.org/abs/2609.19818
作者: Kunyu Feng,Yuxiang Wang,Li Wang,Wan Lin,Zhizheng Wu
类目: ound (cs.SD); Artificial Intelligence (cs.AI)
备注: 5 pages, 2 figures, 3 tables

点击查看摘要

Abstract:Generalizing to unseen attacks remains challenging for audio deepfake detectors, and collecting training data covering all potential attacks is impractical. We explore recurrent refinement in an already-trained SSL-based detector without additional data or changes to its original parameters. However, directly recycling encoder outputs as inputs degrades detection in our diagnostic. We propose CoReLoop, which makes this reuse effective by adapting recurrent inputs to the frozen encoder, controlling state updates, and aligning refined outputs with the frozen classifier. By training only lightweight refinement modules and loop-specific low-rank adapters on the original data, CoReLoop enables additional refinement while preserving the detector’s original first-pass prediction. On 14 cross-domain test sets, the 24-layer model reduces pooled equal error rate (EER) from 4.85% to 3.74% with two passes, with approximately 10M trainable parameters out of 598M. To selectively apply this refinement, an optional halting head chooses the depth for each utterance, achieving 3.73% pooled EER with an average of 1.18 passes.

[AI-84] Contagion on the Trading Floor: How Adversarial Signals Spread in Multi-Agent Trading Systems KDD2026 ECML

链接: https://arxiv.org/abs/2609.19789
作者: Qi Rong Sua,Junhao Dong,Nguyen Duc Thai,Yuqing Wen,Cheston Tan,Yew-Soon Ong
类目: Artificial Intelligence (cs.AI)
备注: Published in ECML PKDD 2026. 25 pages, including 7 pages of supplementary material

点击查看摘要

Abstract:Multi-agent trading systems built on large language models (LLMs) are beginning to appear in quantitative finance, yet their robustness to adversarial inputs is largely unknown. We study the vulnerability of LLM trading stacks to black-box, input-only attacks that enter solely via admissible social-media feeds. We introduce the Generic Multi-Agent Trading System (GMATS), a framework that captures modern multiagent trading architectures and instantiate a class of black-box poisoning attackers that treat an LLM as a post generator and inject budget-constrained, plausibly benign social-media content into the analyst’s evidence stream. We define contagion metrics that trace how adversarial content propagates through the stack, including belief-shift scores at analyst and coordinator layers and attack-clean deltas on standard backtest metrics. Experiments on a safe offline benchmark with historical market and social data show that even simple input-only attackers can materially degrade risk-return profiles, sharply reducing Sharpe ratios. At the same time, we find that suitably designed multi-agent topologies and coordinator prompts can dampen adversarial shocks and improve average robustness under identical poisoning budgets.

[AI-85] Integrating knowledge from case reports: a medical ontology based multimodal information system with structured summary

链接: https://arxiv.org/abs/2609.19775
作者: Shuyu Guo,Lan Huang,Yichen Liu,Hanbin Ma,Tian Bai
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Published medical case reports serve as a crucial medical information carrier, documenting discoveries in rare diseases, diagnostic methods, and innovative treatments. Despite the wealth of clinical knowledge in millions of case reports in the public medicine literature database (PubMed), accessing relevant information efficiently is hindered by the limitations of traditional keyword-based retrieval tools on unstructured and diverse case reports. To address the above issues, we introduce a comprehensive multimodal information system for case reports integrating structured clinical summaries of patients including medical images and biomedical named entities from 52949 open-access case reports published from 2000 to 2021. The multimodal essential information is organized in a well-structured medical ontology. Also, a powerful interface for searching and browsing case reports is designed to assist junior clinicians in retrieving cases effectively and improving the identification and diagnosis of rare diseases.

[AI-86] orchCraft: Unified binder design by inverting an all-atom structure predictor

链接: https://arxiv.org/abs/2609.19770
作者: TorchCraft Team:Yu Liu,Zhouhanyu Shen,Zhengyi Li,Xikun Huang,Jiaqi Liu,Shuxian Gao,Qilin Yu,Xiayan Qin,Yucheng Zhang,Mingchen Chen
类目: Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:All-atom structure predictors model diverse molecular interactions, but using their learned structural priors for binder design remains challenging. Here we present TorchCraft, a unified binder-design framework that optimizes sequence logits through a frozen all-atom predictor. Implemented in TorchFold, TorchCraft combines confidence, contact, geometric, and sequence-prior objectives within a shared optimization procedure for minibinders, framework-conditioned VHHs, cyclic peptides, and ligand-binding proteins. Using pretrained AlphaFold 3 weights, TorchCraft generated representative minibinders and VHHs with experimentally measured binding across four targets in each format, without post hoc sequence redesign. Computational benchmarks further demonstrated the framework’s applicability to cyclic peptides and ligand-conditioned pocket design. TorchCraft extends predictor inversion to multiple binder formats and molecular contexts, providing a common framework for reusing all-atom structural priors in design.

[AI-87] Rethinking Multi-Agent Collaboration: When More Is Less

链接: https://arxiv.org/abs/2609.19759
作者: Yishuo Yuan,Yibo Wu,Yihan Zhang,Minyuan Sun,Shenliang Li,Xinkai Ma,Yifan Li,Jiaheng Liu
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:The rapid advancement of large language models and single-agent harnesses has reshaped the landscape of autonomous systems, raising a critical question of when multi-agent collaboration offers genuine value. As individual agent capabilities continue to scale, multi-agent collaboration faces diminishing returns while incurring growing context overhead. Through systematic analysis, we delineate the capability boundaries of multi-agent collaboration relative to single-agent alternatives, showing that it confers systematic benefits specifically in long-horizon tasks with sparse dependencies, while single-agent harnesses remain superior in tightly coupled, sequential workflows. Building on these insights, we propose SAIGE, a lightweight multi-agent collaboration mechanism based on Semantic-Aware Incremental Graph Evolution. SAIGE models collaboration as a dynamically evolving graph, where nodes are agent instances spawned on demand and edges encode semantic dependencies established through content-based information retrieval. Experiments on long-horizon, complex task benchmarks show that SAIGE achieves a favorable trade-off between context efficiency and task performance, and that scaling the agent pool or deepening the recursion level does not consistently improve outcomes. Our findings suggest that multi-agent superiority is bounded by task structure rather than universal, and that more agents do not necessarily make a system more intelligent.

[AI-88] When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models

链接: https://arxiv.org/abs/2609.19671
作者: Jaejun Shim,HyunJin Kim,Young Jin Kim,JinYeong Bak
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Large Reasoning Models (LRMs) achieve strong performance on complex tasks but exhibit systematic inefficiency: they often overthink easy problems and underthink hard ones. Existing approaches based on uniform length penalties or rigid routing incur an efficiency tax, trading reduced computation on easy instances for accuracy loss on hard instances. We formulate efficient reasoning as an instance-adaptive computation allocation problem and propose When2Think, a post-training framework for hybrid reasoning that dynamically allocates computation based on problem difficulty. Our method introduces Instance-level Difficulty-Aware Control (IDAC), a reward-shaping mechanism that leverages pre-computed reference statistics (accuracy and token usage) to regulate reasoning depth. Combined with verifier-based rewards and batch-wise standardized advantages, IDAC enables stable critic-free optimization without learned reward models or online reference-model queries. When2Think encourages direct answering on easy instances while preserving extended reasoning on hard instances, thereby learning when to use System 1 (NoThink) versus System 2 (Think). Experiments on mathematical benchmarks demonstrate improved accuracy-efficiency trade-offs: on AIME24, Pass@3 increases by 10.0% while token usage is reduced by 27.9% relative to the base model, and on AIME25, When2Think achieves 40.0% Pass@3, outperforming compression and routing-only baselines.

[AI-89] ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI

链接: https://arxiv.org/abs/2609.19644
作者: Jaehyun Nam,Jinsung Yoon,Yanzhou Pan,Yubo Wang,Rui Meng,Parthasarathy Ranganathan,Tomas Pfister
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Scientific discovery is defined by the ability to identify the boundaries of existing knowledge and venture into unexplored territory. The ultimate vision for AI in science is problem-driven autonomous research: given a fundamental challenge by a human expert, the AI independently navigates the scientific landscape, uncovers theoretical and empirical bottlenecks, and systematically expands the frontier of knowledge. In this paper, we introduce ScientistTwo, a fully autonomous multi-agent framework designed to realize this vision. Specifically, ScientistTwo takes an initial problem as input, establishes state-of-the-art baselines, formulates novel hypotheses, and coordinates specialized agents to orchestrate an end-to-end discovery cycle without human intervention. Moreover, the framework rigorously conducts experiments using diverse datasets and metrics, refines methodologies through automated ablation studies, and validates research findings via a closed-loop simulated peer-review rebuttal engine. To evaluate ScientistTwo’s capabilities against the highest standards of human scientific achievement, we benchmark it across papers accepted at top-tier conferences such as ICLR, ICML, and NeurIPS. As a result, ScientistTwo autonomously generates expert-level, publishable papers and fully verified, executable codebases. Its solutions consistently outperform human state-of-the-art models, and achieve higher average review ratings than human-authored papers under automated AI review agents. These results show that ScientistTwo is not merely an assistive tool but an autonomous scientific pioneer capable of pushing the frontiers of human discovery. Project website: this https URL

[AI-90] Reach or Solve? Attributing Agent ic RL Gains with Checkpoint Handoffs

链接: https://arxiv.org/abs/2609.19636
作者: Xuan Liu,Jingbin Qian
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Reinforcement learning now trains language-model agents that act over dozens of steps in live environments. The gains are large, and they are read as better decision-making. An agent in a closed loop writes its own inputs. Each observation follows from its own earlier actions, so the states it meets late in an episode are partly of its own making. An SFT checkpoint and an RL checkpoint are then scored from different states, even on identical tasks. Endpoint success mixes two changes: where the agent arrives, and what it does once it is there. Restricting the comparison to states both policies reach does not separate them. That restriction selects on an outcome, and in our data it flips the sign of the effect. We introduce checkpoint handoff, an evaluation protocol that clones a state one released checkpoint reached and hands it to another, with no retraining. Crossing a reacher role and a solver role over SFT and RL splits an endpoint gain into REACH and SOLVE. REACH is how often a policy arrives at a state the environment confirms is a fixed number of actions from success. SOLVE is how often it finishes from an identical cloned state. Across two benchmarks and two independently released pipelines, the reacher by solver interaction is positive in all five conditions. An RL history is worth more to an RL solver than the same history is to an SFT solver. On ALFWorld, RL improves both terms, and the SFT solver never succeeds where the RL solver fails. Independent REACH and SOLVE gaps predict the aggregate interaction. Handoff asks only that one checkpoint’s history can be replayed under another, so long-horizon evaluation can report arrival and completion beside endpoint success.

[AI-91] acSushi: Tactile-Grounded World-Action Modeling for Dexterous Sushi Manipulation

链接: https://arxiv.org/abs/2609.19613
作者: Haodi Hu,Kaen Kogashi,Toshiaki Koike-Akino
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 8 pages, 5 figures

点击查看摘要

Abstract:Dexterous food manipulation requires control under deformation, occlusion, and uncertain contact. We present TacSushi, a tactile-grounded, Cosmos3-based world-action policy that learns from recorded future consequences while acting on current observations. The backbone encodes current RGB, language, and hand state, and feature-wise gated fusion incorporates fingertip tactile features into the action representation. During training, a decoder conditioned on demonstrated action chunks predicts logged future visual observations, task progress, relative contact risk, and tactile summaries; this decoder is removed at deployment. Failed trials provide consequence supervision, but their actions are excluded from imitation. We train TacSushi on 340 successful and 50 failed real-robot trials and compare six methods in 600 separate rollouts across three in-distribution tasks and two out-of-distribution ingredient variants. To assess food quality beyond a single geometric threshold, we score terminal outcomes using an anchored visual-quality protocol that equally weights five human ratings and three vision-language-model ratings per rollout. Full TacSushi achieves 68.3% average in-distribution success and 37.5% out-of-distribution success, compared with 36.7%/10.0% without future-consequence supervision and 25.0%/17.5% with direct tactile concatenation in place of gated fusion. These comparisons support complementary benefits of feature-wise gated tactile fusion and training-only predictive supervision.

[AI-92] SIMLIFE: Pattern Understanding for Long-Horizon Human-Agent Partnership

链接: https://arxiv.org/abs/2609.19610
作者: Run Peng,Zinnia Nie,Jing Ding,Yinpei Dai,Yichi Zhang,Zengqing Wu,Yao Fu,Ziqiao Ma,Jiayuan Mao,Joyce Chai
类目: Artificial Intelligence (cs.AI)
备注: COLM 2026 Learning from Situated and Embodied Interaction Workshop

点击查看摘要

Abstract:Understanding humans over long horizons requires agents to infer not only what people need in the moment, but also how routines form, why they repeat, and when they change. We introduce SimLife, a scalable platform for simulating long-term household life with rich visual observations, ground-truth action logs, and synthetic dialogues with audio. Built on SimLife, SimLife-BP evaluates long-context pattern understanding: the ability to infer latent behavioral rules from weeks or months of everyday observations. The benchmark contains 106 episodes averaging 15.49 hours and 38.57 in-game days, and 1,439 question-answer pairs. Each task probes direct, counterfactual, noisy, and inverse reasoning under different levels of rule hints. Evaluating frontier models and architectures, we find that current models often achieve surface-level prediction without comprehensive rule understanding, rely on frequency-based heuristics rather than if-then reasoning over evidence, and struggle to adapt when behavioral patterns change. These findings suggest that long-context pattern understanding remains a major bottleneck for future embodied agents, while SimLife opens a broader space for studying memory, personalization, adaptation, and long-horizon planning in everyday human-AI interaction.

[AI-93] DeltaSelect: Affordable A/B Testing for Coding Agents

链接: https://arxiv.org/abs/2609.19607
作者: Nicholas J. Conn
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Coding-agent benchmarks are built for broad and comprehensive comparisons, not frequent development decisions. Individual runs vary, full suites are expensive, and the benchmark harness may differ from the harness used in practice. In a resampling analysis of DeepSWE’s published trials, only 19.5% of tasks (22 of 113) had a fifth-percentile Pearson correlation of at least 0.50 with full-benchmark performance. The paper presents DeltaSelect, an open-source method that identifies tasks whose one-run results consistently track full-benchmark performance using Pearson correlation, maps fractional verifier results to a common score using linear regression, and selects a fixed task set within a dollar budget. DeltaSelect is intended for repeated baseline-versus-candidate comparisons during development, not model rankings. In a gpt-5.6-luna low-reasoning case study, DeltaSelect was used to revise custom skills and instructions. Across 13 evaluations, the recorded cost was USD 27.86 at rates published August 16, 2026. The adopted version cost 58.1% less than the initial version (USD 1.75 versus USD 4.18; p=0.008), while the calibrated score was higher (42.36% versus 36.46%; published-analog variance p=0.326).

[AI-94] Continual Enterprise World Model Discovery in Dynamic Systems

链接: https://arxiv.org/abs/2609.19551
作者: Shambhavi Mishra,David Vazquez,Perouz Taslakian,Marco Pedersoli,Jose Dolz,Issam H. Laradji
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:In an enterprise system, updating one field can set another, create a record, or start an approval. These effects are produced by business rules that are not built into the platform but written by each organization and revised over time. An agent working in such a system cannot predict the result of its own actions without knowing these rules. We study continual enterprise world model discovery, where an agent starts without knowledge of these business rules and discovers them by interacting with records and observing the outcomes. From those observations it builds a world model, which it revises as the rules change. To evaluate this, we introduce EnterpriseWorldShift, built on a live ServiceNow environment with nine tables, 25 hidden rules and 600 evaluation actions. It presents four versions of the same enterprise world, with the tables and records held fixed while a rule is modified, then added, then removed, so that discovery, revision, extension and retirement are each tested in turn. Our Continual Discovery Agent (CDA) builds such a model and carries it from one world to the next. It predicts the effects of the hidden rules more accurately than looking them up for each question, the approach taken by prior work, by up to 8.98 IoU points, and it answers from its own model without querying the running system.

[AI-95] Compressed Active Subspaces for Scalable Bayesian Inference

链接: https://arxiv.org/abs/2609.19539
作者: Thomas Flynn,Sanket Jantre,Byung-Jun Yoon,Kibaek Kim
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Methodology (stat.ME); Machine Learning (stat.ML)
备注:

点击查看摘要

Abstract:Active subspace methods provide a framework for quantifying predictive uncertainty in high-dimensional models by identifying and performing inference along parameter directions that have the greatest influence on the model output. However, the construction of active subspaces requires storing many full-dimensional model gradients, which becomes prohibitive as model size increases. We address this limitation by proposing Compressed Active Subspaces (CAS), a scalable approach that first maps the model parameters to a compressed space using a structured isometric embedding and then constructs the active subspace within this reduced parameterization. Our approach substantially reduces the memory required for active subspace construction and enables Bayesian inference for large models where standard active subspace methods become impractical. We demonstrate the scalability of CAS on neural networks of increasing size while maintaining predictive performance and robust uncertainty estimates.

[AI-96] Detecting Soft Errors in Parallel Software with LLM -tuned Instruction Duplication

链接: https://arxiv.org/abs/2609.19531
作者: Yafan Huang,Guanpeng Li
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI); Programming Languages (cs.PL)
备注:

点击查看摘要

Abstract:We propose PaRID (PaRallel Instruction Duplication), a software-directed soft error detection framework that requires only compile-time effort for multithreading parallel programs. PaRID addresses two key challenges: supporting parallel programs with mixed serial and parallel regions and minimizing performance overhead without relying on costly dynamic profiling. It combines parallel-aware code transformation with LLM-tuned performance modeling, guided by eight generalizable findings from an offline characterization study, to enable fast soft error detection in parallel applications. Evaluation on NPB benchmarks shows that PaRID reduces protection overhead from 162.79% to 59.84% on average and achieves up to 5x speedup while maintaining full error detection effectiveness.

[AI-97] AURORA: A Natural Language-Driven Agent ic Framework for Understanding Reasoning and Orchestrating Reliable Air-Ground Co-Simulation

链接: https://arxiv.org/abs/2609.19527
作者: Keshu Wu,Hao Zhang,Rui Gan,Xiangbo Gao,Xiaopeng Li,Zhengzhong Tu,Yang Zhou
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Air-ground transportation research increasingly relies on co-simulation, yet constructing scenarios remains labor-intensive and difficult to validate. More importantly, a generated scenario may execute successfully while failing to realize the spatial, temporal, communication, or behavioral relationships requested by the user. This paper presents AURORA, a natural-language-driven agentic framework that treats air-ground scenario generation as a process of compilation with verification. Central to AURORA is the Air-Ground Scenario Graph (AGSG), a typed intermediate representation that explicitly connects agents, aerial missions, events, communication links, success conditions, and their cross-domain dependencies. This shared representation enables simulator-grounded parsing, joint road-airspace grounding, temporal planning, pre-execution feasibility checking, trace-based runtime verification, failure localization, and bounded repair within a unified workflow. We further introduce AURORA-Bench to evaluate not only whether generated scenarios execute, but whether they faithfully realize the requested interactions. Experiments across multiple language models show that structured execution substantially improves reliability, while runtime verification exposes silent failures that completion-based evaluation overlooks. Localized repair further resolves many violations without regenerating the entire scenario. The results show that reliable scenario generation requires verifying realized behavior, not merely executable code, and demonstrate the value of explicit intermediate representations for verifiable and repairable language-driven co-simulation.

[AI-98] Self Improvement via Fast Tree-search

链接: https://arxiv.org/abs/2609.19526
作者: Xinghong Fu,Aravinth Kulanthaivelu,Yutaro Yamada
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Coding agents can recursively modify their own implementations, forming a loop of self-improvement. While prior work shows this can boost performance on coding benchmarks, existing approaches are costly and compute-intensive. We introduce a simple, sample-efficient self-improvement framework that significantly improves coding performance under strict budget constraints. We identify evaluation of candidate self-modifications as the main runtime bottleneck since prior approaches estimate their effectiveness by re-running a subset of benchmark tasks with the modified agent, which is time-consuming. We introduce Recursive Self Improvement via Fast Tree-search (SIFT), which augments these downstream task evaluations with an LLM-as-a-judge signal that performs pairwise comparisons between candidate patches, where the win-loss record is aggregated with a regularized Bradley-Terry model, and the resulting strength scores drive rank-based parent sampling inside a lightweight disaggregated tree search. Expensive downstream task evaluations are reserved only for the most promising nodes. Using a fully disaggregated tree search pipeline, the judge scores provide intermediate signal to guide exploration on promising candidate patches without being bottlenecked by slow evaluation runs. SIFT outperforms existing tree-search based self-evolution frameworks on the full Polyglot benchmark with significantly lower resource requirements in terms of CPU hours, wall clock time, and API cost.

[AI-99] A Unified Evaluation Framework for Trustworthy Large Language Models Agent ic AI and Multimodal Systems

链接: https://arxiv.org/abs/2609.19524
作者: Shaina Raza,Ahmed Y. Radwan,Imran Liaquat,Kathryn Hume
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Benchmark scores alone provide an incomplete basis for assessing the trustworthiness of modern artificial intelligence systems. Large language models (LLMs), agentic systems, and multimodal models (MLLMs) require different forms of assessment, yet their evaluation evidence must remain interpretable for development and oversight. We propose a unified framework that connects output-level, trajectory-level, and cross-modal assessment through eight trustworthiness dimensions: capability, robustness, safety, fairness, transparency, governance, oversight, and efficiency. The framework preserves system-specific metrics while mapping native measurements to common performance bands, accompanied by uncertainty estimates and traceable evidence. A meta-evaluation layer examines the validity, reliability, and reproducibility of the evaluation itself. Multidimensional profiles expose strengths and weaknesses, while safety-critical overrides prevent aggregate scores from masking critical failures. Mappings to governance frameworks, international standards, and European Union regulatory requirements connect technical assessment with oversight needs. The framework provides a structured basis for assessing both system performance and the credibility of the evidence supporting it, with empirical validation across deployment contexts remaining an essential next step.

[AI-100] An Architecture for Long-Horizon Agents : Levels Ticks and Cascaded Intelligence

链接: https://arxiv.org/abs/2609.19519
作者: Erik Nijkamp,Anurag Koul,Egor Pakhomov,Bo Pang
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Language-model agents are increasingly asked to carry out work spanning days or weeks, such as an operations remediation or a research programme. Such a task outlives any context window, any process and any interval at which a person can attend. In this paper, we argue that a long-horizon agent must run continually without forgetting before it can learn continually. This ability lies in the harness around the model rather than in the model itself. We derive seven bottlenecks from the long-horizon setting and answer them with a hierarchical architecture of three parts: (i) levels indexed by time scale, each keeping a bounded file summarising the level below; (ii) a clocked tick as the unit of autonomous action; and (iii) cascaded intelligence, where work is escalated to a more capable model only after failing review. We report on a ten-day campaign in which an agent built on this architecture reproduced a published reinforcement-learning result with a human attending once a day, and show (1) the agent kept the thread across every context reset and session boundary of the campaign, (2) operating knowledge written early changed later behaviour with no change to model weights, and (3) where learned components would enter such a system. Overall, our experience suggests continual learning for these agents needs a substrate outliving every context and process, and the checks the harness already runs are where a learner belongs.

[AI-101] LLM -as-an-Improver: Turning Verification into Better Candidates

链接: https://arxiv.org/abs/2609.19515
作者: Akiyoshi Tomihari,Yuma Ichikawa
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Verifier-based selection improves LLM performance by generating multiple candidate solutions and using a verifier to select the most promising one. However, existing methods typically treat verification only as a ranking step and discard its feedback once a fixed candidate pool has been evaluated. In this paper, we ask whether verification can also improve the candidate set itself. To this end, we introduce LLM-as-an-Improver and propose Verify–Repair–Reselect (VRR), which uses verification feedback to generate and reselect improved candidates. VRR retains the initial winner while conditionally generating three complementary alternatives: repaired versions of the winner and runner-up, and a solution based on a new approach. It filters invalid and duplicate candidates using only inference-time information and then reselects the final answer under the original evaluation criteria. Across diverse models and code-generation and reasoning benchmarks, VRR improves over fixed-pool verifier-based selection in many settings and can recover correct solutions even when all candidates in the initial pool are incorrect. These results highlight a broader role for LLMs as improvers: verification feedback can not only select among existing solutions but also construct stronger candidates beyond the initial pool.

[AI-102] QVAC Genesis III: A Large-Scale High-Quality Open Synthetic STEM Corpus for Efficient Language Model Pre-Training

链接: https://arxiv.org/abs/2609.19513
作者: Davide Vitabile,N. Ranjan,Akshay Nambiar,Kamal K. Gupta,Amril Nazir
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:High-quality pre-training data is a critical bottleneck for educational and STEM-specific language models targeting edge AI and on-device deployment where token budgets are tightly constrained. While major organizations train ever-larger models on private corpora, the open ecosystem lacks STEM-focused synthetic datasets that deliver high per-token learning value efficiently for small models. To address this gap, we introduce QVAC Genesis III, a 191.43B-token, STEM-focused multi-domain synthetic corpus covering 19 domains across several difficulty levels and different educational styles. QVAC Genesis III is built via a dual generation strategy that performs targeted teacher distillation using a weak edge-scale student model as signal: the student’s failures are converted into corrective explanations, while its successes are expanded into contrastive option-level reasoning over all answer choices. We further introduce an LLM-as-a-parser evaluation protocol that extracts final answers from free-form outputs and tracks both accuracy and answer validity. To validate the effectiveness of our QVAC Genesis III data, we conduct controlled from-scratch ablations with 1.7B-parameter models, showing that models trained with QVAC Genesis III consistently outperform both models trained with the open-source synthetic corpus Cosmopedia-v2 and the publicly released Cosmo-1B model across ARC, GPQA Diamond, and MMLU STEM benchmarks, achieving up to +28.57% on ARC-E and +21.35% on ARC-C, while reaching a Valid Answer Rate of up to 99.45%.

[AI-103] CoreSense: Traceable Failure Recall and Conflict-Aware Belief Gating for Auditable Robot Decisions

链接: https://arxiv.org/abs/2609.19512
作者: Zoe Li
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: 6 pages, 2 figures, 3 this http URL

点击查看摘要

Abstract:Robots can recall prior failures without knowing whether recalled evidence remains valid, conflicts with current observations, or is sufficient to guide a decision. We present CoreSense, a robot-system integration architecture that combines traceable episodic evidence with a conflict-aware belief gate and bounded, auditable recommendations. The gate checks scope, provenance, time, contradiction, and support before it permits PROCEED, requests re-observation, abstains, or escalates. Evaluation follows three complementary layers without commanding a physical robot: offline public real-robot data, a frozen signal-level simulation, and a live cloud deployment path. On CableTrace-120 and BotFails-200, belief gating reduces protocol-defined unsafe proceeds from 20% and 40% to 0%. A disjointly calibrated raw-video policy also reaches 0% unsafe proceed, but overblocks every nominal episode. On public data, a ViFailback-BotFails visual detector reaches 0.778 AUROC yet remains all-blocking, whereas cycle-disjoint UR3 telemetry for protective stops yields 0% unsafe proceed, 36.1% overblocking, and 61.9% coverage; grip-loss transfer remains a negative result. Controlled physical corroboration yields 3.3%, 0%, and 42.0%, while conflict-aware fusion yields 4.7%, 0%, and 42.8%. Finally, 20/20 cloud recalls validate a CockroachDB Cloud-Amazon Bedrock deployment path. The evidence supports an auditable integration pattern, not autonomous recovery or certified safety.

[AI-104] Efficiently Linking Unstructured Data for Multi-step Reasoning

链接: https://arxiv.org/abs/2609.19491
作者: Jiaming Liang,Haydn Jones,Jacob R. Gardner,Mark Yatskar,Zachary Ives
类目: Databases (cs.DB); Artificial Intelligence (cs.AI)
备注: 23 pages, 9 figures

点击查看摘要

Abstract:Modern LLMs and AI agents increasingly support data engineering workflows that integrate evidence from unstructured sources. Such pipelines typically do data retrieval, integration, and ranking before proceeding to more complex agentic reasoning or actions, e.g., for scientific discovery. The core retrieval problem in these workflows jointly executes multi-attribute filtering, multi-vector search, exact relational joins, and thresholded embedding-similarity joins. Given a planned query and monotone scoring function, our DASE query engine constructs and ranks candidate evidence tuples. It comprises (i) a multi-step reasoning query model over structured predicates, multiple vectors, and relational links; (ii) SemJI, a sparse materialized embedding-similarity join index for rare near-neighbor pairs; and (iii) a co-designed execution layer that combines predicate-aware ANN traversal, batched access, and threshold-based score aggregation. On scientific-discovery workloads, DASE retrieves candidate evidence for multi-step reasoning queries 6x to 46x faster than strong RDBMS, rerank, and vector-database baselines at comparable recall; and for tasks that require semantic-operator post-processing, DASE acts as a high-recall prefilter that makes downstream LLM evaluation both cheaper and more accurate – e.g., on SemBench E-Commerce it improves BigQuery quality from 0.67 to 0.80 while cutting cost from 2.42 to 0.54. Comments: 23 pages, 9 figures Subjects: Databases (cs.DB); Artificial Intelligence (cs.AI) Cite as: arXiv:2609.19491 [cs.DB] (or arXiv:2609.19491v1 [cs.DB] for this version) https://doi.org/10.48550/arXiv.2609.19491 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-105] Compositional Reasoning in Language Models under Reinforcement Learning Post-Training

链接: https://arxiv.org/abs/2609.19465
作者: Yu He,Yingxi Li,Yifei Wang,Ellen Vitercik
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Compositional reasoning is critical for real-world problem solving: since training data is necessarily limited, models must generalize by composing learned skills in new ways. While post-training methods such as reinforcement learning (RL) have substantially improved the reasoning abilities of language models (LMs), their effects on compositional reasoning remain less well understood. We propose a dependency-graph framework to formalize compositional reasoning, yielding three levels of compositionality with increasing complexity. Empirically, we instantiate this framework with data-structure tasks, which provide deterministic reward computation and clear compositional structure. We find a consistent decomposed-to-composed asymmetry: decomposed-skill training does not reliably transfer to composed tasks, whereas composed-task training transfers more readily back to decomposed tasks. We provide theoretical explanation for this asymmetry, and further evaluate compositional generalization under length extrapolation, structural distribution shift, and transfer to tasks requiring unseen skills. Finally, we present a pilot study on real-world tool-calling benchmarks, showing preliminary evidence that the decomposed-to-composed asymmetry can extend to practical settings.

[AI-106] he syntax and semantics of goals

链接: https://arxiv.org/abs/2609.19448
作者: David M. Abel,Mark K. Ho
类目: Artificial Intelligence (cs.AI)
备注: To appear in Topics in Cognitive Science, special issue on Goal-Centric Perspectives in Cognitive Science

点击查看摘要

Abstract:In both cognitive science and computer science, goals are conceptualized as cognitive states that flexibly combine with world knowledge to organize and specify purposeful behavior. In this way, goals are compositional representations whose content relates to rational behavior. We here draw attention to goals as representations and their content because it highlights a parallel with other areas in cognitive science - in particular, the syntax-semantics interface in linguistics and logic - while also foregrounding foundational questions about the expressivity, design, and efficiency of different goal representations. For example, goals are typically taken as fixed and imposing constraints on desirable behaviors, but we can also identify constraints on goal representations themselves, such as whether a particular goal language is sufficiently expressive to capture behaviors of interest, or whether different goal representations capture the same behavior. Here, we synthesize work that aims to characterize the properties of different goal representations and suggest these are points of a broader design space. We close by discussing how distinguishing the form and meaning of goals can elucidate the implicit assumptions we make about goals, inform the study of interactions between higher-level cognition and motivation, and isolate axes of variation for different conceptions of goals.

[AI-107] Predict Before You Deploy: Offline Prediction of Quantization-Induced Task Degradation for World Action Models

链接: https://arxiv.org/abs/2609.19441
作者: Jiuyi Xu,Jinjia Guo,Meida Chen,Jing Du,Yangming Shi
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: 9 pages, 3 figures, and 3 tables

点击查看摘要

Abstract:World action models (WAMs) rely on video-generation backbones, requiring substantial memory and compute for deployment. Post-training quantization reduces memory and can accelerate inference, but bit width, grouping, and quantizer choice define a large configuration space. Identifying configurations that preserve task performance through exhaustive closed-loop evaluation is costly. We propose PreDE (Predict Before You Deploy), a policy-calibrated framework for predicting quantization-induced task degradation from offline action deviations. Using closed-loop outcomes from a small development set, PreDE calibrates two thresholds and accepts, rejects, or defers new configurations using a fixed observation log. Under a within-setting label-ordering hypothesis, the rule issues decisions where all thresholds consistent with the development labels agree. Across five WAMs and four benchmark settings, quantization produces configuration-dependent task losses that cannot be explained by bit width alone or a shared deviation threshold. Across 28 held-out configurations from two policies, PreDE issued 21 decisions before observing closed-loop outcomes (75% coverage), all matching the observed acceptable or degraded labels. Deferred candidates included both acceptable outcomes and a 33-percentage-point loss. In 450 Franka Research 3 trials across two independently fine-tuned policies, all configurations assigned to high-deviation groups before testing showed significant degradation, while low-deviation comparisons showed no statistically significant degradation. On the real robot, W4A4 achieved a 1.37x action-query speedup and approximately 44% lower peak memory. These results support policy-specific behavioral calibration for quantization configuration selection while identifying candidates that require closed-loop evaluation. The code is available at this https URL.

[AI-108] Closed-World Resolution Against Tool Hallucination in LLM Agents

链接: https://arxiv.org/abs/2609.19425
作者: Laxmipriya Ganesh Iyer
类目: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Software Engineering (cs.SE)
备注:

点击查看摘要

Abstract:Tool-augmented large language model (LLM) agents fail in a way no tool-selection or tool-security method addresses: they call tools that do not exist and pass arguments no schema declares. Existing defenses either pick the right tool (selection) or constrain what an agent may do with real tools (gating), both of which presuppose the emitted call refers to a real tool at all. We show this is a structural blind spot: a hallucinated call is by construction not a decision any gate made, so no gate can reject it. This paper is primarily a measurement and benchmark study. We give a five-class taxonomy of tool hallucination (H1-H5) and, as a reference point, the Resolution Rung: a training-free, closed-world resolver (registry membership plus a signature check) whose interest is where it must sit, not what it computes. We prove hallucination defense must precede any causal gate, and characterize the one irreducible residue (borrowed arguments schema-indistinguishable from a valid call). Across ten hosted models under two invocation surfaces we measure 322 genuine hallucinations; fabricated-tool calls concentrate on the unconstrained raw-JSON surface (34 vs. 3), and model scale does not help (a 675B model matches a 7-8B one). We then extend to the Model Context Protocol, where merging several servers into one namespace creates hallucination surfaces a single registry cannot express (a second taxonomy, M1-M5); on the live MCP surface we measure 154 hallucinations, including from frontier models that were clean on the single-registry surface, because collisions and shadowing are structural to the merge. We release the versioned Hallucinated-Tools Benchmark (HTB) so any resolver is comparable across submissions.

[AI-109] From Rollout to Reset: A Graph-Based Harness for Autonomous Long-Horizon Manipulation Evaluation

链接: https://arxiv.org/abs/2609.19413
作者: Jing Jiang,Yue Yang,Xinkai Jiang,Gedas Bertasius,Daniel J. Szafir,Rudolf Lioutikov
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Robot manipulation policies are improving quickly, and real-robot evaluation remains the standard evidence for that progress. It still relies on a human to reset the scene between rollouts, which consumes operator time and leaves the initial state distribution unspecified, so results reproduce poorly. A recent system, AutoEval, automates both reset and scoring, but only for single-step tasks, because a long-horizon rollout can terminate in combinatorially many configurations that no single learned reset policy covers. We present HALTER, a Harness for Autonomous Long-horizon Task Evaluation and Reset, which restores the scene by planning over a library of learned atomic reset skills, so demonstration cost scales with the size of that library rather than with the number of terminal states. HALTER builds a spatial scene graph online from point clouds and vision foundation models, and an LLM reasons over this graph to score the rollout, plan the reset, and verify that the reset succeeded, without collecting labeled success images for any task. On four long-horizon tasks on a Franka arm, HALTER restores the scene in 76% of episodes, against 52% for AutoEval and 65% for a motion-planning reset, and it estimates the completed-skill fraction correctly in 90% of episodes, against 76%. Its reset-verification verdict is correct in 91% of episodes, compared with 78% for AutoEval. It also cuts the operator time of an evaluation campaign by 72% relative to manual reset. We further measure compositional generalization on three held-out tasks, where HALTER resets 74.7% of episodes against 1.3% for a per-task reset policy, and we ablate the scene representation and the graph update rate.

[AI-110] Efficient Nash Equilibrium Computation for Cybersecurity Games

链接: https://arxiv.org/abs/2609.19399
作者: Michael Lanier,David Farmer,Yevgeniy Vorobeychik
类目: Computer Science and Game Theory (cs.GT); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Computing Nash equilibria of simulation-based cybersecurity games with policy-space response oracles (PSRO) is bottlenecked by payoff estimation: every payoff-matrix entry costs Monte-Carlo rollouts of a slow simulator, while policies and restricted-game solves are cheap. We introduce Regret-Weighted Payoff Sampling (RWPS), a budgeted estimator that simulates only the cells an equilibrium is sensitive to and fills the rest with a surrogate trained on every entry simulated earlier in the run. The sup-norm error bound cannot evaluate such an estimator, because it is set by the cells left deliberately inaccurate. We prove an instance-dependent bound that weights error by the opponent’s equilibrium mixture, a certificate computable from simulation data alone, and a coverage result showing that once the deviation-relevant set is simulated, surrogate error cannot affect either player’s regret. On three 21x21 general-sum games, two synthetic and an asymmetric Colonel Blotto, the refined bounds are four to six times tighter on the estimator’s own output, and the coverage result predicts in advance which games are cheap: 18% of the matrix for small-support games against 82% for Blotto. In growing-pool PSRO, RWPS reaches lower exploitability than minimum-regret-first search, information-gain search, and progressive sampling at a matched budget, and on the CyGym and ANSG cyber simulators it is lowest at the smallest budgets.

[AI-111] MAGS: Multi-agent Auto-formalization Guarantees Safety for Agent ic Outputs

链接: https://arxiv.org/abs/2609.19391
作者: Albert Wu,Nicholas Roberts,Tzu-Heng Huang,Haoran Lin,Gil Friedman,Sungjun Cho,Gabriel Orlanski,Frederic Sala
类目: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Software Engineering (cs.SE)
备注:

点击查看摘要

Abstract:LLM coding agents now generate complex programs at a scale that makes thorough human review increasingly difficult, raising the risk of safety and security failures. Common approaches, including fuzz testing, static analysis, and LLM-as-a-Verifier, can detect many failures but struggle to cover all possible edge cases. Formal verification addresses this by providing machine-checkable guarantees over specified properties, but traditionally demands substantial manual specification and proof engineering. We introduce a unified multi-agent framework, MAGS, that generates executable programs with formal safety guarantees, using Dafny as a verification-aware intermediate representation where safety properties can be mechanically checked. MAGS formalizes and freezes human-audited APIs and safety requirements, translates generated code into Dafny, repairs violations using verifier feedback, and compiles verified programs back into executable code. We evaluate MAGS on 100 CUDA kernels, 100 terminal scripts, and 20 robotic-arm tasks. Across all 220 examples, it achieves a 100% success rate in producing programs with non-trivial safety guarantees against frozen specifications. Independent safety and functional evaluations further show strong performance across all three domains, while revealing failures when the auto-formalized semantics do not fully capture the target behavior.

[AI-112] Do AI Agents Understand Computer Architecture?

链接: https://arxiv.org/abs/2609.19387
作者: Ambika Sharan,Grigory Chirkov,Soheil Abbasloo
类目: Artificial Intelligence (cs.AI); Hardware Architecture (cs.AR)
备注: 10 pages, 3 figures, 3 tables

点击查看摘要

Abstract:Agents are increasingly asked to design hardware, and increasingly reported to succeed. Such reports establish that a design improved; they cannot establish why. An agent that improves an accelerator may be reasoning about the machine, or may be searching competently over knobs whose meaning it never recovers – and only the first transfers to the next architecture. Existing evaluations cannot tell the two apart, because they vary the agent while holding the framing of the problem fixed. We do the opposite. AutoTuring hands the same agent the same 15-dimensional accelerator space twice: once as named architectural knobs with simulator counters, once as anonymous variables on [0,1], with the evaluator, the legal space and the reachable optima held identical, so that the only thing that varies is whether the problem means anything. The gap between the two is the measurement. On a nine-kernel FP16 GEMM basket, meaning pays: the architect beats a modeled H200 by 5.4% and its blind counterpart by 12.3% on average, with 70.1% fewer simulator calls. It does not pay uniquely: a critic loop recovers most of that gap for the blind agent and buys the architect nothing, so architectural knowledge and structured critique behave as substitutes rather than as complements. We report these as preliminary findings – five to six runs per condition on a single modeled accelerator – and take the comparison itself, not the accelerator, to be the contribution.

[AI-113] How to Guide Your Language Flow

链接: https://arxiv.org/abs/2609.19356
作者: Rohit Dilip,Tianrong Chen,Yuyang Wang,David Van Valen,Joshua Susskind,Miguel Angel Bautista
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:We introduce a new method to guide flow matching models. Our approach, which we call probe guidance, uses the frozen internal states of an existing diffusion model to construct a guidance signal. This works using a similar principle as autoguidance, but eliminates the need for an additional forward pass at inference time and provides a reliable path to ensure that the weak and strong model share similar dynamics. We apply and benchmark this method on continuous diffusion language models, where probe guidance sets a new state-of-the-art performance on unconditional generation. When applied to a 1.7B diffusion language model, probe guidance consistently improves on multiple choice question answering benchmarks. Using our probes, we study the traditional autoguidance setting where the strong model is a weak checkpoint, and find that the weak model must come from a low-entropy region of training. These findings both provide a practical way to improve diffusion language models and shed light on the actual mechanism behind autoguidance, which is currently poorly understood.

[AI-114] Kinematics-Grounded Agent ic AI for Robotic Additive Manufacturing Process Planning

链接: https://arxiv.org/abs/2609.19347
作者: Jingzhan Ge,Ruimin Chen,Azadeh Haghighi,Jiong Tang,Farhad Imani
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Systems and Control (eess.SY)
备注: 25 pages, 17 figures

点击查看摘要

Abstract:Robotic additive manufacturing (AM) extends material-extrusion printing beyond gantry kinematics but makes process planning robot-dependent. A slicer-generated plan that appears favorable in part coordinates can become infeasible or robotically unfavorable on a manipulator because slicer-process decisions and part orientation determine the generated path, while part orientation and workspace placement affect its kinematic realization. Existing AM tools, large language model (LLM)-based decision-support methods, and digital-shadow systems do not provide integrated pre-execution evaluation of these coupled decisions. This paper presents agentic robotic additive manufacturing (A-RAM), an agent-specialist-tool framework that converts user intent and a part file into traceable, execution-ready plans. The LLM interprets manufacturing objectives and constraints, identifies prescribed and searchable planning variables, and encodes this reasoning in a schema-constrained request; a deterministic Planning Agent instantiates the corresponding search workflow, while domain tools compute quantitative evidence for slicing, placement, inverse kinematics, trajectory timing, Joint-6 jerk, and extrusion. The framework is evaluated on a six-axis robotic-arm AM cell through three case studies covering expert-specified planning, goal-only planning, objective-dependent infill screening, and geometry-dependent orientation-placement selection. Across the evaluated candidate sets, selected plans achieve up to 53.5% lower maximum Joint-6 jerk and 48.3% lower mean absolute Joint-6 jerk than the least favorable valid candidates, while objective-specific infill screening yields motion-plan completion times up to 40.1% shorter and extrusion paths up to 12.7% shorter than the corresponding least favorable screened patterns.

[AI-115] GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning

链接: https://arxiv.org/abs/2609.19315
作者: Ruiyang Wang,Hao-Lun Hsu,Swarajh Mehta,Jiwoo Kim,Zhihao Dou,Miroslav Pajic
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models (LLMs) provide a flexible interface for long-horizon robot planning, but generated plans often fail to respect embodiment constraints, recover from planning errors, or reason effectively under partial observability. We present GAVEL, a framework for verifying and repairing long-horizon LLM planning built around an explicit graph world model. The graph represents relevant object-relations, action pre-conditions and effects, and probabilistic beliefs over unobserved object locations. This model can predict the consequences of LLM-generated actions before execution, detect violations, and repair those whose corrections follow directly from the world model. This method also reserves LLM replanning solely for errors requiring semantic reasoning. For multi-task instructions, GAVEL reasons over distributions of possible object locations to reorder remaining subtasks and minimize expected search cost. We evaluate GAVEL on BEHAVIOR-1K across 100 single long-horizon tasks and 500 multi-task instructions. With Qwen3-8B, GAVEL improves single-task success from 41.2% to 91.8% and multi-task success from 19.9% to 92.6%. Distributional belief reasoning also reduces travel distance by approximately 5.4% compared with a static variant. These improvements show that an explicit graph world model harness can substantially improve the reliability and efficiency of long-horizon embodied planning across compact and frontier hosted LLM capabilities.

[AI-116] Randomized SVD Approximations for Spectral Co-Clustering of Word-Document Matrices ICML

链接: https://arxiv.org/abs/2609.19243
作者: Fateme Mazdarani,Carlos Toxtli
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Accepted to 2026 IEEE International Conference on Machine Learning and Applications (ICMLA)

点击查看摘要

Abstract:Spectral co-clustering is a useful tool for discovering latent structure in word-document matrices, but its reliance on singular value decomposition (SVD) can make standard formulations expensive on high-dimensional data. This paper presents two randomized approximations for normalized spectral co-clustering of bipartite text data when the numbers of document and word clusters may differ. The first method uses randomized SVD through random projection, while the second combines partial SVD with element-wise random sampling. Across real-world and synthetic datasets, both methods reduce runtime relative to the full-SVD baseline, but their behavior depends on matrix sparsity. The random projection method is the more reliable approximation across the tested settings, whereas the sampling-based method is most useful on denser matrices and provides limited benefit on already sparse text data. These results show that randomized approximations for spectral co-clustering should be selected according to the underlying structure of the data.

[AI-117] Robust Conformal Intrusion Detection via Traffic-Aware Calibration and Attack-Orbit Invariance

链接: https://arxiv.org/abs/2609.19241
作者: Zhenpeng Li
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 12 pages, 1 figure, 5 tables. Submitted to IEEE Transactions on Dependable and Secure Computing

点击查看摘要

Abstract:Large language models fine-tuned for network intrusion detection emit single-point predictions without statistical validity guarantees. Conformal prediction supplies a finite-sample coverage guarantee, but a threshold calibrated on clean traffic fails once an adversary perturbs controllable network features. We demonstrate this failure across three intrusion detection benchmarks and propose traffic-aware conformal prediction, which calibrates on traffic drawn from the perturbation mechanism an attacker is expected to use and provably restores coverage whenever that mechanism is known and can be sampled. A stronger, adaptive attacker that queries the target model’s own score can still degrade this matched-calibration guarantee. We address this second threat model by excluding attacker-controllable features and their deterministic descendants from the scored representation, and prove that this yields an exact, pathwise coverage guarantee rather than a probabilistic bound. Across three independently fine-tuned language model architectures, this representation remains completely unchanged under every evaluated attack attempt, at a quantified seven-to-fourteen-point cost in clean accuracy relative to the unrestricted feature set.

[AI-118] PAPC: Platform Mediation for Privacy-Propagation Externalities in AI-Mediated Workflows

链接: https://arxiv.org/abs/2609.19226
作者: Tao Huang,Guosen Wu,Chen Hou,Guolong Zheng
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 25 pages, 1 figure

点击查看摘要

Abstract:AI-mediated platforms coordinate work through LLM agents acting for different principals. In these workflows, privacy loss can be created before a final answer appears: a memory write, shared-workspace update, inter-agent message, or tool event may impose downstream exposure cost on another principal. We model this failure mode as a privacy-propagation externality, where the cost of a raw disclosure depends on topology and fanout as well as content. We present PAPC, a platform-mediated mechanism that intercepts information-moving events before they update shared state or external channels. PAPC combines policy, provenance, topology/fanout, privilege, and content signals to allow an event, release a policy-safe abstraction, quarantine raw content, block a transition, or narrow onward rights. The model explains why final-output control misses intermediate exposure costs and why high-fanout objects amplify propagation. Across retrieval-memory and multi-agent workflow benchmarks, PAPC preserves deterministic task completion and eliminates measured exact raw-value and external raw-value exposure. The results position event-level mediation as a platform-governance primitive for agent-mediated online work.

[AI-119] Layer-wise Curriculum Learning for Efficient LLM Compression EMNLP

链接: https://arxiv.org/abs/2609.19213
作者: Donggeon Lee,Dooyeon Na,Seungmin Oh,Jongbin Ryu
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Accepted by the Conference on Empirical Methods in Natural Language Processing (EMNLP) 2026

点击查看摘要

Abstract:In this paper, we introduce layer-wise curriculum learning for efficient LLM compression. The proposed method facilitates the knowledge transfer from the teacher model to the student model, utilizing a curriculum learning approach that begins with easier optimization tasks and progressively tackles harder ones. In order to adopt the layer-wise learning in LLM compression, we partition the whole model into multiple segments consisting of layers, thereby enabling more computationally efficient knowledge transfer for LLMs. Based on our theoretical analysis of cumulative error phenomenon, layer-wise curriculum learning accelerates convergence while stabilizing the knowledge transfer process. In addition, we present a feature caching method with a multi-threading strategy to efficiently address feature misalignment across layers, maximizing GPU utilization. Consequently, our method exhibits advanced model compression performance, as well as high computational efficiency in terms of minimized memory usage and short training hours. Experiments on multiple datasets show that the proposed method achieves state-of-the-art performance while reducing GPU memory usage and training hours by more than 50% on BERT and GPT-2. Moreover, it outperforms the other pruning methods on LLaMA-family and Qwen models under the same training hours, with a lower GPU memory footprint.

[AI-120] What Do Current Systematic Generalization Tasks Miss? A Reasoning -Centered Analysis

链接: https://arxiv.org/abs/2609.19212
作者: Chengwen Qi,Deheng Ye,Yatao Bian
类目: Artificial Intelligence (cs.AI)
备注: Preprint. Under review

点击查看摘要

Abstract:Systematic generalization, the ability to solve novel problems by recombining known atomic elements, is central to human intelligence but difficult to study rigorously under controlled settings. Existing studies therefore rely on simplifications such as approximately linear action composition, productivity-based tests, and action-explicit goals, which make systematic generalization easier to study but omit some essential aspects of this capability. To characterize what these simplifications miss, we adopt a reasoning-centered lens and introduce TranSGrid, a testbed that brings deductive, inductive, and abductive reasoning together within a unified task. Experiments with seven Transformers on 4,800 TranSGrid instances show that all models perform much worse on TranSGrid than on a held-out test set: the largest model solves 79.6% of the test set, but only 55.3% of TranSGrid and 15.8% of the hardest subset. The gap remains within the training length range, showing that productivity alone is not sufficient to evaluate systematic generalization. Additionally, we reintroduce the other two simplifications into TranSGrid: one variant makes actions compose almost linearly (reducing the inductive demand), the other makes goals action-explicit (reducing the abductive one). In both, solve rates return to roughly the test set level, showing that either simplification alone is enough to reduce TranSGrid to an ordinary held-out test set. Together, our results show that existing tasks reduce either or both of the inductive and abductive demands, and that comprehensively measuring systematic generalization requires a task that involves all three forms of reasoning.

[AI-121] Not All Nodes Are Created Equal: Homophily-Aware Stratification for Stable GNN Evaluation

链接: https://arxiv.org/abs/2609.19210
作者: Naga Venkata Sai Jitin Jami,Thomas Altstidl,Sebastian Hoefler,Jonas Mueller,Dario Zanca,Bjoern Eskofier,Heike Leutheuser
类目: ocial and Information Networks (cs.SI); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Machine Learning (stat.ML)
备注: 10 pages

点击查看摘要

Abstract:Graph neural networks are widely used for transductive node classification, with accuracy typically measured on randomly drawn train/validation/test splits. Reported accuracy has been shown to shift substantially across different random splits of the same dataset, making published comparisons between architectures unreliable. The classical remedy in non-graph settings is stratified k -fold cross-validation, which ensures each test fold reflects the full class distribution of the dataset. We argue that class stratification alone is insufficient for graphs: nodes are not isolated but connected, and folds that differ in their distribution of local neighbourhood homophily expose the model to systematically different relational conditions that directly affect message-passing behaviour. The resulting cross-fold variation reflects the homophily composition of each split, inflating reported variance beyond what model behaviour alone would produce. To address this, we propose \hp, a topology-aware stratification procedure that treats node homophily as the primary stratification axis, aligning folds with respect to local relational consistency alongside the class marginal that standard stratification already controls. Stratifying on homophily alone does not guarantee class balance, so \hp incorporates class label as a secondary axis, preserving class representativeness as a natural consequence of the procedure. We evaluate \hp on a broad benchmark suite comprising 15 node-classification datasets spanning the full homophily spectrum and 7 GNN architectures. \hp achieves a mean stability rank of 1.49 compared to 2.31 for random k -fold, achieving the lowest mean stability rank on 13 of 15 datasets while preserving class balance close to class-stratified splits and substantially better than random. We argue that homophily-aware split construction merits broader adoption for GNN evaluation.

[AI-122] MeshKV: A Network-on-Chip KV Cache Fabric for Scalable Transformer Decoding Accelerators

链接: https://arxiv.org/abs/2609.19207
作者: Dong Liu,Yanxuan Yu
类目: Hardware Architecture (cs.AR); Artificial Intelligence (cs.AI)
备注: Accepted as Conference Paper at the 44th IEEE International Conference on Computer Design. (ICCD 2026)

点击查看摘要

Abstract:Autoregressive transformer decoding is constrained by irregular key-value (KV) cache movement on tiled accelerators. Prior compression and DRAM-placement systems still concentrate traffic on centralized memory paths that bottleneck long-context serving. We present MeshKV, a KV cache fabric that moves blocks as packetized flows over a lightweight NoC. It co-designs (i) TaKV affine striping to spread homes and cut hotspot load, (ii) Mare multicast with verified duplicate suppression, and (iii) Pad, which overlaps prefetch, tile multiply, and streaming softmax behind credit-aligned FIFOs. Together they convert bisection back-pressure into useful KV transfer. On our 8x8 FPGA implementation with LLaMA-2-7B and Mistral-7B at 8K-32K, MeshKV reduces interconnect traffic by up to 58%, improves KV bandwidth utilization by 2.1x, and delivers up to 1.9x multi-stream throughput.

[AI-123] REACT: A Fully Spiking State-Space Model for Real-Time Event-Driven Temporal Perception

链接: https://arxiv.org/abs/2609.19204
作者: Geoffroy Keime,Nicolas Cuperlier,Benoit R. Cottereau
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Robotic systems operating in dynamic environments require visual perception that evolves continuously with the incoming sensory stream. Event cameras provide microsecond temporal resolution and asynchronous sensing, but most learning-based methods accumulate events into frames or temporal bins, introducing an integration delay that can limit fast reaction. Here we propose REACT, a fully spiking state-space model for event-driven temporal perception that processes raw events one by one, without temporal accumulation. REACT uses a complex-valued spiking neuron, C-SiLIF, whose continuous-time dynamics are driven by the physical inter-event interval, allowing its internal state to evolve at the temporal resolution of individual events. We evaluate REACT on gesture recognition and time-to-collision (TTC) estimation from full-field event streams, without a target bounding box or localization input. On EvTTC, REACT achieves a 9.59% relative TTC error with 4.6 ms end-to-end inference latency, within 0.15 percentage points of the best learned method while requiring no target prior. At the dataset’s mean approach speed, this latency corresponds to only 4 cm of vehicle motion, compared with 1 m for the fastest competing learned method. REACT further supports anytime TTC prediction, zero-shot transfer to a different driving sequence, and INT8 quantization, reducing the estimated energy consumption from 18.5 to 2.8 mJ per 32,768 events. These results show that event-driven spiking state-space dynamics can provide low-latency, continuously updated temporal perception for reactive robotic systems.

[AI-124] Code-as-Auditor: Executable Compliance Reasoning via Regulation-to-Code CIKM’26

链接: https://arxiv.org/abs/2609.19199
作者: Jisoo Kim,Taeyoon Kwack,Jinwoo Jang,Woo Kyung Kim,Honguk Woo
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注: 12 pages. Accepted at the 35th ACM International Conference on Information and Knowledge Management (CIKM '26), November 7-11, 2026, Rome, Italy

点击查看摘要

Abstract:Large Language Models (LLMs) are increasingly adopted for compliance and legal reasoning tasks, yet their outputs often lack explicit grounding in legal logic and evidence. We present Code-as-Auditor, an LLM-based framework that extends the model’s reasoning capability toward structured and evidence-grounded compliance assessment. The framework translates regulatory information into (1) formalized checklists and executable decision trees, encoding regulations and conditions as interpretable code structures. During inference, each checklist item is (2) dynamically expanded into factual and counterfactual questions, guiding the model to reason over case-specific evidence and potential violations. This process establishes a reasoning pipeline that proceeds from evidence identification, through rule application, to final decision-making, while a self-verification loop improves the logical consistency of the generated code and the traceability of outcomes. Experiments on privacy and data protection scenarios demonstrate that Code-as-Auditor delivers more accurate and evidence-backed evaluations, enabling automated compliance regulation checking grounded in explicit regulatory criteria.

[AI-125] BioPhys-Bridge: A Benchmark for Interdisciplinary Scientific Reasoning in Physics-Grounded Biological Research

链接: https://arxiv.org/abs/2609.19180
作者: Qingyang Xu
类目: Artificial Intelligence (cs.AI); Quantitative Methods (q-bio.QM)
备注: Empirical Methods in Natural Language Processing 2026, 11 pages, 3 figures

点击查看摘要

Abstract:Language models face unique challenges in analyzing interdisciplinary scientific research literature. In biophysics research, faithful answers require grounding observed data in source evidence, interpreting it through a quantitative physics model, and linking it to a biological mechanism. To address this challenge, we introduce BioPhys-Bridge, a novel benchmark dataset for evidence-grounded scientific reasoning over biophysical literature. Each case contains evidence blocks, stable evidence IDs, quantitative values, units, equations, assumptions, mechanisms, and next decisions as grounding targets for question answering (QA) and retrieval-augmented generation (RAG). The initial release contains 500 cases, 1,517 agent-facing tasks, and covers six biological domains and nine physical model families, including three sparse families reserved for future expansion. We enforce strict quality gates for all cases in schema, evidence-integrity, quantitative-grounding, source-license, duplicate, unit-normalization, with domain expert review and annotation for 81 cases. Preliminary evaluations show that DeepSeek-V4-Flash obtain the highest evidence-ID F_1 score (0.360), followed by Qwen3.7-Max (0.316) and GPT-4o-mini (0.294). BioPhys-Bridge is an interdisciplinary benchmark for evaluating attribution, faithfulness, hallucination reduction, and biological experiment design with complex, multi-step scientific reasoning. Future works will increase the size and complexity of the dataset and perform comprehensive evaluations. Code and data are available in the GitHub repository and on Hugging Face.

[AI-126] Regularized Emphatic Temporal-Difference Learning: Stability under Constant Stepsizes

链接: https://arxiv.org/abs/2609.19170
作者: Xingguo Chen,Zhaohui Wu,Jinguo Ye,Chao Li,Shangdong Yang,Guang Yang,Skylar Liang,Wenhao Wang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Emphatic temporal-difference learning (ETD) stabilizes the expected off-policy TD update and changes its projection geometry, but neither property determines constant-stepsize sampled dynamics. We construct an ergodic two-state counterexample in which the ETD mean map contracts while the sampled product has a positive top Lyapunov exponent. Regenerative-cycle analysis separates this sign from the infinite variance of the follow-on trace. We introduce regularized emphatic TD (RETD), a normalized first-order post-shock repair that leaves the trace and importance ratios unchanged, stores the emphatic TD signal in a leaky scalar state, and releases a delayed correction. RETD’s raw equilibrium is an affine shift of the ETD equilibrium; single- and two-regularization readouts recover the ETD fixed point exactly. We prove almost-sure convergence for harmonic diminishing stepsizes and a conditional constant-stepsize moment-contraction result from a Markovian random-product bound. RETD has certified negative exponents on the two-state construction and one Baird point, whereas the positive Baird ETD sign remains numerical. Paired 10,000-run experiments validate both separations, fixed-point recovery, a nonmonotone stability region, and task dependence. RETD changes post-shock dynamics; it does not reduce the shared follow-on-trace variance.

[AI-127] StableEval Arena: A Cost-Aware Agent ic Benchmark for Stablecoin Price Stability Prediction

链接: https://arxiv.org/abs/2609.18949
作者: Sean Wan,Dongping Liu,Luyao Zhang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); General Economics (econ.GN); Computational Finance (q-fin.CP)
备注:

点击查看摘要

Abstract:We introduce StableEval Arena, a cost-aware benchmark framework for evaluating agentic AI systems on stablecoin peg-risk prediction. StableEval Arena evaluates LLM-backed agentic systems on diagnosing peg stress and forecasting deviations from the one-dollar peg over a hidden seven-day horizon, using leakage-safe historical replay with exchange price-volume data and market-context features. We report two complementary experiment blocks: a 120-case stress-enriched validation block and a 507-case natural-distribution full-arena evaluation block. Across six LLM-backed agent configurations and baselines, StableEval Arena measures prediction quality, calibrated-label behavior, structured-output reliability, latency, token consumption, and estimated inference cost. Rather than ranking agents by accuracy alone, the framework treats trustworthiness as a joint property of forecast quality, operational reliability, and computational cost. The results show a gap between protocol-following reliability and financial-risk reliability: agents reliably produce valid structured outputs at modest measured cost, but still miss most rare severe-stress and sustained-depeg cases. To support auditing and replication, we release the benchmark dataset on Hugging Face and the source code on GitHub.

[AI-128] Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation

链接: https://arxiv.org/abs/2609.20758
作者: Sho Kawano,Zehang Richard Li,Paul A. Parker
类目: Machine Learning (stat.ML); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Applications (stat.AP); Methodology (stat.ME)
备注: 15 pages of main text, 30 pages total, 4 figures

点击查看摘要

Abstract:Evaluating an AI system requires disaggregated assessment, as performance varies across domains such as benchmark task types or conversation types in deployed agents. Exhaustive testing is expensive, so evaluation rests on a sample of labeled units. We treat the evaluation set as a finite population and seek accurate point and interval estimates of each domain mean. Direct estimators, including prediction-powered inference (PPI), use only a domain’s own labels and are imprecise where labels are few. Small area estimation addresses this problem, and we build on it to develop an integrated workflow for estimation and validation. For estimation, we propose prediction-powered smoothing (PP-S), a Bayesian model fit to each domain’s prediction-powered estimate, with an extension that borrows strength across a reporting taxonomy (PP-TS). For validation, we derive a new, approximately unbiased design-based cross-validation score for choosing among direct and smoothed estimators. We study a curated benchmark with verifiable grading and deployed agent traffic graded by humans, each with every outcome observed. In both, the proposed estimators improve on the direct estimators in point and interval estimation, with near-nominal coverage. At the same sampling budget, our score selects as well as an independent validation sample does and estimates the selected estimator’s error far more accurately.

[AI-129] A Mathematical Model of Motivated Emotional Mind - Cognitive Embodied System

链接: https://arxiv.org/abs/2609.20437
作者: Wiesław L. Galus,Janusz A. Starzyk
类目: Neurons and Cognition (q-bio.NC); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:This article presents a mathematical model of the Motivated Emotional Mind cognitive architecture developed for embodied intelligent systems. Such a system learns to maintain its homeostasis through a generalized form of reinforcement learning based on its internal motivations, termed motivated learning (ML). The principal contribution of this article is a rigorous formalization of the re-entrant loop integrating feedforward processing, lateral interactions, and feedback pathways, together with the representational selection mechanisms that govern adaptive system responses. The model specifies how ongoing exteroceptive and interoceptive signals, bodily-motivational context, and memory traces are bound into associative memory structures termed semblions, which compete for access to further processing and top-down reconstruction. The formalization encompasses secondary perception, representational competition, curiosity, procedural gaps, and action selection directed toward limiting allostatic violations. Within this framework, motivated learning is tailored to embodied systems whose dynamics are shaped by needs, affect, and the current regulatory state. Unlike standard reinforcement-learning models, the proposed approach incorporates need thresholds, goal generation and shifting goals, bodily state, resource constraints, and action uncertainty, thereby providing a more adequate account of response selection under regulatory pressure. Global affect functions as a central control signal, modulating the learning rate, representational valence, and the balance between exploration and exploitation. The model presented here is a step toward a more rigorous formalization of cognitive phenomena and may provide a basis for further theoretical analysis, computer simulation, and implementation in artificial-intelligence systems inspired by biological processes.

[AI-130] Human and AI-generated texts between modal logic and statistics

链接: https://arxiv.org/abs/2609.20311
作者: Simone Cuconato,Donato Ferrari
类目: Logic (math.LO); Artificial Intelligence (cs.AI); Statistics Theory (math.ST)
备注:

点击查看摘要

Abstract:We read the geometry of semantic neighbourhood graphs as modal logic and give that reading a statistical form, in order to make precise the structural difference between human and machine-generated text. Texts are the worlds of a finite frame whose accessibility is the k -nearest-neighbour relation of a transformer embedding, and the symmetry, transitivity, Euclideanity and seriality frequencies of the two subcorpora are shown to be degrees of validation of the modal axioms \mathsfB , \mathsf4 , \mathsf5 , \mathsfD . Each degree is at once the proportion of instances of a rule that the subframe licenses in Negri’s labelled calculus \mathsfG3.K and a plug-in estimate of a population probability. A prompt-balanced comparison finds consistently higher artificial degrees for \mathsf4 and \mathsf5 . We add a degree of groundedness and of situatedness, and recast the licensing reading in Cuconato’s one-sided sequent-style tableaux, where each degree becomes a rate of set membership.

[AI-131] Risk-Set Transported Synthetic Control with Difference-in-Differences Adjustment under Staggered Treatment Adoption

链接: https://arxiv.org/abs/2609.20264
作者: Mojtaba Eslami
类目: Applications (stat.AP); Artificial Intelligence (cs.AI); Econometrics (econ.EM); Statistical Finance (q-fin.ST); Methodology (stat.ME)
备注:

点击查看摘要

Abstract:In staggered treatment-adoption designs, later-treated units are valid controls for an earlier-treated cohort only until their own treatment begins, so the admissible donor set contracts with event time. Fixing the donor pool at the longest horizon discards temporarily eligible donors, whereas re-estimating synthetic-control weights independently at each horizon can make the counterfactual unstable as donor composition changes. We propose Risk-Set Transported Synthetic Control with Difference-in-Differences Adjustment (RT-SC-DiD). For each cohort and event-time horizon, the estimator fits weights on the currently untreated donors while shrinking them toward a transported reference that reallocates the weight of exiting donors to similar surviving donors. A DiD baseline correction removes persistent level differences. We characterize distortion from horizon-by-horizon reoptimization, derive the loading change induced by naive deletion and renormalization, and give a conditional recursive bound for error propagation under an explicitly assumed regularity condition on the transport map. We also introduce donor-support diagnostics and a donor-only placebo procedure for selecting the transport penalty. In an 80-replication pilot comparison and a separate 40-replication-per-value sensitivity analysis, intermediate transport regularization reduces average RMSE relative to independent horizon-specific estimation and strong anchoring. This evidence supports the method’s bias-variance motivation but is not a proved guarantee. RT-SC-DiD is intended for settings where later-treated units provide useful short-horizon information and donor support contracts materially over time. Existing staggered synthetic-control and synthetic difference-in-differences methods do not, to our knowledge, explicitly regularize within-cohort weight sequences toward transported references as risk sets contract.

[AI-132] Self-complementary completions on six vertices

链接: https://arxiv.org/abs/2609.20231
作者: Xinan Dai,Wenhao Deng,Yingdong Shi,Tailin Wu,Yuchen Yang
类目: Combinatorics (math.CO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Let (\cthreshold(n)) be the largest integer (q) such that every loopless digraph on (n) vertices with at most (q) arcs is isomorphic to a spanning subdigraph of a self-complementary digraph of order (n). We prove that (\cthreshold(6)=7). The upper bound is witnessed by [ \bK3\dunion (x\longrightarrow y\longrightarrow z), ] and follows from a direct argument with a self-complementing permutation. We also determine the complete eight-arc obstruction layer: it consists of five isomorphism classes, or three after converse digraphs are identified. All five are arc-minimal. Each nevertheless packs with an isomorphic copy of itself, so ordinary packing is strictly weaker than same-order self-complementary completion already at this first failure layer.

[AI-133] Long-horizon autoformalization of a core theorem underlying MIP* = RE

链接: https://arxiv.org/abs/2609.19814
作者: Sirui Lu,Ruixuan Deng,Yanqiao Zhu,Zhengfeng Ji
类目: Quantum Physics (quant-ph); Artificial Intelligence (cs.AI); Logic in Computer Science (cs.LO)
备注: 72 pages. Main text 13 pages with 4 figures and 1 table, followed by supplementary appendices (57 pages, 9 figures, 17 tables) and references. Lean 4 library: this https URL

点击查看摘要

Abstract:Landmark mathematical formalizations have taken specialist teams years to complete. We present FormalFlow, a system that coordinates AI proving agents under human supervision to address statement drift and proof composition in long-horizon formalization. Drawing on software engineering principles and practices, it uses a shared blueprint to guide nested planning, proving and review loops. Agents strengthen verification and review throughout formalization. We completed a machine-checked Lean 4 proof of the quantum soundness of the classical low individual-degree test, a core theorem underlying MIP* = RE. Developing the proof took 63 days; greater parallelism could further reduce this time. The final library contains 126,367 lines of Lean code, all generated by agents. The formalization corrects side conditions and intermediate errors while preserving the published final error bound under corrected assumptions. This work provides a verified foundation for quantum complexity and demonstrates a route to affordable verification of major research proofs by small teams.

[AI-134] Physics-Informed Hemodynamic Modeling for Data-Free Prediction and Sparse-Data Assimilation

链接: https://arxiv.org/abs/2609.19290
作者: Xi Chen,Jianchuan Yang,Hongde Li,Guangxin He,Qiuyu Ye,Qiang Luo,Mao Chen,Wenqi Hu
类目: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Clinical decision-making for coronary intervention relies mainly on angiography and fractional flow reserve (FFR). However, angiography is two-dimensional and lacks depth information for 3D lesion characterization, while FFR provides only a single functional index, offering limited hemodynamic insight. Among existing methods, numerical analysis is computationally expensive, whereas learning-based approaches require extensive supervision and often lack physical consistency. To address these limitations, we propose physics-informed hemodynamic modeling, an integrated deep learning framework for 3D coronary blood flow analysis from dual-view angiography. First, an attention-enhanced CNN reconstructs coronary geometry from angiography. The resulting point clouds are then mapped to a reference domain and Fourier-encoded for joint representation. A decoupled network separately predicts velocity and pressure fields, with embedded physical priors enabling efficient transfer across physiological conditions. Across 32 clinical patients evaluated under four flow conditions, the trans-stenotic pressure-drop mean absolute percentage error was 2.02%, while the velocity and pressure relative-L2 errors were 0.054 and 0.023, respectively. Validation against hospital-measured FFR further achieved 93.8% diagnostic accuracy (30/32; exact 95% CI, 79.2%-99.2%). The framework also supports illustrative revascularization comparisons and sparse-data assimilation, with the full angiography-to-hemodynamics pipeline completed within 20 minutes per patient.

[AI-135] Optimal Transport Metric Learning for Feature Alignment in Partially Supervised Segmentation

链接: https://arxiv.org/abs/2609.19176
作者: Dakini Mallam Garba,Salim Abdou Daoura
类目: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Multi-organ segmentation is often challenged by partially annotated datasets and domain shifts across different imaging sources. To address these limitations, we propose a two-stage learning framework that efficiently leverages partial supervision. In the first stage, the model learns from available annotations to produce accurate segmentations of annotated organs, establishing robust feature representations. In the second stage, we introduce learnable organ prototypes and a Sinkhorn-triplet loss to enforce organ-wise feature consistency across datasets. This encourages latent embeddings of the same organ to remain close, while increasing separation between different organs, even when annotations are missing. Our approach achieves performance comparable to state-of-the-art methods on the BTCV dataset, while remaining computationally efficient. By explicitly aligning feature distributions rather than relying solely on pseudo-labels, the framework effectively mitigates domain shift, making it particularly suitable for medical image segmentation tasks with limited annotation resources.

机器学习

[LG-0] Score Centering Stabilizes Off-policy Reinforcement Learning

链接: https://arxiv.org/abs/2609.20807
作者: Martin Marek,Max Ryabinin
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Reinforcement learning (RL) of large language models is notoriously sensitive to small differences between training and inference engines, often referred to as the training-inference mismatch (TIM). However, completely eliminating TIM is impractical, as it would come at a major cost to rollout efficiency. In this paper, we show that the instability of RL under TIM is primarily caused by drift: a persistent bias between training and inference engines that accumulates with every training step. We derive an additive “score centering” correction term that stabilizes RL under TIM by canceling drift. When training models from 0.6B to 30B parameters, score centering alone matches or outperforms methods based on importance sampling under quantization, with the gap growing as the mismatch becomes more severe. Because the correction is additive, score centering also composes with importance sampling – their composition outperforms pure importance-sampling baselines in our staleness experiments.

[LG-1] PosteriorBench: From Point Estimates to Posterior Matching in Evaluating Generative Inverse Solvers

链接: https://arxiv.org/abs/2609.20794
作者: Jiachen Yao,Zi-Siang Hsu,Xi Deng,Aditi Gupta,Xin Ju,Sally M Benson,Gege Wen,Anima Anandkumar
类目: Machine Learning (cs.LG); Computational Engineering, Finance, and Science (cs.CE)
*备注: 32 pages, 10 figures, 21 tables; the code is available at this https URL

点击查看摘要

Abstract:Generative models are increasingly used to solve scientific inverse problems, but existing evaluations still focus primarily on whether a method can produce a single plausible reconstruction. This is insufficient for ill-posed problems, where multiple solutions may be consistent with the same sparse or noisy observations. In these settings, a method can achieve strong pointwise accuracy while still failing to capture the true posterior through mode collapse, overconfident uncertainty, or averaging incompatible solutions. We introduce PosteriorBench, a benchmark for evaluating the distributional accuracy of generative inverse solvers. PosteriorBench evaluates four physics-based inverse problems: Darcy flow inversion, Poisson source recovery, carbon capture and storage, and light transport material inference. For each task, we construct high-fidelity reference posteriors using computationally heavy but established procedures such as rejection sampling and Markov chain Monte Carlo, enabling direct assessment of whether solvers recover the full set of solutions rather than the single best sample. We pair these references with a five-metric posterior evaluation suite: posterior-mean error, posterior-standard-deviation error, maximum mean discrepancy, sliced Wasserstein distance, and radially averaged power-spectrum error. These metrics assess pointwise accuracy, marginal uncertainty, distributional alignment, and global frequency fidelity. The benchmark spans sparse sensing, low-resolution observations, nonlinear forward models, varying noise levels, and multimodal priors, with a unified pipeline for distribution matching and uncertainty quantification. Our experiments reveal substantial distribution-matching gaps across current solvers, while showing that neural operators improve resolution robustness, and guidance weights and generation noise are key to posterior-variance calibration.

[LG-2] Calibrated RF-Fingerprinting Under Interference With Heterogeneous Transmission Protocols

链接: https://arxiv.org/abs/2609.20765
作者: Tariq Abdul-Quddoos,Xiangfang Li,Lijun Qian
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Radio Frequency(RF)-Fingerprinting is a spectrum monitoring technique that identifies specific transmitters based on hardware impairments imprinted within the emitted signal. Although widely researched, studies almost exclusively consider scenarios where only one transmitter is emitting at a time, limiting real world applicability. In this work, we further the study of RF-Fingerprinting by considering co-channel interference, with multiple emitted signals interfering with each other, overlapping in time and frequency. Specifically, we formulate this problem as a multi-label classification problem and employ a 1D convolutional neural network (CNN). Furthermore, the models are calibrated such that the confidence thresholds for the label probabilities are derived, with guarantees on the upper bound on the average number of False Negatives, providing a degree of confidence in not missing a true spectrum policy violation. The proposed method is validated using real world data from the POWDER 5G testbed on devices transmitting 802.11a(Wi-Fi), 4G LTE, and 5G NR waveforms. The results show accuracy as high as 97% and as low as 73% after calibration depending on channel conditions. Also calibrating for various average false negatives upper bounds achieves micro recall scores of approximately (1 - calibrated false negatives) with the calibration robust to out-of-distribution interference, demonstrating the potential of the proposed method in a realistic high contention wireless environment

[LG-3] Agile-WAM: An Agile Tactile World Action Model for Contact-Rich Robot Control

链接: https://arxiv.org/abs/2609.20761
作者: Hanchu Zhou,Brendan Lynch,Raman Goyal,Dechen Gao,Begum Kasap,Boqi Zhao,Junshan Zhang
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:World Action Models (WAMs) advance beyond conventional visuomotor policies by jointly predicting future world states and robot actions, enabling the policy to learn physical dynamics that support effective control. However, recent tactile WAMs often rely on large-scale pretrained generative backbones to capture contact-rich physical dynamics, which limit their inference efficiency and flexible deployment. In this paper, we present \ABBR, an agile tactile World Action Model for contact-rich robot control. \ABBR encodes visual and tactile observations into a shared latent that serves as the source of a direct vision-tactile-to-action flow-matching process, which can jointly generate latent representations of action chunks and future visual/tactile latents. A key observation is that vision and tactile signals evolve at inherently different timescales: adjacent visual frames are often highly similar, whereas tactile signals can change abruptly upon contact. We therefore introduce multi-horizon multimodal prediction in \ABBR, which provides supervision for visual latent at a larger temporal offset while predicting the tactile latent in the next frame to capture fine-grained contact dynamics. Across nine simulated and five real-world contact-rich manipulation tasks, \ABBR demonstrates strong and robust performance, outperforming the strongest baseline in success rate while maintaining low inference latency. In particular, in five real-world experiments, \ABBR yields a relative gain of \textbf29.4% in overall success rates while achieving inference latency of \textbf11.9 ms . These results demonstrate that multimodal WAM can be achieved with an agile architecture suitable for precise and high-frequency robot control. More details are available on our project page: this https URL.

[LG-4] MILER: Semantic Mid-Level Representation for Sim-to-Real Reinforcement Learning in Unstructured Autonomous Driving ATC WWW

链接: https://arxiv.org/abs/2609.20747
作者: Thomas Steinecker,Denis Trescher,Alexander Bienemann,Thorsten Luettel,Mirko Maehlisch
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注: Evaluation video: this https URL

点击查看摘要

Abstract:Reinforcement learning constitutes a promising approach owing to its potential for superhuman performance and self-learned policies. However, its application to real-world autonomous driving remains scarce, particularly in unstructured environments, because of the challenges associated with sim-to-real transfer for unstructured environments. In this work, we present MILER, an end-to-end policy framework with zero-shot sim-to-real transfer. During offline training, we employ a custom semantic mid-level representation (MLR) simulator and train the policy network using reinforcement learning, with its control outputs applied directly to a bicycle model. During deployment on the real vehicle, camera and LiDAR data are processed by BEVFusion to generate a semantic bird’s-eye-view representation consistent with that of the MLR simulator. The actions generated by the policy network are not applied directly to the real vehicle. Instead, we employ a trajectory-alignment strategy that enables zero-shot sim-to-real transfer of both perception and control. We extensively evaluate the proposed framework on a diverse test track comprising numerous challenges, including various obstacles, hairpin curves, velocities of up to 33.6 km/h, and off-road sections. In total, we drove 17.3 km with two different vehicles on a 3.0 km test track without human intervention, thereby demonstrating the effectiveness of our approach. Furthermore, the entire software stack runs on a Jetson AGX Orin.

[LG-5] Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

链接: https://arxiv.org/abs/2609.20744
作者: Haocheng Xi,Yiming Xie,Hexu Zhao,Yiwen Zhang,Michael Liu,Thomas Creavin,Kurt Keutzer,Xiuyu Li,Zhaoyang Lv,Chenfeng Xu,Haiwen Feng
类目: Machine Learning (cs.LG)
*备注: 20 pages, 10 figures

点击查看摘要

Abstract:Video diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major computational bottleneck. Linear attention offers an appealing alternative and has been widely adopted in recent large language models, but directly applying it to video models often fails to preserve the fine-grained interactions required for high-quality generation. We present Video DeltaNet (VDN), which combines local Softmax attention with bidirectional linear memory for long-range video context. Its linear branch introduces Video Delta Attention (VDA), which updates memory once per frame by jointly incorporating its spatial tokens. Separate output projections and learnable gates calibrate the two branches, while a staged teacher-alignment recipe progressively introduces the new pathway into pretrained models. We instantiate VDN on MiniMax H3, applying the hybrid to video-to-video interactions while retaining Softmax for interactions involving text or audio. With eight-step distillation and an optimized SGLang serving stack, VDN-H3 completes DiT denoising for a 14.3-second, 768p video in 6.70 seconds on eight NVIDIA B200 GPUs, corresponding to a 14.5x speedup over the 50-step dense H3 baseline on the same GPU count.

[LG-6] RISC-V and machine learning: a survey

链接: https://arxiv.org/abs/2609.20677
作者: Shriman Keshri,Apparna Singh,Chinmaya Kumar Palo,Shreya Adya,Subhankar Mishra
类目: Machine Learning (cs.LG); Hardware Architecture (cs.AR)
*备注:

点击查看摘要

Abstract:The intersection of open-source processor architectures and machine learning is driving the demand for customizable, efficient, and accessible hardware. This survey examines the state of the RISC-V ISA in machine learning applications, analyzing current capabilities, challenges, and future directions based on recent research. The analysis covers academic and commercial implementations, software frameworks, and real-world applications. The RISC-V machine learning ecosystem is evaluated, from instruction set extensions and core implementations to compiler optimizations and deployment strategies. Key contributions include a unified taxonomy of RISC-V ML implementations, a comparative analysis of performance and design trade-offs, an evaluation of software toolchain maturity, and the identification of emerging trends in instruction set extensions and specialized accelerators. Findings reveal progress in energy efficiency, specialized instruction development, and framework integration, while highlighting challenges in standardization, verification complexity, and ecosystem fragmentation. The analysis proposes four research directions to address current limitations: specialized neural processing extensions, adaptive and modular processor architectures, security frameworks, and energy-efficient multi-domain architectures. These directions provide a roadmap for advancing RISC-V as a foundational platform for next-generation machine learning systems.

[LG-7] Epidemiological Causal Graph Identification: Challenges Identifiability and Algorithms

链接: https://arxiv.org/abs/2609.20676
作者: Sambit Mishra,Yingying Wang,Christine K. Johnson,Urbashi Mitra
类目: Machine Learning (cs.LG); Methodology (stat.ME); Machine Learning (stat.ML)
*备注: 5 pages, 2 figures, accepted at the 60th Asilomar Conference on Signals, Systems, and Computers 2026

点击查看摘要

Abstract:Causal discovery from observational data is fundamental to statistics and machine learning, yet determining causal direction without interventions necessitates structural assumptions. Existing identifiability research primarily focuses on continuous variables under additive noise models, often neglecting mixed datasets containing ordinal scales, counts, and continuous measurements. This paper investigates causal discovery in Directed Acyclic Graphs (DAGs) where nodes follow either an ordinal distribution (via an ordered logit model) or a regular one-parameter exponential family distribution. We prove that the edge direction between an ordinal and an exponential family node is distributionally identifiable for generic parameter values. Our findings generalize previous Ordinal-Poisson results to the broader exponential family. Computationally, we introduce a score-based exhaustive search and a masked continuous optimization framework using DAGMA for larger graphs. Numerical results validate the theory, recovering edge orientations within a Markov equivalence class that are unidentifiable under classical structural equation models.

[LG-8] Multi-center Medical Data Mining with FL-Net - A One-stop Shop for Federated Learning

链接: https://arxiv.org/abs/2609.20650
作者: Simon Süwer,Julian Klemm,Elisa Acitelli,Mathieu Almeida,Lucia Altucci,Zsolt Bagyura,Michelangela Barbieri,Zsolt-Zoltán Bedő,Rosaria Benedetti,Béla Bihari,Csongor Csalóka,Lucia Dicunta,Stanislav Ehrlich,Bjoern M. Eskofier,Sándor-József Fejér,Georg Fröwis,Walter Hötzendorfer,Alexandra Kautzky-Willer,Jens Johann Georg Lohmann,Marianna Maranghi,Lorenzo Marconi,Rudolf Mayer,Wouter Leonard Megchelenbrink,Monika Moga,Adham Mottalib,Sanjeev Mehta,Madeleine Müller,Thomas Nyström,Balázs-Attila Orbán,Paul O’Toole,Giuseppe Paolisso,Paolo Parini,Matteo Pedrelli,Enrico Petrillo,Philipp Poindl,Niklas Probul,Anastasia Pustozerova,Tanja Šarčević,Lukas Weilguny,Jan Baumbach,Andreas Maier
类目: Machine Learning (cs.LG); Cryptography and Security (cs.CR); Distributed, Parallel, and Cluster Computing (cs.DC)
*备注: 69 pages, 8 figures, includes supplementary material

点击查看摘要

Abstract:Federated learning enables collaborative training without sharing patient-level data, but most studies remain simulations. Based on five requirements derived from the literature, we analyzed 14 FL frameworks and found that none fully satisfied these requirements. We present FL-Net, a novel federated clinical research framework to fulfill all requirements. It integrates modular data harmonization, data discovery, disclosure control, securely built versioned FL-Net-Tools and containerized federated workflow execution into a persistent network. It enables the re-use of harmonized data and workflows across studies. FL-Net’s end-to-end capabilities were evaluated through harmonization, cross-study patient discovery across MIMIC and US-130, and reproducible, audited federated workflows with up to 50 concurrent clients. FL-Net is being developed within the dAIbetes and Microb-AI-ome EU projects and will cover over 800,000 patients across 10 hospitals in 9 countries covering longitudinal and single point in time data, FL-Net provides a practical foundation for interoperable, reproducible, and privacy-preserving multicenter clinical research.

[LG-9] Beyond PINNs: A Unified Gauss–Newton and Petrov–Galerkin Framework for Neural and Hybrid PDE Solvers

链接: https://arxiv.org/abs/2609.20641
作者: Nilo Schwencke,Roland Maier
类目: Numerical Analysis (math.NA); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Physics-informed neural networks and finite element methods provide two different paradigms for the numerical approximation of partial differential equations: the former are commonly trained by minimizing pointwise strong residuals, whereas the latter are naturally built from weak variational formulations and the finite-dimensional systems obtained after discretization. In this work, we introduce a common framework based on the discretization of functional Gauss–Newton problems by finite families of linear measurements. We show that, through an appropriate duality pairing, the linear measurements can be represented by test functions. The resulting Gauss–Newton system is then precisely a Petrov–Galerkin discretization of the linearized functional problem. This perspective recovers pointwise collocation and natural-gradient constructions as particular cases, while making the choice of test functions an explicit algorithmic design choice. We specialize this framework to elliptic problems, where it naturally leads to weak residual formulations and to a hybrid finite element–neural construction acting on complementary approximation spaces. Numerical experiments support the proposed framework and demonstrate the effectiveness of weak Gauss–Newton formulations and hybrid finite element–neural approximations.

[LG-10] Recursive Quantum Long Short-Term Memory for Stable Short-Horizon Temperature Forecasting

链接: https://arxiv.org/abs/2609.20594
作者: Mu-En Lee,Yen-Ku Liu,Samuel Yen-Chi Chen,Yun-Cheng Tsai
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Quantum long short-term memory (QLSTM) models extend recurrent sequence learning with variational quantum circuits, but their optimization behavior can vary substantially across random initializations and temporal contexts. This paper evaluates a recursive QLSTM architecture against a standard QLSTM for one-step-ahead prediction of daily minimum and maximum temperature. Using daily weather observations from Toronto and identical training settings, we compare convergence, predictive accuracy, and generalization across input windows of 8, 16, and 32 days over 20 random seeds. The recursive model consistently reaches a near-optimal test loss earlier, reduces mean absolute error and root mean squared error, and exhibits a smaller generalization gap. These results indicate that recursive quantum feature transformations can improve stability and out-of-sample performance for compact hybrid quantum–classical temporal models.

[LG-11] CrystalMO-TuRBO: Multi-Objective Trust-Region Bayesian Optimization for High-precision Joint Crystal Structure Refinement

链接: https://arxiv.org/abs/2609.20592
作者: Joseph Agada,Yishu Wang,Arpan Biswas
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Crystal structure refinement is a fundamental inverse problem in materials characterization, where structural parameters are optimized to reproduce experimental diffraction data. Conventional approaches, such as least-squares and likelihood-based optimization, rely on local search and often struggle with non-convex, noisy, and highly correlated parameter landscapes, particularly when integrating multiple diffraction modalities. Joint refinement of X-ray and neutron data is especially challenging due to their complementary but competing sensitivities, which are typically combined through scalarized objectives requiring manual weighting and leading to suboptimal solutions. We propose CrystalMO-TuRBO, a multi-objective trust region Bayesian optimization architecture for joint crystal structure refinement. The method models X-ray and neutron discrepancies as separate objectives and transforms the problem into a normalized maximization setting. A two-phase optimization strategy is introduced: Phase 1 performs global exploration using parallel trust-region Bayesian optimization across multiple scalarizations to identify promising regions of the parameter space, while Phase 2 conducts localized refinement within a shrinking region to achieve high-precision solutions. This design explicitly separates global search from fine-grained optimization, addressing the unique accuracy requirements of refinement tasks. We evaluate the proposed method on experimentally collected X-ray and neutron diffraction data from single-crystal Ho2Ti2O7. Results demonstrate improved convergence, robustness, and parameter precision compared to classical refinement methods and Bayesian optimization baselines on refinement of a single-crystal pyrochlore material system.

[LG-12] Parallelism critical windows and separations among diffusion language models

链接: https://arxiv.org/abs/2609.20539
作者: Sitan Chen,Liye Wang
类目: Machine Learning (cs.LG); Data Structures and Algorithms (cs.DS); Statistics Theory (math.ST); Machine Learning (stat.ML)
*备注: 90 pages

点击查看摘要

Abstract:A popular selling point of diffusion large language models (dLLMs) is their capacity for parallelism: the ability to generate sequences of text far more efficiently than autoregressive models, which require one forward pass per token. Yet among the many competing paradigms for dLLMs, from masked to uniform to Gaussian diffusion, principled understanding of how these different proposals compare in parallelism remains limited. In this work, we initiate a fine-grained comparison of the capacity for parallelism among these three leading approaches and prove the following: - Uniform and Gaussian diffusion can sample in a number of forward passes which scales with the dual total correlation of the underlying distribution, a measure of intrinsic complexity which can be much smaller than the context length. Previously, it was only known how to achieve this using masked diffusion. - For a certain family of random empirical measures, we show that \widetilde\Theta(\sqrtd) forward passes are necessary and sufficient to sample using uniform or Gaussian diffusion, yet there exist approximate score oracles for which \widetilde\Omega(d) forward passes are needed for masked diffusion. This establishes the first provable separation in parallelism between the three prevailing dLLM paradigms. Contrary to popular intuition that masked diffusions are harder to parallelize because they must commit to token values, the latter separation instead comes from the fact that the critical windows in masked diffusion sampling are asymptotically narrower than those in uniform and Gaussian diffusion sampling. Comments: 90 pages Subjects: Machine Learning (cs.LG); Data Structures and Algorithms (cs.DS); Statistics Theory (math.ST); Machine Learning (stat.ML) Cite as: arXiv:2609.20539 [cs.LG] (or arXiv:2609.20539v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2609.20539 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Sitan Chen [view email] [v1] Thu, 17 Sep 2026 15:10:18 UTC (123 KB)

[LG-13] When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation

链接: https://arxiv.org/abs/2609.20511
作者: Yuxiao Yang,Tianrun Yu,Shangzhe Li,Kaixiang Zhao,Xuchao Zhang,Chetan Bansal,Huaxiu Yao,Taylor W. Killian,Weitong Zhang
类目: Machine Learning (cs.LG)
*备注: 30 pages, 12 figures, 3 tables, code available at this https URL

点击查看摘要

Abstract:We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even exhaust the generation budget. We identify \emphtermination-token mismatch between base students and post-trained teachers as an important source of this behavior. Across Qwen3, Llama, and Gemma, the two models can place their stopping probability on different EOS tokens, even when their declared stopping sets are identical. This mismatch can suppress the student’s preferred termination action without reliably transferring the teacher-preferred alternative. We show that aligning the decoding stopping set alone is insufficient, while treating functionally equivalent EOS tokens as a shared semantic stopping action substantially mitigates mismatch-induced length inflation across all three model families. To further understand how termination behavior evolves over training, we study OPD across different K2-Horizon training stages. This stage-wise analysis shows that termination preferences can shift substantially during training, while also revealing a distinct length inflation late in the OPD run that persists beyond termination alignment. Together, these results identify termination mismatch as an important, but not exhaustive, source of OPD length dynamics. We release an implementation incorporating the proposed termination-handling corrections.

[LG-14] Radio Frequency Detection and Classification of Microplastics in Water

链接: https://arxiv.org/abs/2609.20507
作者: Jaden Tolbert,Md Saiful Islam,Pingshan Wang
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Micro- and nano-plastic particles (MPs/NPs) are ubiquitous environmental contaminants whose increasing abundance and potential health impacts have created an urgent need for rapid, label-free detection methods. As particle size decreases to the low-micrometer range, conventional optical and spectroscopic techniques become increasingly challenging because of limited throughput and/or complex sample preparation. In this work, we present a machine learning (ML)-assisted radio-frequency (RF) dielectric spectroscopic cytometry (DiSC) platform for the label-free detection and classification of MPs. Eight types of 10 \mum nominal-diameter MP particles suspended in deionized (DI) water were characterized at four frequencies spanning 0.2\text-9\text GHz . The measured alterations in RF scattering parameters (S-parameters), referenced to the carrier medium, were used to train supervised ML models for material classification, including the identification of MPs in mixed samples and saline-water environments. For eight MP classes suspended in DI water, the proposed method achieved macro-average F1-score, precision, and recall values exceeding 0.71 . Furthermore, PET classification performance was largely maintained in saline carrier media containing 3.3% and 6.6% sea salt. These results demonstrate the feasibility of ML-assisted RF DiSC for rapid, single-particle MP classification in aqueous environments. Future work will focus on improving classification performance through enhanced RF calibration, increased spectral coverage, larger training datasets, and validation using environmentally aged and biologically contaminated microplastics.

[LG-15] Distributionally Robust Federated Learning with Multi-Source Data

链接: https://arxiv.org/abs/2609.20501
作者: Yingzhu Liu,Zhongkui Li,Pengcheng You,Ashish Cherukuri
类目: Machine Learning (cs.LG); Optimization and Control (math.OC)
*备注: 11 pages, 2 figures

点击查看摘要

Abstract:Federated learning trains a shared model from private client data. In practice, data-generating distributions may differ, and the true mixture across clients is often unknown, making the underlying group distribution difficult to specify. Existing approaches address cross-client mixture uncertainty by optimizing against the worst-case mixture, yet assume accurate client-wise distribution estimates. However, these estimates can be unreliable when based on finite samples. To handle both cross-client mixture uncertainty and within-client distributional ambiguity, we construct a global ambiguity set as the union of admissible mixtures of local ambiguity sets. The construction allows client-specific ambiguity radii and admits a client-wise separable reformulation. Leveraging this structure, we establish a high-probability out-of-sample performance guarantee. We further develop a federated algorithm for a penalty-based reformulation and prove its convergence under milder regularity conditions. Simulations validate the algorithm’s effectiveness.

[LG-16] Resolution limits for process comparison from event data

链接: https://arxiv.org/abs/2609.20489
作者: Antony R. Lee,Peter Tiňo,Iain B. Styles
类目: Databases (cs.DB); Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 35 pages, 4 figures; 13-page supplementary material as an ancillary file

点击查看摘要

Abstract:One hospital runs bloods and imaging at the same time. Another runs them one after the other, in either order, equally often. Knowing which actually happened, and how it is recorded in data, is critical for all operational managers. In process mining, the standard approach is to construct an event log, and attempt to discover concurrent and sequential processes in a data-driven way. We show this standard approach, built on the stochastic language of an event log, reports only the assumptions of its discovery algorithm, because every such log is explained equally well by a model with no concurrency at all. Further, before any data is acquired, we characterise when data can and cannot distinguish concurrent behaviour. Where it cannot, the distinction is recoverable from evidence the stochastic language discards, such as the times at which activities start and end, or object-centric records that fix an order within an execution. The remedy is therefore a choice of what is recorded, rather than a larger sample. This impacts decision making, as planning resource for truly concurrent services is very different from sequential services.

[LG-17] raining Neural Networks to Approach the Optimum Bayes Estimator in Dense Multi-Emitter Localization

链接: https://arxiv.org/abs/2609.20465
作者: Yi Sun,Mona Sharifi,Muzna Yumman
类目: Machine Learning (cs.LG)
*备注: 28 pages, 5 figures

点击查看摘要

Abstract:We train neural networks on synthesized frames to approach the optimum Bayes estimator for dense emitter localization. The result justifies the future work on training neural networks to achieve high-throughput large-FOV super spatiotemporal resolution SMLM.

[LG-18] Seismic Site Response Prediction from Sparse Observations Using Finite-Element-Pretrained Latent Dynamics

链接: https://arxiv.org/abs/2609.20451
作者: Yi Zhu,Su Chen,Xiaojun Li
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Numerical site-response predictions often deviate from observations, yet correcting these discrepancies is difficult because records are limited in both sensor coverage and number of events. This study proposes the Transfer-Enabled Forced Latent Autoencoder for Response Equations (FLARE-T) to improve these predictions by learning and calibrating low-dimensional latent dynamics that connect the base acceleration input to acceleration outputs at multiple depths. FLARE-T learns a low-dimensional response manifold and input-driven dynamics from dense finite-element simulations. It then trains a sparse encoder to map simulated sensor responses into the learned coordinates and uses limited records to calibrate the dynamics within them. A short response window initializes each prediction, while the complete base motion drives the response. The framework was evaluated using a layered-soil centrifuge test and the Lotung field vertical array. Test-set results show that FLARE-T improved multi-depth acceleration histories and 5%-damped pseudoacceleration response spectra relative to the original finite-element models, reducing errors at every evaluated sensor for motions of different intensities and, at Lotung, for both horizontal components. Two Lotung source models with different constitutive parameters achieved comparable test-set accuracy, indicating reduced dependence on precise prior calibration. FLARE-T therefore provides a data-efficient means of combining dense numerical response information with limited field records to improve future site-response predictions.

[LG-19] he Bias of Nonlinear Two-Time-scale Stochastic Approximation under Constant Step-Sizes

链接: https://arxiv.org/abs/2609.20409
作者: Djamel Rassem Lamouri,Dorian Baudry,Nicolas Gast
类目: Machine Learning (cs.LG); Optimization and Control (math.OC); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Two-timescale stochastic approximation (TTSA) is a fundamental tool for analyzing coupled iterative algorithms in reinforcement learning, optimization, and stochastic control. However, finite-time guarantees for nonlinear two-timescale schemes remain difficult to obtain, especially under constant step-sizes. In this paper, we study nonlinear TTSA with step-sizes \alpha\gg\beta . Under standard stability, regularity, and Markovian noise assumptions, we upper bound the mean-squared error and the bias of both iterates around their limiting equilibria. Our bounds scale as O(\alpha+\beta^2/\alpha^2) , which we prove to be tight when \beta\le\alpha^3/2 . The analysis separates the contributions of initial conditions, fast-timescale tracking error, Markovian dependence, and timescale coupling, thereby clarifying the origin of the \beta^2/\alpha^2 term. Our results reveal qualitative differences from the linear TTSA setting previously studied, showing that nonlinear dynamics introduce additional finite-time effects that are absent in the linear case.

[LG-20] Learning Principal-Agent Contracts for Equitable Smallholder Carbon Farming under Moral Hazard and Adverse Selection

链接: https://arxiv.org/abs/2609.20404
作者: Rishi Bharadwaj,Yadati Narahari
类目: Machine Learning (cs.LG); Computer Science and Game Theory (cs.GT)
*备注: 14 pages, 2 figures

点击查看摘要

Abstract:Agricultural soils are a major untapped carbon sink. Carbon farming is emerging as a promising practice for tapping this potential. Smallholder farmers, who dominate agriculture across South Asia and sub-Saharan Africa, are key to scaling climate mitigation via carbon farming. It is ironic that real-world carbon programs largely fail to reach them. We study this important gap through the lens of contract design. An aggregator offers a single pooled contract to a heterogeneous population of smallholder farmers who have private adoption costs (adverse selection) and exert unobserved effort (moral hazard), with agronomic outcomes evolving over multiple seasons. We formulate this evolving contracting problem as a POMDP and use reinforcement learning to learn a dynamic profit-maximising contract. We analyse the performance of the aggregator under various conditions. We find that a profit-maximising aggregator does not merely inherit the exclusion of smallholders, it amplifies it. On large farms the aggregator realises 87.7% of achievable adoption, against only 8.2% on smallholdings. Per-hectare Measurement, Reporting and Verification (MRV) costs fall as farm size rises, and the aggregator’s pooling contract compounds this gradient rather than offsetting it. A counterfactual that makes MRV costs purely area-proportional eliminates this disparity. Our results and simulation can guide contract and policy design that opens carbon income to smallholders while enabling agricultural soils to contribute to climate mitigation at scale.

[LG-21] Minimax-Optimal Online Contract Design with Unrestricted Bounded Contracts

链接: https://arxiv.org/abs/2609.20353
作者: Rui Ai,David Simchi-Levi,Han Zhong
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We study repeated contract design when a principal observes outcomes but not the actions that generate them. The principal may use any bounded outcome-contingent payment vector, and the agent’s best response can make expected profit discontinuous in those payments. For every fixed number m\ge2 of outcomes, the minimax regret over T rounds is of order T^m/(m+1) , up to logarithmic factors. The upper bound allows arbitrary action spaces and agent heterogeneity, without smoothness or monotone-surplus assumptions. Its key is an effective-dimension reduction that the benchmark can be normalized even when fixed tie-breaking is not shift invariant, after which revealed preference yields a monotone response map in payment-difference coordinates. A learning policy built on a Lipschitz parametrization of this map attains the rate using only observed outcome categories. The lower-bound construction accounts for how incentive losses accumulate across outcome dimensions. It shows that each additional contractible outcome creates a precise and unavoidable increase in the worst-case cost of learning.

[LG-22] COMPASS: Ordered Clustered Routing at 100K Scale

链接: https://arxiv.org/abs/2609.20352
作者: Ido Greenberg,Hugo Linsenmaier,Piotr Sielski,Shie Mannor,Alex Fender,Gal Chechik,Eli Meirom
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Large-scale routing often requires visiting clusters of nodes in a prescribed order, giving rise to the Ordered Clustered Traveling Salesman Problem (OCTSP). Optimizing each cluster independently seems natural, but misses non-local dependencies. We introduce the COMPASS algorithm for OCTSP, which combines search with learning-accelerated routing by orchestrating parallel sub-solvers. COMPASS has no quality ceiling and its solutions keep improving with compute. It exploits the clustered structure, and can reach exact solutions in time exponential in cluster size rather than instance size. Empirically, COMPASS consistently outperforms alternative methods. Unlike common large-scale routing solvers, COMPASS consumes general distance matrices and is not limited to coordinate inputs. We demonstrate scaling to 100K synthetic nodes and to 28.5K real e-commerce nodes. To our knowledge, the latter is the largest reported routing solution over asymmetric distances, 9x beyond established ATSP benchmarks.

[LG-23] Detecting Deceptive Recruitment: A Signal-theoretic Machine Learning Framework for Early Identification of Labour Exploitation

链接: https://arxiv.org/abs/2609.20336
作者: Sajid Siraj,Mahnaz Hosseinzadeh,Amin Vafadarnikjoo,Shuyang Li
类目: Computers and Society (cs.CY); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Deceptive online job advertisements have emerged as a primary pathway into forced labour, yet systematic detection methods remain underdeveloped due to data scarcity and absence of empirically validated indicators. We formalise this detection challenge as a classification problem under signalling theory, where exploiters transmit costless signals mimicking legitimate communications across textual, visual, and structural dimensions. Using 464 verified cases (164 deceptive, 300 legitimate) collected through anti-slavery charities across nine origin countries and 21 industries, we develop multimodal detection models combining computer vision, natural language processing, and semantic embeddings. Through systematic feature ablation experiments and repeated stratified cross-validation, we demonstrate that individual modalities achieve substantial discriminatory power (ROC-AUC: 0.87–0.97), whilst their integration yields modest further gains. SHAP-based analysis reveals that text quality and domain-specific risk language are the primary discriminators, with readability indices, risk keyword density, and visa sponsorship mentions ranking highest, followed by visual colour and texture features. These production quality gaps reflect resource constraints that prevent exploiters from maintaining professional standards across all communication channels simultaneously. We operationalise findings through a proof-of-concept decision support system providing interpretable risk scores for practitioners. This work demonstrates how rigorous analytical frameworks can address complex humanitarian operations challenges characterised by information asymmetry and limited ground-truth data.

[LG-24] Sharp Reconstruction Bounds for Autoencoders Using the Same Forward Map

链接: https://arxiv.org/abs/2609.20333
作者: Patricia Medina,Hy P. G. Lam
类目: Machine Learning (cs.LG); Dynamical Systems (math.DS)
*备注: 9 pages, 1 figure, 2 tables

点击查看摘要

Abstract:We study reconstruction in autoencoders that apply the same forward map before and after setting the observed coordinates to zero. For equal odd input and hidden dimensions d\geq 3 , among orientation-preserving diffeomorphisms whose Jacobian singular values lie in [m,M] , we show that the least uniform reconstruction-derivative error is \max\1-M(M-m)/2,0\ , with affine maps attaining this sharp bound at every prescribed depth. A translated radial rotation can nevertheless reconstruct any prescribed ball exactly with singular values arbitrarily close to one, motivating additional conditions for a finite-data bound. We test this prediction on a 798,452-point terrestrial LiDAR forest scan. At input scale 0.05 , the mean theoretical bound is 0.155 , about 84% of the mean normalized training error 0.185 across four spatial regions, two depths, and three seeds. At this scale, adding one hidden coordinate reduces the mean reconstruction error below 6\times10^-6 .

[LG-25] EviRec: Continual Evidence Learning for Dual Cold-Start POI Recommendation

链接: https://arxiv.org/abs/2609.20313
作者: Rongchao Xu,Lin Jiang,Guang Wang
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Point-of-Interest (POI) recommendation is a core task in location-based services, yet most existing methods assume a fixed user population and POI catalog. Through a large-scale data-driven analysis of 10 U.S. cities, we identify substantial POI churn, user turnover, category drift, and decay in static POI memory, motivating the study of continual dual cold-start POI recommendation. To address this setting, we propose EviRec, a continual evidence-learning framework that estimates how much historical evidence should be trusted separately for each candidate POI. EviRec scores each visible candidate from three complementary views: a matching view based on the user’s recent mobility profile, a transition-memory view that captures repeated mobility routines, and a lifecycle view that reflects candidate maturity. Because a near-zero transition score may indicate either irrelevance or insufficient observation, EviRec qualifies the evidence using each candidate’s observation state and applies a reliability gate to adaptively route between transition-memory and lifecycle evidence. We evaluate EviRec on a full-year, five-city POI check-in dataset containing more than 30,000 users and 684,200 trajectories. Experimental results show that EviRec consistently outperforms state-of-the-art baselines, with the largest gains concentrated on cold-start queries. In particular, EviRec improves NDCG@10 by 20.4% on Dual-New cases over the strongest baseline. In-depth analyses further confirm that these gains arise primarily from candidate-specific reliability gating while largely preserving previously learned mobility routines.

[LG-26] ZeroHAT: Behavior-Conditioned Zero-Shot Human Activity Trace Generation

链接: https://arxiv.org/abs/2609.20310
作者: Rongchao Xu,Dahai Yu,Lin Jiang,Guang Wang
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Human activity traces record individuals’ timestamped visits to points of interest and are essential for applications such as mobility prediction and urban simulation. However, accessing large-scale HATs is challenging due to high collection costs and privacy concerns. Synthetic HAT generation offers a promising way to make such data available and has attracted growing interest from both industry and academia. Although many efforts have been devoted to this topic, most of them rely on real data from a region to generate synthetic data for the same region, which is infeasible for the many regions where real HATs are unavailable. To fill this gap, we propose ZeroHAT, a behavior-conditioned framework that generates synthetic HATs for a target region in a zero-shot manner by transferring behavioral patterns learned from real HATs in source regions and adapting them with publicly available contextual information about the target region. ZeroHAT has three key novel components: (i) a multidimensional consistency-aware intent extractor; (ii) a cross-region behavioral cloning module; and (iii) a behavior-conditioned activity realization module. We evaluate ZeroHAT on a ten-city benchmark, where extensive experiments show that ZeroHAT achieves 4.5-6.4x the normalized downstream utility of the strongest baseline and improves average fidelity by 15.6%-40.8% across target regions.

[LG-27] Hypernetwork-Parameterized Spatially Adaptive Neural Operators for PDE Learning

链接: https://arxiv.org/abs/2609.20309
作者: Jiaquan Zhang,Chaoning Zhang,Shuxu Chen,Meng Ye,Xiaofeng Zhang,Qiang He,Weifeng Huang,Guoqing Wang,Yang Yang,Caiyan Qin
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Spatially heterogeneous partial differential equations (PDEs) exhibit location-dependent dynamics arising from variations in geometry and physical coefficients. Existing neural operators improve localized modeling through multiscale features, attention mechanisms, or domain decomposition, yet their update rules often remain spatially shared. Hypernetwork-based methods adapt parameters across PDE instances but typically generate only one global parameterization per instance. Consequently, shared operators may underfit boundaries and high-gradient regions, with these localized errors accumulating during autoregressive rollout. We propose a spatially adaptive neural operator (SANO), which replaces this spatially shared parameterization with a spatially continuous field of location-dependent operator parameters. SANO uses Fourier-encoded coordinates and a coordinate-conditioned hypernetwork to generate spatial operator-conditioning codes at sampling points. A Hyper-Neural Element (HNE) mechanism interpolates these codes within local subregions, coupling neighboring operators while allowing their update rules to vary across space, and partition-of-unity weights assemble the overlapping local predictions. Experiments on one-, two-, and three-dimensional PDEs and two perforated-domain elliptic benchmarks show that SANO consistently outperforms competitive neural-operator, hypernetwork-based, and physics-informed baselines.

[LG-28] SAGG: Sample-Adaptive Gradient Gating for Robust Multimodal Learning under Heterogeneous Corruption

链接: https://arxiv.org/abs/2609.20302
作者: Wentao Zhang,Yifan Zhu,Yutong Zhang,Wentao Mo
类目: Machine Learning (cs.LG); Multimedia (cs.MM)
*备注: Accepted by ACM Multimedia 2026

点击查看摘要

Abstract:Multimodal gradient balancing methods modulate encoder gradients with a shared scalar per modality, implicitly assuming that corruption is uniform across the training batch. In practice, corruption is sample-heterogeneous: within a single mini-batch, different samples may have different modalities corrupted. We prove that under this heterogeneous corruption model, any batch-level sample-agnostic linear estimator with a shared modulation parameter incurs an irreducible bias with respect to the clean-data gradient, and that sample-level all-or-nothing gating is the unique unbiased strategy within a natural distribution-free estimator class. Motivated by this result, we propose Sample-Adaptive Gradient Gating (SAGG), which makes a binary retain-or-discard decision per sample via an online feature-norm quality test and incorporates a truncation mechanism for variance control. We prove that SAGG-based SGD converges at the standard O(1/sqrt(T)) rate to stationary points of the clean loss without a corruption-dependent error floor, and derive a certified robustness radius for the independent-encoder architecture that connects per-modality Lipschitz constants to the classification margin. Experiments on Kinetics-Sounds and UCF-101 under Gaussian noise injection, partial modality missing, and natural contribution imbalance show that SAGG consistently outperforms ten existing methods, with the largest gains in high-corruption regimes where batch-level bias is most severe.

[LG-29] Improving Generalization and Robustness in Offline Reinforcement Learning via Boundary-Aware Data Augmentation

链接: https://arxiv.org/abs/2609.20300
作者: Gong Gao,Weidong Zhao,Xianhui Liu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Current offline reinforcement learning (ORL) algorithms tend to overfit the training dataset and exhibit poor in-distribution generalization and robustness performance when deployed to real environments, thus compromising their effectiveness. Existing methods typically enhance in-distribution generalization and robustness by leveraging regularization techniques widely used in computer vision. However, due to the high sensitivity of low-level physical signals to distributional shifts, these methods still suffer from notable limitations in in-distribution generalization and robustness, making it difficult to achieve stable performance in complex environments. To address this issue, we theoretically analyze the error bounds of the behavior policy and action-value function trained with random episode interpolation, revealing that the error scales positively correlated with the distance between states. Based on this insight, we propose a method called \bfB oundary- \bfA ware \bfD ata \bfA ugmentation (BADA), which leverages neighboring states to construct interpolation boundaries, enabling the generation of synthetic data that more faithfully preserves the original data distribution. We first conduct qualitative studies in a toy environment, showing that BADA generates mixed samples that preserve desirable policy smoothness while accurately reconstructing multimodal value distributions. Extensive experiments on limited offline datasets further demonstrate that BADA attains state-of-the-art performance across diverse benchmarks.

[LG-30] A Learning Algorithm for Threshold Boolean Networks with Prescribed Fixed Points

链接: https://arxiv.org/abs/2609.20298
作者: Gonzalo A. Ruz
类目: Machine Learning (cs.LG)
*备注: 14 pages, 3 figures, to be published in 20th International Meeting, CIBB 2025, Revised Selected Papers

点击查看摘要

Abstract:We present a learning algorithm for inferring threshold Boolean networks (TBNs) with a prescribed set of fixed points. The proposed method employs a custom differentiable loss function that jointly enforces fixed point preservation, penalizes spurious attractors, encourages binary outputs, and promotes sparsity through L1 regularization. Applied to the FOS-GRN model of Arabidopsis thaliana, the approach achieved perfect reconstruction (i.e., all 10 desired fixed points and no spurious ones) in 5 out of 30 independent runs, recovering on average 8.53 \pm 0.90 correct fixed points with no spurious attractors. In contrast, standard methods such as the Perceptron and Logistic Regression recovered up to 10 fixed points but introduced between 8 and 31 spurious ones. An additional analysis varying the sparsity coefficient ( \lambda ) confirmed that the method’s performance and the structural properties of the inferred networks remain robust within a practical range (up to 0.01) of regularization strengths. Overall, the results demonstrate the effectiveness and stability of the proposed algorithm in capturing meaningful network dynamics under prescribed dynamical constraints.

[LG-31] Intact-to-Amputee Transfer in Surface-EMG Gesture Decoding: Training Source and Calibration Budget

链接: https://arxiv.org/abs/2609.20297
作者: Jethro Odeyemi,W. J. Zhang
类目: Machine Learning (cs.LG); Neurons and Cognition (q-bio.NC)
*备注: 22 pages, 6 figures, 11 tables

点击查看摘要

Abstract:A recogniser trained on one person rarely transfers to the next, and useful performance usually demands a fresh round of labelled calibration from the end user. A systematic review of 1077 studies quantifies where the evidence is thin: amputees appear in about one in six. Here a montage-agnostic cross-user encoder is carried to eleven transradial amputees on a protocol matched to its intact-limb training data. Zero-shot cross-population transfer fails outright: the encoder requires labeled data from the new user before it begins decoding, and it then exceeds the per-user classifier a clinic would fit by 0.190 macro F1 at three repetitions and for every subject in the cohort. Given three labelled repetitions it reaches 0.779 macro-F1 against 0.589 for the per-user pipeline. Training on forty intact subjects produces better transfers to a new amputee than training on ten other amputees, and combining the two produces better transfers than either individually. The prediction pre-registered for this study, which extends the encoder’s baseline-strength account with the premise that amputee EMG is less separable, holds true only after a few repetitions become available and after enriching the source pool with additional amputees. At a single repetition, and at every budget under a source matched to the intact-limb comparison, it fails. Thus, it locates the boundary of the proposed account.

[LG-32] A Table-Free Index for Tapered Memoization Grids: Compact Out-of-Core Evaluation of Functions of Sorted Arguments

链接: https://arxiv.org/abs/2609.20276
作者: Tamal Maharaj
类目: Data Structures and Algorithms (cs.DS); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Many applications must repeatedly evaluate an expensive function f of a sorted score vector whose influence decays with rank: Plackett-Luce choice probabilities, alpha-entmax attention thresholds, and rank-weighted aggregates. Biswas and Regan (TCS 2015) introduced a tapered grid that memoizes such functions, indexed through precomputed node-count tables. We first make explicit that the tapered grid’s key set is exactly the set of multiset combinations, so its index is the classical combinatorial number system: this yields a table-free closed-form O(d) rank that eliminates the O(Bd)-O(Bd^2) preprocessing tables of the original scheme, generalizes it beyond a pinned first coordinate, and supplies the previously missing O(d) unranking, which enables order-free parallel construction and key-free storage. The resulting structure is a values-only flat array: at N=37.4M entries it occupies 5.7x less memory than a hash-map memo and answers queries 1.1-1.8x faster once both structures exceed cache, and it remains operable memory-mapped beyond RAM, where pointer-based alternatives cannot reside. We give design guidance for choosing the taper: the optimal per-level refinement ratio equals the influence-decay ratio, and we derive a finite-epsilon closed form for the size penalty of a mismatched ratio – accurate to a few percent where the classical asymptotic rate overstates the penalty by 16-43%. End-to-end, memoizing the Plackett-Luce normalization – whose exact evaluation is an iterative transcendental root-find – is 25-55x faster than Newton’s method at 2.6e-3 mean error, against the approximately 10x reported originally; we also report a negative result, alpha-entmax thresholds, where the exact solver’s tiny active support makes it 2.6x faster than any table, and distil the scoping rule this implies.

[LG-33] Placement Is Free Composition Is Not: The Latin Square as a Provably-Balanced Construction for Heterogeneous Sequence-Mixer Stacks

链接: https://arxiv.org/abs/2609.20269
作者: Taebong Kim,Youngsik Hong,Minsik Kim,Sunyoung Choi,Jaewon Jang,Minseo Kim
类目: Machine Learning (cs.LG)
*备注: 18 pages, 5 figures

点击查看摘要

Abstract:Since GPT, most Transformers have repeated the same attention mechanism at every layer. Yet this design is largely a convention rather than a tested conclusion. When multiple sequence mixers are combined in one stack, improvements may arise from mechanism choice, placement, or both, making causal attribution difficult. We introduce Aether-7B-5Attn, a 6.59B-parameter mixture-of-experts model ( \approx 2.98B active) whose 49 layers contain seven sequence-mixing mechanisms arranged as a 7\times7 Latin square. Because each mechanism appears exactly once in every row and column, the design guarantees balanced exposure across depth while eliminating placement confounds. To evaluate this principle, we build a parameter-matched proxy with four mechanisms arranged as a 4\times4 Latin square over sixteen layers, matched to 700.9M parameters and trained with eight seeds per arm. The results reveal a clear dissociation. Rearranging a distributed heterogeneous stack into a balanced periodic cycle changes validation loss by only 0.16%, indicating that exact placement has little effect. In contrast, clustering the same mechanisms into contiguous depth bands incurs a 0.59% penalty, while replacing the heterogeneous stack with a homogeneous one incurs a 1.68% penalty. These results indicate that performance depends primarily on heterogeneous composition distributed across depth rather than on any particular permutation. We confirm this finding at 2.16 \times larger scale (1.514B parameters), where the homogeneous-stack penalty increases to 2.63% and removing the SSM-family mechanism produces a 3.20% degradation. We further report per-mechanism cost profiles, English and Korean evaluations, and a causal-safety audit of all 49 layers. We release model weights, training recipes, training code, logs, and architecture source code.

[LG-34] Improving Online Reinforcement Learning via Bidirectional Behavior Prior Distillation

链接: https://arxiv.org/abs/2609.20268
作者: Gong Gao,Xiao Lai,Jiaji Shen,Ning Jia,Xianhui Liu,Weidong Zhao
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Online reinforcement learning (RL) algorithms frequently exhibit poor sample efficiency and unstable learning dynamics, stemming from systematic critic estimation errors that are exacerbated by greedy policy updates. Existing behavior-prior reinforcement learning methods attempt to alleviate this issue by relying on offline pre-training to learn behavior models from fixed datasets and using policy priors to constrain online policy updates. However, the limited quality of offline datasets often hinders the ability to provide high-value policies that can effectively guide policy updates. The absence of expert trajectories significantly impairs online policy learning, leading to low sample efficiency and suboptimal performance. To address these challenges, we depart from conventional behavior prior approaches and propose a Bidirectional Behavior Prior Distillation (B2PD) algorithm. B2PD leverages action-value priors to guide a conditional variational autoencoder (CVAE) in generating a high-value behavior support set. The resulting expert behavior priors are further distilled into the agent, effectively reducing inefficient exploration and enabling stable policy optimization, while establishing a bidirectional knowledge flow mechanism. Empirical evaluations on both state- and pixel-based tasks verify that B2PD substantially improves sample efficiency while maintaining stable policy optimization. More broadly, this work shows that enforcing high-quality behavioral support during online learning effectively mitigates critic-induced error amplification, enabling structured behavior priors to guide policy updates in a principled and sample-efficient manner.

[LG-35] How Far Can Sub-3B Open Language Models Go in Zero-Shot Essay Scoring on an 8 GB Consumer GPU?

链接: https://arxiv.org/abs/2609.20250
作者: Nguyen Dung Son,Dang Quang Minh,Nguyen Huu Loi,Truong Viet Vu,Nguyen Thai Anh
类目: Machine Learning (cs.LG)
*备注: 15 pages, 4 figures

点击查看摘要

Abstract:Zero-shot essay scoring with large language models is usually demonstrated with proprietary API models, yet the settings where automated scoring is most needed, such as public schools grading thousands of essays under strict privacy rules, are often those where sending student writing to a third-party API is unacceptable. We ask how much capability survives when the model must be a sub-3B open model running fully locally in FP16, with a controlled study of four instruction-tuned models from two families (Qwen2.5 at 0.5B/1.5B/3B, SmolLM2 at 1.7B) on all eight ASAP-AES prompts on a single 8 GB consumer GPU, with bootstrap confidence intervals, Holm-corrected paired tests, and deployment-realistic variants of the key design choices. Three findings emerge. (i) Rubric-decomposed prompting beats holistic prompting for every model under batch min-max aggregation (though Qwen2.5-3B drops significantly on one prompt), and under mean aggregation two unrelated families land within 0.01 at the 1.5-1.7B scale. (ii) Mapping trait scores into the prompt range is fragile to grader calibration: one model compresses traits into a narrow low band (2-4 on 0-10) and naive mean aggregation collapses, while the min-max normalization of Multi-Trait Specialization repairs it (macro QWK 0.204 to 0.388) and stays within 0.03 when its statistics are frozen on 30 held-out essays. (iii) Signed error falls with essay length in eleven of twelve configurations, opposite to the verbosity bias reported for large LLM judges; normalized rubric decomposition largely flattens this slope for well-calibrated models. We anchor results honestly: the best local configuration (0.388) remains far below both the human inter-rater ceiling (0.769) and a length-only baseline (0.523), so we position sub-3B local models strictly for formative, human-supervised feedback.

[LG-36] Accuracy Is Not Enough: A Cross-Architecture Audit of Demographic Bias in Deep Knowledge Tracing

链接: https://arxiv.org/abs/2609.20249
作者: Dang Quang Minh,Nguyen Dung Son,Nguyen Huu Loi,Truong Viet Vu,Nguyen Thai Anh
类目: Machine Learning (cs.LG)
*备注: 15 pages, 3 figures

点击查看摘要

Abstract:Deep knowledge tracing (DKT) models implicitly decide which students an adaptive system believes have mastered a skill, yet almost all evidence on their demographic fairness comes from Bayesian knowledge tracing; the deep models that power modern systems have received no comparable cross-architecture audit. We close this gap: four architectures (DKT, DKVMN, SAKT, AKT) trained under three regimes (standard, reweighting, adversarial) on two public datasets with demographic metadata, Eedi (15.9M interactions) and OULAD (167k after preprocessing), evaluated with ABROCA, student-level bootstrap confidence intervals, and permutation tests addressing recent critiques of fairness-metric instability. Three findings emerge. (i) Bias is real but context-dependent: every architecture shows a significant socioeconomic ABROCA on Eedi (0.018-0.023, p0.005 ), with per-group AUC lower for economically disadvantaged students, while gender bias is significant on OULAD for three of four architectures after multiplicity correction yet negligible on Eedi. (ii) The most accurate architecture is the most biased: AKT gains about 4 AUC points from item-level Rasch embeddings and shows the largest socioeconomic ABROCA, exceeding every other architecture under a paired bootstrap ( p\leq0.002 ); ablating only the Rasch embeddings removes the accuracy gain and the excess bias together. (iii) Standard mitigation is unreliable: reweighting and adversarial debiasing leave ABROCA essentially unchanged in every configuration that preserves accuracy, even though the adversary is pinned at chance at full reversal strength and a weak-strength positive control rules out a dead probe.

[LG-37] Explaining spatial information flow in short-term traffic forecasting models using a gated graph attention network

链接: https://arxiv.org/abs/2609.20217
作者: Yue Li,Shujuan Chen,Ying Jin
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Short-term traffic forecasting supports real-time monitoring and control of road networks, and graph attention networks (GAT) are the standard means of representing spatial dependence in these models. GAT layers are widely described as capturing the influence of neighbouring locations, but this is seldom verified, because the attention weights offered in support cannot be compared against any measured quantity. That leaves two questions open, how the model should be explained and which of its components are necessary. We address this by adding a gate to the GAT layer which learns, at every sensor and every time step, what share of a sensor’s updated state is drawn from its neighbours rather than from itself. Regularising the gate withdraws neighbour information progressively and thereby provides a graded form of ablation. We apply the gated GAT to ST-MetaNet, whose encoder and decoder each place one GAT layer between two recurrent layers, and train it on one calendar year of records from 498 loop detectors on the strategic road network of England. The gate assigns a larger share of neighbour information to sensors carrying heavier traffic and follows the daily and weekly cycle of travel, consistent with adjacent locations being more strongly coupled when busy. Mild regularisation improves accuracy slightly, and accuracy declines at higher strengths as the penalty withdraws information the model needs. The encoder gate closes before the decoder gate, but direct ablation qualifies that ordering. Removing either GAT layer alone leaves accuracy at least as good as keeping both, whereas removing both degrades it substantially, so the two layers are largely redundant rather than either being indispensable. The gated GAT therefore yields a modest accuracy gain, an explanation of where and when spatial information flows, and evidence on which layers the architecture requires.

[LG-38] Subdomain-aware representation compression for pretrained image embeddings

链接: https://arxiv.org/abs/2609.20213
作者: Poowanut Niamluang,Jittat Fakcharoenphol
类目: Machine Learning (cs.LG)
*备注: Appeared in JCSSE2026

点击查看摘要

Abstract:Dimensionality reduction is a well-known technique for improving space efficiency, typically applied uniformly across an entire dataset. This paper investigates the possibilities of using dimensionality reduction techniques for subdomain representation compression. We explore standard techniques such as Principal Component Analysis (PCA) and Linear discriminant analysis (LDA) in image domains. The results not only demonstrate the expected improvements in space and computation complexity crucial for edge-device ML applications but also show improvements in accuracy over direct full-embedding procedure. One possible explanation is that dimensionality reduction effectively extracts subdomain features. We also performed experiments to demonstrate transfer learning capabilities using the compressed representations.

[LG-39] Evaluating Financial Sentiment in the Age of AI

链接: https://arxiv.org/abs/2609.20198
作者: Arslan Bisharat,Oudom Hean
类目: Machine Learning (cs.LG)
*备注: 19 pages, 5 figures, 3 tables, preprint

点击查看摘要

Abstract:Financial sentiment measures are widely used in empirical finance, but it remains unclear whether general-purpose large language models (LLMs) improve on existing finance-specific methods. This paper evaluates twelve sentiment models, including dictionary-based methods, finance-specific transformers, and open-source LLMs, using two criteria: linguistic validity and economic validity. We find that general-purpose LLMs achieve classification performance comparable to finance-specific transformer models without task-specific fine-tuning. However, higher classification accuracy does not translate into stronger economic relationships. Several models produce sentiment measures that are significantly associated with earnings surprises, but none is significantly associated with next-day stock returns. Model performance is strongest for announcements with large earnings beats or misses and substantially weaker for announcements with more moderate earnings surprises. These findings suggest that financial sentiment captures information about firms’ economic performance but has limited ability to explain short-run market reactions

[LG-40] When Does Retrieval Help Time-Series Forecasting?

链接: https://arxiv.org/abs/2609.20193
作者: Mert Onur Cakiroglu,Elham Buxton,Mehmet Dalkilic,Hasan Kurban
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Retrieval plug-ins supply a deep forecaster with information its lookback window cannot carry. Published evaluations report consistent gains, and each credits its own mechanism. We show that the benefit belongs instead to the operating point: the relation between window length S and dominant seasonal period L , an axis the standard protocol never varies. Stratifying the evaluation by that relation exposes the regime. At S=12 , a simple control that repeats the last observed period beats the six standard backbones, in aggregate, on four of seven benchmarks by 8% to 44% of MSE. It beats the strongest plug-in we run on ETTm1 and matches it on ECL. It is worse by up to 25% on the three datasets whose training-split spectra lack a concentrated, shared period. A controlled synthetic sweep of horizon, period, and window shows the benefit boundary tracks the period (correlation +0.71 ), not the horizon ( -0.23 ). A paired control with no phase to recover nearly erases the effect, consistent with phase starvation. Zero-shot pretraining does not escape it: a foundation model trails trained backbones by 22% to 50% on the periodic benchmarks. Within our instrument, exact lookup matches graph diffusion: the payoff is consulting the record, not the machinery on top. Two interpretable statistics, a trend test and a staleness rate, predict the sign of the per-cell benefit at 0.76 accuracy under leave-one-dataset-out evaluation, a suggestive margin over the 0.69 majority rule, where a 22-feature stack manages 0.57 . We propose no new plug-in. The contribution is the regime map, the protocol that reveals it, and two statistics that screen it before deployment. Code: this https URL.

[LG-41] Robust Federated Q-Learning with Almost No Communication

链接: https://arxiv.org/abs/2609.20174
作者: Sreejeet Maity,Aritra Mitra
类目: Machine Learning (cs.LG); Systems and Control (eess.SY)
*备注: Accepted at the 2026 American Control Conference (ACC 2026)

点击查看摘要

Abstract:We consider a federated reinforcement learning setting involving M agents, all of whom interact with a common Markov Decision Process (MDP). The agents exchange information via a central server to learn the optimal value function. Our goal is to understand to what extent one can hope for collaborative sample-complexity speedups in such a setting, when a small fraction of the agents are adversarial and can act arbitrarily. To that end, we propose Robust Fed-Q, a federated Q-learning algorithm that blends ideas from both model-based and model-free RL, along with the median-of-means device from robust statistics. We prove that despite corruption, with high-probability, Robust Fed-Q (i) guarantees exact convergence to the optimal value function in the limit of infinite samples, and (ii) enjoys near-optimal finite-time rates that benefit from collaboration. In addition, our approach requires just \tildeO(1) rounds of communication to achieve each of the above guarantees, a feature of independent interest in FL where communication is the major bottleneck.

[LG-42] Support Thresholds Not Algorithms Limit Rare-Association Recovery in Co-Purchase Networks

链接: https://arxiv.org/abs/2609.20171
作者: Xiao Han,Zhen Zhang,Xin Zhao,Jiechun Lei,Moxuan Zheng,Youting Wang
类目: Machine Learning (cs.LG); Computational Engineering, Finance, and Science (cs.CE)
*备注:

点击查看摘要

Abstract:The support threshold of the Apriori algorithm involves a trade-off in conducting market basket analysis: the associations that occur frequently are noted with high threshold; however, the low ones lead to generating the large amount of rules. The paper compares five methods for co-purchase edge filtration on two grocery datasets: i.e., Instacart (3.2 million baskets) and Dunnhumby (208 thousand baskets), including Apriori, Apriori + lift post-filtering, top- K ranking based on lift, and two methods based on networks, noise-corrected (NC) and disparity filter (DF). The top- K method ensures the maximum average lift, while the NC achieves similar lift level by means of a single value of the significance parameter ( \alpha ). These two methods recover substantially more rare high-lift associations than Apriori (80-100% against 22-28%). NC and top- K select meaningfully different edges (18-29% non-overlapping): NC retains statistically validated pairs, while top- K retains rare pairs with high lift but low statistical significance. A rolling-origin holdout evaluation shows that top- K edges recur at higher rates at every split, but NC edges are ~12 pp more likely to remain statistically significant in the held-out network.

[LG-43] Small Enough to Know Everything: The Fully-Enumerable Transformer as an Instrument for the Science of Delayed Generalization

链接: https://arxiv.org/abs/2609.20166
作者: Yoshiyuki Ootani
类目: Machine Learning (cs.LG)
*备注: 9 pages, 2 figures. Code and per-seed run records will be released publicly

点击查看摘要

Abstract:Tiny transformers trained on fully-enumerable tasks occupy an unusual position in the study of grokking: every input can be evaluated, every generalization ceiling can be computed exactly, and hundreds of seeds cost minutes. We argue this regime is a scientific instrument with four capabilities that approximate settings cannot offer: (a) exact, falsifiable generalization ceilings; (b) task surgery that manipulates one structural variable while provably fixing all others; © direct observation of every weight; and (d) survival-time statistics over many seeds that recast “does not grok” as a censored observation. The obvious objection is that laws characterized at 10^4 parameters may not mean anything beyond them. We answer it with a preregistered conservation study: three task-side laws established at 12K parameters – a recoverability-ceiling law, a role-conflict delay law, and a weight-decay response law – are re-measured under an identical from-scratch protocol at 12K, 1M, and 50M parameters (a 4,000x span; 360 runs plus a 44-run control arm). The ceiling law and the delay law are conserved (0/144 Holm-corrected ceiling violations; Spearman rho = 0.75 at every scale, permutation p 1e-4), while the weight-decay law deforms systematically, steepening with scale. Preregistered controls show the 50M role-conflict deficit survives learning-rate adjustment and a tripled budget. Conservation was tested against criteria frozen before data collection, and one law’s deformation shows the test could have failed. These results license the fully-enumerable transformer as a model organism for the task-side laws of delayed generalization: what it measures exactly, larger models largely obey – and where they deviate, the deviation is itself lawful and measurable.

[LG-44] LEO Satellite Internet of Things: Architecture Technology and On-Orbit Verification

链接: https://arxiv.org/abs/2609.20165
作者: Ming Ying,Xiaoming Chen,Qiao Qi,Yichao Xu,Jiajun Pan
类目: Information Theory (cs.IT); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Low Earth orbit (LEO) satellite constellations are poised to become a cornerstone of the sixth-generation (6G) Internet of Things (IoT), providing truly global coverage and ubiquitous connectivity. This article presents a holistic two-dimensional system architecture for 6G LEO satellite IoT that incorporates composition and functional perspectives to facilitate the seamless integration of LEO satellites and terrestrial networks. Building upon this architecture, we evaluate three pivotal enabling technologies targeting the uplink, downlink, and inter-satellite links (ISLs). Specifically, we analyze massive grant-free random access for efficient uplink connectivity, investigate deep learning-based multibeam precoding for robust downlink transmission, and examine distributed cooperative routing for resilient ISL data delivery. Furthermore, we present an on-orbit verification platform that validates the real-world feasibility and performance of the proposed solutions. Finally, we outline key open challenges and future research directions to guide the realization of future LEO satellite IoT.

[LG-45] A Noise Optimum in Rehearsal-Free Continual Learning: Isolation Mechanism and Scope

链接: https://arxiv.org/abs/2609.20162
作者: Gunner Levi Howe
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Injecting stochastic noise into a consolidation rule can improve a network’s retention of earlier tasks up to an optimal level, then degrade it – an inverted-U in retention vs. noise. This paper isolates what produces that optimum and maps where it holds, entirely in simulation. (1) Phenomenon: the retention inverted-U appears on several related-task continual-learning benchmarks (Split-MNIST, FashionMNIST, continual Yin-Yang). (2) Isolation: a magnitude-matched ladder shows the effect requires coherent restoring toward the consolidated weights – a random-direction force of identical magnitude produces no optimum, and a coherent force toward the wrong target actively hurts. (3) Active ingredient: most of the optimum is recovered by coupling the anchor gain to the injected-noise variance sigma^2 – a one-line rule that neither Ornstein-Uhlenbeck Adaptation (fixed gain) nor MESU (posterior-variance gain) implements. A forced Ornstein-Uhlenbeck calculation derives the rising flank and predicts that the optimal noise rises with per-task interference g – confirmed out-of-sample in direction against pre-existing measurements (the exponent is unresolved at our grid). The barrier-conditioning of the originating Doob h-transform is a low-sigma safety net that bounds forgetting where the coupled gain is too weak. (4) Scope: the optimum requires shared task structure – it is absent on permuted-MNIST, and a controlled rotated-vs-permuted comparison localizes the boundary to task structure; the precise governing quantity is left open. (5) Length: at matched severity the advantage persists but attenuates with task count, and we show no rotation family can attribute the trend (a compact-group identity). A single-seed BrainScaleS-2 demonstration of the originating rule is reported separately (Howe, arXiv:2607.06924); this paper makes no hardware claim.

[LG-46] Fast-varying Natural Frequencies and Damping Ratio Identification for Linear Time-Varying System

链接: https://arxiv.org/abs/2609.20138
作者: Melisa Bozaci,Alice Cicirello
类目: Machine Learning (cs.LG)
*备注: Preprint submitted to Mechanical Systems and Signal Processing

点击查看摘要

Abstract:This work proposes a physics-enhanced machine learning approach for the system identification of Linear Time-Varying (LTV) systems under time-varying operating conditions in terms of fast-varying natural frequencies and damping ratios by combining a long short-term memory network with an Extended Kalman Filter (EKF). The proposed approach uses vibration data (displacement and velocity measurements), domain knowledge of modal damping ratios, and a physics-based model that can yield an approximate natural frequencies time-dependency model. The approach is validated using synthetic data generated from a finite element model of a 2-blade offshore wind turbine under realistic environmental and operating conditions. This system displays fast time-varying frequencies due to operating conditions, whose identification is particularly challenging because of the wind and wave loading. The robustness of the proposed approach is assessed under assumed incorrect system information (e.g. damping ratio). The proposed approach is evaluated across different environmental and operating conditions to show its applicability to different operating regimes. The results show the approach can accurately identify the selected fast-varying natural frequency, 1st Fore-Aft (FA-1) mode, with a maximum root mean square error of 0.0012 Hz. The results demonstrate that the model trained on EKF estimates depends on accurate damping values, whereas the model trained on physics-based data exhibits robustness to incorrect damping assumptions. The approach is extended to damping ratio identification for the selected mode by estimating the root mean square error between models trained on EKF estimates and physics-based data. The results show that the approach can yield a good approximation of the FA-1 mode damping ratio using grid search, offering an improvement over covariance-driven stochastic subspace identification.

[LG-47] QoS-Aware Federated Learning for Multimodal In-Cabin Interaction in Smart Vehicles

链接: https://arxiv.org/abs/2609.20123
作者: Baran Can Gül,Mert Nakıp,Nasser Jazdi,Michael Weyrich
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Modern smart vehicles leverage multimodal sensors, ranging from high-bandwidth vision systems to low-rate physiological monitors, to provide personalized in-cabin services. However, integrating high-fidelity multimodal fusion with collaborative training is often hindered by the heterogeneous and time-varying Quality of Service (QoS) constraints of vehicular networks. Standard Federated Learning (FL) approaches enforce rigid synchronous rounds that fail to account for these resource asymmetries, leading to safety-critical timing violations and energy exhaustion. In this paper, we propose FedQoS, a novel asynchronous, event-triggered FL framework that decouples local computation from global communication via a two-phase gating mechanism. First, we introduce a resource-aware training gate that initializes local learning only when sensing buffers and energy reserves meet safety thresholds, preventing ML tasks from compromising core vehicle mobility. Second, a QoS-aware transmission policy gates uplink updates based on an efficiency score that balances model novelty against instantaneous latency and energy costs. Locally, clients optimize an objective featuring a staleness-aware proximal term that dynamically adjusts the global anchor strength based on update age. Extensive experiments on multimodal vehicular datasets demonstrate that FedQoS achieves competitive personalized accuracy with only marginal performance loss compared to FedAvg, while substantially reducing QoS violations, cutting communication overhead by 76.7%, and lowering latency cost by 26.0%, demonstrating a highly favorable accuracy and efficiency balance for real-world vehicular deployments.

[LG-48] CARE-VI: Conservative Adaptive Reliability Estimation for Value Improvement in Off-Policy Actor-Critic Learning

链接: https://arxiv.org/abs/2609.20098
作者: Xiang Zou,Shengzhu Shi,Junqi Gao,Zhichang Guo
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Reliable temporal-difference targets are central to off-policy actor-critic learning. Direct value improvement refines the next-state target with alternative actions, but the reliability of this refinement depends on how candidate actions are ranked, reviewed, and weighted. Noisy rankings may force premature candidate commitment, reusing selection scores may bias target valuation, and fixed enhancement weights may amplify weak evidence. To address these risks, we develop Conservative Adaptive Ranking and Screening (CARS), which retains an ordered candidate prefix within a preset budget and narrows it only when the observed boundary gap exceeds a disagreement-scaled uncertainty radius. Selector-Evaluator Value Assessment (SEVA) uses selector critics to order candidates and a separately parameterized evaluator critic to review the selected value, then caps the reviewed value at the selector reference. Dynamic Adaptive Risk-aware Enhancement (DARE) then regulates each residual correction using candidate reliability, the gap between selector and evaluator signals, and a finite stage factor. Together, CARS, SEVA, and DARE form CARE-VI, an evidence-regulated target construction framework that preserves the backbone interfaces for critic regression and actor updates. The analysis bounds the CARS boundary error, the SEVA selected-value overestimation, and the one-sided deviation of the DARE residual displacement from its population counterpart, and establishes fixed-policy recovery after the finite-stage perturbation ends. Experiments with SAC, TD3, and TD7 on four MuJoCo tasks show that CARE-VI achieves the highest mean return in all twelve settings. Grouped ablations and scalar diagnostics support the roles of the three components in improving target reliability.

[LG-49] SETTer: Sparse-Encoder Transformer for Long-term Multivariate Time Series Forecasting

链接: https://arxiv.org/abs/2609.20086
作者: Abraham Ezema,Chijioke Eze,Ferdinanda Ponci,Antonello Monti
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Long-term multivariate time series plays a significant role in many application areas such as power systems, trading, etc. However, their accurate prediction is quite difficult for conventional forecasting methods as they often exhibit high dimensionality and complex relationships. Recent works show that transformer-based approaches are quite effective for long-term forecasting thanks to their attention mechanism. However, in the presence of complex high-dimensional inputs, they show evidence of oversmoothing, limited capacity, and opacity. To this end, this paper introduces SETTer, a transformer-based model that addresses these challenges by incorporating novel techniques for decoupled self-attention and hybrid masking. The proposed techniques enable SETTer to effectively capture the dominant short- and long-term patterns across the temporal and channel dimensions. In addition, we enrich the model layers with simple explainable structures that indicate the discriminative pattern of SETTer. We show that with a single-layer transformer architecture, SETTer can effectively model long-term dependencies in the presence of varying data complexities. Extensive experiments on real-word benchmark datasets for long-term multivariate time series forecasting demonstrate that SETTer outperforms state-of-the-art models in 88% of the scenarios.

[LG-50] Evaluating Explanation Methods by the Predictors They Induce

链接: https://arxiv.org/abs/2609.20058
作者: Jacob Selbæk,Hugo L. Hammer
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Explanations of machine learning models are usually judged by criteria that are hard to compare. We propose a simpler test: if an explanation really describes how a model uses its features, it should be possible to rebuild the model’s predictions from it. We turn each explanation into a predictor by reading each feature’s effect and adding them up, and measure how well that predictor reproduces the model on unseen data. Nothing is fitted, so the score reflects the explanation itself. The test applies to any explanation that can be written as a function of the features; we demonstrate it on partial dependence plots (PDP), accumulated local effects (ALE), SHAP and LIME. We prove that summing partial dependence curves gives the best possible additive summary of a model when its features are independent, and that this fails when they are dependent. Across 13 real datasets and 9 synthetic designs and four model families, which method scores best depends entirely on feature dependence: where features are independent SHAP is slightly worse than PDP, exactly as the theory predicts; on dependent real data SHAP leads. Some widely used quality metrics even prefer a damaged explanation to an intact one.

[LG-51] CellRFT: Reinforcement Fine-Tuning for Single-Cell Perturbation Modeling

链接: https://arxiv.org/abs/2609.19970
作者: Jie Yan,Li Liu,Hanze Guo,Jiaxin Hu,Houxin He,Xiaoning Qi,Haoran Wang,Cong Li,Zhong-Yuan Zhang,Yong Wang
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Predicting cellular responses to perturbations supports the study of gene function, disease mechanisms, and therapeutic strategies. Despite advances in single-cell perturbation modeling, existing models typically optimize surrogate losses that do not directly reflect the biological criteria used for evaluation, so better data fitting need not yield better biological predictions. To address this mismatch, we introduce \textbfCellRFT, a reinforcement fine-tuning framework that uses biological evaluation as direct training feedback. CellRFT uses policy-gradient optimization to learn from non-differentiable evaluations of generated cell populations and integrates multiple biological rewards through hierarchical reward aggregation. Comprehensive experiments demonstrate CellRFT’s applicability across different pretrained models and effectiveness in improving perturbation prediction, reveal that optimizing one biological criterion can help or hinder others, and show that complementary rewards can improve criteria beyond those directly optimized, offering a way to probe how biological metrics shape model behavior, with the potential to inform evaluation design. Code will be made available.

[LG-52] Graph-Based Stochastic Power-UCT: Monte-Carlo Graph Search with Power Mean Estimation

链接: https://arxiv.org/abs/2609.19956
作者: Tung Tran,Viet Bao Mai,Hoang Ta,Tuan Dam
类目: Machine Learning (cs.LG)
*备注: No

点击查看摘要

Abstract:Tree-based Monte-Carlo Tree Search (MCTS) duplicates the same state when it is reached through different trajectories, which can waste simulations in stochastic MDPs. We introduce Graph-Based Stochastic-Power-UCT (GS-Power-UCT), which shares states reached at the same planning depth while keeping separate values for states reached at different depths. This design applies to general stochastic MDPs, including problems with cycles. We prove that for a fixed planning horizon, the root estimate converges to the finite-horizon value at rate O(n^-1/2) , matching tree-based Stochastic-Power-UCT while reusing samples across shared states. We also study two full-state variants: GS-Power-UCT-F, which stores one node per physical state to increase sample sharing but may mix values from different remaining horizons, and GS-Power-UCT-F ^+ , which uses an adaptive horizon to control this bias. The latter converges to V^\star(s_0) , the optimal infinite-horizon discounted value at the root state s_0 , when the remaining cross-depth gap vanishes. Experiments on stochastic planning benchmarks show improved sample efficiency over tree-based and graph-based baselines.

[LG-53] One Intervention per Component is Enough: Towards Identifiability in Linear Stochastic Dynamics from Steady State

链接: https://arxiv.org/abs/2609.19955
作者: Saber Salehkaleybar
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We study the problem of recovering the parameters of a multivariate Ornstein-Uhlenbeck (OU) process from steady-state observational and interventional data. In many applications, such as large-scale gene perturbation experiments, only stationary “snapshot” measurements are available, making standard stochastic differential equation estimation methods that rely on time-series trajectories inapplicable. We first establish an identifiability result: one intervention per strongly connected component (SCC) of the drift graph suffices to recover all OU process parameters generically up to a global scaling factor. This holds provided that the SCC condensation graph is connected with a single root and certain spectral nondegeneracy assumptions hold. We propose a recursive learning algorithm that orders SCCs topologically and, for each component, isolates its marginal dynamics and solves a linear system derived from the steady-state moment equations, leveraging parameters recovered for upstream components. Building on this theoretical foundation, we propose a regularized least-squares estimator that jointly minimizes residuals of the steady-state mean and covariance equations across observational and interventional data. Experimental results validate our theoretical findings in recovering parameters of the underlying OU process.

[LG-54] Stringological sequence prediction III: layered ziplines and a tradeoff between efficiency and expressivity

链接: https://arxiv.org/abs/2609.19940
作者: Vanessa Kosoy
类目: Formal Languages and Automata Theory (cs.FL); Data Structures and Algorithms (cs.DS); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:In previous papers, we began the study of sequence prediction algorithms adapted to stringological word complexity measures. In particular, we defined a complexity measure called Arithmetic Repetition Complexity (ARC) which admits a polynomial-time prediction algorithm with a mistake bound quasilinear in the complexity. Here, we show a weaker complexity measure related to ARC that admits an especially efficient prediction algorithm: an algorithm that runs in quasilinear time and polylog space for appropriate highly-structured sequences. The complexity measure is defined via a restricted class of “zipline programs” (a variant of straight-line programs), which we call layered. We thus get a less expressive measure with a more efficient algorithm (compared to our results for ARC), demonstrating a possible tradeoff.

[LG-55] he Life of a Token: from Words to Bits on the Wire

链接: https://arxiv.org/abs/2609.19924
作者: Davide Avesani(CEDRIC - ROC),Pengwenlong Gu(CEDRIC - ROC),Sotiris Skaperas(Cnam),Stefano Secci(CEDRIC - ROC)
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG); Networking and Internet Architecture (cs.NI); Performance (cs.PF)
*备注:

点击查看摘要

Abstract:Large Language Models (LLMs) transform vast collections of unstructured text into semantic patterns used for language generation and reasoning tasks. Behind their ease of use lies a complex process: words become tokens, tokens become vectors, and vectors ultimately give rise to streams of bits that flow through High-Performance Computing (HPC) systems. As modern LLMs grow to billions or trillions of parameters, this path increasingly unfolds across thousands of interconnected accelerators, making the underlying communication fabric a critical and often opaque component of model training. This tutorial aims to walk the reader through the journey from words to network traffic, shedding light on how language is translated into communication flows within HPC training systems. Using concrete examples from Dante’s Divine Comedy, we illustrate how model architecture, tokenization, embeddings, and parallelization strategies shape the volume, structure, and timing of data exchanged across the network. We combine architectural analysis with analytical traffic models and numerical examples to characterize the communication requirements of LLM training. We try to demystify how words travel across the network and provide practical insights into the network requirements needed to support the journey from text to trained model.

[LG-56] Amortizing Physics-Informed Neural Solvers via Graph Hypernetworks

链接: https://arxiv.org/abs/2609.19915
作者: Cheng Jing,Abhishek Verma,Kallol Bera,Yixuan He,Kookjin Lee
类目: Machine Learning (cs.LG)
*备注: Accepted at the Learning on Graphs Conference (LoG), 2026

点击查看摘要

Abstract:Amortizing physics-informed neural networks (PINNs) across related PDEs requires describing each equation to a reusable solver. Coefficient vectors encode numerical parameters in predefined slots, leaving operator and cross-field assignments implicit. We make these relationships explicit in an operator graph, with nodes for fields, derivatives, terms, and residuals and coefficients retained as term attributes. A graph hypernetwork generates diagonal codes that initialize a meta-trained factorized PINN for each target equation. Meta-training and target-specific adaptation use governing equations and prescribed conditions without solution labels. We compare coefficient-vector, DeepSets-based term-set, and graph conditioning by solution accuracy within a fixed adaptation budget. In scalar convection-diffusion-reaction problems, both term-based descriptors improve high-reaction accuracy, with similar performance. In two-field Fisher-KPP, meta-training sees uncoupled and one-way systems; after 3,000 adaptation steps on unseen two-way coupling, the graph’s mean final error is 35.7% below the term set and 67.7% below the coefficient vector. In a fixed-structure capacitively coupled plasma model, the coefficient vector performs best. These results support extending coefficient conditioning with explicit equation relationships for physics-based solver adaptation.

[LG-57] Digital Twins for Opinion Dynamics: A Generative LLM Framework for Social Networks

链接: https://arxiv.org/abs/2609.19913
作者: Omran Berjawi,Giuseppe Fenza,Rida Khatoun,Sherali Zeadally
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:The study of opinion dynamics in social networks is one of the key challenges in computational social science with direct relevance to understanding political polarization, misinformation, and health responses. Current approaches focus on simplified mathematical models that ignore linguistic and contextual factors related to belief updates or use Large Language Model (LLM)-based simulations that have not been validated against real data. We present a framework based on the concept of a digital twin to simulate opinion dynamics in social networks. The approach fills the gap by cloning a real-world Twitter network, assigns a set of attributes for agents (such as persona, emotions, centrality, stubbornness, and influence), and employs Mistral-7B to perform opinion update based on memory and social exposure. To evaluate the proposed approach, we validate it against two real Twitter datasets (COVID-19 discourse and U.S elections 2020). The results show that the capability of the proposed framework reproduces opinion trajectories and reduces individual prediction error by more than 50% compared to the best-performing classical baseline (Mistral-7B achieves Mean Absolute Error (MAE) = 0.150 and 0.121 on the COVID-19 and US Election 2020 datasets, respectively). We observe similar improvements in structural alignment (Delta_r = 0.120 and 0.180) and polarization dynamics (Delta_Var = 0.106 and 0.115) on the two datasets, respectively. Additionally, the ablation studies confirm that agent attributes, memory, and social exposure all contribute to the framework’s predictive fidelity in reproducing opinion trajectories, with agent attributes being the most critical contributor. Overall, our results demonstrate that grounding Mistral-7B within empirically cloned interaction networks produces a realistic simulation framework capable of reproducing complex social dynamics.

[LG-58] REARL: A Closed-loop Autonomous Driving Simulation Enhancement Framework with Real Traffic Data and Large Language Models

链接: https://arxiv.org/abs/2609.19903
作者: Xiaojun Bi(1),Jun Jiang(1),Yiwen Sun(2 and 3),Quanyi Ou(1),Ke Cheng(4),Mingjie Bi(3),Yexin Li(3) ((1) Minzu University of China, Beijing, China, (2) Peking University, Beijing, China, (3) BIGAI, Beijing, China, (4) Beihang University, Beijing, China)
类目: Machine Learning (cs.LG)
*备注: 14 pages, 8 figures, 3 tables. Corresponding author: Yiwen Sun. This work was supported by the National Natural Science Foundation of China (Grant No. 62503015)

点击查看摘要

Abstract:Accurate simulation is crucial for autonomous driving development, yet capturing real-world traffic complexity remains challenging. Existing simulators that rely on predefined rules or static data playback struggle with dynamic traffic. CRITICAL uses real traffic data and a large language model (LLM) to adjust the initial simulation configuration, but the simulated distribution still diverges from real traffic as the rollout evolves. We propose REARL, a closed-loop simulation enhancement framework that integrates real traffic data with LLMs. Real traffic data are clustered, and each cluster center is used as a representative scenario that provides typical real-world traffic patterns for the LLM. A timed sliding-window detector then monitors discrepancies in vehicle speed distribution and mean spacing between pairs of vehicles. If a metric exceeds a threshold, the LLM adjusts vehicle decision-making; otherwise the existing controller is kept. The LLM also selects a matching real vehicle from a traffic snapshot and modulates the simulated vehicle with reference to that real action. In a controlled HighD highway setting, compared with the CRITICAL baseline and a PPO-based learning baseline, REARL reduces the Hellinger distance for speed distributions to 0.3067 and the MAPE for mean spacing to 0.8371, while achieving a time headway (THW) of 22.8575 and a lane change rate of 0.0708.

[LG-59] Delphi Scanner: efficient and interpretable static malware detection via API sequence modeling

链接: https://arxiv.org/abs/2609.19900
作者: Bijied Brahimi,Vincent Cohadon,Gabriel Glazman,Rayan Al Mohaize,Omran Berjawi,Rida Khatoun
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Static malware detection for Windows Portable Executable files demands a careful balance between detection effectiveness, computational efficiency, and analytical interpretability. This paper introduces Delphi Scanner, a static malware detection system for Windows PE files that balances efficiency with behavioral interpretation. It uses a convolutional neural network (CNN) to model Windows API sequences to classify PE and a decoupled interpretation layer based on a rule-based layer to categorize APIs into high-level malicious capabilities. Evaluated on over 190,000 Windows PE files, the system achieves 95.35% accuracy with a 1.53~MB model footprint. Robustness experiments on 5,647 out-of-distribution MalwareBazaar samples, paired packed and unpacked executables, and three adversarial manipulation strategies confirm generalization beyond the training distribution and resistance to functionality-preserving evasion techniques. Overall, these results demonstrate that API sequence-based static analysis offers a practical, interpretable, and efficient foundation for malware triage in local deployment scenarios.

[LG-60] Online Adaptive Kernel Mixing for Gaussian Process Decision Making

链接: https://arxiv.org/abs/2609.19891
作者: Kavin Aravindan,Mani Tej Sriram,Gautam Dasarathy,Tejas Bodas
类目: Machine Learning (cs.LG)
*备注: 35 pages, 9 figures. Accepted as a full paper at IFIP Performance 2026

点击查看摘要

Abstract:Gaussian Processes (GPs) are widely used as surrogates for black-box functions in sequential decision-making problems such as Bayesian optimization (BO), level set estimation (LSE), and Bayesian active learning (BAL). GP performance critically depends on kernels, and standard kernels can lead to suboptimal decisions under misspecification. To address this, we introduce HACK GPs (Hedge Adaptive Cumulative Kernels), a method that views kernel selection as an online learning with expert advice problem. HACK treats each candidate kernel as a GP “expert” and updates a distribution over experts online using AdaHedge, based on a loss received as a proxy for their ability to fit the function and align with the task objective. We provide two variants of HACK: (i) Mixture of Gaussians (MoG) and (ii) categorical sampling. We establish general guarantees showing that, under a loss-gap condition, the weight concentrates on the best kernel and the resulting acquisition function is close to that of the best expert. Empirically, we observe robust performance across BO, LSE, and BAL compared to standard kernels such as Squared Exponential and Matern-5/2, as well as simple ensemble baselines.

[LG-61] AURA: Adaptive Uncertainty-Routed Analysis for Email Threat Detection

链接: https://arxiv.org/abs/2609.19873
作者: Omran Berjawi,Walid fahs,Rida Khatoun
类目: Machine Learning (cs.LG)
*备注: Under review

点击查看摘要

Abstract:Email spam and phishing attacks remain a critical security threat. Adversaries increasingly exploit large language models to craft contextually convincing malicious messages, and existing spam detection systems often struggle to keep pace. Generalization across diverse and evolving attack scenarios is limited, which reduces effectiveness once these systems are deployed in practice. This paper introduces Adaptive Uncertainty-Routed Analysis (AURA), a multimodal email threat detection system that analyzes both the content of an email and its embedded URLs. AURA is built around two layers: the first quantifies prediction uncertainty from a URL classifier, and only ambiguous messages are escalated to a fine-tuned transformer encoder for semantic analysis. The system is evaluated on eight heterogeneous training corpora together with two held-out real-world corpora spanning a decade of adversarial campaigns. AURA reaches a macro F1-score of 0.9858 in-distribution, and on NazPhish-Eval and GuenterTrap-Eval it maintains 0.9502 and 0.9436, respectively, which is evidence of robust generalization under genuine distribution shift.

[LG-62] Pretrained Medical Representations for the Practical Screening of Drug Repositioning Candidates ICML2026

链接: https://arxiv.org/abs/2609.19865
作者: Yuhei Fujioka,Daitaro Misawa,Shingo Fukuma
类目: Machine Learning (cs.LG)
*备注: Accepted at ICML 2026 AI for Science Workshop

点击查看摘要

Abstract:Representation learning from medical code sequences in electronic health records and medical claims data has been successful in various clinical applications, such as those regarding disease prediction. However, significant challenges remain in extending this approach to the discovery of scientific hypotheses. One reason is that many existing BERT-based models fail to adequately capture the hierarchical structure of medical codes and the complex interactions between diagnoses and treatments. To address these limitations, we propose a new unified pre-training framework that explicitly integrates hierarchical sub-token aggregation, partial masking, and cross-reference mechanisms. The proposed model consistently outperformed existing methods on both pre-training objectives and downstream clinical event prediction tasks, including the onset of dementia and hospitalization. We also conducted an in silico drug repositioning case study targeting Alzheimer’s disease. In the hypothesis generation step, our approach successfully rediscovered known promising drugs in a data-driven manner without relying on such external knowledge sources as the literature. Subsequently, in the hypothesis prioritization step, we introduced a Task-Adaptive Representation Approach to alleviate the over-encoding of historical prescription information within diagnostic vectors, enabling the robust prioritization of generated hypotheses. This study establishes an exploratory screening workflow for hypothesis generation and prioritization based on observational associations. Importantly, this framework is not intended to provide causal evidence, but rather to identify promising candidates for subsequent rigorous causal inference. Overall, this study demonstrates that domain-informed representation learning combined with task-adaptive representation control can enable a practical hypothesis discovery workflow.

[LG-63] Expected Hypervolume Maximization for Multiobjective Optimization under Uncertainties

链接: https://arxiv.org/abs/2609.19858
作者: Victor Trappler(Mines Saint-Étienne MSE, LIMOS, FAYOL-ENSMSE, FAYOL-ENSMSE)
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:The problem of multiobjective optimization under uncertainties is often approached by taking the expectation of each objective. In this work, we propose instead to formulate this as a Bayesian decision problem and to rely on the expected value of the hypervolume, which is to be maximized with respect to a finite set of input points. We show that this can be performed using methods based on gradients in a stochastic optimization framework, provided that care is taken with respect to dominated points. Moreover, in the absence of readily available differentiable code, we propose to use Gaussian Processes as differentiable surrogate models, in order to perform the optimization. An additional contribution in this work are some active learning strategies, through acquisition functions which helps construct a surrogate model well-designed for the multiobjective optimization problem at stake. These strategies are compared on simple analytical problems to assess their performances.

[LG-64] Beyond Flattened Tokens: Structure-Preserving EEG Decoding with Reusable TriDim Blocks

链接: https://arxiv.org/abs/2609.19842
作者: Shiyue Su,Song Wang,Zekai Zhan,Junjie Zeng,Ziling Lu,Zongsheng Li,Xinyuan Ye,Zhiyuan Ma,Xinke Shen,Quanying Liu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Effective EEG decoding requires representations that preserve organization among channels, local waveform dynamics, and long-range temporal context. Existing EEG architectures often capture these structures using separate specialized modules or collapse them into a single token sequence, making it difficult to maintain their distinct roles and coordinate their interactions throughout the backbone. We propose TriDim, a reusable block that preserves the representation shape and keeps three EEG axes explicit: channel, sample position within each patch, and patch position across the recording. These axes correspond to spatial, short-term temporal, and long-term temporal information, respectively. Each TriDim block applies feed-forward transformations along individual axes and cross-axis attention to coordinate information exchange among them. By stacking TriDim blocks with a multi-level tri-axis readout, we construct TriDimEEG, a standalone EEG decoder. Under strict cross-subject evaluation on eight datasets spanning clinical diagnosis, sleep staging, motor imagery, and emotion recognition, TriDimEEG achieves the best overall performance among fifteen evaluated models, with a 4.3% relative improvement in average accuracy over the second-best model. Replacing Transformer blocks in three EEG foundation models with TriDim blocks yields an average relative improvement of 7.4% in downstream accuracy while reducing parameter counts by 17.0% to 47.3%. These results establish TriDim as an effective and reusable building block and TriDimEEG as a strong standalone EEG decoder. Code and parameters of TriDimEEG are available at this https URL.

[LG-65] DeliveryGym: An RL Environment for Long-Horizon Embodied Agent Planning with Adaptive Curriculum

链接: https://arxiv.org/abs/2609.19801
作者: Haoqiang Kang,Yiming Zhang,Yiyang Guo,Chuying Li,Jianzhi Shen,Tianruo Rose Xu,Xiaokang Ye,Lianhui Qin
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Executable environments enable LLM agents to learn from the consequences of their actions. For embodied agents, those consequences extend beyond whether the current task succeeds: completing a delivery can consume the time, energy, or money needed for later work. Learning to plan therefore requires environments that preserve these dependencies and turn them into feedback across a complete trajectory. We introduce DeliveryGym, a 3D environment for evaluating and training agents on continuous courier shifts. It couples multimodal tool interaction with persistent world dynamics and computes trajectory rewards from simulator events, making the costs of an agent’s decisions available for reinforcement learning (RL). The environment also adapts future training shifts to the policy’s observed weaknesses while keeping evaluation fixed. Across six models and 13 city maps, evaluation exposes a gap between reliably executing assigned deliveries and choosing and sequencing work over a shift. On the fixed test suite, RL improves Qwen3-VL-4B’s net income by 54.3%, showing that learning from complete shifts improves performance under these coupled constraints. Adapting the training environment improves test income by 16.5% over uniform sampling at the same rollout budget, indicating that which situations an agent practices also matters. DeliveryGym provides an executable setting for studying how agents learn to coordinate deliveries and preserve resources for later orders within an episode.

[LG-66] Learning-Based Reconstruction of Optical Properties in Bilayered Media from Single-distance Time-Resolved Reflectance Measurements

链接: https://arxiv.org/abs/2609.19786
作者: Caterina Amendola,Giulia Maffeis,Lorenzo Buffoni,Lorenzo Chicchi,Francesco Coghi,Duccio Fanelli,Raffaele Marino,Fabrizio Martelli,Riccardo Paoli,Lorenzo Pattelli,Lorenzo Spinelli
类目: Machine Learning (cs.LG); Optics (physics.optics)
*备注:

点击查看摘要

Abstract:The inverse problem of reconstructing optical properties, specifically absorption and scattering coefficients, in layered biological media from time-domain reflectance measurements remains a significant challenge for traditional analytical models. Inverse solvers based on the diffusion equation often struggle with structural heterogeneity, frequently yielding poor accuracy for superficial absorption and deep-layers scattering. In this work, we propose a machine learning framework as an alternative approach to reconstruct the optical properties of a bilayered medium, benchmarking its efficiency and accuracy against model-based algorithms. To overcome the intrinsic approximations of diffusion theory and inverse reconstruction, we generated a robust synthetic dataset of forward DTOF using exact Monte Carlo simulations at multiple source-detector distances. A machine learning pipeline was then trained on this dataset and validated against state-of-the-art model-based reconstruction methods. Besides the significant reconstruction speed-up, the machine learning approach achieves higher accuracy than model-based inverse solvers, further providing an estimate of the parameter space dimensionality without requiring any a priori information about the number of layers in the investigated geometry. Further enhancements in the reconstruction accuracy can be expected in future extensions of this work, by training the pipeline over multiple DTOF curves from the same medium, in a joint multi-distance reconstruction approach.

[LG-67] PhyRestore: Physics-Structured Latent-Factor Restoration

链接: https://arxiv.org/abs/2609.19776
作者: Ahmed Shafee,Chayan Lahiri
类目: Machine Learning (cs.LG)
*备注: 12 pages, 2 figures, 2 tables

点击查看摘要

Abstract:Estimating temporal soil-loss change is challenging when physically meaningful input factors are noisy or corrupted, particularly because substantial changes are rare relative to the large number of locations exhibiting little change. We study this problem through the Revised Universal Soil Loss Equation (RUSLE) and introduce PhyRestore, a physics-structured latent-factor restoration framework. Rather than directly predicting soil-loss change or correcting a degraded physical estimate, PhyRestore restores corrupted physical factors and reconstructs temporal change through the known physical relationship. We evaluate PhyRestore in a watershed-scale bitemporal raster setting under isolated and simultaneous corruption of rainfall erosivity and cover management, comparing it with the degraded RUSLE estimate and Direct RF, XGBoost, MLP, and CNN models. Factor restoration improves high-magnitude recovery when the corrupted factors remain identifiable, but its advantage weakens under joint corruption, sparse positive extremes, and factor values outside the training support.

[LG-68] OceanMoE: Structured Conditional Sparse Computation for Long-Horizon Multivariate Ocean Forecasting

链接: https://arxiv.org/abs/2609.19768
作者: Yishun Zhu,Jian Wang
类目: Machine Learning (cs.LG)
*备注: 15 pages, 4 figures. Yishun Zhu and Jian Wang contributed equally to this research

点击查看摘要

Abstract:Multivariate ocean forecasting must exploit shared evolution in a coupled ocean system while adapting to the heterogeneous statistical and dynamical characteristics of different prediction variables and locations. Fully shared models may lack the flexibility to handle this heterogeneity, whereas fully independent models discard the common ocean context shared across variables. The key question is how to retain shared context in a unified model while allowing computation to specialize according to the prediction target and local state. We propose OceanMoE, a structured conditional sparse Mixture-of-Experts framework that combines sharing and specialization for multivariate ocean forecasting. OceanMoE fuses cross-variable information to construct target-specific local representations and uses them to perform content-conditioned sparse routing at each spatial location, with the number of active experts adapted to router confidence. In the decoder, routing is augmented with a learned geographic bias parameterized by spherical-harmonic spatial bases, while shared residual and seasonal pathways provide common cross-variable and month-dependent context. Experiments on long-horizon autoregressive ORAS5 forecasting show that OceanMoE lowers aggregate forecasting error in both evaluated settings and maintains lower geometric-mean normalized RMSE than the corresponding baselines over most later rollout months. Routing analyses further show that expert allocation varies with prediction targets and spatial locations. These results support structured conditional computation as a modeling strategy for balancing shared ocean context with adaptive specialization.

[LG-69] Alliance Beats Isolation: Unifying Heterogeneous Allied Datasets Improves Classifier Performance

链接: https://arxiv.org/abs/2609.19748
作者: Girish Keshav Palshikar
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:In many application domains, such as student dropout, insurance fraud, loan approval, and machine failures, several labelled public datasets are available where (i) data is about the same type of objects but the set of actual underlying objects are disjoint; and (ii) the class labels are same; and (iii) the feature spaces of the datasets are largely distinct (heterogeneous), with a few shared features. We call such datasets as allied. A single classifier cannot be trained on both datasets together, and one classifier trained on one dataset cannot be tested on the other. In this paper, we propose a method to merge the feature-spaces into a single feature-space for a pair of given allied heterogeneous datasets. We then use a matrix completion method to create a unified dataset based on the merged feature-space. The hypothesis is that the merged representation facilitates the transfer of classification knowledge from one dataset to another. We conduct experiments on several pairs of allied, heterogeneous datasets and several classifiers to demonstrate that any classifier trained on the unified representation always outperforms classifiers separately trained on the constituent allied datasets on several pairs of allied datasets. This work provides an easy way to substantially improve classifier performance by unifying and using multiple allied datasets together.

[LG-70] ALIBI: Adversarial Legitimacy Injection in Binary Input against LLM Malware Analyzers

链接: https://arxiv.org/abs/2609.19722
作者: Hyeongjun Choi,Wonyoung Jung,Haehoon Seo,Sungyup Nam
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注: 12 pages, 4 figures, 4 tables

点击查看摘要

Abstract:Large language models are being integrated into malware triage workflows as reasoning components that summarize static evidence and produce analyst-facing verdicts. This paper shows that the same reasoning capability introduces a new attack surface. We present ALIBI, a semantic cover story attack against frontier LLM-based malware analyzers. ALIBI adds a small, non-executed read-only section to a compiled binary, containing a coherent but false security product narrative, without altering imports or executable behavior. Instead of issuing direct instructions to the model, it reframes suspicious evidence as expected behavior of a benign endpoint security tool. On a frozen PE set of 50 malicious samples, the payload flips 30 of the 35 baseline-malicious samples to benign on Gemini 2.5 Pro, while GPT-5.5 Pro and Claude Opus 4.7 produce substantial severity downgrades with significant confidence reductions even when verdict labels are preserved. The attack transfers to ELF binaries, where Gemini flips 16 of 40. A verification-guided defense prompt roughly halves the benign verdicts, but 42.9 percent of malicious samples still reach benign. LLM malware analyzers therefore require provenance checks that separate verified facts from attacker-controlled claims, not narrative trust.

[LG-71] Odds-Ratio Thompson Sampling: A Specification and Design Guide for Contrast-Based Multi-Armed Bandits

链接: https://arxiv.org/abs/2609.19709
作者: Sulgi Kim
类目: Machine Learning (cs.LG)
*备注: 25 pages, 7 figures, 2 tables in the main text; 2 figures and 3 tables in the supplement. Reference implementation at this https URL

点击查看摘要

Abstract:Batched multi-armed bandits update on a service’s own schedule, and the usual implementation carries each arm’s absolute reward rate from one update to the next. When the shared level moves between batches, that memory goes stale even though the comparisons between arms may not have. Odds-Ratio Thompson Sampling (OR-TS) instead carries the joint posterior over log-odds contrasts and fits the common level afresh in every batch, marginalizing it out. This paper specifies that update, places it inside a Bayesian bandit agent with two controls, decay for how much past evidence survives an update and aggressiveness for how sharply belief becomes allocation, and evaluates it against absolute-rate memory. Across 86 public A/B series the level varies about twenty-five times more than the contrast. In prespecified synthetic environments a moving level costs absolute-rate memory five times the regret and leaves the best arm below a majority of traffic in 7 of 20 runs, against none for OR-TS. In a policy simulation built from 71 real experiments, where the contrasts are too small to resolve, expected-click differences stay within 0.1% for 58 of them, yet contrast memory still ends on the better arm more than twice as often. Where the contrasts themselves move, the bet fails, and that case is reported too.

[LG-72] Opinion Dynamics-based Coalition Formation for Federated Learning in Heterogeneous IoT Systems

链接: https://arxiv.org/abs/2609.19695
作者: Mohammed El Hanjri,Anas Abouaomar,Hamidou Tembine,Abdellatif Kobbane
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Federated learning (FL) enables privacy-preserving, on-device training across heterogeneous Internet-of-Things (IoT) deployments such as smart-city water-metering networks, where each smart meter observes a household-specific consumption time series. Under such statistical heterogeneity, the standard Federated Averaging (FedAvg) aggregation averages dissimilar local models into a single global model that may fail to capture client-specific patterns. We address this by forming client coalitions directly in the local-weight space and aggregating at the coalition level. Extending a prior weight-driven coalition-formation scheme, we model coalition formation as a Hegselmann-Krause (HK) bounded-confidence opinion-dynamics process acting on the local weights, and develop variants of the HK interaction based on Euclidean-distance and cosine-similarity confidence criteria. The framework is applied to short-term water-consumption forecasting with local Long Short-Term Memory (LSTM) models and evaluated against FedAvg, Per-FedAvg, FedProx, and FedAvg with Euclidean-distance or cosine-similarity coalition formation. Experiments on a real smart-metering dataset of water consumption show that the proposed HK-based coalition formation produces stable, endogenous coalition structures within at most ten inner iterations, incurs no additional client-side computation or communication compared to FedAvg, and reduces the average MAE by up to 54% relative to FedAvg, 39% relative to FedProx, and 24% relative to Per-FedAvg, while achieving the highest global accuracy (83-85%).

[LG-73] UniExo: Unified Multi-Skill Policies for Musculoskeletal Locomotion and Co-Adaptive Exoskeleton Control

链接: https://arxiv.org/abs/2609.19690
作者: Yifei Yuan,Jakob Wolf,Ghaith Androwis,Xianlian Zhou
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注: 9 pages, 8 figures

点击查看摘要

Abstract:Daily locomotion encompasses diverse activities and frequent transitions between them, yet most exoskeleton controllers are designed for a single activity or a narrow set of related movements. Changes in activity therefore typically require explicit mode switching and separately tuned or retrained controllers. Simulation-based learning reduces the need for hardware-based tuning but generally retains this limitation. Here we present UniExo, a framework that first constructs a multi-skill musculoskeletal human policy and then jointly trains an exoskeleton control policy with it. Four single-skill imitation experts for walking, turning, running and backward walking are distilled into a single network structured by a skill latent and subsequently fine-tuned through reinforcement learning on transition sequences. The resultant unified human policy achieves a mean tracking success rate of 94.7% on unseen clips of the four skills and exhibits greater robustness to perturbations than its constituent experts. A single hip exoskeleton controller (UniExo) is initialized from hip moment prediction of the human policy and co-adapted with it through multi-agent reinforcement learning across the four skills. This co-adaptation shifts the timing of the assistance torque and raises the fraction of positive work delivered to the hip. When deployed on a custom hip exoskeleton, the controller generalizes across four treadmill speeds in six participants and assists one participant through a continuous route of all four skills and their transitions, without skill labels or explicit mode switching. UniExo thus provides a step towards replacing activity-specific controllers with unified, user-specific controllers that support diverse locomotor activities and the transitions between them.

[LG-74] Conservation Buys Stability and Factoring Buys Counterfactuals in Physical World Models

链接: https://arxiv.org/abs/2609.19674
作者: Yufeng Wang,Parivesh Priye,Lu Wei,Haibin Ling
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:A learned simulator can reproduce its training conditions accurately yet fail in two distinct ways once those conditions change. Over long rollouts, small errors accumulate until the trajectory drifts away from physically plausible behavior; under an intervention on a physical parameter, the model may continue to follow the law seen during training rather than the intervened one. We show that these two failures require different structural remedies. Evolving a learned energy with a symplectic integrator preserves the geometry of the conservative dynamics and keeps rollouts bounded and physically meaningful for up to 100\times the training horizon, while equal-capacity predictors, an energy-regularized predictor, and a tuned neural ODE diverge. By contrast, encoding the physical coupling through an explicit linear factorization enables the model to follow a never-seen sign of that coupling, whereas an unrestricted parameterization remains locked to the training law. Crucially, the two mechanisms are separable: removing the structure responsible for long-horizon stability leaves counterfactual transfer intact, while removing the factorized coupling destroys counterfactual transfer without eliminating stability. This double dissociation, established with matched controls that remove or replace one structural component at a time, persists beyond the headline three-body system and remains visible when the physical state must be inferred from pixels rather than provided directly. The result is a concrete design principle for physical world models: long-horizon stability and changed-law generalization arise from distinct structural commitments, and each can be imposed deliberately without requiring the other.

[LG-75] CoRe: Coherence and Relational Alignment for Multivariate Time Series Forecasting ICONIP2026

链接: https://arxiv.org/abs/2609.19670
作者: Xiaoyu Lin,Huiran Duan,Yining Liu,Zhixiang Wu,Chu Lin,Lin Lu
类目: Machine Learning (cs.LG)
*备注: Accepted at the International Conference on Neural Information Processing (ICONIP 2026)

点击查看摘要

Abstract:Direct forecasting has become a standard paradigm for multivariate time-series forecasting because it predicts the full future horizon in a single pass. However, its training objective is often still decomposed into pointwise errors such as MSE. Such objectives provide stable supervision, but they do not explicitly preserve the structure of the future trajectory: temporal coherence within each variable and relational consistency across variables can both be weakened. We propose CoRe, a model-agnostic learning objective for direct multivariate forecasting. CoRe replaces pointwise supervision with two output-space constraints: a frequency coherence loss that aligns predicted and target spectra, and a low-rank relational graph loss that matches sampled pairwise differences in a target-derived PCA subspace. The resulting objective introduces no trainable parameters and can be applied to existing forecasting backbones by changing only the loss. Experiments on standard benchmarks show that CoRe improves strong baselines, compares favorably with recent forecasting objectives, and remains effective across different backbones, datasets, and hyperparameter settings overall consistently.

[LG-76] EmbodiedMind: Adaptive Data Curation and Prefix-Tree Reinforcement Learning for Efficient Embodied Intelligence

链接: https://arxiv.org/abs/2609.19659
作者: Feifan Wang,Zongbing Zhang,Yu Zhang,Lingfeng Wang,Yurui Zhu,Jin Deng,Mingliang Zhang,Zhengguang Gao,Yongcheng Wang,Jin Xu,Ri Yang
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Training embodied foundation models typically requires massive-scale datasets and extensive computational resources, yet often suffers from three critical limitations: (1) inefficient sample utilization due to low-informative samples; (2) imbalanced gradient contributions across heterogeneous tasks; and (3) severe credit assignment problem in long-horizon planning, where trajectory-level rewards indiscriminately penalize all tokens. To address these issues, we propose an efficient training paradigm that achieves state-of-the-art average performance through strategic data selection and hierarchical policy optimization. Our approach consists of three synergistic stages. First, Rejection Sampling-based Fine-Tuning (RSFT) filters out low-informative samples to establish robust behavioral priors while preventing distributional collapse. Second, Iterative Rejection GRPO (IR-GRPO) employs task-specific queues stratified by difficulty to keep datasets balanced across reinforcement learning iterations, coupled with a hybrid reward mechanism for precise cross-task feedback. Third, to enhance long-horizon task planning, we introduce Trie-GRPO, a novel reinforcement learning algorithm based on action prefix trees, which enables step-level advantage estimation. This resolves the credit assignment problem by isolating intermediate correct decisions from downstream errors, while effectively balancing exploration efficiency and depth compared to conventional search trees. As a result, EmbodiedMind achieves a state-of-the-art average performance of 70.02% across 18 benchmarks, and significantly outperforms other embodied foundation models in long-horizon task planning accuracy. Our project will be released for reproducibility.

[LG-77] PrefixBench-H100: Characterizing Prefix Reuse and Time-to-First-Token in H100 LLM Serving

链接: https://arxiv.org/abs/2609.19657
作者: Omkar Shewale,Deepak Kumar,Divakar Kumar Yadav
类目: Performance (cs.PF); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Repeated prompt prefixes are increasingly common in LLM serving workloads, appearing in system prompts, templated retrieval-augmented generation pipelines, agent frameworks, and multi-turn conversations. Modern inference runtimes such as vLLM and TensorRT-LLM provide mechanisms for reusing previously computed KV-cache state across requests, yet it remains unclear when prefix reuse materially improves serving performance on contemporary accelerators and when its benefits are limited by scheduling, cache granularity, concurrency, or memory pressure. This paper presents PrefixBench-H100, a reproducible benchmark and measurement framework for characterizing prefix reuse on a single NVIDIA H100. PrefixBench-H100 combines controlled synthetic traces with chat-style and retrieval-style workloads, and evaluates two widely used LLM serving runtimes under matched workload conditions. The benchmark varies shared-prefix length, suffix diversity, request arrival pattern, concurrency, output length, and cache configuration, while collecting time-to-first-token, inter-token latency, end-to-end latency, throughput, cache-hit statistics, GPU memory usage, and selected profiling traces. The goal of PrefixBench-H100 is not to introduce a new caching algorithm, but to expose the practical operating envelope of prefix reuse for H100-class LLM serving. The study identifies the regime where prefix reuse provides substantial first-token latency reductions and the regime where cache pressure erodes them, while showing that cache effectiveness itself is largely insensitive to concurrency and output length; the cross-runtime differences that remain arise above the cache, in the scheduling layer. Subjects: Performance (cs.PF); Machine Learning (cs.LG) Cite as: arXiv:2609.19657 [cs.PF] (or arXiv:2609.19657v1 [cs.PF] for this version) https://doi.org/10.48550/arXiv.2609.19657 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-78] A Policy Profile for Croissant: Refusal as a Property of the Dataset

链接: https://arxiv.org/abs/2609.19640
作者: Alexander Chernov
类目: Machine Learning (cs.LG); Cryptography and Security (cs.CR); Databases (cs.DB)
*备注: 23 pages. Reference implementation and conformance corpus archived at doi: https://doi.org/10.5281/zenodo.22018156 and doi: https://doi.org/10.5281/zenodo.22016112

点击查看摘要

Abstract:Croissant is the de facto machine-readable descriptor for ML datasets: JSON-LD over this http URL. Since version 1.1 it also carries data use conditions, recommending DUO and ODRL for them. What no version specifies is how any of them is evaluated: no decision procedure, no bound on evaluation cost, no outcome for a condition an implementation cannot evaluate, no record of what was checked, and nothing on composition with caller-side authority. We supply that half. An additive profile lets a dataset declare the operations it admits and the conditions under which it admits them, over a closed set of five operators whose decision procedure is given in full, so a gate decides from the descriptor alone and records what it checked. Two corpora evaluate it and their evidence is kept apart. Three descriptors that gated a real nf-core pipeline give the deployment result: decisions from a profile document match the gate’s native descriptor record for record, stripping the layer leaves a valid Croissant document, and the added cost is 11.7 \mu s against a 119 \mu s decision. A corpus generated from the profile’s grammar gives the breadth, covering every operator, refusal class and conformance clause. Across its valid cases, 552 complete decision records agree three ways – native descriptor, profile terms, and the same policy as ODRL in usageInfo. The carrier is therefore not the contribution; the evaluation semantics is. Finally, caller-bound and data-bound policies range over non-overlapping state spaces, so neither permit set contains the other.

[LG-79] he Complexity Kink: A Prompt-Side Structural Complexity Index for Code-Generation Reliability

链接: https://arxiv.org/abs/2609.19616
作者: Michael Hernandez,Tian Zhao
类目: Machine Learning (cs.LG); Software Engineering (cs.SE)
*备注:

点击查看摘要

Abstract:Complexity measured from generated code is failure-dependent: a difficult prompt can yield a short failing program and be assigned low output complexity. We introduce a six-dimension prompt-side structural-complexity index scored before generation and kept separate from correctness. We select 5,000 Python prompts across six bands of a preliminary single-rater rubric. Four out-of-panel LLM raters rescore the locked prompts, giving 19,997 score rows; composite inter-rater reliability is ICC = 0.872 on the 4,998 prompts with all four ratings. We evaluate 21 models per prompt, yielding 105,000 generations. In the unadjusted mean-pooled analysis, pass rate has a nonmonotone breakpoint at composite 13.75, with 79.9% at or below and 87.6% above. This is not a universal failure cutoff. Task-type fixed effects shift the breakpoint to 10.75 and cut the regime gap from 7.6 to 2.1 points. A construction-frame control shifts it to 8.50 with a raw gap of -3.5 points, and neither frame alone reproduces the pooled +7.6-point change. Model-specific fits include 16 upward and five downward changes. A 365-prompt audit-clean extension matches the original five-model estimates at bins 15 and 16 but adds only 14 prompts above bin 16. Among zero-pass generations with computable Lizard complexity, 28.5% pair a prompt composite above 8 with output complexity at most 10. Human agreement is moderate and rater-dependent on a disagreement-enriched calibration set; paraphrase and cross-language rescoring preserve score ordering. Overidentification tests reject the joint restrictions on the six dimensions, so we treat the composite as an index and make no causal interpretation of the 2SLS estimates. The contribution is a pre-generation measurement framework and a bounded observational analysis of reliability regimes.

[LG-80] he Output-Space Hypothesis: Enumerative Equivalence Checking for Tensor Programs

链接: https://arxiv.org/abs/2609.19611
作者: Paul Biberstein,Joseph Devietti,Mayur Naik
类目: Programming Languages (cs.PL); Machine Learning (cs.LG); Software Engineering (cs.SE)
*备注:

点击查看摘要

Abstract:Tensor programs, as used in deep learning models, are a prime target for optimization, as small performance improvements can have a large impact across training or inference workloads. However, such optimizations are complicated and can produce subtle bugs. Traditionally, correctness is assumed when differential testing against a reference on random inputs fails to reveal bugs. However, the inputs to these programs are massive tensors, and finding bugs can require generating extremely low likelihood inputs with precise relationships among their values. We propose a novel way to find bugs more consistently by flipping the quantifiers. Rather than generating a single input and checking all output tensor locations for equivalence, what if you could check a single output tensor location’s equivalence for all inputs? We implement this idea in a system, \dirigo, by using a novel symbolic execution strategy. We demonstrate that \dirigo can find bugs effectively in a public dataset of 6,988 AI-written CUDA kernels that are all marked correct by differential testing. Of these, \dirigo finds 600 kernels that are actually buggy, and finds 97.3% of those bugs within two minutes. Subjects: Programming Languages (cs.PL); Machine Learning (cs.LG); Software Engineering (cs.SE) Cite as: arXiv:2609.19611 [cs.PL] (or arXiv:2609.19611v1 [cs.PL] for this version) https://doi.org/10.48550/arXiv.2609.19611 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-81] FedFIbOS: Fisher Importance based Optimal Submodelling for Heterogeneous Federated Learning

链接: https://arxiv.org/abs/2609.19559
作者: Yasmeen Afzal,Jeremiah D. Deng,Haibo Zhang
类目: Machine Learning (cs.LG)
*备注: 9 pages, 4 figures

点击查看摘要

Abstract:Heterogeneous federated learning requires clients with diverse computational capacities to collaboratively train a global model, where each client trains a capacity-constrained submodel. Existing methods select submodel parameters using heuristic importance measures—most prominently parameter magnitude—without theoretical justification for why these measures support convergence. We identify a fundamental gap: existing parameter selection criteria lack theoretical grounding in the convergence framework, partial client participation introduces additional estimation effects in the Fisher scores. We propose \textbfFedFIbOS: Fisher Importance-based Optimal Submodelling for heterogeneous federated learning, using Fisher Information in a principled criterion derived from minimizing submodel masking error. %We formally establish when magnitude selection is equivalent to Fisher selection fail under non-IID heterogeneous federated learning. We theoretically formulate submodel selection through a Fisher-weighted quadratic masking surrogate and show that the raw Fisher top- k rule implemented by FedFIbOS solves this surrogate under a Fisher-dominant ranking condition. The resulting method retains the convergence structure of the underlying masked federated optimization bound. Fisher scores are efficiently estimated from empirical diagonal Fisher information using squared gradients, enabling stable and adaptive parameter selection without additional optimization overhead. Experiments on CIFAR-10, CIFAR-100, and AGNews under pathological and Dirichlet non-IID settings show FedFIbOS achieves \approx10% higher accuracy than the state of the art, with improvements becoming more pronounced under stronger heterogeneity.

[LG-82] LSTM-UT and Recurrent-Depth Transformers on Cellular Automata

链接: https://arxiv.org/abs/2609.19521
作者: Aras Kavuncu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Recurrent-depth Transformers apply shared computation repeatedly, but differ in how they retain information across steps. We compare a Block Universal Transformer (BUT), which carries only its current hidden state; CoTFormer, which also retains an expanding attention cache; and a new LSTM Universal Transformer (LSTM-UT) with bounded gated memory. On Rule 30 cellular automata, BUT extrapolates to unseen recurrent depths more reliably than CoTFormer, although its accuracy eventually degrades. State and cache interventions show that CoTFormer’s failure depends on their interaction: correcting the current state can temporarily restore accuracy, while retained history can undermine that correction. In a delayed-recall task, BUT also outperforms CoTFormer despite lacking direct access to past states; CoTFormer does not reliably select the requested cached representation. LSTM-UT improves both depth extrapolation and delayed recall over these baselines. The results support bounded gated memory as an effective inductive bias for repeated computation and later retrieval in these tasks.

[LG-83] Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling DATE

链接: https://arxiv.org/abs/2609.19499
作者: Mobina Kashaniyan,Ali Jannesari
类目: Machine Learning (cs.LG); Distributed, Parallel, and Cluster Computing (cs.DC); Performance (cs.PF)
*备注: 8 pages, 4 figures, 9 tables. Experiments evaluate Phi-3-mini and Qwen2.5-1.5B on GSM8K and SciQ using NVIDIA A100 and V100 GPUs. Studies LLM test-time scaling, candidate-generation scheduling, latency, throughput, GPU-hours, and GPU-device energy

点击查看摘要

Abstract:Test-time scaling can improve large language model reasoning by generating and combining multiple candidate responses. In sampling-based methods, the inference budget is often described by the number of generated candidates, N. However, N tells us how many candidates are generated, not how they are executed. The same candidate budget can be produced in one batched generation call or split across several sequential calls with smaller batch sizes. We first study the effect of increasing N on reasoning accuracy using Phi-3-mini and Qwen2.5-1.5B on 500 GSM8K prompts. As expected, increasing N from 1 to 8 improves accuracy by 8.4 percentage points for Phi-3-mini and 18.4 points for Qwen2.5-1.5B. However, accuracy alone does not show the systems cost of using a larger candidate budget. We therefore fix N = 8 and compare four generation schedules: 1x8, 2x4, 4x2, and 8x1, where axb denotes a generation calls with b candidates per call. We measure latency, throughput, GPU-hours, and gross GPU-device energy while keeping the total candidate count fixed. On A100 GPUs, eight serial calls use 4.64-4.86x as much gross GPU-device energy and have 5.77-6.12x the P95 latency of one batched call with eight candidates. The same pattern appears across three independently scheduled A100 nodes per model and in short-output SciQ/V100 experiments. These results show that candidate count alone is not enough to describe the systems cost of multi-candidate test-time scaling. When candidates are independent and memory allows it, fewer generation calls with larger batch sizes are more efficient. Evaluations should therefore report not only candidate count and accuracy, but also generation schedule and GPU-level systems metrics.

[LG-84] Search at the Cost of Sampling: Nearly-Instant Latent Space Bayesian Optimization

链接: https://arxiv.org/abs/2609.19476
作者: Donney Fan,Colin Doumont,Aleksandra Kalisz,Paul Duckworth,Jacob R. Gardner,Henry Moss,Geoff Pleiss
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Generative models are increasingly central to many de novo discovery pipelines, in which designs are generated at scale and filtered through virtual screens to determine a set of candidates to experimentally validate. While Bayesian optimization (BO) is a natural fit for this setting, as it uses past evaluations to guide future proposals, the computational overhead required for its sequential decision-making becomes a bottleneck when virtual screens are relatively cheap. We make BO practical in this regime by exploiting the unique combination of a linear model constrained to a spherical domain where high-dimensional latents concentrate. We build off recent work justifying the use of linear surrogates, while deriving nearly closed-form solutions to the surrogate modelling and acquisition problems that exploit spherical symmetry. The result is at least a 100x speedup over state-of-the art baselines, with matching or improved performance across molecular and image generation benchmarks. Altogether, our method makes BO a practical drop-in for de novo pipelines where it was previously too slow to consider.

[LG-85] Enhanced Agriculture-informed Neural Network by Domain Knowledge

链接: https://arxiv.org/abs/2609.19466
作者: Ci Lin,Futong Li,Rose Chong-Wu,Tet Yeap,Iluju Kiringa
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Accurate prediction of nitrous oxide (N2O) emissions from agriculture is important for assessing environmental impacts and supporting sustainable farming. However, prediction remains difficult because N2O emissions result from complex interactions among soil properties, climate, biochemical processes, and management practices, while high-quality observations are limited. Deep learning models can capture nonlinear relationships but often lack physical interpretability and may generalize poorly across environmental conditions. We propose the Knowledge-enhanced Agriculture-informed Neural Network (KAINN), a hybrid neural-mechanistic framework that extends the Agriculture-informed Neural Network by incorporating domain knowledge about fertilizer diffusion, soil respiration, and water-filled porosity. We evaluate KAINN using CNN, LSTM, and Transformer architectures across multiple growing seasons and input-feature configurations. The results show that KAINN generally provides lower root mean square error and mean absolute error and higher R-squared values than purely data-driven models and the original AINN. Analysis of the learned interfaces also shows smoother and more physically consistent parameter trajectories with reduced uncertainty. These findings demonstrate that incorporating environmental knowledge into neural networks can improve the reliability, interpretability, and generalization of agricultural N2O-emission predictions.

[LG-86] Sharpness-Aware Minimization (SAM) Improves Classification Accuracy of Bacterial Raman Spectral Data Enabling Portable Diagnostics ICLR2024

链接: https://arxiv.org/abs/2609.19453
作者: Kaitlin Zareno,Jarett Dewbury,Siamak K. Sorooshyari,Hossein Mobahi,Loza F. Tadesse
类目: Machine Learning (cs.LG); Quantitative Methods (q-bio.QM); Machine Learning (stat.ML)
*备注: Accepted for oral presentation at the ICLR 2024 Workshop on Practical ML for Low Resource Settings (PML4LRS)

点击查看摘要

Abstract:Antimicrobial resistance is expected to claim 10 million lives per year by 2050, and resource-limited regions are most affected. Raman spectroscopy is a novel pathogen diagnostic approach promising rapid and portable antibiotic resistance testing within a few hours, compared to days when using gold standard methods. However, current algorithms for Raman spectra analysis 1) are unable to generalize well on limited datasets across diverse patient populations and 2) require increased complexity due to the necessity of non-trivial pre-processing steps, such as feature extraction, which are essential to mitigate the low-quality nature of Raman spectral data. In this work, we address these limitations using Sharpness-Aware Minimization (SAM) to enhance model generalization across a diverse array of hyperparameters in clinical bacterial isolate classification tasks. We demonstrate that SAM achieves accuracy improvements of up to 10.5% on a single split, and an increase in average accuracy of 2.7% across all splits in spectral classification tasks over the traditional optimizer, Adam. These results display the capability of SAM to advance the clinical application of AI-powered Raman spectroscopy tools.

[LG-87] GLAMDRING: Gait Learning And Morphology co-Design via Reinforcement LearnING of CPGs

链接: https://arxiv.org/abs/2609.19452
作者: Amogh Joshi,Kaushik Roy
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Robots are moving out of the structured factory floor and into unstructured environments such as disaster sites, planetary surfaces, and agricultural fields, for which the right robot often does not yet exist. We present GLAMDRING, a framework that synthesizes the optimal robot for a locomotion task and, jointly, learns the controller that drives it. For the given specifications of forward-velocity bounds, a per-actuator power budget, an actuator library, and a payload requirement, GLAMDRING returns a matched quadruped morphology (link geometry and per-joint actuators) and a Hopf-oscillator Central Pattern Generator (CPG) gait policy. We rank feasible designs against a target design objective, viz., maximum speed, minimum Cost of Transport (CoT), or max Payload Margin. Because body and locomotion are coupled, the optimal morphology dictates how a robot is driven, while optimal gait depends on the physical body. We train a small number of CPG policies by reinforcement learning across the space of candidate morphologies, co-learning the gait with the underlying robot hardware. Link lengths and actuators are then resolved post-hoc from the policy’s logged operating envelope, reducing synthesis cost to a small, fixed number of reinforcement-learning runs instead of one per candidate. Our experiments show three key findings: co-designing body and gait is necessary to satisfy locomotion constraints; actuator-envelope feasibility, rather than locomotion success alone, determines realizable payload capacity; and canonical animal gaits emerge naturally in most designs from morphology and constraints alone. A real-world demonstration further highlights the efficacy of our work.

[LG-88] Bayesian Optimization with Rich Auxiliary Information via LLM s

链接: https://arxiv.org/abs/2609.19437
作者: Tejus Gupta,Efe Mert Karagözlü,Rohit Sonker,Barnabás Póczos,Jeff Schnieder
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Bayesian Optimization (BO) is widely used for optimizing expensive black-box functions, yet many real-world optimization problems contain substantially richer information than function evaluations alone. Examples include training curves in hyperparameter optimization, expert notes and images in scientific experimentation, and prior knowledge about where optima may lie. We show that large language models (LLMs) can effectively leverage such rich auxiliary information to guide optimization. Motivated by these findings, we develop three methods for incorporating auxiliary information into BO using LLMs. Across hyperparameter optimization benchmarks and a real-world nuclear fusion optimization task, our methods consistently outperform both standard BO and existing LLM-based optimization approaches. Our results demonstrate the effectiveness of LLMs for leveraging rich auxiliary information in BO.

[LG-89] Demystifying Linear Operator Learning for Control Systems

链接: https://arxiv.org/abs/2609.19428
作者: Max Beier,Nicolas Hoischen,Sandra Hirche,Petar Bevanda
类目: ystems and Control (eess.SY); Machine Learning (cs.LG)
*备注: accepted to the 65th IEEE Conference on Decision and Control; authors’ version

点击查看摘要

Abstract:This paper proposes a structured approach to learning linear operators for control systems from data. We address both structural and learning-theoretic aspects of the problem. To derive structural assumptions, we propose using the well-established framework of (semi)groups for evolution equations, as operators in control systems are of the same type. Further, we propose analyzing learning algorithms through the lens of the inverse problems framework. This reveals how a learned model depends on the data via error decompositions, convergence guarantees, and optimal regularization – enabling us to compare existing methods and derive provably advantageous algorithms. In order to obtain these results, we restrict our scope to bounded operators on Hilbert spaces. Although this may appear restrictive, existing approaches often make this assumption implicitly to obtain matrix-like representations. We demonstrate the power of using these frameworks by deriving a convergent estimator for time-varying systems.

[LG-90] Improving Offline Goal-Conditioned Reinforcement Learning via Selective Reward Stimulation

链接: https://arxiv.org/abs/2609.19414
作者: Jing Zhang
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Goal-conditioned reinforcement learning aims to learn policies that reach specified goals, but remains challenging in offline settings with sparse rewards and long-horizon dependencies. In such settings, goal-completion information can be temporally distant from the early decisions that enable success, while offline value estimation introduces additional error. We study this issue from a reward-propagation perspective and show, in a stylized delayed-goal setting, how goal-directed value separation can become small relative to local estimation error. Motivated by this analysis, we propose Reward Stimulation Implicit Q-Learning (RSIQL), a simple non-hierarchical method that introduces additional reward signals at progress-making intermediate states in offline trajectories. RSIQL uses an auxiliary goal-conditioned value function to identify intermediate states estimated to make progress toward the goal and applies reward stimulation to provide less-delayed training supervision. Unlike hierarchical methods, RSIQL does not learn a separate high-level subgoal policy. Experiments on D4RL goal-reaching benchmarks and OGBench show that RSIQL improves over goal-conditioned IQL on average and achieves performance competitive with hierarchical offline goal-conditioned methods, while retaining a simple flat policy structure.

[LG-91] FCx: An algorithm for finding Feasible Counterfactual Explanations

链接: https://arxiv.org/abs/2609.19383
作者: Kleopatra Markou,Vana Kalogeraki,Dimitrios Gunopulos
类目: Machine Learning (cs.LG)
*备注: 20 pages, 9 figures

点击查看摘要

Abstract:Counterfactual (CF) explanations identify changes that alter an input’s classification. While existing methods produce realistic and low-cost CFs, they often fail to ensure feasibility, by suggesting non-constructive modifications or incompatible with future changes (e.g., changing an individual’s race to secure a job offer). We introduce a refinement of CF explanations that explicitly enforces feasibility. Our approach is the first to efficiently generate CFs that are realistic, low-cost and feasible. We accommodate both hard feasible constraints, specified by domain knowledge users, and soft feasible constraints, inferred automatically via causal inference from the dataset. Our method, Feasible Counterfactual Explanations (FCx), is based on a modified Variational Autoencoder (VAE) optimized with a multi-factor loss function. We measure the cost of a change based on the absolute change in values (proximity) as well as the number of features changed (sparsity) while realism is measured based on the LOF for density estimation, guaranteeing that CFs reside in densely populated regions. Extensive experiments on four public datasets show that our approach matches state-of-the-art performance across multiple metrics while guaranteeing feasibility.

[LG-92] Machine-Learning Assessment of the Predictive Value of Inflammatory Biomarkers for Cognitive Impairment in an Older Hispanic Adult Cohort

链接: https://arxiv.org/abs/2609.19374
作者: Antony Garcia,Gabrielle Britton,Alcibiades Villarreal,Diana Oviedo,Giselle Rangel,Xinming Huang
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Small clinical tabular datasets require interpretable machine learning because deep learning is often impractical and ensemble models can be difficult to inspect. A key pitfall is that statistical significance does not necessarily imply predictive utility. Using data from the Panama Aging Research Initiative–Health Disparities (PARI-HD) cohort (n=165), we implemented a leakage-safe threshold-likelihood Bernoulli/Categorical Naive Bayes (BNB/CNB) classifier. Within every training fold, each continuous predictor was reduced to a supervised chi-square-derived state, while income entered the model through a categorical likelihood. All data-dependent steps were performed within repeated stratified 10-fold cross-validation with 30 repeats. The demographic baseline achieved a ROC-AUC of 0.630 +/- 0.017. I-309 (CCL1) was the dominant incremental feature, increasing AUC by 0.110, with paired DeLong tests yielding p0.05 in 100% of repeats. In the pre-specified primary analysis, I-309 produced a fixed-partition DeLong p=0.0018, with robustness assessed across 200 random partitions, where the median p-value was 0.0011. Within the exploratory family of 18 candidate markers, I-309 achieved a Benjamini-Hochberg-adjusted q=0.032 on the frozen partition and satisfied q0.05 in 85% of random partitions, whereas no other marker demonstrated reliable incremental predictive value. Because the fitted model is an inspectable table of thresholds and class-conditional probabilities, these results identify I-309/CCL1 as an interpretable candidate feature for tabular prediction of cognitive impairment, pending external validation.

[LG-93] Stiefel Attention: When the Geometry of Transformer Projection Matrices Dominates Optimizer Choice—and When It Does Not

链接: https://arxiv.org/abs/2609.19363
作者: Rubén Darío Guerrero
类目: Machine Learning (cs.LG); Numerical Analysis (math.NA)
*备注: 16 pages, 2 figures

点击查看摘要

Abstract:The query and key projections \WQ,\WK in attention are almost always trained by Euclidean optimizers with no constraint on their geometry. We constrain them to the Stiefel manifold and optimize them there with a Riemannian Adam that carries one scalar second moment per frame, caps its step by a trust region, and retracts polarly. Four propositions prove this update is steepest descent in the embedded metric, independent of gradient scale, well conditioned, and exactly \mathrmO(d) -equivariant, each certified numerically in \textttfloat64. A fifth supplies the mechanism: weight decay has \emphidentically zero Riemannian gradient on \St(d,r) , since W = W I_r lies in the normal space, so the learned attention geometry survives the collapse cycles that decay drives through the rest of the model. On modular arithmetic grokking, a single run holds 97.0% validation accuracy at epoch 20,000 against the baseline’s 61.1% —an unstable endpoint we report as evidence for the mechanism rather than as an effect size. On CIFAR-10 patches the same rule gains \mathbf+8.98 ,pp over 12 paired starts ( t=60.6 , 12/12 ), and the gap widens with data rather than eroding. The step rule earns this: a fixed-step Riemannian update is degree one in the gradient, so it moves 24 – 40\times less per step than an identically shaped AdamW matrix—its frames barely leave their initialization, and freezing them outright costs only 0.28 ,pp. An ablation credits the whole gain to making the step scale free, and nothing measurable to the projector or to equivariance. A negative result sharpens the account: gauge removal cannot motivate the method, because a direction along which the loss is invariant carries no gradient at all.

[LG-94] Smart Insole Human Activity Recognition for Continuous Monitoring in Elderly Care

链接: https://arxiv.org/abs/2609.19359
作者: Edwin Rios,Antony Garcia,Fengpei Yuan,Xinming Huang
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Falls in older adults are often preceded by changes in mobility, balance, and postural transitions. This paper presents a wireless smart insole platform and machine-learning workflow for recognizing sitting, standing, walking, and unstable walking from plantar-pressure and inertial signals. Each insole integrates 16 active pressure-sensing locations and a six-dimensional IMU stream consisting of tri-axial acceleration and angular velocity. Data were collected from 15 healthy adults at 80~Hz and segmented into overlapping windows. Window length and candidate model families were first screened with stratified 10-fold cross-validation; the primary performance estimate was then obtained with participant-independent 5-fold Stratified Group cross-validation, ensuring that all windows from a participant remained in a single fold. Under this protocol, Histogram-Based Gradient Boosting (HGB) achieved macro-F1 scores of 0.954 and 0.959 for the left and right feet, respectively, and 0.980 with bilateral sensing. A compact 1D-CNN evaluated with the same participant-independent folds did not significantly outperform HGB ( p=0.0625 ). The results show that low-profile footwear sensing can infer activity state from pressure and IMU measurements for participants unseen during training, establishing a basis for activity monitoring and fall prevention in elderly care.

[LG-95] Personalized Federated Hierarchical Gaussian Processes for Privacy-Preserving Modeling of Heterogeneous Distributed Systems

链接: https://arxiv.org/abs/2609.19337
作者: Xianjian Xie,Hao Yan
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We present Personalized Federated Hierarchical Gaussian Processes (pFedHGP) for probabilistic regression and classification when data are distributed across heterogeneous clients. Each client’s latent function decomposes into (i) a shared global component, (ii) a client-specific deviation that shares the global kernel structure, and (iii) a flexible local residual. Sparse inducing-variable approximations and federated variational inference keep raw data local while the server synchronizes only low-dimensional statistics for the shared component. Full predictive distributions support uncertainty-aware decisions. In application studies, pFedHGP attains perfect fault classification in press tonnage monitoring using 13.77% of labeled cycles and recovers geographic zones in federated air-quality modeling without centralizing station-level time series. An Instantaneous Linear Mixing Model viewpoint links the hierarchy to multi-output Gaussian processes for correlated sensors.

[LG-96] Learning-Induced Dynamical Transition in Recurrent Neural Networks

链接: https://arxiv.org/abs/2609.19288
作者: Varun Vaidya
类目: Machine Learning (cs.LG); Disordered Systems and Neural Networks (cond-mat.dis-nn); Chaotic Dynamics (nlin.CD)
*备注: 16 pages, 7 figures

点击查看摘要

Abstract:Learning in recurrent neural networks can fundamentally reshape their underlying dynamics, transforming initially chaotic activity into stable task-dependent behavior. We develop a non-equilibrium dynamical mean-field theory(DMFT) to describe this transition during learning. We show that a slow feedback-driven learning process generates an evolving effective feedback strength that drives the network through a transition from chaotic to stable dynamics defined by a bifurcation of the DMFT solution. By deriving the two-time correlation function throughout learning, we identify a critical feedback strength and a corresponding learning rate dependent critical time separating these regimes. The transition arises from the progressive deformation of an effective dynamical landscape by the growing learned feedback structure. Starting from the untrained state, the theory predicts the time evolution of the network output during training and shows quantitative agreement with numerical simulations.

[LG-97] Radio-Frequency Convolutional Neural Networks

链接: https://arxiv.org/abs/2609.19279
作者: Zhihui Gao,Shi-Yuan Ma,Yiran Chen,Dirk Englund,Tingjun Chen
类目: Machine Learning (cs.LG); Emerging Technologies (cs.ET); Signal Processing (eess.SP); Applied Physics (physics.app-ph)
*备注: 15 pages, 4 figures. Supplementary Information: 50 pages, 31 figures, 1 table

点击查看摘要

Abstract:Running artificial intelligence (AI) models directly on edge devices such as smartphones, wearables, and drones offers low latency, pervasive scalability, and data privacy, but these devices rarely carry the computing capability that modern neural networks demand. Edge accelerators have been developed in response, yet each adds computing hardware to devices already constrained in size, weight, power, and cost (SWaP-C). An alternative lies in what these devices already carry: the frequency mixer in every wireless radio multiplies signals in time, natively performing convolution in the frequency domain. Here we introduce radio-frequency convolutional neural networks (RF-CNNs), which repurpose existing communication hardware for CNN inference. Multi-channel convolutions are mapped onto frequency tones for a passive mixer to execute in a single pass. We experimentally demonstrate that RF-CNN runs deep CNNs up to 26.4 million parameters and nine layers from classification of wireless signals and images to controllable image generation, close to full-precision performance. Because the weights arrive over the air and the analog hardware is shared with communication, the edge device spends energy only on data preparation and readout-down to 0.72 femtojoules per multiply-accumulate, two orders of magnitude less than it would cost on an added digital processor. These results suggest that deployed wireless infrastructure can bring efficient, state-of-the-art AI inference to the billions of devices it already connects.

[LG-98] Block Parallelism For Efficient Distributed Long-Context Diffusion Language Model Training

链接: https://arxiv.org/abs/2609.19242
作者: Tarun Suresh,Pranshu Chaturvedi,Hangoo Kang,Parth Shroff,Ishan S. Khare,Hermann Kumbong,Azalia Mirhoseini
类目: Machine Learning (cs.LG)
*备注: 25 pages, 6 figures

点击查看摘要

Abstract:Block diffusion language models (BDLMs) combine autoregressive dependencies across blocks with parallel denoising within blocks, but long-context training is constrained by distributed attention communication and activation memory. Conventional context parallelism (CP) shards the combined clean-plus-corrupted sequence by position, communicating shared clean K/V together with block-specific corrupted K/V and their gradients. We observe that the BDLM objective separates over target blocks. We introduce block parallelism (BP), a new distributed parallelism dimension that assigns each corrupted-block computation to one rank. To scale BP to long contexts, we introduce context-sharded block parallelism (CSBP), which also shards the shared clean sequence across those ranks. CSBP keeps corrupted K/V and gradients local, avoids replicated clean prefixes, and preserves BDLM training semantics. On 16 H200 GPUs at 256K context, CSBP improves throughput over the best baseline by 1.18-1.45x for supervised fine-tuning and 1.27-1.33x for conversion of autoregressive models to BDLMs, while matching or reducing peak HBM. Full-model speedup reaches 1.61x at 512K. On eight H100 GPUs, CSBP accelerates DFlash2 speculative-decoder training by 2.48x at 512K and 7.59x at 1M. In matched 12-hour DiffusionGemma 26B-A4B SFT runs, CSBP achieves higher pass rates at every trained checkpoint on SWE-bench Verified and Terminal-Bench Lite. Code: this https URL

[LG-99] Generative Query Suggestion via Intent Coverag e and Query-Level Credit Assignment

链接: https://arxiv.org/abs/2609.19209
作者: Xinpeng Liu,Lu Ma,Jiayi Qiao,Mengyu Zhou,Linglong Li,Xiaofeng Bian,Haonan Chen,Xiaoxi Jiang,Guanjun Jiang
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Generative query suggestion aims to enhance user engagement by anticipating user intents and recommending relevant follow-up queries. A central challenge is to generate slates whose individual queries are useful while the slate covers distinct intents. We propose an Intent-Driven Query Suggestion Framework with dual-stage optimization. First, intent-aware diversity modeling constructs intent-aligned supervised fine-tuning (SFT) data and uses an Intent-Aware Diversity Reward to optimize intent coverage. Second, query-level credit assignment routes individual quality signals to the corresponding query tokens while sharing a slate-level diversity signal across the slate. Experiments on a large-scale production dataset, including online A/B testing and offline evaluation, show improvements in click-through rate, query quality, and intent coverage.

[LG-100] How Does Distribution Shift Shape Pretraining Gains in Neural PDE Surrogates? NEURIPS2026

链接: https://arxiv.org/abs/2609.20814
作者: Pochinapeddi Sai Bhargav,Nithin Somasekharan,Rohit Sunil Kanchi,Sicheng He,Shaowu Pan
类目: Computational Physics (physics.comp-ph); Machine Learning (cs.LG); Fluid Dynamics (physics.flu-dyn)
*备注: 15 pages, 5 figures. Representations for the Physical Sciences Workshop, NeurIPS 2026

点击查看摘要

Abstract:Pretraining a neural PDE surrogate can reduce the amount of new CFD data needed when geometry or modeled physics changes. However, it remains unclear how different components of distribution shift affect this benefit. We pretrain a surrogate on 254,909 RANS solutions from one airfoil family and fine-tune it on a new family under two target settings with matched freestream ranges: the same Spalart-Allmaras (SA) modeling and SA with added e^N transition modeling. At N=1000 , the pretrained model matches the accuracy of a model trained from scratch on 3.25\times as many samples for the same-SA target, but 2.58\times as many for the transition-modeled target. By N=5000 , this ordering reverses ( 1.56\times versus 1.86\times ). At N=1000 , sampling more distinct airfoils lowers error on both targets, but only for the same-SA target is the gain increase larger than the observed draw-to-draw variation ( 3.3\times to 4.0\times ). These results show that pretraining value depends jointly on target-data budget, target-data coverage, and whether source and target differ in modeled physics.

[LG-101] Stable Movement for Nondual Lipschitz Convex Optimization: Efficiency and Nearly Optimal Oracle Rates

链接: https://arxiv.org/abs/2609.20701
作者: David Martínez-Rubio,Cristóbal Guzmán
类目: Optimization and Control (math.OC); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We study efficient algorithms for realizing the first-order oracle complexity of optimization of G -Lipschitz convex functions with respect to the \ell_q -norm over an \ell_p -ball of radius R , where 1\leq p,q\leq \infty . For pq , we obtain error \widetildeO_p,q(GR/T^1/p-(1/q-1/2)+) after T oracle queries, efficiently realizing the nearly optimal rates of (MBG+26), thereby resolving the nonsmooth end of the COLT 2015 open problem (Guz15b). In particular, the rate is \widetildeO(GR/T) for Euclidean Lipschitzness over an \ell_1 -ball of radius R ( p=1,q=2 ). Our solution consists of reducing convex Lipschitz optimization to the chasing nested convex sets problem in sublevel sets of an evolving bundle (LNN95; BBE+20): at each query we either find a point with low function value or we produce a deep cut in the current sublevel of the bundle, that we chase. The dichotomy between stability of selectors and forced movement by deep cuts bounds the number of iterations of the algorithm near optimally. For nested subsets of R B_p^d , we introduce a novel notion of stable center whose movement is bounded by \widetildeO_p,q(RT^1-1/p+(1/q-1/2)+) in the \ell_q -norm after T steps, which we show is nearly optimal in high dimensions. A Monte Carlo average of the proposed selector achieves near-optimal rates with high probability and can be implemented in polynomial time for our optimization algorithm in the real-arithmetic model.

[LG-102] risCNN for interpretable detection of phases of matter from experimental quantum simulator data

链接: https://arxiv.org/abs/2609.20693
作者: Kacper Cybiński,Björn van Zwol,James Enouen,Guillaume Bornet,Thierry Lahaye,Antoine Browaeys,Antoine Georges,Anna Dawid
类目: Quantum Physics (quant-ph); Disordered Systems and Neural Networks (cond-mat.dis-nn); Machine Learning (cs.LG)
*备注: 34 pages, 25 figures

点击查看摘要

Abstract:Detecting phases of matter in general relies on identifying the correct order parameter - a task that remains notoriously difficult for unknown transitions and traditionally is guided by physical intuition and educated guess. Neural networks have recently offered an alternative route by locating phase transitions in known models without any a priori physical knowledge. Yet these approaches remain black boxes and only identify phases without elucidating their properties. Moreover, they often struggle when confronted with realistic, noisy experimental data, which constitute the ultimate testbed for automated methods in physics. Here, we bridge these perspectives by introducing TetrisCNN, a convolutional architecture with parallel branches of differently shaped filters, reminiscent of Tetris blocks, that learns sparse, interpretable latent representations directly in terms of spin correlators. Applied to experimental snapshots of two-dimensional Ising and XY quantum simulators measured in multiple bases, the network not only detects phase transitions and crossovers but also expresses its latent representation and decision boundaries as symbolic formulas built from experimentally measurable spin correlators. This framework opens the way to integrating interpretable neural networks with quantum simulators to uncover and understand new phases of matter.

[LG-103] he First-Order Oracle Complexity of Lipschitz Convex Optimization in Nondual Settings

链接: https://arxiv.org/abs/2609.20687
作者: David Martínez-Rubio,Brian Bullins,Cristóbal Guzmán,Mathieu Molina
类目: Optimization and Control (math.OC); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We study first-order black-box convex optimization over an \ell_p -ball for objectives Lipschitz in the \ell_q -norm, solving in the affirmative the nonsmooth version of the COLT open question (Guz15b) on whether the geometry of a smaller feasible set ( p q ) can improve convergence rates in convex optimization, and matching prior lower bounds up to logarithmic factors. Our rates include (\widetilde O(1/T)) for convex Euclidean-Lipschitz optimization over the \ell_1 -ball, improving on the O(1/\sqrtT) classical rate under general assumptions. The key technical device is a new online learning game, where the comparator is evaluated using the maximum of affine losses observed so far. We bound the value of this game above and below in terms of a combinatorial online learning quantity: the sequential fat-shattering dimension, which we characterize for the \ell_p / \ell_q case. Our results generally apply when the feasible set X and the set of possible subgradients H are convex, centrally symmetric, and admit a type of minmax theorem, advancing on a fundamental question by Sridharan [Sri12, Section 10.1.2, Q3]. As a geometric consequence of our analysis, of independent interest, we obtain estimates for the expected distance of a convex hull of samples to their mean in several Banach geometries, a version of the celebrated Wendel’s theorem (Wen62), but quantitative and for bounded general distributions as opposed to centrally symmetric ones.

[LG-104] AP Accuracy Below the Fluctuation Scale and Universal Posterior Geometry in Spherical Linear Models

链接: https://arxiv.org/abs/2609.20577
作者: Jingbo Liu,Zhiyuan Yu
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We study the Bayes-optimal spherical linear model as the ambient dimension and sample size grow proportionally, under a quantitative Marchenko–Pastur spectral-regularity condition on the design. This condition is satisfied by normalized i.i.d. designs with standardized entries of finite fourth moment, but does not require entrywise independence or impose conditions on the singular vectors. Under this condition, we prove a quantitative all-temperature TAP approximation and characterize the posterior geometry. For the natural finite-aspect-ratio TAP functional, the normalized spherical free energy and the TAP optimum differ by O_P(p^-1) . Each is within O_P(p^-1/2) of its explicit deterministic equivalent, and this fluctuation scale is sharp. Uniformly over all global TAP maximizers, the normalized squared Euclidean distance to the spherical posterior mean is O_P(p^-1) . We also prove that the posterior mass outside a data-dependent band determined by the ridge estimator has sharp exponential order. More precisely, uniformly over sufficiently small band widths \varepsilon , the logarithm of this mass is at most -cp\varepsilon^2+O_P(1) . For every fixed geometrically admissible width, a spherical-cap construction gives a matching exponential-order lower bound on this mass. For every deterministic sequence of widths \varepsilon_p\gg p^-1/2 , the corresponding bands capture asymptotically all posterior mass.

[LG-105] Noise-Robust Quantum State Characterization for Remote State Preparation with Deep Learning

链接: https://arxiv.org/abs/2609.20523
作者: Bo Tang,Zixuan Liao,Hao Li,Yilin Yang,Jiani Lei,Zengya Li,Jing Qiu,Zhaohui Dong,Zhengyang Mao,Yuanhua Li,Yuanlin Zheng,Xianfeng Chen
类目: Quantum Physics (quant-ph); Machine Learning (cs.LG); Optics (physics.optics)
*备注:

点击查看摘要

Abstract:Quantum communication underpins secure information processing and scalable quantum networks. In particular, remote state preparation (RSP) enables efficient quantum state transfer, but accurately estimating target states under complex noise remains challenging. Here, we propose a Transformer-based Quantum State Characterizer (TQSC) model for noisy RSP experiments. Our model reconstructs experimentally prepared pure and mixed photonic polarization states from noisy measurements in complex scattering environments, while its attention patterns provide physically grounded insights into correlations among the measured observables. The method achieves a mean estimator-target fidelity exceeding 99.999% under complex scattering and dynamic Gaussian noise, while its robustness and generalization are further examined using Qiskit-simulated Bloch-ball this http URL, in a practical MNIST image transmission task with held-out states, the decoded bit error rate is reduced from 50.34% to zero after TQSC post-processing. The TQSC model enables accurate tomographic characterization under dynamic noise and provides physically grounded post-hoc insights, holding promise for intelligent quantum information processing applications.

[LG-106] runcated automatic sparse differentiation for machine learning interatomic potentials

链接: https://arxiv.org/abs/2609.20510
作者: Marcel F. Langer,Adrian Hill,Michele Ceriotti
类目: Chemical Physics (physics.chem-ph); Materials Science (cond-mat.mtrl-sci); Machine Learning (cs.LG)
*备注: 19 pages, 5 figures, 4 tables (7 pages main text). Additional information at this https URL

点击查看摘要

Abstract:Machine learning interatomic potentials (MLIPs) learn the mapping from atomic positions to potential energy. The forces, the negative gradient of this energy, drive molecular dynamics and are readily obtained using automatic differentiation. Higher-order derivatives, most notably the Hessian, describe collective motion and allow the direct prediction of experimental observables, but are considered computationally inaccessible for large systems. We suggest a solution: in physical systems, interactions decay with distance, and most MLIPs build on this locality through message passing up to a finite receptive field. This implies both sparsity of higher-order derivatives and their decay with distance. This structure can be exploited using automatic sparse differentiation (ASD). We explain how to compute the sparsity pattern for MLIP derivatives and demonstrate that, for multiple foundation MLIPs, ASD computes full Hessians of large porous materials exactly, but with modest speedups at best. The larger gains come from truncated ASD: discarding small, but nonzero, Hessian entries between distant atoms yields order-of-magnitude speedups with negligible impact on predicted observables.

[LG-107] Correlation-Free Transition Path Sampling through Shooting Point Generation Guided by Committor Learning

链接: https://arxiv.org/abs/2609.20461
作者: Maximilian Negedly,Sebastian Falkner,Alessandro Coretti,Christoph Dellago
类目: Computational Physics (physics.comp-ph); Statistical Mechanics (cond-mat.stat-mech); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Studying the dynamical behavior of a system often depends on characterizing how it transitions between long-lived states. Because such transitions are rare, observing them usually requires specialized enhanced sampling techniques. Transition Path Sampling (TPS) is a well-established method for generating reactive trajectories, which is simple to implement and does not require the definition of a preconceived reaction coordinate. However, its efficiency is limited by its sequential nature and the resulting correlations between sampled paths. Previous work addressed this limitation by combining TPS with a sampling scheme based on conditioned Boltzmann Generators, a generative machine learning model capable of sampling a given target probability distribution. This approach produces uncorrelated transition paths but relies on an accurate reaction coordinate, which is rarely known in advance. Building on recent advances in committor learning, specifically on the Artificial Intelligence for Molecular Mechanism Discovery (AIMMD) method, in this work we introduce GenAIMMD, an iterative algorithm that actively and self-consistently learns the ideal reaction coordinate (the committor) and trains a conditioned Boltzmann Generator to sample from arbitrary bias windows along it. GenAIMMD thereby provides a correlation-free and fully parallelizable path sampling scheme that does not require prior knowledge of the system’s transition mechanism. We apply GenAIMMD to a two-dimensional toy model and a higher-dimensional polymer system. In both cases, GenAIMMD succeeds in training the Boltzmann Generator and learning the committor. Benchmark results show a substantial increase in performance compared to standard TPS.

[LG-108] Online Supervised Dimension Reduction with Random Features: Diagnostics and Computational Trade-offs

链接: https://arxiv.org/abs/2609.20454
作者: Zhenlin Yao,Wei Xiong
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Computation (stat.CO)
*备注: 41 pages, 4 figures, 18 tables; includes core supplementary material

点击查看摘要

Abstract:Accurate optimization of a supervised spectral objective need not produce an accurate population subspace or a better predictive representation. We investigate these distinctions for Online Kernel Supervised Principal Component Analysis (OKSPCA), which combines a centered cross-moment in finite random-feature coordinates with an Adam-style orthonormal basis update for an established objective. Fixed-map consistency, concentration and perturbation results describe the estimator and its exact subspace; same-target comparisons then assess the practical iterate separately. Across six predictive benchmarks, performance depends on the declared pipeline: replacing the tracker with the exact empirical target leaves the two regression deficits largely unchanged. Direct classification-rank models capture nearly all terminal objective energy on average, but a saved intermediate state exhibits substantial geometric deviation; a controlled sample-size study further separates empirical accuracy from population recovery. In distinct numerical-service workloads, exact on-request computation is faster in the tested classification settings, whereas Adam saves time relative to the tested full thin-SVD service for some dense wider-regression requests, alongside persistent geometric error. These diagnostics limit explanations based solely on terminal optimization accuracy and distinguish numerical cost from quality, rank coverage and freshness; they establish neither practical-tracker convergence nor predictive or deployment benefits from basis availability.

[LG-109] Model-based Bootstrap for Offline Policy Evaluation in Tabular Reinforcement Learning

链接: https://arxiv.org/abs/2609.20389
作者: Weiwei Wang,Yuqiang Li,Xianyi Wu,Bingyi Jing
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Offline policy evaluation (OPE) is crucial in high-stakes reinforcement learning applications, where new policies must be assessed reliably before deployment. In such settings, point estimates alone are insufficient; principled uncertainty quantification, such as confidence intervals and variance estimates, is essential for safe and risk-aware decision-making. A comprehensive way to unify these tasks is to estimate the sampling distribution of the evaluation error. Existing approaches, however, often suffer from limited robustness, scalability, or finite-sample validity. In this paper, we propose a model-based bootstrap framework for uncertainty quantification of OPE in finite-horizon, time-inhomogeneous Markov decision processes (MDPs). Unlike classical bootstrap methods that rely on resampling complete episodes, the proposed method regenerates trajectories from an estimated MDP and can therefore accommodate a much broader range of offline data formats, including complete trajectories, transition-level observations, and trajectory fragments. This flexibility further improves finite-sample statistical efficiency. We establish bootstrap distributional consistency, asymptotically valid confidence intervals, and consistent variance estimation for the target policy value. Extensive simulations show that the proposed method accurately captures the sampling distribution of the OPE estimator, yielding tighter confidence intervals and more accurate variance estimates in most settings.

[LG-110] Near-Optimal Pure Single-Loop Extrag radient Method for Strongly Convex–Strongly Concave Minimax Optimization

链接: https://arxiv.org/abs/2609.20327
作者: Minhao Zhang,Zi Xu
类目: Optimization and Control (math.OC); Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:We study smooth strongly convex–strongly concave minimax optimization with general nonlinear coupling in the deterministic unconstrained setting. We propose a pure single-loop damped extragradient method with fixed parameters and two new full-gradient evaluations per iteration after one initialization query. The method uses an auxiliary feedback recursion and requires no inner solves, accuracy schedules, or staged restarts. We establish last-iterate linear convergence and show that reducing the squared Euclidean distance to the saddle point to an \varepsilon fraction of its initial value requires O(\sqrt\kappa_x\kappa_y\log(2\kappa_x\kappa_y/\varepsilon)) full-gradient queries, where \kappa_x=L/\mu_x and \kappa_y=L/\mu_y . This bound attains the optimal condition-number order up to logarithmic factors through fixed explicit updates. Numerical experiments demonstrate the effectiveness of the method.

[LG-111] QEncodeBench: Can Large Language Models Encode Classical Problems into Verified Quantum Oracles?

链接: https://arxiv.org/abs/2609.20319
作者: Xujun Che,Hanhan Wu,Yuchen Yuan,Chenyang Yu
类目: Quantum Physics (quant-ph); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Grover search, amplitude amplification, and quantum counting all rely on the same reusable subroutine, a phase oracle, whose construction the algorithms literature takes as given: the classical predicate is assumed to be already encoded as a correct, resource-bounded circuit. We turn this assumption into a measured capability. QEncodeBench tasks large language models (LLMs) with encoding classical constraint problems as phase oracles and scores the generated circuits with an adversarially self-validated verifier that decides full solution-set equivalence up to a global phase, with ancillas restored and resource budgets enforced. Sampled basis-state tests, we show, systematically overestimate this ability. Measured this way, models separate sharply: code models without a reasoning mode solve essentially nothing, and enabling native reasoning on identical weights improves accuracy by an order of magnitude. The failures are overwhelmingly semantic rather than syntactic. Two architectures, a unit-verified constraint agent and a neuro-symbolic compilation pipeline, close most of the remaining gap by delegating correctness-critical composition to deterministic procedures. Ablations quantify the contribution of each component, and resource gating exposes an architecture-dependent trade-off between circuit width and depth. Finally, controlled difficulty escalation reveals architecture-specific responses to difficulty structure: different difficulty axes degrade different methods, while the neuro-symbolic pipeline passes every evaluated instance. Code and data are available at this https URL.

[LG-112] Counterexamples and Sufficient Conditions: Comments on “Optimally-Transported Generalized Method of Moments”

链接: https://arxiv.org/abs/2609.20260
作者: Masahiro Kato
类目: Econometrics (econ.EM); Machine Learning (cs.LG); Statistics Theory (math.ST); Methodology (stat.ME); Machine Learning (stat.ML)
*备注: Comment on arXiv:2511.05712

点击查看摘要

Abstract:We comment on the optimally-transported generalized method of moments (OTGMM) estimator proposed by Schennach Starck (2026a) and give counterexamples to Theorems 2-6 under their stated assumptions. First, the assumptions used in the small-error analysis are insufficient for consistency in Theorem 2 and asymptotic normality in Theorem 3. Next, we consider the large-error analysis, in which Theorem 4 states that the OTGMM estimator is equivalent to a GMM estimator with modified moments. We show that in a scalar model, Theorem 4 selects a value that differs from the unique OTGMM minimizer and violates the OTGMM sample moment restriction. In an overidentified model satisfying the assumptions used in Theorems 5 and 6, the first component of the Lagrange multiplier has different probability limits under the OTGMM estimator and the GMM estimator with modified moments. Under misspecification, the population value selected by OTGMM depends on the transport metric and on which variables may be adjusted. We give a sufficient condition under which solutions of the modified moment equations also solve the original constrained problem at a given parameter value, and separate conditions for consistency and asymptotic normality of the OTGMM estimator. We also show that Assumption 16 does not imply the matrix bound used in the supplemental proofs and replace Assumption 16 with a matrix condition that yields the bound.

[LG-113] ransformer fault diagnosis using an efficient simulation-driven variational quantum classifier with domain-aware feature encoding

链接: https://arxiv.org/abs/2609.20214
作者: Huy Hoang Le,Ba Tu Phung,Dai Huynh,Kim-Anh Nguyen
类目: Quantum Physics (quant-ph); Machine Learning (cs.LG)
*备注: This paper has been published in Alexandria Engineering Journal. Please cite the published version

点击查看摘要

Abstract:Early transformer fault diagnosis is challenged by nonlinear dissolved-gas interactions, overlapping fault signatures, and limited labeled data, while practical deployment further requires reliable performance under realistic computational constraints. This paper presents a simulation-driven modeling framework for dissolved gas analysis-based transformer fault diagnosis, in which a carefully engineered variational quantum classifier (VQC) is employed as the computational core and systematically analyzed through simulation. The framework integrates domain-aware feature modeling derived from Duval geometry with a lightweight two-qubit quantum representation, enabling nonlinear gas-interaction effects to be captured within a shallow parameterized circuit. A hybrid ZX-YY quantum feature map is designed to model non-commuting feature interactions, while a full-entanglement EfficientSU2 ansatz provides adequate expressive capacity under strict resource limits. Model behavior is evaluated using a comprehensive simulation pipeline including noise-aware circuit emulation, cross-dataset validation, and limited hardware-in-the-loop execution, allowing key effects of circuit depth, noise, and optimization strategy to be examined. Simulation results on benchmark dissolved-gas-analysis datasets demonstrate high diagnostic accuracy, strong generalization capability, and robustness to realistic noise levels with minimal quantum resources. The results highlight the effectiveness of simulation-informed modeling for practical transformer diagnostic applications, offering a reproducible and resource-efficient pathway for evaluating quantum-enhanced fault diagnosis methods.

[LG-114] Special Lagrangian cones in Deep Learning

链接: https://arxiv.org/abs/2609.20159
作者: Tejas Kotwal,Govind Menon
类目: Differential Geometry (math.DG); Machine Learning (cs.LG); Symplectic Geometry (math.SG)
*备注: 16 pages

点击查看摘要

Abstract:We introduce a matrix generalization of the cone of Harvey and Lawson and prove that it is an exact special Lagrangian manifold. We further show that it belongs to a family of exact special Lagrangian manifolds that foliate the balanced manifold arising in deep learning.

[LG-115] Quantum Graph Convolutional Networks: Implementation and Trainability Analysis

链接: https://arxiv.org/abs/2609.19983
作者: Paul San Sebastian Sein,Theodor Iosif,Tilen G. Limbäck-Stokin,Kin Ian Lo,Yidong Liao
类目: Quantum Physics (quant-ph); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Graph Neural Networks (GNNs) achieve state-of-the-art performance on graph-structured data, but training and inference on large graphs are often bottlenecked by memory constraints and sparse linear-algebra workloads. Quantum computing offers an alternative set of primitives that may improve scalability for graph learning. Building on the quantum graph neural network (QGNN) framework of Liao \textitet al., this work implements two representative architectures — the Simplified Graph Convolution (SGC) and Linear Graph Convolution (LGC) models — and evaluates them on open benchmark graph datasets and semi-supervised learning tasks using quantum simulation. We compare predictive performance and optimization behavior against classical baselines, showing that the quantum models achieve competitive performance with fewer parameters. Finally, we present a cost gradient analysis that identifies the tasks for which the models showcased are trainable. This is followed by a classical simulability study to find regimes in which the proposed circuits remain robust during training.

[LG-116] Error bounds in Sobolev norms for approximations with norm constrained ReLU neural networks

链接: https://arxiv.org/abs/2609.19937
作者: Xianjun Li,Yunfei Yang
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Recent studies have shown that smooth functions can be well approximated by ReLU neural networks with path norm constraint on the weights. We extend these results from uniform approximation to approximation in Sobolev norm. Specifically, we analyze how well Sobolev functions in W^n,p can be approximated by neural networks with width W , depth L and path norm bounded by K , when the approximation error is measured in the W^1,p -norm. For shallow networks with depth L=1 , we derive the approximation error bound \mathcalO(\max\W^-(n-1)/d, K^-(n-1)/(s-n)) , when the smoothness index satisfies ns=(d+3)/2 and the input is d -dimensional. For deep networks, we remove the restriction on the smoothness by showing that the approximation bound \mathcalO(K^-(n-1)/(d+d/p+1)) holds if the width W and depth L are sufficiently large.

[LG-117] Self-Replicating Neural Cellular Automata: Quantifying Emergent Phenotypic and Genotypic Diversity in an OpenEnded Substrate

链接: https://arxiv.org/abs/2609.19902
作者: Sanyam Jain,Felix Simon Reimers,Stefano Nichele
类目: Populations and Evolution (q-bio.PE); Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE)
*备注:

点击查看摘要

Abstract:We study an in-silico substrate in which every pixel of a two-channel cellular-automata grid carries a tiny neural network (an agent) that senses its Moore neighborhood. A cell persists only by self-replication: a living neighbor is cloned and its weights are mutated by a uniform perturbation, so that phenotype (cell state) is driven entirely by genotype (network weights). From a handful of seeded founders the system grows into a spatially organized ecosystem of coexisting, competing and dominating species. Our main contribution is a battery of coarse-grained diversity metrics that make such growth measurable at two scales: four phenotypic tools based on cellular-type frequency, entropy and cell variance, and two genotypic tools that colour each agent by a hash of its full weight vector versus a sparse random-weight probe. Across a five-fold sweep of 1680 small runs and 24 long (1000-generation, 200 x 200) runs, the substrate is persistent and self-maintaining in 20 of the 24 long configurations and exposes a clear phenotype-genotype diversity trade-off: raising phenotypic diversity collapses genotypic diversity and vice versa. Full-genome hash colouring further reveals lineage structure that a random-weight probe systematically misses. Code, data and animations are released as supplementary material.

[LG-118] Well-posedness of neural turbulence closures and tangent dissipation

链接: https://arxiv.org/abs/2609.19647
作者: Zhen Zhang,George Em Karniadakis
类目: Fluid Dynamics (physics.flu-dyn); Machine Learning (cs.LG)
*备注: 20 pages, 4 figures

点击查看摘要

Abstract:A neural turbulence closure defines a new boundary-value problem, R(U)=N(U)+F(U)=0 , with a coupled Jacobian J(U)=N’(U)+F’(U) , where N is the original mean-flow operator and F the learned closure. We establish two consequences of global tangent dissipation. For a monotone original operator, a positive uniform margin supplied by the original operator and closure together guarantees existence, uniqueness and a global inverse-sensitivity bound relating a posteriori solution error to the a priori residual. For a general original operator, a dissipative closure cannot worsen tangent dissipation, but this alone does not guarantee uniqueness. Tangent dissipation depends on both diffusion and reaction. We study two complementary ways to promote it: (1) an exact-integral construction enforcing non-negative tangent diffusion while leaving reaction unconstrained, and (2) a penalty on tangent-reaction violations at sampled states. Tangent diffusion enters the Jacobian, and non-negative secant eddy viscosity alone does not control its coercivity. We conduct tests with channel flow at Re_\tau=180 – 5200 , which provides a strongly monotone baseline. Both constrained closures reach accurate solutions in all 50 training-seed/Reynolds-number cases. At Re_\tau=1000 , we conduct tests with 10,000 starts for one fixed network per closure and we find one root for each constrained closure and multiple roots for the other closures. Although this does not prove uniqueness, it provides strong empirical evidence for uniqueness of the tested constrained closures. At Re_\tau=5200 , the construction and penalty reduce the reported inverse sensitivity relative to the original operator by approximately 372\times and 11\times , respectively.

[LG-119] Next-token functional estimation

链接: https://arxiv.org/abs/2609.19529
作者: Milind Nakul,Vidya Muthukumar,Ashwin Pananjady
类目: Machine Learning (stat.ML); Information Theory (cs.IT); Machine Learning (cs.LG); Probability (math.PR); Statistics Theory (math.ST)
*备注:

点击查看摘要

Abstract:Suppose we observe the first n points of a sequence of random variables having length n+1 , and wish to estimate a functional of the unobserved final point and the empirical measure of the n observed training points. Such next-token functionals include the probability that the next token is novel (also known as the surprise probability), the tail probability of the minimum distance between the next token and training points, and the test error of a classifier trained on the observed points. All of these quantities are classically estimated by the leave-one-out method, which is inconsistent under temporal dependence. We propose a leave-a-window-out estimator, which deletes a window of length \tau after each index before forming the empirical measure and reduces to leave-one-out at \tau = 1 . Under natural assumptions, we show that the error of our estimator decays at a parametric rate for any stationary \beta -mixing process that also admits a Marton coupling. Our results thus cover several natural functionals on a large class of stochastic processes. We complement these upper bounds with a sharp minimax lower bound for estimating the surprise probability on mixing Markov chains. Simulations on Markov chains, moving-average processes, and autoregressive processes show that our estimator succeeds in many scenarios where leave-one-out and add-constant baselines fail.

[LG-120] Null importance: Disentangling relevance for interpretable machine learning

链接: https://arxiv.org/abs/2609.19511
作者: Garvesh Raskutti,Kris Sankaran,Jiaxin Ye
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: 29 pages, 9 files. Submitted to Statistical Science

点击查看摘要

Abstract:Feature importance is central to interpretable machine learning, but the term “importance” encompasses several fundamentally different notions of relevance. We develop a unified perspective based on null importance: a population-level characterization of when a feature is irrelevant under a specified notion of relevance. We consider standard notions of null importance arising from marginal and conditional statistical relevance, predictive risk, functional invariance, and causal effects, and show how these notions answer different scientific questions. We illustrate the framework in two applications in which the distinction is particularly consequential: algorithmic fairness, where common fairness criteria correspond to different notions of null importance, and genomic perturbation modeling, where different notions of relevance lead to different conclusions about what a prediction model has learned. The framework connects three aspects of feature analysis: the scientific question defining relevance, the data and model assumptions that shape how different null notions relate, and the methods used to assess importance. We establish sufficient conditions under which null notions coincide and give counterexamples showing how they diverge when those conditions fail. We then characterize which nulls different method families target and when their zero-importance statistics identify those targets. Finally, simulations spanning feature dependence, redundancy, nonlinearity, hidden features and other standard phenomena, along with case studies on image and multiomics data, provide empirical evidence for these theoretical distinctions and their practical consequences. Taken together, these results provide a common statistical language for relating scientific questions, data-generating assumptions, and algorithms, and clarify the conclusions that feature-importance analyses can support.

[LG-121] Stable Policy Learning

链接: https://arxiv.org/abs/2609.19418
作者: Harvey Barnhard,Giacomo Opocher,Rahul Singh
类目: Econometrics (econ.EM); Machine Learning (cs.LG); Statistics Theory (math.ST); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:In evidence-based policymaking, typically one experimental sample is observed, then a learned policy recommendation is implemented at scale. Policies learned from the experimental data can perform well in expected welfare, yet random sampling in the experiment can produce recommendations with poor welfare outcomes. In this paper, we ask: how should policy learning algorithms balance expected welfare against sampling risk? Our main contribution is to show that algorithmic stability plays a central role in characterizing and navigating the tradeoff. Intuitively, if a policy learning algorithm’s recommendation remains stable when one experimental unit is replaced, then that algorithm has limited sampling risk. We propose a method for policy learning called policy-vote bagging, which learns treatment decisions on many subsamples then averages their votes into treatment probabilities. Relative to using one subsample, averaging across subsamples preserves expected welfare and improves expected utility for a risk-averse researcher. We derive sharp bounds linking estimation accuracy, subsample size, and welfare variation, including an exact guarantee under CARA utility.

[LG-122] Deep Learning Detection of Beyond-General-Relativity Deviations in Gravitational-Wave Signals: A Detection-Threshold Study with Real LIGO Noise

链接: https://arxiv.org/abs/2609.19416
作者: Muhammad Adnan Shahzad
类目: General Relativity and Quantum Cosmology (gr-qc); Instrumentation and Methods for Astrophysics (astro-ph.IM); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We study machine-learning detection of controlled beyond-General-Relativity (beyond-GR) deviations in gravitational-wave signals, using both synthetic aLIGO-PSD noise and real LIGO H1 detector strain. Three deviation families are applied to General-Relativistic inspiral-merger-ringdown waveforms: amplitude modulation, phase modulation, and frequency modulation, each parameterized by a dimensionless strength coefficient \beta . A hybrid classifier combining a one-dimensional convolutional neural network with ten hand-crafted waveform statistics is trained on GR and modified waveforms and tested on a deviation type excluded from training. The central result is a quantitative detectability curve as a function of \beta . Using the real GW150914 strain as a template and real H1 detector noise, we find a detection threshold at \beta \approx 0.25 , with accuracy rising smoothly from chance at \beta \leq 0.2 to perfect classification at \beta \geq 0.5 . The threshold value is specific to the quadratic-in-time modulation form adopted here and should not be interpreted as a generic constraint on beyond-GR parameters. We nevertheless argue that the negative result at small \beta is informative: it establishes a quantitative limit on machine-learning-only beyond-GR searches in real detector noise, in the absence of matched-filter signal extraction.

[LG-123] Learning Submanifolds for Subsequent Inference on Random Dot Product Graphs Part 1: Theory

链接: https://arxiv.org/abs/2609.19357
作者: Michael W. Trosset,Carey E. Priebe
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: 23 pages

点击查看摘要

Abstract:We propose a framework for restricted inference on random dot product graphs whose latent positions lie on an unknown low-dimensional support manifold. For general decision problems, we propose semisupervised decision rules that use auxiliary data to learn the support manifold. Specifically, our rules use the Isomap manifold learning procedure to construct a low-dimensional Euclidean representation of the observed graph, in which space an isometrically invariant function maps configurations of points to actions. We study the behavior of the proposed rules as the quantity of auxiliary data sampled from the unknown support manifold increases. We show that, as the auxiliary sample size increases, the risk of the semisupervised rule converges to the risk of an oracle rule that relies on the maximal amount of low-dimensional Euclidean structure that can be extracted from the support manifold. Examples, applications, and simulation studies are deferred to a sequel.

[LG-124] Federated Soft Clustering via Generalized Total Variation Minimization ICASSP’27

链接: https://arxiv.org/abs/2609.19202
作者: Shamsiiat Abdurakhmanova,Alexander Jung
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: submitted to ICASSP '27

点击查看摘要

Abstract:We study federated soft clustering over federated learning (FL) networks of devices that each hold a private local dataset and fit a personalized Gaussian mixture model (GMM). Generalized total variation minimization (GTVMin) couples the local maximum likelihood problems through a graph regularizer that penalizes a discrepancy between the models of connected nodes. The choice of discrepancy measure is a key design decision: we compare a squared Euclidean distance between model parameters, which requires component matching, with two measures that compare the local model distributions directly and hence need no matching: a Monte-Carlo approximated Kullback-Leibler (KL) divergence and a closed-form maximum mean discrepancy (MMD). All three resulting GTVMin instances are optimized by synchronous projected gradient updates; for the smooth MMD instance we provide a convergence guarantee to stationary points. We characterize their computational cost and evaluate their robustness to data heterogeneity.

附件下载

点击下载今日全部论文列表