本篇博文主要内容为 2026-09-21 从Arxiv.org论文网站获取的最新论文列表,自动更新,按照NLP、CV、ML、AI、IR、MA六个大方向区分。

说明:每日论文数据从Arxiv.org获取,每天早上12:30左右定时自动更新。

提示: 当天未及时更新,有可能是Arxiv当日未有新的论文发布,也有可能是脚本出错。尽可能会在当天修复。

目录

概览 (2026-09-21)

今日共更新647篇论文,其中:

  • 自然语言处理87篇(Computation and Language (cs.CL))
  • 人工智能125篇(Artificial Intelligence (cs.AI))
  • 计算机视觉100篇(Computer Vision and Pattern Recognition (cs.CV))
  • 机器学习152篇(Machine Learning (cs.LG))
  • 多智能体系统7篇(Multiagent Systems (cs.MA))
  • 信息检索8篇(Information Retrieval (cs.IR))
  • 人机交互33篇(Human-Computer Interaction (cs.HC))

多智能体系统

[MA-0] Bayesian Belief Layer for Controllable Opinion Dynamics in LLM Agents EMNLP2026

【速读】:该论文旨在解决大语言模型(LLM)代理在社会模拟中隐式修正观点所导致的可解释性与可控性问题:代理对说服的开放程度无法显式设定或验证,且集体结果受模型训练先验的隐性影响。其解决方案的关键在于提出贝叶斯编年史代理(Bayesian Chronicle Agents, BCA),构建一个最小化的信念层,将代理的内在信念(belief)与外在表达(speech)明确分离。每个立场以概率形式表示,并在每接收一次话语后通过一次贝叶斯更新进行调整;仅通过一个先验强度参数κ即可编码代理的固执程度,该参数借鉴了弗里德金-约翰森(Friedkin–Johnsen, FJ)意见动态模型中的作用机制。通过调节κ,可按需生成三种典型的意见动态模式(共识、持续分歧、少数派坚定影响),其中持续分歧情形下的拟合优度R²达到0.93–0.99,与FJ模型的解析解高度一致。此外,研究证明κ在语言交互的“往返”过程后仍可被准确恢复,所有四类模型均实现完美的排序级恢复。显式的信念结构还增强了模拟的可审计性,使各模型潜在的立场偏差得以显现,而这些偏差在端到端模拟中会被无声吸收。

链接: https://arxiv.org/abs/2609.21997
作者: Hafsa Akbar,Daniel Platnick,Marjan Alirezaie,Hossein Rahnama
机构: MIT Media Lab, Massachusetts Institute of Technology(麻省理工学院媒体实验室); Flybits Labs, Creative AI Hub; Toronto Metropolitan University(多伦多都会大学)
类目: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI)
备注: Accepted to The 2nd Workshop for Research on Agent Language Models (REALM) at EMNLP 2026

点击查看摘要

Abstract:LLM agents in social simulation revise their opinions implicitly, in context: how open an agent is to persuasion can neither be specified nor verified, and collective outcomes inherit the model’s training prior. We introduce Bayesian Chronicle Agents (BCA), a minimal belief layer separating \emphwhat an agent believes from \emphhow it speaks. Each stance is a probability, updated by one Bayesian step per utterance heard. A single prior-strength parameter \kappa encodes stubbornness, modeled after its role in Friedkin–Johnsen (FJ) opinion dynamics. We then sweep this parameter to yield three canonical regimes of opinion dynamics on demand (consensus, persistent disagreement, committed-minority influence), with persistent disagreement matching the FJ closed-form fixed points at R^2!=!0.93 – 0.99 . We further show that prescribed \kappa remains recoverable after the language round-trip, with perfect rank-order recovery across all four models. Explicit belief also makes simulation auditable: the layer surfaces systematic per-model stance biases that end-to-end simulation would silently absorb.

[MA-1] CityLearn v3: A Configurable Simulation and Evaluation Framework for Realistic Control Studies of Renewable Energy Communities

【速读】:该论文旨在解决可再生能源社区(Renewable Energy Communities, RECs)在实际运行中面临的动态参与、设备可用性变化、服务截止时间约束、数据质量波动等复杂条件下的控制策略评估难题。现有控制器研究常因过度简化这些现实因素,导致低电价或峰值负荷的优化结果掩盖了服务未履约或电力请求不可行的问题。为此,本文提出CityLearn v3——一个可配置的仿真与评估框架,用于在真实复杂的动态环境下对REC控制策略进行系统性研究。其关键在于:通过统一仿真环境集成成员与资产的动态变化、柔性负荷的截止时间、需求响应请求、本地能源共享以及数据或设备故障;引入建筑与相位功率限值以约束可调度请求,并采用声明式时间步长确保功率-能量核算一致性;同时记录控制器输入,明确区分请求动作与实际执行动作。该框架还提供基准控制器、面向服务与约束的性能指标及轨迹导出功能,支持跨社区与跨场景比较。通过软件校验与应用案例,验证了其在服务交付、电气约束满足、结算机制和情景演变分析中的有效性,并以合成高频数据重放实证表明聚合可能隐藏短时峰值但不改变年总能耗,从而实现宏观性能与微观服务失败、动作削减及参与者层面结果的协同解读。

链接: https://arxiv.org/abs/2609.21570
作者: Tiago Fonseca,Luis Lino Ferreira,Armando Sousa,Ava Mohammadi,Zoltan Nagy
机构: Unknown
类目: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI)
备注: 34 pages, 14 figures

点击查看摘要

Abstract:Renewable energy communities (RECs) coordinate buildings, photovoltaic generation, batteries, electric vehicles and flexible loads. Controller studies often simplify changing participation, equipment availability, service deadlines and data quality, so lower cost or peak demand can conceal missed services or infeasible power requests. This paper presents CityLearn v3, a configurable simulation and evaluation framework for REC control studies under these conditions. It represents changing members and assets, flexible-load deadlines, demand-response requests, local energy sharing, and data or equipment failures within one simulation environment. Building and phase power limits constrain controllable requests, while a declared timestep preserves consistent power-to-energy accounting. The framework records controller inputs and distinguishes requested actions from those applied to the simulated equipment. Reference controllers, service- and constraint-aware performance indicators, and trajectory exports support comparisons within and across communities. Software checks and application examples examine service delivery, electrical constraints, settlement and changing scenarios; a synthetic high-frequency trace replay illustrates how aggregation can conceal short peaks without changing annual energy. Together, these records allow aggregate performance to be interpreted alongside service failures, action reductions and participant-level outcomes.

[MA-2] AI-GRACE: A Use-Case Operationalization Framework for Agent ic AI: From Organizational Objectives and Obligations to Deployment Capabilities and Architecture

【速读】:该论文旨在解决组织在部署生成式AI(Generative AI)代理系统时,如何将治理、风险、保证、控制与证据(Governance, Risk, Assurance, Controls, and Evidence, GRACE)要求有效落地到具体应用场景中的核心挑战。现有方法往往缺乏从战略目标到技术实现的系统性映射机制,导致治理措施与实际运行脱节。其解决方案的关键在于提出AI-GRACE框架,通过设计科学方法论构建可操作的治理路径,并借助情境化方法工程(Situational Method Engineering)实现框架的定制化适配与复用。该框架首先明确业务目标与合规义务,识别七个关键风险领域(包括使命与价值实现),进而推导出部署前的保障需求、运行期控制机制及证据生成要求,形成能力评估、差距分析与逻辑架构设计的基础。其中,代理操作边界(Agent Operating Envelope)定义了允许行为与升级条件,而风险对齐独立性等级(Risk-Aligned Independence Levels, RAIL)则量化了授权自主程度。通过一个虚构的零售银行应用案例验证了该框架的可行性,其核心贡献在于提供了一套可追溯的决策依据,帮助组织清晰界定必须实施、已有支持及仍待解决的治理要素,从而提升部署决策质量、执行效率与方法复用性。未来需通过实证研究进一步验证其有效性。

链接: https://arxiv.org/abs/2609.21192
作者: John Cuneo,David Chun,Gaurav Khanna
机构: 未知
类目: Artificial Intelligence (cs.AI); Computers and Society (cs.CY); Multiagent Systems (cs.MA)
备注: 31 pages, 2 figures, 6 tables

点击查看摘要

Abstract:Organizations deploying agentic artificial intelligence must determine more than whether a model is trustworthy; they must establish what to validate, control, and observe for a use case to deliver its intended outcome while meeting applicable obligations. This paper proposes AI-GRACE (Agentic Intelligence-Governance, Risk, Assurance, Controls, and Evidence) as a use-case operationalization framework connecting organizational governance with technical implementation. The proposal draws on professional observations and a purposive synthesis of standards and literature, using design science to frame the method contribution and situational method engineering to guide contextual tailoring and reuse. The framework establishes objectives and obligations and then assesses risks in seven proposed domains, including mission and value realization. It derives requirements for assurance before deployment, runtime controls, and evidence, which guide capability qualification, gap assessment, and a logical architecture. An Agent Operating Envelope specifies permitted actions and escalation conditions, while Risk-Aligned Independence Levels (RAIL) summarize the authorized independence. A fictional retail banking application illustrates the method. The contribution is a traceable basis for deciding what an organization must implement, what it already supports, and what remains unresolved. Empirical evaluation must establish whether it improves deployment decisions, efficiency, and reuse.

[MA-3] Loopjacking: Hijacking Human-in-the-Loop Approval

【速读】:该论文旨在解决生成式智能体(Generative AI Agent)在执行关键操作前,人类审批环节中存在的“审批绑定失效”问题,即所谓的“环路劫持”(Loopjacking)。其核心问题是:人类用户批准的指令(操作A)与实际被执行的操作(操作B)之间存在不一致,导致用户在不知情的情况下授权了与预期实质不同的行为,从而破坏了人机协作的安全边界。解决方案的关键在于确保审批时所呈现的操作与最终执行的操作在语义和状态层面完全一致,具体包括两个核心机制:一是实现完整的规范性审批渲染(complete canonical approval rendering),保证审批界面展示的内容与系统内部定义的操作完全一致;二是通过精确的使用时比对(exact use-time comparison)或禁止未授权的待执行状态修改(preventing unauthorized pending-state mutation),防止在审批后发生状态替换或表示错位。实验验证表明,采用序列化延续(serialized continuation)机制的OpenAI Agents SDK 0.22.0/0.22.2能有效阻断攻击,而其他多个主流智能体系统版本则暴露于不同形式的环路劫持风险中。该研究将此贡献独立于传统的误导对话、会话走私、动作绑定及授权连续性等已有工作,强调了在动态执行环境中维持操作可追溯性和语义一致性的重要性。

链接: https://arxiv.org/abs/2609.21081
作者: Adithyan Arun Kumar
机构: 未知
类目: Cryptography and Security (cs.CR); Multiagent Systems (cs.MA)
备注: 17 pages, 3 figures, 3 tables. Evidence archive: this https URL

点击查看摘要

Abstract:Human approval is often treated as the last security boundary before an agent executes a consequential operation. That boundary is only meaningful if the operation presented for review is the operation later authorized or released. We call failures of this binding Loopjacking: a human approves what they understand as operation A, while the implementation uses that decision for a materially different operation B. We distinguish two variants. In a representation-based attack, B is already encoded but omitted or misrepresented at approval time; in a post-approval state-substitution attack, the human sees the correct A and mutable workflow state later replaces it with B. We evaluate a purposive set of released agent products. We reproduce post-approval substitution in seven tested Agno AgentOS releases ending at 3.0.9 and in 12 tested versions of a conditional in-memory LangGraph Agent Server composition ending at 0.14.0. We reproduce representation mismatch in OpenClaw 2026.2.23 and its rejection in 2026.2.24. OpenAI Agents SDK 0.22.0 and 0.22.2 provide a negative control: serialized continuation preserves exact per-call binding and rejects mutated B. These results do not estimate ecosystem prevalence. They show that complete canonical approval rendering and exact use-time comparison, or preventing unauthorized pending-state mutation, block the tested attacks while preserving legitimate execution. We separate this contribution from established work on misleading dialogs, session smuggling, action binding, and authorization continuity. Comments: 17 pages, 3 figures, 3 tables. Evidence archive: this https URL Subjects: Cryptography and Security (cs.CR); Multiagent Systems (cs.MA) Cite as: arXiv:2609.21081 [cs.CR] (or arXiv:2609.21081v1 [cs.CR] for this version) https://doi.org/10.48550/arXiv.2609.21081 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[MA-4] Proxifield: Decentralized Multi-Agent Communication through Semantic Proximity

【速读】:该论文旨在解决大规模多智能体系统中因传统通信协议结构僵化而导致的协调瓶颈问题,尤其是在智能体数量增加时性能显著下降的挑战。其核心解决方案是提出一种无需模型训练或中心化规划器的去中心化多智能体通信协议——Proxifield,该协议通过在推理阶段动态生成四类路由信号(直接指派、信息需求、计划对齐与信息互补性),基于智能体间语义相似性的演化构建稀疏通信图。这一机制实现了对通信关系的语义自适应调整,显著提升了系统的可扩展性与容错能力。实验结果表明,在无人机搜救与集体推理基准测试中,随着智能体规模扩大,Proxifield相较中心化星型协议(Star)和去中心化共享上下文协议(Shared Context)展现出更优的任务奖励表现,且在极端故障场景下仍能保持73.6%的原始任务奖励,远超其他基线方法,验证了其在复杂环境下高效、鲁棒的协同能力。

链接: https://arxiv.org/abs/2609.20889
作者: Pradyumna Tambwekar,Yenchia Feng,Deep Patel,Karime Maamari
机构: Distyl AI(迪斯蒂尔人工智能)
类目: Multiagent Systems (cs.MA)
备注:

点击查看摘要

Abstract:As LLM capabilities have expanded, multi-agent communication has emerged as an increasingly active area of research. Prevailing protocols often adopt rigid structures that introduce coordination bottlenecks and can degrade as the number of agents increases. We introduce Proxifield, a round-adaptive multi-agent protocol with decentralized agent decision-making that constructs sparse communication graphs from the evolving semantic proximity of agents. Without model training or a centralized planner, Proxifield connects agents using four routing signals derived at inference time: direct address, information needs, plan alignment, and information complementarity. We compare Proxifield with two representative coordination baselines, a centralized Star protocol and a decentralized Shared Context protocol, across two domains: Drone Search and Rescue and the collective-reasoning benchmark HiddenBench. We first ablate base-model capability and find that, in both domains, the performance of Proxifield improves with model size (35B - 397B parameter model) and Proxifield outperforms all baselines at the largest scale. As team size increases, Proxifield’s task-reward advantage over Star widens from 5.4% at (N=5) to 53.0% at (N=25) and 59.5% at (N=50), while Shared Context consistently underperforms both protocols. Proxifield is also substantially more robust to permanent agent failure, retaining 73.6% of its no-failure task reward under the most severe condition, compared with 58.3% for Shared Context and 38.8% for Star. These results demonstrate that decentralized, semantically adaptive routing can improve the scalability and fault tolerance of multi-agent systems.

[MA-5] BirdsongChat: A Hybrid Multi-Agent Framework for Multimodal Embodied Behavior Simulation EMNLP2026

【速读】:该论文旨在解决多模态具身系统中人类意图到跨异构模态可解释且协调行为的映射问题,尤其针对现有方法依赖隐式表征所导致的可控性差与跨模态一致性不足的瓶颈。其解决方案的关键在于提出一种统一参数表征(Unified Parameter Representation, UPR),通过基于大语言模型(LLM)的推理代理将多模态输入转化为显式的UPR,该表征编码了行为状态与可解释的控制参数,进而驱动仿真代理同步生成三维运动、空间化声景及环境行为。实验以BirdsongChat为原型系统,在鸟类交互行为模拟任务中验证了该框架在文本与图像引导下的跨模态连贯性(94.4%)、情感一致性(100%)及生成一致性(92.6%)表现优异,证明显式中间表征能有效衔接语义推理与物理执行,显著提升多模态协同的可控性与同步性,为具身人工智能系统中语义到物理的跨模态协调提供了一种可泛化的设计范式。

链接: https://arxiv.org/abs/2609.20887
作者: Callie C. Liao,Duoduo Liao,Ellie L. Zhang
机构: Stanford University (斯坦福大学); George Mason University (乔治梅森大学); IntelliSky
类目: Multiagent Systems (cs.MA)
备注: Paper contents accepted by EMNLP 2026 REALM

点击查看摘要

Abstract:Multimodal embodied systems require translating human intentions into interpretable and coordinated behaviors across heterogeneous modalities. However, existing multimodal agents often rely on implicit representations, limiting controllability and cross-modal consistency. We present a hybrid multi-agent framework for interactive multimodal behavior simulation that bridges semantic reasoning and physical execution through a Unified Parameter Representation (UPR). LLM-based reasoning agents transform multimodal inputs into UPR, which encodes behavioral states and interpretable control parameters for simulation agents generating synchronized 3D motion, spatialized soundscapes, and environmental behaviors. We develop BirdsongChat as a prototype implementation of the proposed framework, using interactive avian behavior simulation as a testbed that tightly couples motion, vocalization, and environmental context. BirdsongChat is evaluated on text- and image-guided scenarios involving species, behaviors, affective states, environments, and multi-bird interactions. The system achieves normalized scores of 94.4% for cross-modal coherence, 100% for affective consistency, and 92.6% for generation consistency. These results demonstrate that an explicit intermediate representation effectively bridges semantic reasoning and physical execution, improving controllability and multimodal synchronization. The proposed framework thus offers a generalizable design principle for embodied AI systems requiring interpretable semantic-to-physical coordination across modalities, with potential applications in bio-inspired ecoacoustics, swarm robotics, virtual environments, and creative multimedia.

[MA-6] Guiding Agents of Quantum Games to Equilibrium using Matrix Exponential Fixed-Point Iteration

【速读】:该论文旨在解决多智能体量子博弈中均衡策略计算困难的问题,其核心挑战在于联合希尔伯特空间维度随参与者局部维度的乘积呈指数增长,导致传统方法在计算上不可行。解决方案的关键在于提出一种扩展的古托斯基-瓦特罗斯(Extended Gutoski-Watrous, EGW)博弈模型,并通过推导收益函数及其梯度的张量收缩表达式,避免显式构造全联合密度矩阵及其与收益算符的高成本矩阵乘法。基于所得的有效哈密顿量,本文进一步提出了带有退火机制的矩阵指数定点迭代算法(Matrix Exponential Fixed-Point Iteration with Annealing, MEFPIA),用于高效搜索博弈均衡点。实验结果表明,相较于矩阵乘法权重更新(Matrix Multiplicative Weights Update, MMWU)算法,MEFPIA在相同参数设置下收敛至相同的策略组合与收益,且以更少的迭代次数实现更低的相对误差,验证了其在求解多智能体量子博弈均衡中的有效性与优越性。

链接: https://arxiv.org/abs/2609.21944
作者: Alireza Habibi,Luis F. Abanto Leon,Setareh Maghsudi
机构: Ruhr University Bochum (鲁尔大学波鸿分校); Ruhr University Bochum (鲁尔大学波鸿分校)
类目: Quantum Physics (quant-ph); Computer Science and Game Theory (cs.GT); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
备注:

点击查看摘要

Abstract:In recent years, quantum game theory has gained significant attention as a framework for studying decision-making in multi-agent systems using quantum principles. However, computing equilibrium strategies is challenging because the dimension of the joint Hilbert space grows as the product of the players’ local dimensions. In this paper, we consider an extended Gutoski-Watrous (EGW) game in which each player’s quantum strategy is represented by a local density matrix. We derive tensor-contraction expressions for the payoff functions and their gradients, thereby avoiding the explicit construction of the full joint density matrix and its computationally expensive multiplication by the payoff operators. Building on the resulting effective Hamiltonians, we propose the Matrix Exponential Fixed-Point Iteration with Annealing (MEFPIA) algorithm to search for equilibrium points in EGW games. We compare MEFPIA with the Matrix Multiplicative Weights Update (MMWU) algorithm in terms of convergence. For the tested instances and parameter settings, both algorithms approach the same strategy profiles and payoffs, while MEFPIA achieves lower relative error in fewer iterations. These results indicate that MEFPIA is a promising numerical method for equilibrium search in multi-agent quantum games. Our findings provide important insights into the quantum game theory’s potential for addressing complex decision-making processes, as well as opening up new paths for future research and exploration in multi-agent quantum systems.

自然语言处理

[NLP-0] Cross-sector generalization of accident-process role classification in occupational accident narratives

【速读】: 该论文旨在解决跨行业职业事故叙述中事故过程角色分类的泛化能力问题,即现有自动化编码系统在不同行业和组织间因术语与写作风格差异而难以有效迁移的问题。其核心解决方案在于通过构建一个由专家标注的法语职业事故叙述语料库,将事实单元划分为四类角色:工作情境(A0)、明确报告的不利条件(A1)、事故事件或偏差(B)以及报告后果(C),并在仅基于建筑行业6,040篇叙述中的42,244个事实单元训练模型的前提下,评估其在冶金和化学—塑料行业及独立收集的企业数据集上的跨领域表现。研究对比了冻结预训练表示、任务特定微调与监督表示学习策略,结果表明任务特定适应显著优于冻结表示,三种最优适配策略在多次训练中于三个目标语料库上实现了85.6%至85.8%的平均平衡准确率,证明了可迁移的辅助编码系统具备一致结构化异质事故叙述的能力,为专家审查与跨行业风险预防分析提供了可靠支持。

链接: https://arxiv.org/abs/2609.22081
作者: Aho Yapi,Pierre Latouche,Arnaud Guillin,Yan Bailly
机构: Laboratoire de Mathématiques Blaise Pascal (UMR 6620 CNRS), Université Clermont Auvergne (克莱蒙奥弗涅大学数学实验室(法国国家科学研究中心6620联合研究单位)); LYF SAS (LYF SAS); Institut Universitaire de France (法国高等教育与科研部大学研究院)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Occupational accident narratives contain valuable information about work situations, unfavourable conditions, accident events, and their consequences. Automatically structuring these narratives can facilitate large-scale accident analysis and support occupational risk prevention. However, the terminology and writing styles used to describe accidents vary considerably across sectors and organisations, raising questions about the ability of automated coding systems to generalize beyond their training domain. In this paper, we evaluate the cross-sector generalization of accident-process role classification in French occupational accident narratives. We construct an expert-annotated corpus in which factual units are classified into four roles: work situation (A0), explicitly reported unfavourable condition (A1), accident event or deviation (B), and reported consequence ©. The role classifiers are developed and selected exclusively on 42,244 factual units extracted from 6,040 construction-sector narratives and are then evaluated on unseen corpora from the metallurgy and chemistry–plastics sectors, as well as on an independently collected company corpus, without retraining or target-domain tuning of the role classifier. We compare frozen pretrained representations with task-specific fine-tuning and supervised representation-learning strategies. The results show that task-specific adaptation consistently improves cross-domain transfer over frozen representations. Across repeated training runs, the three leading task-adapted strategies achieved average balanced accuracies between 85.6% and 85.8% across the three target corpora. These findings support the development of transferable assisted-coding systems capable of consistently structuring heterogeneous occupational accident narratives for expert review and cross-sector prevention analysis.

[NLP-1] An Interpretable Memory Decision Controller for LLM Agents Based on Three-Signal Complementarity: Decoupling Confidence and Consistency

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在使用检索增强生成(Retrieval-Augmented Generation, RAG)技术时,因记忆系统缺乏对检索内容可信度判断而导致的幻觉(Hallucination)问题。尤其当记忆库中存在相互矛盾的信息时,传统RAG方法会无差别地注入所有检索结果,从而放大错误信息传播。其核心解决方案是提出一种零参数的“记忆决策层”(Memory Decision Layer, MDL),作为检索与生成之间的可控中介。MDL的关键在于基于前额叶皮层记忆信号机制设计的三信号互补编码器,通过QR分解实现正交子空间投影,融合相关性、可靠性与任务风险三重信号,并引入元工作记忆信号,生成可解释的可信度表示。该架构实现了置信度与一致性解耦,支持风险反转与显式拒绝响应。实验表明,MDL在多数场景下将冲突记忆下的幻觉率降低约56.04%,在高风险场景中几乎消除幻觉。该控制器为全白盒设计,仅依赖几何运算,无需训练参数,单次决策延迟仅约0.14毫秒,远快于嵌入-检索步骤(约50倍更快)及大模型自评估调用(快四至五个数量级)。

链接: https://arxiv.org/abs/2609.22043
作者: Yiming Zhang,Jinghong Zhang,Haoran Zhao,Yiren Ma,Chunlei Zhao
机构: Tianjin University of Technology (天津理工大学)
类目: Computation and Language (cs.CL)
备注: 17 pages, 6 figures, 10 tables

点击查看摘要

Abstract:Memory systems for large language models have focused predominantly on efficient retrieval, whereas the decision of whether retrieved memories should be trusted has received comparatively little attention. When the memory store contains conflicting positions, standard retrieval-augmented generation (RAG) blindly injects memories and amplifies hallucinations: in models susceptible to memory injection, the RAG hallucination rate under conflicting memories is markedly higher than that of a memory-free baseline. Inspired by memory signaling mechanisms in the prefrontal cortex, we propose the Memory Decision Layer (MDL), a zero-parameter memory decision controller situated between the retrieval and generation stages. Its core is a three-signal complementary encoder that fuses relevance, reliability, and task risk through QR-based orthogonal subspace projection and a meta-working-memory signal into an interpretable decision representation that quantifies the trustworthiness of retrieved memories. Building on this encoder, MDL explicitly decouples confidence from consistency and introduces risk inversion and explicit abstention. Evaluations on mainstream large language models and multiple open-source datasets show that MDL reduces the hallucination rate under conflicting memories by about 56.04% in general scenarios and approaches zero hallucination in high-risk scenarios. The controller is fully white-box: it relies purely on geometric operations, requires no trained parameters, and adds only about 0.14 ms per decision – roughly 50x faster than the embedding-retrieval step that precedes it and four to five orders of magnitude faster than an LLM self-evaluation call.

[NLP-2] QuranicMMLU: A Cognitively-Aware Benchmark for Evaluating Generative AI Solutions on Quranic Linguistic Knowledge

【速读】: 该论文旨在解决现有生成式人工智能(Generative AI)在《古兰经》阿拉伯语评估中缺乏多维度语言复杂性探测的问题。当前的《古兰经》基准测试主要聚焦于通用问答与语义检索任务,未能深入考察特定语言能力,也未根据认知需求层级和经文难度进行分层设计。为此,研究提出一个五支柱的《古兰经》语言学分类体系,涵盖语音学(Phonology)、形态学(Morphology)、句法学(Syntax)、语义学(Semantics)和语用学(Pragmatics),共包含31个细分语言现象,从诵读规则(tajwīd)与词根-模式构词到启示背景及跨章节连贯性。针对每个细分类别,构建了基于布卢姆认知层次(Bloom’s cognitive level)和经文困惑度(verse perplexity)分层的问题,并采用大语言模型(LLM)作为独立评分者对每道题进行回答与打分,再将结果路由至人工审核。最终形成包含980个经人工审核的问题数据集,每个问题均以开放式与选择题两种形式呈现。在12个系统上的基准测试显示,伊斯兰专精模型表现最优,但所有系统在选择题准确率(平均84%)上显著高于开放式回答质量(平均60%),两者排名高度一致(Kendall’s τ=0.73),然而选择题评分掩盖了在去除选项后暴露的模型失败。因此,QuranicMMLU提供了一个严谨、语言学基础坚实的框架,用于评估阿拉伯语自然语言处理在《古兰经》领域的实际能力。

链接: https://arxiv.org/abs/2609.22038
作者: Rawan El Ghali,Umm Kulsoom,Anas Madkoor,Dima Faris Alsaudi,Roaa Abdelmagid,Roaa Ibrahim,Raghad Mousa,Hamza Aljaji,Abdullah Khanafer,Abdallah Alkanani,Salah Feras Alali,Rawan Khaled Mohamed,Ehsaneddin Asgari
机构: Qatar Computing Research Institute, Hamad Bin Khalifa University (卡塔尔计算研究研究院,哈马德·本·哈利法大学); Carnegie Mellon University Qatar (卡内基梅隆大学卡塔尔分校); University of Doha for Science and Technology (多哈科学与技术大学); Qatar University (卡塔尔大学); American University of Beirut (贝鲁特美国大学); University of Birmingham (伯明翰大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:We introduce QuranicMMLU, a benchmark for evaluating generative AI on Quranic Arabic across multiple dimensions of linguistic complexity. Existing Quranic benchmarks center on general question answering and semantic retrieval, without probing specific linguistic competencies or stratifying by cognitive demand and verse difficulty. We construct a five-pillar Quranic taxonomy spanning Phonology, Morphology, Syntax, Semantics, and Pragmatics, with 31 leaves covering phenomena from tajwīd and root-and-pattern morphology to occasions of revelation and inter-surah coherence. For each leaf we generate questions stratified by Bloom’s cognitive level and verse perplexity, then have LLM as a judge to independently answer and score every item and route the annotations to manual review. The resulting dataset comprises 980 human-reviewed questions, each issued in both open-ended and multiple-choice form. We benchmark 12 systems on these items and find that the Islamic-specialized model leads, yet every system scores higher on multiple-choice accuracy (average 84%) than open-ended answer quality (average 60%): the two rankings agree closely (Kendall’s \tau=0.73), but multiple-choice scoring hides failures that surface only once answer choices are removed. QuranicMMLU thus offers a rigorous, linguistically grounded framework for evaluating Arabic NLP in the Quranic domain.

[NLP-3] DiaVLo: Diagnosing Behaviours of Vision-Language Models EMNLP2026

【速读】: 该论文旨在解决视觉语言模型(Vision-Language Models, VLMs)在实际部署中难以验证其行为是否符合预期、是否存在有害行为的关键问题。由于现有方法在识别和诊断VLM行为方面仍显不足,导致模型的可靠性与安全性难以保障。为此,论文提出DiaVLo诊断框架,其核心解决方案在于结合人工标注与VLM自身的生成能力,构建出期望行为与实际观测行为的规范描述,从而揭示潜在的行为错位(misalignment)。此外,DiaVLo进一步提供因果估计,识别对模型行为最具影响力的语义概念,实现对行为驱动机制的可解释性分析。实验结果表明,DiaVLo生成的行为标签与模型性能高度相关,并能为性能评估提供上下文依据;同时,该框架成功识别出明显对齐与错位的行为模式,揭示了VLM在概念感知、组织与优先级排序方面的内在规律。

链接: https://arxiv.org/abs/2609.22008
作者: Lorenzo Corti,Jie Yang
机构: Delft University of Technology (代尔夫特理工大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 34 pages. To appear in EMNLP 2026 (findings)

点击查看摘要

Abstract:Vision-language models (VLMs) rely on storing and transferring appropriate information across their sub-components. Verifying that the VLMs exhibit desired behaviours, while avoiding harmful ones, is central to their reliable deployment. Yet, methods that identify VLM behaviours remain scarce. We present DiaVLo, a diagnostic framework that leverages human curation and VLMs’ generation capabilities to construct specifications of desired and observed VLM behaviours, surfacing potential misalignments. Beyond this, DiaVLo also provides causal estimates to identify the most influential concepts steering VLM behaviours. We evaluate DiaVLo on several open-source VLMs under both classification and generation conditions. Our experiments show that DiaVLo produces behaviour labels that correlate with model performance and provide context for measured performance. DiaVLo surfaced behaviours that are clearly aligned and misaligned, alongside patterns in how VLMs perceive, organise, and prioritise concepts.

[NLP-4] Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention

【速读】: 该论文旨在解决生成式语言模型预训练中注意力机制(attention)路径优化的争议问题,即为何门控(gating)机制能提升模型性能,而现有研究对此存在分歧。其核心解决方案的关键在于揭示门控机制引入了两个传统softmax注意力所不具备的重要功能:回避(abstention)噪声过滤(noise filtering)。其中,回避允许注意力头在特定情况下输出零值,从而摆脱注意力权重必须归一化的约束;噪声过滤则通过在值路径上引入门控机制,抑制残差流中叠加特征带来的干扰。在参数规模从1000万到3500万的对比实验中,作者分别通过学习每个头的“下沉对数偏置”(sink logit)实现回避,并通过值路径上的门控实现噪声过滤。实验发现:第一,回避带来的性能增益随模型规模增长而下降,而噪声过滤的增益随规模上升,在350M模型中,噪声过滤贡献了绝大部分改进;第二,所有规模下最优模型均同时具备两种机制;第三,通过向值路径注入可控干扰,验证了门控确实能有效移除干扰,并揭示了两种门控形式各自存在特定盲区。该方案仅增加极少参数,且与键值缓存(key-value cache)兼容,具有良好的可扩展性与实用性。

链接: https://arxiv.org/abs/2609.22005
作者: Richard Zhe Wang
机构: St. John Fisher University (圣约翰费舍尔大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 21 pages (8 pages main text plus appendices), 5 figures, 12 tables

点击查看摘要

Abstract:Gating the value pathway of attention reportedly improves language model pretraining, and prior studies disagree on why. We argue and provide experimental evidence that such gates supply two different things that softmax attention lacks: abstention and noise filtering. The first is abstention, which allows an attention head to output nothing, bypassing the requirement that attention weights must sum to one. The second is noise filtering, which allows the value pathway of an attention head to suppress interference from superposed features in the residual stream. In our experiments in matched models from 10M to 350M parameters, we supply abstention through a learned per-head sink logit in the softmax and noise filtering through a gate on each value. We report three empirical findings. First, the benefit of abstention, measured as the reduction in validation loss relative to a matched baseline, declines as models grow, whereas the benefit of noise filtering increases with scale. In particular, abstention accounts for nearly all of the gain from gating at 10M and filtering for most of it at 350M. Second, the best model at every scale is the one with both primitives built in. Third, injecting controlled interference into the values a head reads confirms that the gate removes such interference, and reveals that each of the two gate forms we study has a characteristic blind spot. Supplying both primitives adds negligible parameters and remains compatible with the key-value cache.

[NLP-5] RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents

【速读】: 该论文旨在解决当前计算机使用代理(Computer-use Agents, CUAs)在图形化交互与代码驱动的软件开发之间割裂的问题,即现有系统通常将二者作为独立路径处理,而真实数字工作场景要求两者以交错方式协同执行。其核心挑战在于如何构建能够自主决策何时探索界面、编写代码、运行程序并视觉验证输出结果的混合型代理。解决方案的关键在于提出RecreationWorld框架——一个基于“复刻”任务范式的五平台统一环境(涵盖Ubuntu、macOS、Windows、Android及Web),支持跨平台可复现的实验设置;该框架通过提供运行中的参考实例作为行为“预言机”(oracle),为代理提供基于执行结果的奖励信号,并配备原生GUI控制与编码工具的统一调度器。研究通过大规模高质量开源应用生成轨迹数据训练模型,使代理在五个分布外的编码与混合计算机使用基准上表现提升,且更频繁地对生成输出进行视觉验证,表明其具备超越复刻任务本身的泛化能力。为进一步评估,论文引入RecreationBench基准测试集,包含250个跨领域、跨平台的多样化任务,采用参考基线的程序化与视觉断言进行多层次交互结果验证,并经人工审核后冻结用于自动化评分。实验结果显示,尽管GPT-6 Astra在总体得分上达到58.1%领先,但仅在2.8%的任务中完全通过所有程序化测试,反映出代理在动态交互和计算输出复现方面仍存在显著不足,而静态界面结构的复现相对更可靠,生成的应用也普遍比参考实现更小且更单体化。研究已公开发布基准、环境与测试套件,推动混合型计算机使用代理的发展。

链接: https://arxiv.org/abs/2609.22000
作者: Shuai Bai,Jiayong Deng,Yikun Fu,Chang Gao,Xuhao Hu,Mianqiu Huang,Yizhen Jiang,Yuheng Jing,Dehui Kong,Keliang Li,Ning Li,Wanli Li,Dayiheng Liu,Dunjie Lu,Changwei Luo,Que Shen,Zheyuan Wang,Zijian Wang,Jie Wu,Gao Wu,Zhihui Xie,Rui Xie,Haiyang Xu,An Yang,Jiakang Yuan,Yanming Zhang,Jiajun Zhang,Xi Zhang,Zhenru Zhang,Zhuo Zhen,Mingkang Zhu,Bowen Zhou
机构: Alibaba Token Hub(阿里巴巴令牌中心); Alibaba Group(阿里巴巴集团)
类目: Computation and Language (cs.CL); Software Engineering (cs.SE)
备注:

点击查看摘要

Abstract:Computer-use agents (CUAs) have advanced along two separate lines: graphical interaction and software development through code and the command line. Real digital work requires both, interleaved rather than stacked end to end. We study hybrid CUAs that autonomously decide when to explore an interface, implement software, and run and visually verify their artifacts. We introduce RecreationWorld, a five-platform framework built around recreation: given a running reference, an agent must discover its behavior and build a faithful implementation with no prescribed workflow. RecreationWorld provides reproducible environments on Ubuntu, macOS, Windows, Android, and Web, plus a unified harness with native GUI control and coding tools. The running reference serves as an oracle for hidden behavioral tests, providing execution-grounded rewards. We scale trajectory generation with high-quality open-source applications. Models trained on these trajectories improve across five out-of-distribution coding and hybrid computer-use benchmarks and more frequently verify their rendered outputs, providing evidence of transfer beyond recreation. For held-out evaluation, we introduce RecreationBench, comprising 250 diverse tasks across domains and platforms. Reference-grounded programmatic and visual assertions cover action-conditioned outcomes at multiple interaction depths; each is validated on the reference and by human reviewers before the suite is frozen for automatic scoring. GPT-6 Astra leads at 58.1% overall, but passes all programmatic tests on just 2.8% of tasks. Agents reproduce static interface structure more reliably than interactions and computed outputs, while generated applications remain smaller and more monolithic than their references. We release the benchmark, environments, and test suites.

[NLP-6] Moral Entropy: Auditing Bias and Uncertainty in Moral Judgment EMNLP2026

【速读】: 该论文旨在解决当前计算伦理(computational ethics)研究中对标注者分歧(annotator disagreement)处理方式的固有缺陷:现有方法通常将标注分歧视为噪声,简单通过多数投票或“任意标注者规则”(any-annotator rule)进行聚合,忽视了其背后蕴含的道德判断不确定性。这种做法不仅掩盖了真实分歧,还可能导致模型训练和评估引入系统性偏差。论文提出的核心解决方案是引入道德熵(Moral Entropy)——一种基于贝叶斯框架的建模方法,能够保持对真实标签的完整后验分布,并将其熵分解为两类不确定性:随机性不确定性(aleatoric uncertainty,不可消除的道德分歧)认知性不确定性(epistemic uncertainty,由标注不足或噪声引起)。该框架允许通过交叉熵、KL散度、Brier评分及期望校准误差等熵度量方法,对现有的共识规则进行可解释、可审计的校准评估。在三个语料库和十五个话语领域上的实证分析表明,标准聚合规则存在显著偏差:任意标注者规则在约30%的样本上与校准后的后验不一致,且主要表现为几乎全部的假阳性;而更严格的多数投票和双人同意规则则漏检了63%-83%的真实正例。这一发现揭示了现有数据处理流程中未被报告的系统性偏误,强调了对标注不确定性进行建模与量化的重要性。

链接: https://arxiv.org/abs/2609.21992
作者: Maciej Skorski
机构: University of Luxembourg(卢森堡大学)
类目: Computation and Language (cs.CL); Computers and Society (cs.CY); Machine Learning (stat.ML)
备注: accepted to UncertaiNLP @ EMNLP 2026

点击查看摘要

Abstract:Most work in computational ethics treats annotator disagreement on moral content as noise to be voted away, collapsed into majority vote or the more permissive any-annotator rule the moment a single annotator flags an item. We argue this uncertainty should instead be modeled and learned from. We introduce Moral Entropy, a Bayesian framework that keeps a full posterior over the true label and decomposes its entropy into aleatoric uncertainty (irreducible disagreement about the moral content) and epistemic uncertainty (from insufficient or noisy annotation) – and lets any heuristic consensus rule be audited against a calibrated ground truth via entropy methods such as cross-entropy/KL, Brier score, and expected calibration error. Across three corpora and fifteen discourse domains, auditing the standard aggregation rules against this posterior reveals bias that no current pipeline reports: the any-annotator rule disagrees with the calibrated posterior on roughly 30% of items – pooled, almost entirely false positives, though the errors invert at the foundation level (19.9%/38.9% mean FPR/FNR on MFTC) – while the stricter majority and two-vote rules miss 63-83% of true positives. Comments: accepted to UncertaiNLP @ EMNLP 2026 Subjects: Computation and Language (cs.CL); Computers and Society (cs.CY); Machine Learning (stat.ML) MSC classes: 62F15, 94A17, 91E10 ACMclasses: I.2.7; G.3; J.4 Cite as: arXiv:2609.21992 [cs.CL] (or arXiv:2609.21992v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2609.21992 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[NLP-7] NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities

【速读】: 该论文旨在解决开放域全双工语音对话系统中多模态能力集成与实时交互行为保持之间的矛盾问题,即如何在保证自然对话时序特性的前提下,实现语音识别、语言理解、工具调用与语音生成的统一高效协同。其解决方案的关键在于提出一种基于流式架构的端到端语音到语音模型——NemotronLabs VoiceChat,通过并行设计专用输出通道以分别处理代理文本输出与结构化函数调用,结合流式语音编码器、解码器仅语言模型(decoder-only language model)、辅助RNN-T分支实现增量用户语音转录,以及流式文本转语音(TTS)解码器,构建了一个集听、转录、推理、工具调用与发声于一体的统一流式系统。该架构在保持低延迟和高响应性的同时,实现了对用户打断的100%接管率及高质后中断响应,显著提升了全双工对话的真实感与可用性。

链接: https://arxiv.org/abs/2609.21967
作者: Jagadeesh Balam,Travis Bartley,Edresson Casanova,Sanjay Chauhan,Chen Chen,Zhehuai Chen,Zijia Chen,Francesco Ciannella,Slyne Deng,Mikyas Desta,Harishchandra Dubey,Slim Essid,Nourchene Ferchichi,Boris Ginsburg,Mariana Graterol Fuenmayor,Negar Habibi,Kevin Hu,Anand Joseph,Viraj Karandikar,Myungjong Kim,Viacheslav Klimkov,Seelan Lakshmi Narasimhan,Lily Lee,Jason Li,Eileen Long,Ameya Mahabaleshwarkar,Aditya Malte,Adi Margolin,Sasha Meister,Valentin Mendelev,Oluwatobi Olabiyi,Ankita Pasad,Yifan Peng,Elena Rastorgueva,Jayda Ritchie,Jason Roche,Nikhil Srihari,Yuanhang Su,Yoshi Suhara,Viet Anh Trinh,Jinhan Wang,Piotr Zelasko,Hui Wang,Puhui Meng,Chaosen Zhang,Yunsheng Liu,Shawn Wang,Wenjing Li,Zhonglei He
机构: NVIDIA(英伟达)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:We introduce NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool-calling capabilities. NemotronLabs VoiceChat combines a streaming speech encoder and decoder-only language model with parallel specialized output streams for agent text and structured function calls, an auxiliary RNN-T branch for incremental user transcription, and a streaming TTS decoder. This design enables the model to listen, transcribe, reason, invoke tools, and speak within a unified streaming architecture while preserving the temporal behavior required for natural conversation. On Full-Duplex-Bench 1.0, NemotronLabs VoiceChat achieves the lowest pause-handling takeover rates among evaluated open-weight systems, 100% takeover following user interruptions, and a 4.33/5 post-interruption response-quality score. On Full-Duplex-Bench 1.5, it resumes its response after user backchannels in 93% of cases. NemotronLabs VoiceChat obtains a 55.1 normalized average on VoiceBench and, on Full-Duplex-Bench 3.0 (FDB 3.0), achieves 82.5% tool-selection F1, while argument accuracy and end-to-end tool execution remain areas for improvement. These results demonstrate that full-duplex interaction, speech recognition and generation, general language capabilities, and external tool use can be integrated in a single open speech-to-speech model without sacrificing real-time conversational behavior.

[NLP-8] Detecting Pretraining Data in Large Language Models from a Free-Energy Perspective

【速读】: 该论文旨在解决大语言模型中预训练数据检测的难题,即高似然性可能源于模型对训练数据的暴露(membership)或其强大的泛化能力,导致现有基于似然性的检测方法在判断成员身份时易产生误判。其核心挑战在于如何有效区分由训练暴露引起的高似然与由强泛化能力带来的高似然。解决方案的关键在于提出一种具有倾斜边界的新型检测机制,该机制通过将预测损失(prediction loss)相对于预测熵(predictive entropy)进行归一化,实现对成员信号的熵校正。这种校正不仅能保留预期的成员信号,还能显著降低信号方差,从而提升成员与非成员之间的标准化分离效果。进一步地,研究将均值-方差分析拓展至存在非零熵间隙的更一般情形,并发现该熵调整后的评分可被解释为亥姆霍兹自由能(Helmholtz free energy),由此引出能量转移检测(Energy Transfer Detection, ETD)方法,从宏观残余自由能转移的角度重新审视预训练数据检测问题。大量实验表明,ETD在多种设置下均表现出最优的平均检测性能,平均AUROC提升达3.5%,在5%假阳性率下的真阳性率(TPR@5%FPR)最高提升5.1%,且具备良好的鲁棒性。

链接: https://arxiv.org/abs/2609.21888
作者: Chenye Ke,Zirui Liu,Qi Liu,Yan Zhuang,Jintao Zhang,Zhenya Huang,Shijin Wang
机构: University of Science and Technology of China (中国科学技术大学); Hefei Comprehensive National Science Center (合肥综合性国家科学中心); Nanjing University of Aeronautics and Astronautics (南京航空航天大学); iFLYTEK AI Research (Central China), iFLYTEK Co., Ltd (科大讯飞人工智能研究院(华中), 科大讯飞股份有限公司)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Detecting pretraining data in large language models is challenging because high likelihood can reflect either training exposure or strong generalization. In the joint space of prediction loss and predictive entropy, a likelihood-only detector uses a horizontal boundary and can mistake predictable non-members for members. Motivated by this, we introduce an inclined boundary that evaluates prediction loss relative to predictive entropy. Our analysis shows that entropy correction can preserve the expected membership signal while reducing its variance, thereby improving standardized member–non-member separation. We further extend the mean–variance analysis to the more general setting with a nonzero mean entropy gap. Interestingly, this entropy-adjusted score admits a Helmholtz free-energy interpretation, leading to Energy Transfer Detection (ETD), which views pretraining data detection from a macroscopic residual free-energy transfer perspective. Extensive experiments show that ETD achieves the best average detection performance, improving average AUROC by up to 3.5% and TPR@5%FPR by up to 5.1%, while remaining robust across diverse settings.

[NLP-9] rialAtlas: Multi-Agent Research Organization for Clinical Trial Design and Optimization

【速读】: 该论文旨在解决药物临床开发中高失败率与开发决策高度依赖人工、主观判断的问题。当前制药企业虽通过临床开发规划(CDP)和成功率评估来管理风险,但此类过程仍严重依赖跨领域专家在异构证据(如文献、竞品试验、监管先例等)上的协同获取、整合与推理,效率低且易受个体偏见影响。其解决方案的关键在于提出一个基于记忆增强的多智能体研究组织——TrialAtlas,该系统模拟真实研发团队协作流程,通过专业化智能体分别执行文献综述、竞品试验情报分析、监管先例研判以及综合性的试验设计与风险推理。此外,TrialAtlas能够从历史临床试验及监管结果(包括既往新药申请,NDAs)中学习,使决策具备经验基础。为验证其在真实监管场景下的有效性,研究构建了TrialAtlasBench基准数据集,涵盖291份美国FDA完整回应函,涵盖缺陷识别、设计改进建议和技管成功预测三项任务。实验表明,TrialAtlas在缺陷检测上达到50.0% F1值,优于最强基线6.1个百分点;在技术与监管成功预测方面,平衡准确率达85.3%,F1值达84.7%,较最优基线提升6.7个百分点(平衡准确率)和12.0个百分点(Cohen’s kappa)。专家评估显示,86.4%的TrialAtlas生成问题被判定为有效,显著高于OpenAI DeepResearch(83.1%)和Gemini DeepResearch(59.3%),证明其在复杂研发决策支持中的卓越性能与可靠性。

链接: https://arxiv.org/abs/2609.21859
作者: Jiacheng Lin,Zifeng Wang,Zheng Chen,Erick Scott,Ziwei Yang,Fanyang Yu,Sheng Zhong,Jimeng Sun
机构: Siebel School of Computing and Data Science, University of Illinois Urbana-Champaign(伊利诺伊大学香槟分校西贝尔计算与数据科学学院); Keiji AI Inc(凯吉人工智能公司); Institute of Scientific and Industrial Research, The University of Osaka(大阪大学产业科学研究所); University of Pennsylvania(宾夕法尼亚大学); AbbVie Inc(艾伯维公司); Carle Illinois College of Medicine, University of Illinois Urbana-Champaign(伊利诺伊大学香槟分校卡勒医学院)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Nearly 90% of drugs entering clinical development ultimately fail, despite billions of dollars in investment. Pharmaceutical companies therefore rely on clinical development planning (CDP) and probability of technical and regulatory success assessment to anticipate development risks, yet these decisions remain labor-intensive and subjective, requiring experts across clinical science, statistics, regulatory affairs, and competitive intelligence to jointly acquire, synthesize, and reason over heterogeneous evidence. Here, we introduce TrialAtlas, a memory-augmented multi-agent research organization for CDP that mirrors this collaborative process by coordinating specialized agents for literature synthesis, competitive trial intelligence, regulatory precedent analysis, and integrated reasoning over trial design and development risk. TrialAtlas further learns from historical clinical trials and regulatory outcomes, including prior New Drug Applications (NDAs), to ground its decisions in accumulated development experience. To evaluate these capabilities in an authentic regulatory setting, we introduce TrialAtlasBench, constructed from 291 FDA Complete Response Letters and spanning three practical tasks: detecting trial design deficiencies, recommending actionable design improvements, and predicting technical and regulatory success. TrialAtlas achieves an F1 score of 50.0% for deficiency detection, outperforming the strongest baseline by 6.1 points, and reaches 85.3% balanced accuracy and 84.7% F1 for prediction of technical and regulatory success, improving over the best baselines by 6.7 points in balanced accuracy and 12.0 points in Cohen’s kappa. In expert evaluation, 86.4% of TrialAtlas-generated concerns were judged valid, compared with 83.1% for OpenAI DeepResearch and 59.3% for Gemini DeepResearch.

[NLP-10] Do Personality-Tuned LLM s Make Better Social Agents ?

【速读】: 该论文旨在解决大语言模型(LLM)在社会模拟中生成对话时存在的“人格一致性不足”与“可控性差”的问题,即尽管模型能够模仿人类行为,但仍表现出一种难以忽视的“陌生感”(alienness)。其核心解决方案是通过人格感知的微调(personality-aware fine-tuning),利用包含人格标注的社会媒体文本与对话语料对小型开源模型(Qwen2.5-7B-Instruct 和 Ministral-8B-Instruct)进行训练,以构建更具人格一致性的对话引擎。关键在于引入结构化的人格标签数据,增强模型在多场景社会交互中对特定人格特征的稳定响应能力。然而,实验结果表明,微调后的模型在人格扮演表现上并未显著优于基线模型,且评估中存在较低的评价者间一致性(inter-rater agreement),削弱了结论的可信度;尽管如此,微调在提升Qwen模型的语言多样性方面显示出积极效果。研究指出,未来工作应更注重训练数据的质量与领域适配性,以实现更精准的人格角色扮演。

链接: https://arxiv.org/abs/2609.21857
作者: Tim Krabbe,Xiaodan Shi
机构: Stockholm University (斯德哥尔摩大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:LLMs are increasingly used in social simulations for socially interactive agents and robots, offering more flexibility than rule-based systems. However, even though they mimic human behaviour very well, there is a persistent alienness to them. This work investigates whether personality-aware fine-tuning can reduce this gap by improving the consistency and controllability of personality-conditioned dialogue generation compared with instruction prompting alone. We fine-tune two small open-weight LLMs, Qwen2.5-7B-Instruct and Ministral-8B-Instruct, using a corpus that combines personality-labelled social media posts and dialogues to create a personality-based dialogue engine for social simulation. The resulting models are evaluated across multiple social interaction scenarios using three independent LLM judges, which assess personality fidelity and provide evidence-based behavioral interpretations. We additionally quantify inter-rater agreement and lexical characteristics of the generated dialogue. Results indicate that fine-tuned models are not better at role-playing different personalities than their respective baseline models. However, low inter-rater agreement limits the confidence with which these results can be interpreted. Concerning the quality of generated texts, fine-tuned models are mostly comparable to the baselines, with fine-tuning improving the linguistic diversity of the Qwen models. While the results appear generally usable and the baseline models offer the best overall performance, future studies should place greater emphasis on the quality and domain alignment of training data for accurate personality role-playing.

[NLP-11] Reusing Latent Speech Representations for Query-Conditioned Topic Localization in Transcripts EMNLP2026

【速读】: 该论文旨在解决长语音转写文本(long transcripts)在下游自然语言处理(NLP)任务中因包含冗余上下文而导致的计算成本高与信息相关性低的问题。其核心问题为:如何在给定主题标题查询(topic-title query)的情况下,精准定位语音转录文本中与该主题最相关的句子片段(sentence span),即实现查询条件下的主题定位(query-conditioned topic localization)。解决方案的关键在于:复用自动语音识别(ASR)编码器的输出状态作为句子级语义表示,并将其与文本嵌入(textual embeddings)进行融合,从而构建轻量级的片段定位器(span locator)。该方法使定位器能够有效利用语音模态信息,而无需额外运行独立的音频编码器,显著提升了定位精度,尤其在严格边界匹配评估标准下表现突出。实验结果表明,在两个公开数据集上均优于纯文本基线模型;跨数据集分析进一步揭示,该方法在结构化或半结构化语音场景中收益最大,而在自发性口语(spontaneous speech)中的提升则有限且不一致。

链接: https://arxiv.org/abs/2609.21844
作者: Steffen Freisinger,Philipp Seeberger,Thomas Ranzenberger,Tobias Bocklet,Korbinian Riedhammer
机构: Technische Hochschule Nürnberg Georg Simon Ohm (诺伊堡应用技术大学)
类目: Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
备注: Accepted at EMNLP 2026 Main Conference

点击查看摘要

Abstract:Long transcripts are costly inputs for downstream NLP systems and often contain irrelevant context. We study query-conditioned topic localization: predicting the sentence span in a transcript that best addresses a topic-title query. To improve span localization, we reuse ASR encoder states as sentence-level representations and fuse them with textual embeddings. This lets lightweight span locators exploit speech information without running a separate audio encoder. Experiments on two public datasets show consistent gains over text-only baselines, especially under strict boundary-matching criteria. Cross-dataset experiments further indicate that the benefits are strongest for structured or semi-structured speech, while gains on spontaneous speech are limited and mixed.

[NLP-12] RheoSampling: Resolving the One-Hot Dilemma in Stochastic Dynamic-Tree Speculative Decoding

【速读】: 该论文旨在解决生成式 AI(Generative AI)在大语言模型(LLM)推理过程中,动态树结构(dynamic-tree)方法在随机采样(stochastic decoding, T0)场景下因概率分布耦合导致接受率严重下降的问题。现有动态树方法如EAGLE-3虽在贪心解码下表现优异,依赖确定性Top-K扩展与全局剪枝保持上下文感知的拓扑结构,但在随机采样时,其将同一概率分布同时用于树构建与令牌验证,导致草案分布坍缩为独热(one-hot)概率,破坏了采样的随机性并引发接受率骤降。这一矛盾本质上源于树构建与验证任务对同一概率分布的冲突性使用,使得随机性难以直接注入。本文提出RheoSampling方法,通过解耦这两个角色实现突破:在树构建阶段,将从草案分布中采样的令牌以代理概率(proxy probability)置于确定性Top-K候选中;而在验证阶段,则使用其真实采样概率进行校验。该机制首次实现了兼具上下文感知的动态拓扑构建与真正随机采样的动态树方法,且保证无损性(losslessness)。作者通过等价类分析(equivalence-class analysis)将复杂的随机树空间压缩为可处理的类别,建立理论上的无损性保障;结合基于最优传输(OT-based)的验证策略与稀疏草案机制,确保理论增益可转化为实际推理加速。实验表明,RheoSampling在多种大语言模型与基准测试上均显著提升接受率与推理速度,超越当前最先进的动态树方法,为随机树结构的设计与分析提供了通用框架。

链接: https://arxiv.org/abs/2609.21827
作者: Qiao Hu,Yepeng Weng,Bo Zhang,Takehisa Yairi
机构: National Center for Mathematics and Interdisciplinary Sciences (NCMIS), AMSS, CAS; The University of Tokyo; SKLMS and AMSS, Chinese Academy of Sciences; School of Mathematical Sciences, University of Chinese Academy of Sciences; Lenovo AI Technology Center
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Speculative decoding accelerates LLM inference by drafting multiple tokens in parallel, with tree-based methods further improving efficiency through hierarchical structures. Dynamic-tree methods such as EAGLE-3 perform well under greedy decoding via deterministic top-K expansion and global pruning. However, in stochastic decoding (T0), this mechanism collapses the draft distribution into one-hot probabilities, causing a severe drop in acceptance rate. This creates a dilemma: dynamic-tree methods sacrifice stochastic sampling to preserve context-aware topology, while static-tree methods preserve stochastic sampling with context-agnostic structures. The issue arises because the same probability distribution is used for two conflicting tasks: constructing the tree and verifying tokens. This coupling makes direct injection of randomness challenging due to the resulting stochastic process. We resolve this by decoupling these roles: RheoSampling assigns a token sampled from the draft distribution a proxy probability for tree expansion and pruning alongside its true sampling probability for verification. Specifically, we inject a sampled token among the deterministic top-K slots and treat it with different probabilities during construction and verification, making RheoSampling the first dynamic-tree method with both context-aware top-K construction and stochastic sampling while maintaining losslessness. We establish the lossless guarantee through an equivalence-class analysis that compresses the stochastic tree space into tractable classes. An OT-based verification strategy and a sparse draft mechanism ensure that theoretical gains translate into practical efficiency. Experiments across LLMs and benchmarks demonstrate improvements in acceptance rate and speedup over state-of-the-art dynamic tree methods. This framework may provide a template for analyzing stochastic tree structures.

[NLP-13] CASCADE Against Jailbreaks: Combination Across Stages with Controlled Attack-Defense Evaluation EMNLP2026

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在面对越狱攻击(jailbreak attacks)时,防御策略在不同处理阶段(如输入修改或输出防护)的部署选择与组合优化问题。现有研究因攻击成功率定义不统一及实验设置差异,导致防御措施多被孤立评估,缺乏系统性指导。为此,本文首次在一致的直接性、黑盒、单轮攻击威胁模型下,系统性地研究了跨阶段及阶段内防御机制的组合效果。其解决方案的关键在于提出一个决策框架,通过规范化攻击成功率的计算方式并控制查询预算,结合明确的公平性规则实现标准化评估。基于19种攻击与15种防御的全面实验,研究发现单一防御无法普适最优,但经过精心设计的防御组合可在保障显著安全性的同时,将性能损失降至最低,从而为构建分层防御体系提供了可落地的实践建议。

链接: https://arxiv.org/abs/2609.21793
作者: Jiale Luo,Eric Han
机构: National University of Singapore(新加坡国立大学)
类目: Cryptography and Security (cs.CR); Computation and Language (cs.CL)
备注: Accepted to Findings of EMNLP 2026

点击查看摘要

Abstract:Defenses against jailbreak attacks on Large Language Models (LLMs) operate at different pipeline stages, such as input modification or output guard, but it remains unclear which defenses to deploy at each stage and how to combine them. Prior empirical studies, fragmented by inconsistent attack-success-rate definitions and experimental settings, have evaluated defenses largely in isolation. Here we present the first systematic study, to our knowledge, of defense combinations both within and across pipeline stages, under a consistent threat model of direct, black-box, single-turn attacks. Our decision framework standardizes evaluation through a principled attack-success-rate formulation with controlled query budgets, together with explicit fairness rules. Across 19 attacks and 15 defenses, we find that no single defense is universally best, but well-chosen combinations achieve substantial safety with minimal utility degradation, yielding practical recommendations for layered defense pipelines.

[NLP-14] Per-Aetiology Contrastive Severity Embeddings with Phonological Pseudo-Labelling for Multilingual Dysarthric Speech

【速读】: 该论文旨在解决多语言构音障碍严重程度评估中普遍存在的标签空间设计问题,即现有系统通常采用单一病因-语言组合进行训练,或简单地将不同病因数据混合归入同一标签空间,这种做法可能引入混淆与偏差。其核心解决方案是通过构建共享主干网络、统一训练流程、统一语料库注册和严格隔离的测试集,对四种匹配的HuBERT-base对比嵌入模型进行受控比较:一个混合病因基准模型与三个针对脑性瘫痪(CP)、帕金森病(PD)和肌萎缩侧索硬化症(ALS)的特异性病因模型。关键创新在于结合临床标注语音与一种无需训练的音系学轮廓分析方法生成的序数伪标签,以增强数据多样性与泛化能力。实验结果表明,在去泄漏的说话人独立测试子集上,各病因特异性模型在所有目标病因上均显著优于混合基准模型(如CP:宏平均F1 0.829 vs 0.676,相对提升22.6%;PD:0.715 vs 0.511,+40.0%;ALS:0.788 vs 0.596,+32.3%)。此外,在CP任务中引入144名SAP和44名CDSD伪标签样本,使宏平均F1从0.786提升至0.829,进一步验证了伪标签的有效性。研究强调了伪标签校准、数据划分规范性和置信度阈值部署等关键局限,为未来多病因多语言构音障碍评估系统的标签空间设计提供了重要启示。

链接: https://arxiv.org/abs/2609.21789
作者: Bernard Muller,Antonio Armando Ortiz Barrañón,LaVonne Roberts
机构: The Scott-Morgan Foundation (英国斯科特-摩根基金会); Tecnológico de Monterrey (墨西哥蒙特雷理工学院); SMF Labs (法国SMF实验室)
类目: Computation and Language (cs.CL); Sound (cs.SD)
备注: Accepted at IEEE SLT 2026, 13-16 December 2026, Palermo, Sicily

点击查看摘要

Abstract:Most multilingual dysarthria-severity systems either train on a single aetiology-language pair or pool heterogeneous aetiologies into one label space. We test that pooling assumption with four matched HuBERT-base contrastive embedding models under a shared backbone, training recipe, corpus registry and held-out evaluation: one mixed-aetiology baseline and three aetiology-specific models for cerebral palsy (CP), Parkinson’s disease (PD) and amyotrophic lateral sclerosis (ALS). Training combines clinically labelled speech with ordinal pseudo-labels from a training-free phonological profiling method [1], [2]. On speaker-disjoint, leakage-filtered held-out subsets, the per-aetiology models outperform the mixed baseline across all three target aetiologies: CP (macro F1 0.829 vs 0.676, +22.6 % relative), PD (0.715 vs 0.511, +40.0 %) and ALS (0.788 vs 0.596, +32.3 %). On CP, adding 144 SAP and 44 CDSD pseudo-labelled speakers lifts macro F1 from 0.786 to 0.829 over a clinical-only CP model (+4.3 percentage points). Training data span three to seven languages per aetiology. We position this as a controlled comparison of label-space design choices and discuss pseudo-label calibration, split hygiene, and confidence-thresholded deployment as important limitations for future work.

[NLP-15] World Modeling in Transformers

【速读】: 该论文试图解决的问题是:尽管基于变压器(Transformer)的模型在行为上表现出“缺乏世界模型”(world model)的现象,但其内部可能已习得对环境的忠实表征。研究以TaxiGPT为例,该模型通过曼哈顿随机游走训练,其导航失败常被解释为缺乏一致的内部地图,但本文通过机制分析与因果干预揭示,该模型实际上能够表征交叉路口和道路、跟踪自身位置,并利用目标罗盘进行导航。其行为失败的关键原因在于:不同交叉路口特征的叠加导致内部表征干扰,从而破坏了定位精度。解决方案的核心在于“可及性打包”(affordance packing),即通过将具有相同合法移动路径的交叉路口表征进行分组,以限制此类错误的传播范围。此外,研究提出了一套机制指标,用于比较不同训练阶段模型的世界建模能力,发现世界模型能力并非在某一固定时间点突然出现,而是随训练逐步涌现。这一发现推动研究范式从“模型是否具备世界模型”的二元判断转向对世界建模机制的深入剖析,即考察模型如何通过相互作用的表征能力来构建并利用环境认知以指导行为。

链接: https://arxiv.org/abs/2609.21748
作者: Pierre Beckmann,Matthieu Queloz,Andre Freitas
机构: EPFL(洛桑联邦理工学院); IDIAP Research Institute(IDIAP 研究所); MATS; University of Bern(伯尔尼大学); University of Manchester(曼彻斯特大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Behavioral failures can make a transformer appear to lack a world model even when it has learned faithful representations of its environment. We demonstrate this in TaxiGPT, a transformer trained on random walks through Manhattan whose failures have been interpreted as evidence of an incoherent internal map. Through mechanistic analysis and causal interventions, we show that the model represents intersections and streets, tracks its position, and uses a goal compass to navigate. We trace its failures to interference between superposed intersection features, which disrupts localization within the internal map. Affordance packing, which groups representations of intersections with the same legal moves, helps limit the consequences of these errors. Finally, we propose mechanistic indicators that we use to compare models and show that world-modeling capacities emerge at different stages of training. Our findings motivate a shift from asking whether a model has a world model to mechanistically studying its world modeling: the interacting capacities through which it represents its environment and uses those representations to guide behavior.

[NLP-16] CIBuzzBench: A Benchmark for Cross-Lingual Understanding of Chinese Internet Buzzwords

【速读】: 该论文旨在解决大语言模型(LLM)在跨语言情境下对中文网络流行语(Chinese internet buzzwords)的非字面意义理解能力不足的问题,尤其关注其在中译英过程中能否准确传递蕴含于本地文化与语用背景中的隐含含义。由于中文网络流行语常依赖谐音、婉转表达、反讽或编码语言等文化特异性手段,其真实意图可能被表层语义掩盖,进而影响内容安全评估的有效性。因此,论文提出的关键解决方案是构建首个面向中英跨语言理解的基准测试集——CIBuzzBench,该数据集包含3,001个经过标注的中文网络流行语,涵盖英文释义、英文对应词、类别标签及有害性标签,并据此设计了三项核心评估任务:语义解释、跨语言等价匹配与文化根基型有害性检测。实验结果表明,当前先进的专有及中文大模型在非字面细粒度理解、对抗扰动下的等价匹配鲁棒性以及有害性判断的校准能力方面仍存在显著短板,揭示了多语言大模型在处理文化嵌入式语言现象时面临的持续挑战,也为未来安全导向的跨语言评估提供了重要依据。

链接: https://arxiv.org/abs/2609.21722
作者: Yifan Wang,Junyu Lu,Qifan Wang,Shun Zhang,Chaozhuo Li,Jiahao Liu,Zhijun Cao,Lingbin Bu,Fanliang Bu
机构: People’s Public Security University of China (中国人民公安大学); Dalian University of Technology (大连理工大学); Meta AI (Meta人工智能); Beijing University of Posts and Telecommunications (北京邮电大学); Meituan (美团); Beijing Police College (北京警察学院)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Chinese social media has generated a vast and continually evolving lexicon of internet buzzwords whose meanings are often non-literal and deeply rooted in local cultural and pragmatic contexts. Existing research has primarily focused on interpreting these buzzwords within Chinese, leaving largely unexplored whether LLMs can transfer such culturally grounded knowledge across languages and accurately convey the intended meanings in English. This cross-lingual capability is also critical for safety, as harmful expressions may obscure their offensive content through culture-specific homophony, euphemism, irony, or coded language. In this paper, we investigate the ability of advanced LLMs to understand Chinese internet buzzwords across languages. To this end, we introduce CIBuzzBench, the first benchmark for cross-lingual Chinese-to-English understanding of Chinese internet buzzwords. CIBuzzBench comprises 3,001 Chinese internet buzzwords annotated with English meaning explanations, English equivalents, category labels, and harmfulness labels. Based on these annotations, we design three evaluation tasks: Meaning Explanation, Cross-lingual Equivalent Matching, and Culturally Grounded Harmfulness Detection. We evaluate representative state-of-the-art proprietary and Chinese LLMs under both English- and Chinese-prompting settings. Our results show that LLMs continue to struggle with the cross-lingual understanding of Chinese internet buzzwords, particularly in fine-grained non-literal interpretation, robust equivalent matching under option perturbations, and calibrated harmfulness detection. These findings highlight the persistent challenges posed by culturally grounded language phenomena for multilingual LLMs and safety-oriented evaluation. The dataset and code are available at this https URL.

[NLP-17] Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation ECCV

【速读】: 该论文旨在解决对话系统中缺乏对听者即时行为反馈的动态建模问题,进而提升生成语句在情感与语调(prosody)上的自然性与情境适配性。其核心解决方案是提出一种两阶段框架ReACT-TTS,利用听者前1秒的面部表情序列作为时间上下文线索,在语音生成前进行情绪与语调的规划。关键创新在于引入时间动态建模(Temporal conditioning),通过分析听者的实时非语言行为,实现对响应风格的精准预测;实验表明,相较于仅依赖文本输入的方法,该方法在十次随机种子测试中均显著提升了平均宏F1分数和情感-激活-支配(VAD)一致性指标,且准确率保持稳定。消融实验进一步验证了时间建模优于其他视觉变体,并证明无需显式区分早期与晚期反应差异;同时,正确匹配听者反应的策略表现优于循环错配。在20位语音研究者的上下文恰当性评估中,76%的评价偏好时间建模方案,凸显其在真实对话场景中的优势。最终,该预测结果被接入Grad-TTS骨干网络,实现了端到端语音合成。研究证实,听者预响应动态可作为对话响应规划的互补性线索,有效增强生成语音的情境感知能力。

链接: https://arxiv.org/abs/2609.21683
作者: Yunji Chu
机构: Sogang University (首尔世宗大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD)
备注: 15 pages, 2 figures, 2026 ECCV Workshop (11th ABAW) Best Student Paper Award

点击查看摘要

Abstract:Conversational speech depends on dialogue context and the listener’s immediately preceding behavior. We propose ReACT-TTS, a two-stage framework that uses a one-second pre-response listener facial sequence to plan the next utterance’s emotion and prosody before speech realization. On a strict dyadic MELD protocol, Temporal conditioning yields higher mean macro-F1 and VAD concordance than Text-only across ten seeds, while accuracy remains essentially unchanged. Ablations show that temporal modeling performs best among the visual variants and that an explicit early-to-late difference is unnecessary; correct listener reactions also outperform cyclic mismatches on average. In a contextual-appropriateness study with 20 speech researchers, 76% of judgments prefer Temporal, 9% Text-only, and 15% report no preference. We further connect the predicted response style to a Grad-TTS backbone for end-to-end speech realization. Overall, the results support pre-response listener dynamics as complementary cues for conversational response planning. The source code is available at this https URL.

[NLP-18] PRISM-BN: A Controlled Corpus and Benchmark for Text-to-Parameterized Bayesian Network Extraction

【速读】: 该论文旨在解决生成式人工智能在构建符号化概率图模型(Probabilistic Graphical Models, PGMs)时缺乏大规模、高质量的文本到参数化贝叶斯网络(Bayesian Network, BN)配对数据集的问题。现有方法受限于无法获取足够规模的文本-贝叶斯网络标注资源,导致训练端到端的文本到参数化BN系统面临挑战。为此,作者提出PRISM-BN,一个包含5054个基于贝叶斯网络的文本描述及其对应的离散参考贝叶斯网络的可控语料库,覆盖五个领域,每个实例均包含变量、状态、有向边、根节点先验以及完整的多父节点条件概率表(Conditional Probability Distributions, CPDs)。这些数据由50个维基百科种子骨架衍生而来,其概率参数为内部构造的基准目标而非外部验证的因果估计。解决方案的关键在于PRISM框架——一种以边缘分布优先的流水线方法,通过提取边缘分布与局部联合分布,解析恢复归一化的条件概率表,并构建局部重参数化的子图结构。研究定义了涵盖语义节点与状态对齐、条件结构评分及严格全CPD评估的基准测试体系。实验表明,六种大语言模型(LLM)提取器在节点识别上的F1值为0.56–0.83,条件边识别F1值为0.90–0.97,而条件概率表的KL散度则在1.11–3.14之间,显示出结构恢复能力较强但全量条件概率表精确匹配仍具挑战性。该趋势在独立生成的GPT-5.5参考样本中依然存在,且人类预实验验证了结构可恢复性与概率解释的一致性。因此,PRISM-BN实现了对结构恢复与概率参数估计任务的分离评估,为神经符号人工智能中的概率建模提供了可信赖的基准数据集与评估标准。

链接: https://arxiv.org/abs/2609.21673
作者: Amartya Bhattacharya,Nikhil Singh,Neeti Pokhriyal,Soroush Vosoughi
机构: Dartmouth College (达特茅斯学院); RAND Corporation (兰德公司)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Probabilistic Graphical Models (PGMs), especially Bayesian Networks (BNs), expose directed structure and probabilistic parameters, making them natural symbolic targets for neurosymbolic AI. Yet training text-to-parameterized-BN systems requires paired text-to-BN resources unavailable at scale. We introduce PRISM-BN, a controlled corpus of 5054 BN-grounded descriptions paired with discrete reference BNs containing variables, states, directed edges, root priors, and full multi-parent CPDs across five domains. The instances are derived from 50 Wikipedia-seeded backbones, and their probabilities are internally constructed benchmark targets rather than externally validated causal estimates. PRISM-BN is built with PRISM, a marginal-first pipeline that elicits marginal and local joint distributions, analytically recovers normalized CPDs, and constructs locally reparameterized subgraphs. We define a benchmark with semantic node and state alignment, conditional structural scoring, and strict full-CPD evaluation. Across six LLM extractors, Node F1 ranges from 0.56 to 0.83, conditional Edge F1 from 0.90 to 0.97, and CPD-KL from 1.11 to 3.14. Conditional state and edge recovery remain consistently strong, whereas strict full-CPD agreement remains challenging. These trends persist with independently generated GPT-5.5 references, and a human pilot corroborates structural recoverability and similar probabilistic interpretations. PRISM-BN supports separate evaluation of structural recovery and probabilistic parameter estimation.

[NLP-19] Accelerating Dense LLM s via L0-regularized Mixture-of-Experts

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在推理过程中存在速度慢、成本高的问题,同时克服现有加速方法导致明显性能下降以及混合专家模型(Mixture-of-Experts, MoE)对计算资源需求过高的局限性。其解决方案的关键在于提出一种轻量级的MoE框架——L0-MoE,通过引入L0正则化实现稀疏激活,从而在几乎不损失模型性能的前提下显著提升推理效率。该方法结合领域感知的数据集精炼策略(基于聚类混淆矩阵)与动态批处理技术,有效优化训练过程并降低冗余计算。实验结果表明,L0-MoE相比密集型LLM可实现最高2.5倍的加速比,且性能优于现有的主流LLM加速基线。

链接: https://arxiv.org/abs/2609.21672
作者: Zhenyu Zhang,Jiudong Yang,Zhaowen Tao,Meng Chen
机构: YZW(智元机器人); FuTu AI(富途AI); Wise AI(智慧人工智能)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Large language models (LLMs) achieve strong performance but suffer from slow and costly inference. Existing acceleration methods often lead to noticeable performance degradation, while Mixture-of-Experts (MoE) models require extensive computational resources. In this paper, we propose L0-MoE, a lightweight MoE approach using L0-regularization to accelerate dense LLMs nearly without performance loss. Our method introduces a cluster confusion matrix for domain-aware dataset curation and applies dynamic batching for efficient training. Experiments show that L0-MoE achieves up to 2.5x speedup over dense models while maintaining competitive performance, outperforming existing LLM acceleration baselines.

[NLP-20] Rethinking Human-Aligned Evaluation: An Analysis of Semantic Metrics Beyond WER

【速读】: 该论文旨在解决传统自动语音识别(ASR)评估中词错误率(Word Error Rate, WER)的局限性问题,即其将所有词汇偏差视为同等代价,忽视了语义是否发生改变,从而可能与人类对转录质量的主观判断不一致。为此,作者提出了HATS-en,一个面向人类中心的英语ASR评估数据集,并在此基础上对比多种词汇级评估指标(包括不同配置的BERTScore和语义距离,SemDist)与人类判断的一致性。研究发现,WER在所有测试指标中与人类判断的相关性最低;表现最佳的SemDist配置在整体一致性上优于词错误率(CER)和BERTScore;且无单一语言模型在所有设置下均最优。尽管如此,CER因其计算简单、成本低,仍表现出与最优SemDist配置相近的性能。研究结果支持将评价体系从WERS转向更易解释、成本更低的CER,尤其适用于英语及音节型书写系统,并建议以SemDist作为补充评估手段,以更准确反映实际评估需求。

链接: https://arxiv.org/abs/2609.21663
作者: Hritika Sharma,Thibault Bañeras-Roux,Alessandra Pinto,Petr Motlicek,Hyunggu Jung,Esaú Villatoro-Tello,Somang Nam
机构: Google(谷歌); Meta(元); Stability.AI(稳定人工智能); Anthropic(Anthropic); Character.ai(字符人工智能); Claude(克劳德)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Word Error Rate (WER), the most commonly used metric for Automatic Speech Recognition (ASR), treats every lexical deviation from the reference as equally costly, regardless of whether it changes meaning. This raises the question: does WER actually track how humans judge ASR transcript quality? We introduce HATS-en, an English dataset for human-centered ASR evaluation. Using this dataset, we benchmark lexical metrics against several configurations of BERTScore and SemDist, varying the language model, layer, and pooling strategy. We find that WER agrees least with human judgment among all metrics tested, that the best-performing SemDist configurations achieve the highest overall agreement, ahead of CER and BERTScore, and that no single model is best across settings. CER, despite its simplicity and low cost, remains remarkably close to these best configurations. In line with prior recommendations, our results support shifting ASR evaluation toward CER both for English and for morphosyllabic writing systems as it is a more interpretable and low-cost metric for what evaluation should actually capture, and using SemDist as a complementary evaluation.

[NLP-21] When Steering Fails in Latent Reasoning : A Latent-to-Language Transition Gap

【速读】: 该论文旨在解决生成式语言模型在隐式链式思维(latent chain-of-thought, latent CoT)推理过程中,通过隐空间调控(activation steering)难以有效影响后续语言生成的问题。尽管隐式思维的隐藏表示被移动了与显式链式思维(explicit CoT)相当的程度,但其对最终输出的控制效果显著减弱。研究发现,任务信息在连续思维中仍可被识别,由此提出“隐空间到语言的转换间隙”(latent-to-language transition gap)这一核心假设:隐空间中的干预效应无法充分传递至语言生成阶段。两个关键证据支持该假设:一是输出分布在转换边界处发生突变;二是任务相关方向在隐式CoT中表现出远弱于显式CoT的双向控制能力。该研究揭示了隐空间调控方法的瓶颈在于转换接口,强调未来设计高效隐式调控策略必须聚焦于优化此界面。

链接: https://arxiv.org/abs/2609.21662
作者: Gaoxiang Huang,Lei Qi
机构: HKUSTGZ(香港科技大学(广州)); SEU(东南大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Activation steering has become a widely used approach for controlling language models during explicit chain-of-thought (CoT) reasoning, motivating its extension to latent CoT. However, we find that steering continuous thoughts produces substantially weaker effects on subsequent language generation than steering explicit CoT, even when the hidden representations are moved by comparable amounts. We first show that task information remains identifiable in continuous thoughts. Hence, we hypothesize a \textbflatent-to-language transition gap, in which an intervention effect in latent space fails to transfer to language generation. Two further results support this hypothesis: the output distribution changes abruptly at the transition boundary, and task-related directions exert much weaker bidirectional control in latent CoT than in explicit CoT. These findings identify the transition interface as a central target for evaluating and designing future latent-steering methods.

[NLP-22] Analysing the Linearity of Linguistic Relations in Language Model Embedding Spaces ICLR2026

【速读】: 该论文旨在解决语言模型嵌入空间中不同语义关系的线性编码强度问题,即探究各类语言关系(如屈折、派生、词汇学及百科知识类关系)在嵌入向量空间中以多大程度上可被线性表征。其解决方案的关键在于提出一种基于约束线性近似的形式化框架,通过对比相关词对与无关词对之间的嵌入差异,量化评估各类关系在嵌入空间中的线性可分解性。研究基于扩展的BATS数据集,覆盖了多种语言关系类型,并在GloVe、RoBERTa和ModernBERT三种模型上进行实验。结果表明,屈折和派生关系具有接近完美的线性编码,而词汇学和百科知识类关系(尤其是多对一、多对多关联)则表现出显著更高的误差;同时,RoBERTa与ModernBERT在关系线性编码方面优于GloVe。该框架能够有效揭示不同模型中哪些关系结构更易于线性表示,为探查和比较不同模型的语义几何特性提供了一种高效且紧凑的工具。

链接: https://arxiv.org/abs/2609.21655
作者: Vasudevan Nedumpozhimana,Fathima Thekkekara,John Kelleher
机构: ADAPT Research Centre; Trinity College Dublin (都柏林三一学院); Indian Institute of Technology Bombay (印度理工学院孟买分校); Mumbai (孟买)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 6 pages. Accepted at the Workshop on Scientific Methods for Understanding Deep Learning (Sci4DL) at ICLR 2026

点击查看摘要

Abstract:We propose a framework to analyse how strongly different linguistic relations are linearly encoded in language model embedding spaces. We formalise linear encoding via a constrained linear approximation over related and unrelated word pairs and apply this to an extended BATS dataset covering inflectional, derivational, lexicographic, and encyclopedic relations in GloVe, RoBERTa, and ModernBERT. Our experiments show near-perfect linear encodings for inflectional and derivational relations, but substantially higher errors for lexicographic and encyclopedic relations, especially for one-to-many and many-to-many associations. We also find that RoBERTa and ModernBERT generally encode relations more linearly than GloVe. These results indicate that our framework can reveal which relational structures are most linearly accessible in embeddings, offering a compact tool for probing and comparing relational geometry across models.

[NLP-23] Configurable Multi-Stage Vision Pipeline for Crop Disease and Pest Diagnosis

【速读】: 该论文旨在解决生成式 AI 在实际农业应用中面临的三大核心挑战:图像质量不可控、作物与病虫害识别准确率不足,以及系统缺乏可配置性和灵活性。具体而言,用户上传的田间照片常存在光照差、模糊、角度不稳等问题,且无文本描述,导致现有服务无法有效判断图像是否可用、无法准确识别作物种类或诊断问题类型,甚至出现大量“无诊断”情况。此外,当前系统为黑箱模型,不可调整阈值、无法新增作物或病害类别、缺乏置信度裁剪机制,严重限制了其在复杂多变的真实场景中的部署能力。为此,论文提出分阶段解决方案:将任务分解为三个模块——图像质量筛选(M0)、作物检测(M1)和病虫害分类(M2)。相较单一语言视觉大模型(如Qwen3-VL-4B)的“一体式”方案,研究采用专业化小模型架构(DaViT + YOLOv26),实现更高效、高精度的端到端处理。实验表明,基于轻量级MobileNetV3的质量门控模型在86.9% F1得分下仅需12毫秒完成判断;而层级化专用模型(DaViT-Base)在作物识别上达到95.41%准确率,显著优于原生产基线(91.46%),且在诊断覆盖度和响应稳定性方面全面领先,同时保留了请求用户提供更优图像的能力。关键创新在于通过模块化设计实现了性能提升与系统可扩展性、可维护性的统一。

链接: https://arxiv.org/abs/2609.21651
作者: Naga Ganesh,Chandrashekar M S,Lakshmi Pedapudi,Aakash Singh,Vineet Singh
机构: Digital Green
类目: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 14 pages, 26 Tables, 12 Figures

点击查看摘要

Abstract:this http URL is Digital Green’s farm advisory service for smallholder farmers. When something looks wrong with a crop, the farmer takes a photograph and sends it, and that photograph is the whole question: no symptom described, no crop named, often no text at all. The service has to determine whether the picture can be used, what crop it shows, and what is wrong with it, from images taken on cheap phones in a field, in poor light and with a moving camera. The system doing this today cannot be adjusted. It has no adjustable thresholds for photograph rejection, crops and problems cannot be added, and there is no confidence cut-off to set. We study about 1.16 million photographs sent to this http URL from Ethiopia, India, Kenya and Nigeria. The production quality gate rejected 46.8% of the images it judged, over a quarter of those reaching diagnosis returned no crop name, and 35.8% of the labelled problems filed under “disease” are pests, identifiable without the crop. We therefore split the work into three stages: a quality gate (M0), a crop detector (M1), and a disease or pest detector (M2). Route A fills all three with one fine-tuned vision-language model (Qwen3-VL-4B) answering in a single call. Route B fills each with a small specialist model (DaViT, YOLO26). We replace our production GPT-4o quality gate with a small MobileNetV3 gate at 86.9% F1 in 12 ms. On one test set scored the same way for every system, a hierarchical DaViT-Base achieves 95.41% crop accuracy against 91.46% for the production baseline. It also leads on diagnosis and never declines to answer, while every language model in the comparison leaves a large share of rows with no diagnosis. The fine-tuned model retains two capabilities the specialists do not have: one call for all three stages, and a request for a better photograph when the image cannot support an answer. Comments: 14 pages, 26 Tables, 12 Figures Subjects: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG) Cite as: arXiv:2609.21651 [cs.CV] (or arXiv:2609.21651v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2609.21651 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[NLP-24] Chinese Competitive Debating Dataset and Benchmark

【速读】: 该论文旨在解决现有研究中缺乏细粒度辩论对话文本与真实竞赛环境下专业裁判评判结果相结合的数据集问题,尤其关注中文竞技性辩论中论点演进的动态追踪。其核心解决方案是构建一个大规模、多层级标注的中文辩论数据集与基准评测体系,涵盖148场正式比赛、2,698个辩论阶段及20,542个交流单元,所有内容均经人工校验并保留原始评分、投票结果、最佳辩手票选以及裁判评析理由。通过定义匹配级(winner-tendency prediction)、阶段级(stage-score prediction)和发言者级(best-debater prediction)三类任务,该研究为评估大语言模型对交互式论证的理解能力提供了标准化测试平台。零样本评估结果显示,最优模型在胜方预测准确率上达到66.2%,阶段评分与人类平均评分间的皮尔逊相关系数达0.250,最佳辩手预测准确率为56.8%,表明当前模型对辩论复杂性的理解仍存在显著提升空间,但该数据集为深入研究模型与专业裁判间的一致性提供了坚实基础。

链接: https://arxiv.org/abs/2609.21637
作者: Zongrui Yang,Haoyuan Li,Zhongsheng Wang,Zhirui Zeng,Pengqian Han,Yi Zhou,Yuting Wang,Jiamou Liu
机构: University of Auckland, New Zealand; South China Agricultural University, China; Sanming University, China
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 25 pages, 2 figures

点击查看摘要

Abstract:Debate adjudication requires tracking how arguments develop through interaction, yet existing datasets rarely combine fine-grained debate transcripts with professional judgments collected during real competitions under a shared rubric. We introduce a dataset and benchmark for evaluating large language models’ understanding of competitive Chinese-language debate at the match, stage, and speaker levels. We organized 182 matches and recruited 120 professional judges, with each match independently adjudicated by three judges using a predefined rubric. After excluding matches with incomplete records, the dataset contains 148 matches, 2,698 stages, and 20,542 exchange units, with manually verified transcripts and segmentation. It preserves original stage scores, match votes, best-debater ballots, and adjudication rationales. We define three tasks: winner-tendency prediction, stage-score prediction, and best-debater prediction. Zero-shot evaluation of multiple large language models yields a highest winner-prediction accuracy of 66.2%, a highest Pearson correlation of 0.250 between model stage scores and mean human ratings, and a highest best-debater prediction accuracy of 56.8%. The dataset and benchmark provide a testbed for studying large language models’ understanding of interactive argumentation and their agreement with professional judges.

[NLP-25] Steering LLM s Responses Towards Moral Foundations on the Norwegian MFQ-30 WWW

【速读】: 该论文旨在解决生成式人工智能(Generative AI)在道德与价值观表征方面缺乏稳定性与可调控性的问题,即现有基于人类心理量表的测评工具是否能稳定反映大语言模型(LLM)的道德基础特征,以及这些特征能否被有效引导至目标人类群体的分布。其解决方案的关键在于通过两种可控干预手段实现对模型道德基础(moral foundations)的定向调节:一是提示层级的“角色扮演”(persona steering),采用未参考人类样本分布信息的中性北欧受访者人格设定,使参与测试的模型在马哈拉诺比斯距离(Mahalanobis d²)上相对于挪威人群体均值的偏差缩小44%-77%;二是激活层级的“单对激活添加”(ActAdd),在固定中间层施加扰动,虽未能单独调整特定道德基础,但整体上实现了对模型道德谱系的扁平化处理。值得注意的是,某一模型在使用同一角色人格后不仅表现出显著的道德谱系迁移,还首次展现出原本缺失的互动参与行为,揭示了潜在的“认知幻影”(cognitive phantoms)现象,提示当前方法可能引发非预期的、类人类的虚假认知表现。

链接: https://arxiv.org/abs/2609.21636
作者: Hans Andersen,David Dichas
机构: University of Oslo (奥斯陆大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
备注: 13 pages, 4 figures, 7 tables. Awarded best Paper Award at WNNLP 2026 (University of Oslo). Proceedings: this https URL

点击查看摘要

Abstract:Recent work applies human psychometric questionnaires to large language models to elicit moral and value profiles, but it is not clear whether these instruments measure anything stable in models or whether the resulting profiles can be moved toward a target human population. We administer the Norwegian Moral Foundations Questionnaire (MFQ-30) to six open-weight LLMs and compare their foundation profiles to a sample of N = 1,282 Norwegian respondents. We test two steering interventions, prompt-level persona steering and activation-level ActAdd. Half the models engage with the questionnaire under our attention check. The other half default to flat or central-tendency outputs that look near-human on average without tracking item content. A neutral Nordic-respondent persona, written without any distributional information from the human sample, brings the engaging models 44-77% closer to the Norwegian mean in Mahalanobis d^2 . One-pair ActAdd at a fixed mid-layer flattens the foundation profile rather than steering individual foundations. For at least one model the same persona that shifts the profile also induces engagement that was absent at baseline, a concrete instance of the cognitive phantoms that Peereboom et al. (2025) warn about.

[NLP-26] Evaluating In-Context Learning and Retrieval Strategies for Devanagari Post-OCR Correction

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在印地语和马拉地语等天城文(Devanagari)脚本的后光学字符识别(post-OCR)纠错任务中缺乏系统性评估的问题,尤其关注无需训练即可实现高效纠错的上下文学习(in-context learning)方法的有效性。其核心挑战在于如何在不进行任务特定微调的前提下,利用少量示例有效提升对天城文文本中常见OCR错误的纠正能力。解决方案的关键在于提出一种名为CharBM25的新颖示例检索策略,该策略基于字符n-gram的BM25相似度,在OCR输入与候选示例间进行密集语义匹配,以捕捉目标句子与训练示例之间的共享错误模式。实验表明,使用字符三元组(trigram)的CharBM25在多个新闻领域基准上显著优于随机选择与密集语义检索,分别在印地语和马拉地语上提升2.8–4.0个百分点和2.9–3.8个百分点的词错误率(WER),且在仅需GPU资源极少的情况下达到甚至超越密集检索性能。此外,研究发现模型规模是决定纠错效果的关键因素:参数量超过12B的通用型大模型结合CharBM25可实现稳定、可靠的纠错表现,而小于8B的模型则无法持续改进,部分小模型甚至导致更多错误。马拉地语因更高的形态复杂性,纠错难度始终高于印地语,进一步凸显了方法设计的必要性。综上,该研究确立了CharBM25作为一项高效、低成本的检索机制,并证明将之与12B以上参数的通用大模型结合,可实现无需训练的高质量天城文后OCR纠错。

链接: https://arxiv.org/abs/2609.21595
作者: Abhishek Bhandari,Gaurav Harit
机构: 未知
类目: Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:In-context learning using Large Language Models (LLMs) offers a compelling path to training-free post-OCR correction, yet its effectiveness for Devanagari script remains entirely unexplored. We present the first systematic evaluation of LLMs (3B-32B) for post-OCR correction in Hindi and Marathi, comparing three in-context example retrieval strategies: domain-random selection, dense semantic retrieval, and our proposed CharBM25, which retrieves examples by character n-gram BM25 similarity over OCR inputs to target shared error patterns with the test sentence. Across a 20,000-sentence benchmark spanning five news domains, retrieval strategy is the decisive factor in correction quality: CharBM25 outperforms domain-random selection by 2.8-4.0pp absolute WER on Hindi and 2.9-3.8pp on Marathi, using character trigrams, which consistently outperform bigrams and unigrams. Scale dominates performance: Gemma-3-27B achieves WER reductions of 55.0% for Hindi and 33.3% for Marathi under CharBM25-5. Few-shot gains are capacity-gated: models below 8B do not reliably improve over the OCR baseline, and on Marathi the smallest models (3B) degrade more sentences than they improve. Marathi is persistently harder to correct than Hindi across all scales, reflecting its greater morphological complexity. These findings establish CharBM25 as an effective, GPU-free retrieval strategy that matches or exceeds dense retrieval at negligible computational cost, and show that combining it with a general-purpose LLM of 12B+ parameters delivers reliable, training-free Devanagari post-OCR correction without task-specific fine-tuning. Dataset: this https URL

[NLP-27] GameLogicBench: Evaluating Coding Agents on Runtime Game Logic with Tick-Level State Assertions

【速读】: 该论文旨在解决现有游戏开发基准测试在评估生成式AI(Generative AI)代码代理时,无法全面、动态地验证游戏规则执行一致性的问题。当前方法仅依赖固定示例回放、视频评分或外部模型评判,缺乏对多样化场景下游戏逻辑全程合规性的严格检查,且难以保证结果的可复现性。为此,研究提出GameLogicBench,一个基于Godot项目的72个游戏逻辑任务基准,通过自动化评估器在每一模拟周期(simulation tick)实时校验游戏规则。该基准包含403个手工设计的场景,结合参数扰动生成1,451个测试用例,并引入变异体(mutant)验证机制,确保评估器仅接受符合功能需求的正确实现,拒绝移除必要能力的错误变体。关键创新在于构建了具备行为导向判别能力的评估框架,不仅关注输出结果,更强调对实现路径的鲁棒性与语义正确性的双重约束。实验表明,尽管最佳模型仅能解决52.78%的任务,但随着任务复杂度从单一机制扩展至多系统交互及项目级特征,所有模型性能均下降;同时发现代理在项目级任务中更频繁调用工具并审查代码。失败案例多数可运行但逻辑错误,且无变异体验证时,错误提交可通过评估。此外分析揭示开放网络访问下代理存在从公共仓库复制代码的行为。因此,可靠评估不仅依赖于测试用例对错误行为的排斥能力,也取决于对外部代码访问的严格控制。

链接: https://arxiv.org/abs/2609.21562
作者: Xinyu Che,Yunfei Ge,Shihao Li,Yanchen Liu,Hang Yan,Xinping Lei,Yanghai Wang,Zixuan Dong,Yifan Yao,Qianqian Xie,Letian Zhu,Jiaheng Liu
机构: Nanjing University (南京大学)
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 36 pages, 9 figures, 13 tables. Xinyu Che, Yunfei Ge, Shihao Li, Yanchen Liu, Hang Yan, and Xinping Lei contributed equally. Jiaheng Liu is the corresponding author. Code and benchmark: this https URL

点击查看摘要

Abstract:Coding agents can modify and test code across large software projects. Game development is a domain where agents must implement gameplay rules. A game can end in a valid state even after violating its rules during the run. Current game-development benchmarks replay fixed examples, score videos, or ask another model to judge the result. However, no existing benchmark checks game rules throughout execution across varied evaluator-selected scenarios while ensuring exactly reproducible verdicts. We introduce GameLogicBench, a benchmark of 72 gameplay-logic tasks in Godot projects. An automated evaluator checks each game’s rules at every simulation tick. Across 403 hand-designed scenarios, seeded parameter variations produce 1,451 test cases. To ensure that the evaluator measures behavior rather than implementation choice, it must accept different correct implementations for each task while rejecting mutants, implementations with one required capability removed. The tasks span isolated mechanics, multi-system interactions, and repository-scale features. Across 20 combinations of language models and scaffolds, the best observed run solves 52.78% of tasks. Under Claude Code, all twelve models solve fewer tasks as task scope expands from isolated mechanics, through interacting systems, to repository-scale features. Agents inspect code more often and make more tool calls on repository-scale tasks than on isolated-mechanic tasks. Most unsuccessful submissions are runnable, but implement some required game behavior incorrectly. We compared versions of our benchmark evaluator built with and without validation using mutants. Without this validation, incorrect agent submissions passed. A separate analysis finds agents copying code from public repositories when network access is open. Reliable evaluation thus depends both on what the tests reject and on what external code agents can access.

[NLP-28] MIRAG E: Multi-Perspective Creative Language Model Reasoning with Reinforcement Learning Guidance ICML2025

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在处理复杂数学、科学及逻辑任务时表现不佳的问题。其核心挑战在于模型缺乏类似人类的认知灵活性,难以动态调整思维视角以应对多维度推理需求。为此,论文提出了一种名为MIRAGE(Multi-perspective Inference-time Reasoning via Agent-Guided Exploration)的推理阶段创造性思维框架,其关键创新在于引入一个**选择器(Selector)与一个求解器(Reasoner)**协同工作机制:选择器负责在推理过程中优先筛选出最具潜力的概念视角(如代数、概率等),而求解器则基于这些视角进行序列化任务求解,若无法获得置信解,则聚合多个视角的结果以提升鲁棒性。该方法在GSM8K、MATH500、MMLU-Pro和Game-of-24等多个基准测试中显著优于Chain-of-Thought及多样化提示集成等现有方法,在几乎无额外推理开销的前提下实现了准确率的大幅提升,为实际应用提供了可扩展的解决方案。

链接: https://arxiv.org/abs/2609.21554
作者: Arash Lagzian,Srinivas Anumasa,Dianbo Liu
机构: National University of Singapore(新加坡国立大学)
类目: Computation and Language (cs.CL)
备注: 18 pages, 5 figures. Accepted at the ICML 2025 Workshop on Multi-Agent Systems in the Era of Foundation Models: Opportunities, Challenges and Futures (MAS-2025)

点击查看摘要

Abstract:Recent advances in Large Language Models (LLMs) have revolutionized artificial intelligence and how human interact with AIs. Despite impressive advancements, LLMs struggle with complex mathematical, scientific, and logical tasks. Inspired by human cognitive flexibility - our ability to dynamically switch mental perspectives - we propose MIRAGE (Multi-perspective Inference-time Reasoning via Agent-Guided Exploration), a novel inference-time creative thinking framework. MIRAGE includes a Selector that prioritizes effective conceptual perspectives (e.g., algebraic, probabilistic) and a Reasoner that sequentially solves tasks until a confident solution emerges, otherwise aggregating multiple perspectives. Tested on GSM8K, MATH500, MMLU-Pro, and Game-of-24 benchmarks, MIRAGE consistently outperforms methods like Chain-of-Thought and diverse prompting ensembles, significantly boosting accuracy with minimal inference overhead, providing a scalable solution for practical applications.

[NLP-29] Benchmarking Gender Bias in Machine Translation Evaluation Metrics across Occupations EMNLP2026

【速读】: 该论文旨在解决机器翻译(Machine Translation, MT)中持续存在的性别偏见问题,尤其关注在源文本未明确指明人物性别时,翻译结果及自动评估系统对性别形式的系统性偏好。其核心问题是:当源语言中性别信息缺失时,翻译系统与评估指标可能无根据地倾向使用男性或女性代词/称谓,从而导致评价偏差。解决方案的关键在于通过构建一个职业平衡的测试集(GAMBIT+的子集),包含每种目标语言1,308对男女性翻译样本(覆盖436个ISCO-08职业类别),并针对英源至阿拉伯语、捷克语、希腊语、冰岛语、俄语和乌克兰语以及德语共七种语言对进行评估。研究不仅分析评分预测与错误标注的表现,还深入考察性别差异在评分方向、幅度和频率上的分布特征,揭示出整体上男性化翻译得分更高,且不同职业领域呈现符合社会刻板印象的性别偏好,但这种偏见程度和一致性在不同语言和评估者间存在显著差异。因此,论文强调,要全面捕捉评估中的性别偏见,必须超越单一聚合指标,引入多维度的评估者行为分析。

链接: https://arxiv.org/abs/2609.21490
作者: Orfeas Menis Mastromichalakis,Giorgos Filandrianos,Wafaa Mohammed,Giuseppe Attanasio,Chrysoula Zerva
机构: Instituto de Telecomunicações, Lisbon, Portugal; National Technical University of Athens, Greece; University of Amsterdam, Netherlands
类目: Computation and Language (cs.CL)
备注: Accepted for publication at the 11th Conference of Machine Translation (WMT26), co-located with EMNLP 2026

点击查看摘要

Abstract:Gender bias remains a persistent concern in machine translation (MT), affecting both generated translations and their automatic evaluation. When a source text leaves a person’s gender unspecified, translations may realize that person using masculine or feminine forms, and both MT systems and evaluation metrics may exhibit systematic preferences between these alternatives despite the source providing no basis for such a distinction. We study this behavior in the WMT 2026 Automated Translation Quality Evaluation Systems Shared Task using an occupation-balanced subset of GAMBIT+. We consider seven English-source language pairs, six from the original dataset, targeting Arabic, Czech, Greek, Icelandic, Russian, and Ukrainian, and extend the original resource with German. The subset contains 1,308 masculine/feminine translation pairs per target language, with three examples for each of the 436 ISCO-08 occupational groups. We evaluate shared-task submissions and baselines for score prediction and error annotation, examining the direction, magnitude, and frequency of gender-related differences. We find an overall tendency for masculine translations to receive higher scores, as well as differences per occupation following stereotypical gender representations, although the strength and consistency of this preference vary considerably across evaluators and languages. Our results show that gender bias remains present in MT evaluation, but that capturing its extent requires looking beyond a single aggregate measure to complementary dimensions of evaluator behavior.

[NLP-30] alking Past the Machine: Morality Politeness and Alignment in Human-AI Dialogue

【速读】: 该论文旨在解决当前对话式人工智能(Conversational AI)系统在人际交互中是否真正实现合作性沟通,抑或仅模拟其表层特征这一核心问题。研究聚焦于道德性(morality)、礼貌性(politeness)与对齐度(alignment)三个构成合作对话的核心维度,通过分析15,881条人机对话与10,784条人际对话的多轮交互数据,运用混合效应模型识别影响回合间对齐的关键因素。研究发现,尽管AI能够生成合作性沟通的表面特征,但其背后的社会结构基础并不存在:道德表达呈现预设而非协商,情感温暖缺乏对“面子”敏感性的回应,语言趋同性持续下降。尤为关键的是,人类互动中促进包容的缓和与软化策略,在AI产生时反而与更低的对齐度相关;而人类中代表分歧的纯粹性框架(purity framing)在与AI互动时却导致用户向AI靠拢。研究进一步表明,主体性(agency)——即赋予用户塑造交流过程的空间——是两类交互中对齐度最稳定的预测因子,而近期模型降低道德强势性并未带来更优的合作表现。这些结果揭示,当前对话式AI仅复现了合作的表象,而缺失了人与人之间基于相互适应的真实合作机制;更令人意外的是,维系人类协作的机制在与AI互动时甚至出现方向逆转,提示仅从回合层面分析交互可能不足以评估其整体有效性。

链接: https://arxiv.org/abs/2609.21401
作者: Marina Mitiaeva,Lu Xiao
机构: Amazon.com(亚马逊); Arizona State University(亚利桑那州立大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Accepted at the 60th Hawaii International Conference on System Sciences (HICSS-60)

点击查看摘要

Abstract:Conversational AI systems produce fluent, socially appropriate responses, yet whether they participate in cooperative communication or merely simulate its surface forms remains unclear - a question central to how these systems are evaluated, trusted, and designed. This study investigates how morality, politeness, and alignment - three dimensions central to cooperative dialogue - function in human-AI interaction compared to human-human conversation. We analyze 15,881 human-ChatGPT and 10,784 human-human multi-turn dialogues, using mixed-effects models to identify which features predict turn-to-turn alignment. We observe a consistent dissociation: AI produces the surface features of cooperative communication without the underlying social architecture. Moral output appears preconfigured rather than negotiated; warmth is generated without face sensitivity; linguistic convergence declines persistently. Most strikingly, the cooperative mechanisms themselves reverse direction: hedging and softening associated with greater accommodation between humans are associated with reduced alignment when produced by AI, and purity framing associated with human divergence coincides with users converging toward the AI. Agency - giving users room to shape the exchange - is the most consistent predictor of alignment across both interaction types, while lower moral assertiveness in more recent models is not accompanied by better cooperation. Together these patterns suggest that AI reproduces the surface of cooperation without the mutual adaptation that grounds it between humans - and, more surprisingly, that mechanisms sustaining human accommodation can run in reverse with AI, suggesting a turn-level view may be insufficient for interaction-level success.

[NLP-31] Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction

【速读】: 该论文旨在解决多模态交互中用户意图理解的深层挑战,即现有模型虽注重响应质量,却忽视了从复杂多模态交互中准确推断用户潜在需求的能力。核心问题是:在语音表达不完整、存在语义模糊或语言不流畅、且受噪声环境干扰的情况下,模型能否基于视觉、听觉与对话历史等多模态上下文正确识别并理解用户的实际需求?尤其当用户以类似请求的言语表达但并非真实需求时,易引发误触发。为此,论文提出“全场景需求理解”(Omni Demand Understanding, ODU)这一新型多模态上下文推理任务,要求模型在给定交互流的基础上判断是否存在用户需求,并从多模态与对话上下文中推断其意图。解决方案的关键在于构建了一个基于挑战驱动的分类体系、由代理生成引导的视频合成方法以及真人录制的真实交互数据集——ODU-Bench,结合媒体对齐标注与人工验证,实现高保真度评估。实验表明,14个主流多模态大模型(MLLMs)中最强者Gemini 3.1 Pro仅能恢复44.7%的需上下文推断的关键信息,且11个模型在非需求场景下的误触发率超过50%,揭示当前模型在上下文感知型需求理解方面存在系统性能力短板。

链接: https://arxiv.org/abs/2609.21392
作者: Qi Chen,Yunfei Chu,Haolin He,Yifan Yang,Zihan Liu,Yuxuan Wang,Ziyang Ma,Ruiyang Xu,Meng Gao,Yinsong Yan,Ling Wang,Hui Wang,Wen Huang,Yiheng Chen,Guanrou Yang,Qiuqiang Kong,Jin Xu,Xie Chen
机构: 未知
类目: Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Sound (cs.SD)
备注:

点击查看摘要

Abstract:Natural audio-visual interaction is emerging as an important interface for AI assistants, allowing users to communicate through speech and vision rather than carefully composed text prompts. However, existing benchmarks of interactive capabilities still focus primarily on response quality, leaving a more fundamental question underexplored: can a model correctly infer the user’s underlying demand from complex multimodal interaction? Real-world user demands are often underspecified in speech and must be inferred from multimodal cues and dialogue history. This inference is further complicated by ambiguous or disfluent expression and noisy acoustic environments. Conversely, request-like speech may not constitute a demand to the assistant, leading to false triggers. We establish Omni Demand Understanding (ODU) as a distinct multimodal contextual inference problem: given an interaction stream, a model must detect whether a user demand is present and infer intent from multimodal and conversational context. ODU evaluates this capability along five dimensions, covering both single-turn and multi-turn interactions. We construct ODU-Bench using a challenge-driven taxonomy, taxonomy-guided agentic video generation, and human-recorded interactions, followed by media-grounded annotation and human verification. We evaluate 14 native MLLMs. Even the strongest, Gemini 3.1 Pro, recovers only 44.7% of key information that must be inferred from visual, acoustic, or conversational context. Moreover, 11 of the 14 models exhibit false-trigger rates above 50% on non-demand scenarios. These results reveal a systematic capability gap in current MLLMs’ ability to infer contextual user demands. We hope ODU can establish the evaluation of a previously underexplored yet essential capability in multimodal interaction: correctly understanding user demands before generating an appropriate response.

[NLP-32] Offline Multimodal Large Language Models for Decision Support in Air Operations

【速读】: 该论文旨在解决在通信受限且安全要求严格的空域作战环境中,分析人员无法依赖外部计算资源时,如何高效、准确地获取并应用作战条令知识的问题。其核心挑战在于:在离线条件下,如何通过自然语言交互实现对技术手册中文本与图像信息的精准检索与推理,同时确保知识溯源可追溯。解决方案的关键在于提出一种模块化、基于检索增强(Retrieval-Augmented)的离线大语言模型架构,该架构支持从技术手册中输入文本和图像,并在无互联网连接环境下运行,从而实现对作战条令知识的实时调用与智能辅助决策。初步试点研究显示,该系统在巴西空军四名图像分析师参与的目标识别任务中,性能与人类水平相当(8/10),且处理时间显著缩短至7.1分钟(相比人工平均26.5分钟),同时揭示了手动生成侦察任务报告(REMIR)存在极高的认知负荷(心理负荷6.0/7,努力度5.0/7)。研究为后续开展人机协同工作流的系统性评估奠定了基础。

链接: https://arxiv.org/abs/2609.21390
作者: Joao P. A. Dantas,Jelton A. Cunha,Gabriel Dietzsch
机构: Instituto de Estudos Avançados (高级研究所), São José dos Campos, São Paulo, Brazil (巴西圣保罗州桑若斯杜斯坎波斯)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Air operations rely on complex rules, established procedures, and time-critical analysis under limited connectivity and strict security constraints. In such environments, analysts must combine written doctrine with images, often without access to external computing resources. This paper studies offline large language models as decision support tools, deployed in isolated and restricted environments to give analysts access to doctrinal knowledge that remains traceable to its original sources through natural language interaction. We describe a modular retrieval-augmented architecture suitable for operation without Internet connectivity, supporting both text and image input from technical manuals. As a first step toward evaluating this architecture, we report a pilot study with four image analysts of the Brazilian Air Force, combining (i) a doctrinal knowledge assessment based on their electronic-target identification doctrine, comparing human and proposed system performance on the same test, and (ii) a measurement of the cognitive workload involved in manually producing a reconnaissance target report (Relatório de Missão de Reconhecimento - REMIR) without AI assistance. The results show a demanding manual task, especially in terms of mental demand (6.0/7) and effort (5.0/7), while the proposed system matches the human score (8/10) and completes the assessment in 7.1 minutes (compared to a human average of 26.5 minutes), establishing a baseline for future AI-assisted evaluation. Finally, we describe a future evaluation protocol to systematically compare manual and AI-assisted workflows.

[NLP-33] Consistent Relexicalization of Clinical Documents using Graph-Based Approach

【速读】: 该论文旨在解决临床自然语言处理(Clinical NLP)中重词化(Relexicalization)技术在保持数据结构完整性、关系一致性及时间连贯性方面的关键挑战。现有方法多依赖于独立实体替换,导致纵向临床记录中出现临床不一致,削弱了重词化数据集在下游科学分析中的有效性。其解决方案的关键在于提出G-RELIC(基于图的上下文重词化,以提升一致性),通过结合大语言模型(LLM)与图结构建模,构建基于图的映射机制,确保原始实体与替代实体之间的一一对应关系;同时引入确定性的时间重定位算法,有效维持时间顺序的一致性。实证评估表明,G-RELIC在关系完整性上提升30.4个百分点(从62.1%增至92.5%),在时间连贯性上提升45.9个百分点(从46%增至91.9%),且不牺牲临床数据集公认的隐私保护标准,显著提升了重词化数据的分析价值并降低了再识别风险。

链接: https://arxiv.org/abs/2609.21387
作者: Dipankar Das,Atri Mandal,Sandeep Singh,Tushar Shandhilya
机构: Oracle Health AI(Oracle健康AI)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Accepted for presentation at the Sixth International Conference on AI ML Systems (AIMLSystems 2026), Lake Como, Italy, October 6-9, 2026

点击查看摘要

Abstract:Relexicalization is a pivotal technique in clinical NLP, as it facilitates robust masking of sensitive information while synthesizing datasets that retain high-fidelity, real-world characteristics. However, preserving structural integrity, relational coherence, and temporal consistency during transformation remains a significant challenge. Existing approaches frequently rely on independent entity replacement, which results in clinical inconsistencies across longitudinal records. This reduces the value of such relexicalized datasets for downstream scientific analysis. To address these limitations, we introduce G-RELIC (Graph Based Contextual Relexicalization with Improved Consistency) which combines the power of LLMs with graphs. G-RELIC implements a graph-based mapping mechanism which optimizes for one-to-one correspondence between original and surrogate entities. It also introduces a deterministic temporal repositioning algorithm to preserve temporal consistency. Empirical evaluations on diverse, real-world clinical datasets validate that G-RELIC significantly outperforms state-of-the-art baselines. G-RELIC yields a 30.4 percentage point improvement in relational integrity (62.1% to 92.5%) and 45.9 percentage point improvement in temporal coherence (46% to 91.9%) without compromising on the recognized privacy benchmarks for clinical datasets. This maximizes the analytical utility of relexicalized datasets while minimizing re-identification risk.

[NLP-34] Prediction Dynamics in Depth-Recurrent Language Models

【速读】: 该论文旨在解决深度循环语言模型(depth-recurrent language models)中一个关键现象:尽管中间预测结果的得分持续变化,但其输出却能逐渐趋于与最终答案一致。这一问题的核心在于理解为何在有限深度下,模型的中间答案能够稳定地与最终答案对齐,即使其置信度分数仍在动态调整。解决方案的关键在于提出一种“尖锐间隔”(sharp margin)表征,将幅度约束的保守性分解为三个可解释成分:公共平移(common translation)、胜者方向相对性(direction relative to the winner),以及每个竞争者更新与其得分差距之间的配对关系(competitor pairing)。通过在Huginn-3.5B和Ouro-1.4B模型上的实证分析发现,考虑更新方向与竞争者配对关系,可在完全答案-文本评分下使平均最早合格深度比仅去除公共平移时进一步减少22.5%–34.4%。此外,研究还揭示了在共享预测分布下,公共运动(common motion)与对比运动(contrast motion)可正交分离,并分别由候选集质量(candidate-set mass)和组内集中度(within-set concentration)刻画。两者能量衰减速率可不同,从而允许偏好变化占比上升的同时,绝对更新量持续减小。这一几何与组合结构解释了有限深度下答案保持稳定的内在机制。

链接: https://arxiv.org/abs/2609.21383
作者: Xinyue Luo,Fei Yu
机构: Ant Group(蚂蚁集团)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Depth-recurrent language models refine predictions through repeated latent updates. Why can intermediate answers agree with the endpoint while their scores continue to change? We derive a sharp margin characterization that decomposes the conservatism of a magnitude bound into common translation, direction relative to the winner, and the pairing of each competitor’s update with its score gap. Across Huginn-3.5B and Ouro-1.4B, accounting for update direction and competitor pairing reduces the mean earliest qualifying depth by a further 22.5-34.4% of the total depth beyond translation removal under full answer-text scoring. This retrospective comparison uses completed trajectories. Substantial contributions also occur under label scoring. For shared predictive distributions, we separate common and contrast motion orthogonally and express the common component through candidate-set mass and within-set concentration. Common and contrast energies can attenuate at different rates, allowing a growing preference-change share to coexist with shrinking absolute updates. These findings explain finite-depth answer preservation through the geometry and composition of observed score changes.

[NLP-35] ArenaFlow: From Trajectory Ranking to Hierarchical Credit Propagation for Open-Ended Agent RL

【速读】: 该论文旨在解决生成式 AI 代理在开放性任务中难以应用强化学习的问题,核心挑战在于此类任务解空间多样且难以获取可靠标量奖励信号。现有成对评估方法虽通过相对偏好替代点态评分缓解了奖励歧视崩溃问题,但仍将丰富的比较反馈压缩为单一轨迹级奖励,导致关键中间步骤信息丢失,阻碍可复用行为模式的提炼与固化。为此,本文提出 ArenaFlow,一种面向开放性任务的分层信用传播框架。其关键创新在于:首先,采用基于锦标赛的相对排序机制生成轨迹级奖励信号;其次,在每次比较中引入结构化反思评估,揭示三类监督信号——决定性成功步骤、可复用策略技能以及检索技能的使用归属。在步骤层面,ArenaFlow 根据锦标赛存活深度,将轨迹级优势精准传播至高置信度的关键步骤,实现对局部推理行为的靶向优化;在技能层面,通过群体级使用归属估计技能效用,并结合效用感知的更新、剪枝与检索机制维护全局技能记忆,使高价值技能可作为未来探索的策略先验。实验结果验证了 ArenaFlow 在开放性任务中的有效性。

链接: https://arxiv.org/abs/2609.21378
作者: Qiang Zhang,Ruixue Ding,Fanrui Zhang,Xi Chen,Boli Chen,Shihang Wang,Yinfeng Huang,Yi Zheng,Pengjun Xie,Kaipeng Zhang,Jiawei Liu,Zheng-Jun Zha
机构: Alibaba Token Hub, Alibaba Group; Amap, Alibaba Group
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Reinforcement learning has substantially improved large language model (LLM) agents in verifiable domains, but remains difficult to apply to open-ended agent tasks, where solutions are diverse and reliable scalar rewards are hard to obtain. Recent pairwise evaluation methods alleviate reward discrimination collapse by replacing pointwise scoring with relative preferences. However, they still compress rich comparative feedback into a single trajectory-level reward, obscuring decisive intermediate steps and preventing successful behaviors from being consolidated into reusable skills. We propose ArenaFlow, a hierarchical credit propagation framework for open-ended agent reinforcement learning. ArenaFlow leverages tournament-based relative ranking to derive trajectory-level reward signals. Each comparison is further equipped with structured reflective evaluation, which reveals three types of supervision: pivotal success steps, reusable strategy skills, and usage attribution of retrieved skills. At the step level, ArenaFlow propagates trajectory-level advantages to high-confidence pivotal steps according to tournament survival depth, enabling more targeted optimization of local reasoning behaviors. At the skill level, ArenaFlow estimates skill utility from group-level usage attribution and maintains a global skill memory through utility-aware updating, pruning, and retrieval. The resulting high-utility skills further serve as policy priors for future exploration. Extensive experiments validate ArenaFlow’s effectiveness on open-ended agent tasks.

[NLP-36] Beyond Atomic Tokens: Factorizing Syllables for Language Model Pretraining

【速读】: 该论文旨在解决传统分词器(tokenizer)在处理越南语和汉语时,因仅基于字符或统计学推导的子词单元而忽略音节内部音韵结构的问题,同时避免因词汇表过大导致的效率低下。其核心解决方案是提出音位分词器(Phonemic Tokenizer),将每个音节转化为国际音标(IPA)并将其分解为三个音韵成分:声母(onset)、韵腹韵尾(rime)和声调(tone),这三个成分共同占据一个上下文位置,既保持了音节级序列长度,又实现了音韵相关音节间的表示共享。该方法采用确定性设计,无需依赖语料库进行词汇学习,从而分别获得仅112个词条(中文)和256个词条(越南语)的小型词汇表。实验表明,该分词器在两种语言中均表现出更高的瑞尼效率(Rényi efficiency),能以恰好为1的“生育率”(Fertility)覆盖标准越南语音节词典中的所有条目,并生成比现有预训练分词器更短的越南语序列。进一步地,作者构建了基于此分词器的音位BERT(PhonemicBERT),通过因子化成分嵌入与三个预测头联合重建被掩码的完整音节,在受控的中文预训练设置下,PhonemicBERT-Zh在多种语言理解任务中表现与字符、子词及SubChar模型相当甚至更优;PhonemicBERT-Vi在越南语及多语言预训练模型中也达到竞争力或更优性能。这些结果证明,音位因子化是一种紧凑、高效且可解释的替代原子或统计分割文本表示的新范式。

链接: https://arxiv.org/abs/2609.21362
作者: Nghia Hieu Nguyen,Thai Bao Huynh,Binh-An Dinh-Le,Phu Gia Hoang,Dat Tien Nguyen,Kiet Van Nguyen,Ngan Luu-Thuy Nguyen
机构: University of Information Technology, Vietnam National University, Ho Chi Minh city, Viet Nam; Independent Researcher, Germany; Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE
类目: Computation and Language (cs.CL)
备注: under review

点击查看摘要

Abstract:Conventional tokenizers represent text as characters or statistically derived subwords, overlooking the internal phonological structure of syllables and often requiring large vocabularies. We introduce \textbfPhonemic Tokenizer, a linguistically motivated tokenizer for Vietnamese and Chinese that converts each syllable into IPA and factorizes it into three phonological components: onset, rime, and tone. The three components jointly occupy one contextual position, preserving syllable-level sequence length while enabling representation sharing across phonologically related syllables. Non-phonological and unsupported units are handled through character-level fallback. This deterministic design requires no corpus-dependent vocabulary learning and yields vocabularies of only 112 entries for Chinese and 256 for Vietnamese. Intrinsic evaluation shows that the tokenizer achieves substantially higher Rényi efficiency in both languages, represents every entry in a standard Vietnamese syllable dictionary with a Fertility of exactly one, and generally produces shorter Vietnamese sequences than existing pretrained tokenizers. We further instantiate the tokenizer in \textbfPhonemicBERT, which combines factorized component embeddings and reconstructs complete masked syllables using three prediction heads. Under a controlled Chinese pretraining setup, PhonemicBERT-Zh is competitive with or outperforms character, subword, and SubChar alternatives across diverse language-understanding tasks. PhonemicBERT-Vi also achieves competitive or superior results to established Vietnamese and multilingual pretrained models. These results establish phonemic factorization as a compact, efficient, and interpretable alternative to atomic and statistically segmented text representations.

[NLP-37] From Memory to Behavior: A Behavior-Aware Role-Playing Framework for Social Media Influencers EMNLP2026

【速读】: 该论文旨在解决真实个体在社交平台上的角色扮演(role-playing)中,大语言模型(Large Language Models, LLMs)难以忠实模拟其行为反应的问题。现有基于上下文学习(In-Context Learning, ICL)的方法无法有效捕捉个体在不同情境下的动态反应模式,且对于非知名人物的评估存在困难。为此,作者提出“情境—内在状态—行为人格”(Situation–Internal state–Behavior Persona, SIBP)方法,通过引入情境依赖的行为策略,增强模型对个体行为一致性的建模能力。其解决方案的关键在于将个体的行为建模为由情境与内在心理状态共同驱动的动态过程,并构建一个包含目标个体参考信息的评估协议,使大语言模型能够更准确地进行自我评估。实验结果表明,该方法在新构建的社交媒体回复生成数据集上优于现有的ICL基线模型,且评估协议与人工判断具有中等程度的相关性;此外,在虚构角色基准测试中的表现也验证了该方法在跨场景应用中的有效性。研究结果表明,整合行为信息可显著提升对真实个体或虚构角色在社交语境中角色扮演的逼真度。

链接: https://arxiv.org/abs/2609.21349
作者: Ji-Lun Peng,Yi-Zhen Zhang,Chun-Nan Chou,Yun-Nung Chen
机构: CMoney Technology Corporation(台湾财金科技公司); National Taiwan University (国立台湾大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Accepted by EMNLP 2026 Findings

点击查看摘要

Abstract:Large language models have shown strong potential as role-playing agents for real individuals, yet faithful impersonating remains challenging. Existing in-context learning-based methods fail to capture how individuals react under different situations. In addition, LLM-based evaluation is difficult for obscure individuals. To address these challenges, we propose Situation–Internal state–Behavior Persona method to incorporate situation-dependent behavioral strategies. We further design an evaluation protocol that provides LLM evaluators with references about the impersonated individual. We evaluate our approach on a newly constructed dataset for the task of generating replies on social media. Experimental results show that our proposed method outperforms state-of-the-art ICL-based baselines, while our evaluation protocol achieves moderate correlation with human judgment. Besides, experiments on fictional-character benchmarks demonstrate that our proposed method is applicable beyond the social media setting. These findings suggest that incorporating behavioral information broadly improves the fidelity of role-playing for real individuals on social media or fictional characters.

[NLP-38] Conformal Privacy Auditing: Calibrated Re-identification Attacks with Statistical Guarantees

【速读】: 该论文旨在解决自然语言文本发布后因生成式AI(Generative AI)与辅助知识结合而引发的个体身份泄露问题,尤其针对基于大语言模型(LLM)的再识别攻击缺乏可量化、可验证的释放时风险评估手段。现有审计方法多依赖特定攻击路径的成功率报告,缺乏有限样本下的统计保证;而训练阶段的差分隐私等防护机制难以直接转化为对单个文本释放的决策支持。为此,论文提出同分布置信隐私审计(Conformal Privacy Auditing, CPA),其核心创新在于构建一种无需分布假设的校准框架,为每份发布的文本提供在交换性假设下具有用户自定义置信水平的再识别风险统计证明。CPA输出一个符合置信度要求的候选身份模糊集(conformal ambiguity set),并以集合大小作为可解释的泄露代理指标,同时兼容对数几率访问与仅采样型攻击者,实现对开源模型与专有API模型的统一审计。实验表明,CPA在多个发布基准和攻击配置下实现了校准覆盖,并揭示了辅助知识、LLM增强及发布机制变化所引起的认证可识别性显著转变,从而为跨攻击场景、数据集与发布机制的释放时关联风险提供了坚实的统计基础。

链接: https://arxiv.org/abs/2609.21340
作者: Shuo Huang,Gholamreza Haffari,Xingliang Yuan,Ting Yu,Lizhen Qu
机构: Monash University (蒙纳什大学); The University of Melbourne (墨尔本大学); Mohamed bin Zayed University of Artificial Intelligence (穆罕默德·本·扎耶德人工智能大学)
类目: Cryptography and Security (cs.CR); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Empirical identity leakage from released text is increasingly driven by attackers that combine large language models (LLMs) with auxiliary knowledge to link documents to individuals. Existing audits typically report success rates for specific attack pipelines but lack finite-sample statistical guarantees, while training-time protections such as differential privacy are difficult to translate into release-time decisions for individual natural-language documents. We introduce Conformal Privacy Auditing(CPA), a distribution-free calibration framework that provides a statistical certificate of re-identification risk for each released document against LLM-empowered adversaries. CPA outputs a conformal ambiguity set of candidate identities that is guaranteed to contain the true identity with user-chosen confidence under exchangeability, together with an interpretable leakage proxy derived from set size. CPA supports both logit-access and sampling-only attackers, enabling audits of open-source models and proprietary API models in a unified framework. Across multiple release benchmarks and attacker configurations, CPA achieves calibrated coverage and reveals sharp shifts in certified identifiability as auxiliary knowledge, LLM augmentation, and release mechanisms vary, providing a statistically grounded basis for reporting and comparing release-time linkage risk across attacker configurations, datasets, and release mechanisms alike.

[NLP-39] FairLMs: A Turnkey Library for Fairness in Language Models

【速读】: 该论文旨在解决大语言模型公平性评估中多工具异构性导致的整合难题,即现有评估工具在模型接口、证据格式、访问限制和结果类型上存在差异,难以实现方法对比与流程复用。其解决方案的关键在于提出一个名为FairLMs的Python库,通过显式声明模型能力与输入需求,统一不同组件间的交互规范;该库集成了33个内在与外在评估指标、14个覆盖四类干预策略的缓解组件、14种数据集与评分工具诊断方法,并提供对三种Transformer架构及主流托管生成API的适配器,支持基准数据集加载。所有操作前执行声明验证,结果附带完整配置信息,确保组件可兼容组合、方法可在统一协议下比较,且工作流可扩展至新模型与数据集,从而构建可复现、可扩展的公平性研究框架。

链接: https://arxiv.org/abs/2609.21296
作者: Jiale Zhang,Michael Larionov,Zichong Wang,Zhipeng Yin,Wenbin Zhang
机构: Indiana University (印第安纳大学); Carnegie Mellon University (卡内基梅隆大学); Florida International University (佛罗里达国际大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Fairness research on language models involves measuring bias, applying mitigation methods, and examining the evidence on which an evaluation rests. Existing tools offer complementary functionality through different interfaces, so combining them requires reconciling model interfaces, evidence formats, access constraints, and result types before applicability can be checked or methods compared. We introduce \textbfFairLMs, a Python library that connects these activities through explicit declarations of model capabilities and input requirements. It provides 33 intrinsic and extrinsic metrics, 14 mitigation components spanning four intervention categories, 14 dataset and scoring-instrument diagnostics, adapters for the three Transformer architectures and supported hosted completion APIs, and benchmark loaders. Declarations are checked before execution and results carry the configuration under which they were obtained, so that compatible components can be combined, methods compared under a common protocol, and workflows extended to new models and datasets. The source code is available at: this https URL.

[NLP-40] How Many Humans Is a Judge Panel Worth?

【速读】: 该论文旨在解决生成式语言模型(Generative AI)在人工判断(human judgment)模拟中的有效性评估问题,核心在于量化一个语言模型评判小组(panel of language models)在多大程度上可代表真实人类判断群体。其关键解决方案是提出一种基于谱分析(spectral analysis)的新型有效规模度量方法,通过匹配归一化残差格拉姆矩阵的参与度比率与条件独立的人类参考样本,定义出反映分类任务中人类判断多样性(nu_H);同时通过匹配分布平方误差,定义出反映分布恢复能力的指标(nu_MSE)。研究发现,在三个ChaosNLI任务中,相同32名模型组成的评判小组表现出显著差异:nu_H在4.24–6.50之间,而nu_MSE仅为2.30–3.75,表明模型组在分类一致性与分布逼近能力之间存在解耦现象。进一步通过谱身份(spectral identity)揭示了决定误差的关键因素——特征值、成员能量及平均方向权重的分离性。实证结果表明,更高的谱多样性并不必然带来更好的分布恢复,即使成员能量相等且相关性非负。此外,共识方向所解释的中心残差方差占比(gamma_co)在MNLI-m和SNLI上分别为43.8%和33.7%,说明平均化策略保留了部分共享变异。因此,有效规模并非通用指标,而是任务依赖的测量结果,强调谱多样性与分布恢复不应被混为一谈,亦不可作为通用的人类替代率标准。

链接: https://arxiv.org/abs/2609.21277
作者: Chao Li,Yingying Yu,Yunfeng Li
机构: Tsinghua University (清华大学); University College London (伦敦大学学院); Autonavi (高德软件)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 18 pages, 10 figures, and 8 tables. Code and data: this https URL

点击查看摘要

Abstract:How many human judgments does a panel of language models represent? The answer depends on what is matched. We audit categorical judge panels against empirical human label distributions, retaining disagreement that binary errors relative to one gold label collapse. We measure spectral residual diversity by matching the participation ratio of a normalized residual Gram matrix to conditionally independent human-reference draws, giving nu_H. We separately match distributional squared error, giving nu_MSE. Across three ChaosNLI tasks, the same 32-judge panels have nu_H=4.24–6.50 but nu_MSE=2.30–3.75. A spectral identity separates the eigenvalues, member energies, and averaging-direction weights that determine error. Realizable hard-label panels show that greater spectral diversity can accompany worse distribution recovery even with equal member energies and nonnegative correlations. In the observed panels, within-size ranking agreement varies sharply by task; some member additions produce conflicting changes that persist across two item halves. The consensus-direction share of centered residual variance is gamma_co=43.8% on MNLI-m and 33.7% on SNLI, quantifying shared variation retained by averaging. We provide aligned votes and analysis protocols for auditing these distinctions. Effective size is therefore a target-specific measurement: spectral diversity and distribution recovery should not be treated as interchangeable measures of panel quality or as general human-replacement rates.

[NLP-41] When Does Reasoning Help in Machine Translation? A Hierarchical Analysis of LRM Reasoning Traces EMNLP2026

【速读】: 该论文旨在解决生成式机器翻译(Machine Translation, MT)中推理痕迹(reasoning traces)的使用效果不明确的问题,即当前大型推理模型在翻译任务中引入中间推理过程时,其有效性在不同模型、语言、领域和数据集之间存在不确定性。核心问题在于:何时推理有助于提升翻译质量,何时反而可能产生负面影响。解决方案的关键在于提出一种可扩展的分层元摘要框架(Hierarchical Meta-Summarization, HMS),该框架无需预设分类体系即可自动识别推理痕迹中的粗粒度与细粒度结构。通过分析发现,最优推理语言具有模型特异性,推理长度与翻译质量呈非单调关系,且推理过程普遍呈现出重复的功能模式——包括理解/规划、翻译/起草、以及精炼/验证三个阶段。研究揭示出,应基于模型特性、推理长度及功能模式对推理行为进行差异化控制,而非对所有情况统一鼓励推理,从而实现更高效、更可控的生成式机器翻译。

链接: https://arxiv.org/abs/2609.21247
作者: Yuxiang Liu,Jiaming Luo,Eleftheria Briakou,Colin Cherry
机构: University of Illinois at Urbana-Champaign (伊利诺伊大学香槟分校); Google DeepMind; Google DeepMind; Google DeepMind
类目: Computation and Language (cs.CL)
备注: Accepted to EMNLP 2026 Main

点击查看摘要

Abstract:Large Reasoning Models increasingly use intermediate traces for machine translation, but it remains unclear when such reasoning helps or hurts. We analyze reasoning traces across models, languages, domains, and datasets, focusing on reasoning language, length, and structure. We find that the best reasoning language is model-specific, reasoning length has a non-monotonic relationship with quality, and traces exhibit recurring functional patterns. To uncover these patterns, we introduce Hierarchical Meta-Summarization (HMS), a scalable framework that induces coarse- and fine-grained reasoning structures without predefined taxonomies. HMS reveals a shared organization–understanding/planning, translating/drafting, and refining/verifying–alongside domain-specific variation. Our results suggest that MT reasoning should be controlled in a model-aware, length-aware, and pattern-aware manner rather than uniformly encouraged.

[NLP-42] Beyond Reference-Based Evaluation: Reward Models for Meta-Evaluation of Grammatical Error Correction

【速读】: 该论文旨在解决传统基于参考答案的语法错误修正(GEC)评估指标(如M²和ERRANT)存在的局限性,即这些指标假设参考集涵盖了所有合法的修正方式,因而会错误地惩罚那些语法正确且语义保持不变但表达方式不同的修正结果。其核心解决方案是提出RM-EVAL,一个基于SEEDA数据集上人类偏好标注训练的奖励模型,作为无参考的元评估器,能够从完整序列和部分序列两个层面预测类人质量判断。该方法不仅实现了与人工排序高度一致的评估性能,还可进一步通过奖励引导文本生成(RGTG)机制,以冻结基础GEC模型为前提,进行在线、基于奖励驱动的解码优化,从而在不依赖黄金参考答案的前提下,构建出一个统一的GEC系统评估与增强框架。

链接: https://arxiv.org/abs/2609.21231
作者: Ruotian Wu,Bill E. Johnson,Gene Saunders,Osama Hamzeh,Ankit Vadehra,Pascal Poupart
机构: University of Waterloo(滑铁卢大学); Vector Institute(向量研究所); Scribendi Inc.(斯克里本迪公司)
类目: Computation and Language (cs.CL)
备注: 5 pages

点击查看摘要

Abstract:Reference-based metrics for Grammatical Error Correction (GEC) such as M ^2 and ERRANT assume that the reference set enumerates all valid edits, and therefore often penalize corrections that are grammatical and meaning-preserving but phrased differently. We introduce RM-EVAL, a reward model trained on human preference data from SEEDA, as a reference-free meta-evaluator that predicts human-like quality judgments at both full-sequence and partial-sequence levels. Beyond evaluation, we show that the same reward model can be used as a learning signal to improve GEC generation via Reward-Guided Text Generation (RGTG), which keeps a base GEC model frozen and performs online, reward-driven decoding. Across SEEDA, RM-EVAL achieves strong agreement with human rankings, and RGTG yields consistent gains in reward and external validation, demonstrating a unified framework for both assessing and enhancing GEC systems without relying on gold references.

[NLP-43] Hallucination-R1: Robustness-Oriented Paraphrase Generation for Factual Consistency

【速读】: 该论文旨在解决生成式模型在面对语义等价但表述不同的问题时,出现事实性幻觉(factual hallucination)的不一致性问题,即模型对原始问题能正确回答,但在其语义等价的改写版本下却产生错误答案。这一现象揭示了模型在语义不变性下的潜在事实稳定性缺陷。现有通用改写方法难以作为有效的鲁棒性监督信号:近似复制的改写提供弱信号,而过度差异化的改写则破坏语义等价性。为此,论文提出HALLUCINATION-R1框架,其核心在于通过两阶段优化机制,学习生成既保持语义忠实又具备挑战性的改写样本,以暴露下游问答模型的事实一致性退化问题。该方法首先确保改写在语义上多样且保真,随后强化那些能有效揭示模型鲁棒性失效的改写样本。实验表明,HALLUCINATION-R1在SimpleQuestions、PopQA和TruthfulQA等多个数据集与模型架构上均实现了良好的一致性-多样性权衡,并揭示出非表面性、非语义漂移的深层事实不稳定性。此外,轻量级微调实验证明,由该框架生成的数据可有效提升模型在改写变体下的鲁棒准确率,验证了其在鲁棒性训练中的实用价值。

链接: https://arxiv.org/abs/2609.21227
作者: Wenhan Yu,Wenxin Wu,Hao Wang,Lei Sha
机构: Beihang University(北京航空航天大学); Beijing, 100191, China
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Factual hallucination is commonly defined by incorrect factual outputs. We study a paraphrase-induced hallucination setting, where a model answers a factual question correctly in its original form but generates an incorrect answer under a semantically equivalent paraphrase. Such inconsistencies expose latent factual instability under semantic invariance. However, general-purpose paraphrases are often insufficient as robustness-oriented supervision: near-copy paraphrases provide weak signals, while overly diverse paraphrases may break semantic equivalence. In this paper, we propose HALLUCINATION-R1, a robustness-oriented paraphrase generation framework that learns to produce semantically faithful yet robustness-challenging paraphrases for factual consistency. Through two-stage optimization, it first stabilizes meaning-preserving and diverse paraphrasing, then rewards paraphrases that reveal factual consistency degradation in downstream QA models. Experiments on SimpleQuestions, PopQA, and TruthfulQA show that HALLUCINATION-R1 achieves a strong consistency–diversity trade-off and exposes robustness failures across multiple model families and datasets. Further analyses indicate that these failures are not reducible to surface-level artifacts or semantic drift, but reveal non-trivial factual instability under meaning-preserving variation. A lightweight fine-tuning study also shows that HALLUCINATION-R1-generated data improves robust accuracy under paraphrase variations, suggesting its utility for robustness-oriented training. Our code and models are publicly available at this https URL.

[NLP-44] When Better Turns Do Not Make Better Agents : Diagnosing the Gap Between Next-Turn Metrics and Workflow Success EMNLP2026

【速读】: 该论文旨在解决当前代理模型(Agent models)评估范式中存在的一大关键问题:基于“黄金交互历史”(gold interaction history)的逐轮决策评估是否能够有效预测模型在自主工作流执行(autonomous workflow execution)中的实际表现。研究发现,尽管监督微调(SFT)在逐轮评估中显著提升了文本轮次成功率和下一动作预测准确率,但这些改进并未转化为真实场景下自主执行多轮客户支持任务的能力提升。其核心问题在于,现有的局部评估指标无法充分反映模型在复杂、动态工作流中的端到端协调与纠错能力。解决方案的关键在于提出应分离报告多个维度的性能指标,包括文本质量、局部动作正确性、工具调用有效性以及端到端任务完成率,以建立更全面、可靠的评估体系,从而避免将局部优化误判为整体性能提升。

链接: https://arxiv.org/abs/2609.21187
作者: Md Tahmid Rahman Laskar,Xue-Yong Fu,Gundeep Singh,Karol Chang,Kevin Sanders,Shi Zong,Tania Habib,Julien Bouvier Tremblay,Shayna Gardiner,Harsh Saini,Matthias Lee,Elena Khasanova,Quinten McNamara,Shashi Bhushan TN
机构: Dialpad Inc.
类目: Computation and Language (cs.CL)
备注: Accepted to the REALM Workshop at EMNLP 2026

点击查看摘要

Abstract:Agent models are frequently evaluated one decision at a time, where the model predicts the next action based on the gold interaction history, which is scored against a reference. We investigate whether improvement under this protocol is predictive of improved autonomous workflow execution. We study pre-SFT and supervised fine-tuned (SFT) Qwen3 models at 4B and 14B parameters and Gemma 3 models at 4B and 12B parameters on multi-turn customer-support workflows. We find that SFT consistently improves text-turn success, and that overall next-turn success increases for every model under gold-history evaluation. However, these improvements do not transfer to autonomous workflow execution. Tool-specific gains also vary across metrics and models. None of the four SFT models succeeds under holistic workflow evaluation, with strict trajectory completion reaching at most 10.4% workflow success. Our results show that next-turn evaluation is not a reliable proxy for workflow success, motivating separate reporting of text quality, local action correctness, tool execution, and end-to-end task completion.

[NLP-45] Ill Keep an Ear Out: Teaching AudioLLM s Proactive Audio Assistance INTERSPEECH2026

【速读】: 该论文旨在解决音频大语言模型(AudioLLM)在实际应用中反应被动的问题,即仅在用户主动查询时才作出响应,无法满足聋哑及听力障碍人群在可穿戴设备场景下的实时、主动式辅助需求。为此,论文提出一种名为“中断与静默建模”(Interrupt and Silent Modeling, ISM)的模型无关范式,其核心在于通过引入两个特殊标记——\texttt{interrupt} 和 \texttt{silent},将主动决策机制嵌入到大语言模型的解码过程之中,从而实现对音频流中关键事件的自主感知与判断。该方法能够捕捉四种状态:事件起始检测(onset detection)、持续相关性触发(sustained-relevance triggering)、无关内容抑制(irrelevance suppression)以及重复信息去重(de-duplication)。实验表明,基于Qwen2-Audio-7B模型的ISM在ESC-50数据集上实现了99.6%的中断F1分数和完美的去重召回率;在嘈杂的Epic-Sounds厨房音频数据上,无需领域特定训练即可达到最高中断F1值,且有效避免了过度触发或过度抑制问题。流式评估验证了其在真实场景下的可行性,平均延迟仅为3.5秒,证明了该方案具备实时部署潜力。

链接: https://arxiv.org/abs/2609.21183
作者: Amit Kumar Singh Yadav,Ritvik Shrivastava,Xuan Zhang,Seungwhan Moon,Shashank Jain,Pinar Donmez,Babak Damavandi
机构: 未知
类目: ound (cs.SD); Computation and Language (cs.CL)
备注: Accepted at Interspeech 2026

点击查看摘要

Abstract:Audio large language models (AudioLLMs) operate reactively, responding only when queried. We introduce proactive audio assistance, where an AudioLLM monitors an audio stream and autonomously decides when to alert the user from a single natural-language intent, motivated by wearable applications for Deaf and Hard of Hearing users. We propose Interrupt and Silent Modeling (ISM), a model-agnostic paradigm that embeds proactive decisions into LLM decoding via two special tokens: \textttinterrupt and \textttsilent, capturing four states: onset detection, sustained-relevance triggering, irrelevance suppression, and de-duplication. Applied to Qwen2-Audio-7B, ISM achieves 99.6% interrupt F1 and perfect de-duplication recall on ESC-50. On noisy Epic-Sounds kitchen audio, ISM achieves the highest interrupt F1 without domain-specific training, the only method maintaining strong onset detection without over-triggering or over-suppression. Streaming evaluation confirms real-time viability with 3.5-second average latency.

[NLP-46] Not All Irregularity Is Equal: Causally Isolating a Rare Failure Mode in Japanese Morphological Inflection EMNLP2026

【速读】: 该论文旨在解决神经形态生成系统在基准数据集上虽表现出高整体准确率,却掩盖了在罕见形态子类中集中出现系统性错误的问题。其核心解决方案在于提出一种基于音写法(orthography)的诊断方法,将平假名(hiragana)视为不仅承载音位转写,更蕴含形态音位结构表征的系统。研究通过两种字符级Transformer架构在五个随机种子下的实验发现,尽管整体准确率超过97%,但仅占数据不足1%的、词干以/e/结尾且需在过去时后缀前进行促音化(gemination)的特定不规则类型,却贡献了30–43%的残余错误,并导致误差量达其出现频率的34–48倍。进一步的控制消融实验表明,仅移除该子类即可带来比移除所有不规则动词更大的准确率提升。这一结果揭示,神经形态学习中的错误集中并非源于不规则性本身,而是由极低频形态模式与特定正字法过程之间的交互所驱动。论文主张在形态评估中引入细粒度子类分析,并讨论了对数据高效、符合语言发展规律的语言模型预训练的启示。

链接: https://arxiv.org/abs/2609.21179
作者: Wen Zhang
机构: 未知
类目: Computation and Language (cs.CL)
备注: BabyLM 2026 Workshop @ EMNLP 2026 CR

点击查看摘要

Abstract:Neural morphological generation systems often achieve high aggregate accuracy on benchmark datasets, yet such performance can conceal systematic errors clustered in rare morphological subclasses. We present an orthography-aware diagnosis of Japanese past-tense verb inflection, treating hiragana not merely as a transcriptional medium but as a representational system that encodes morphophonological structure. Using two character-level Transformer architectures evaluated across five random seeds, we show that although both systems exceed 97% aggregate accuracy, a single structurally specific irregular subtype, verbs whose stems end in /e/ and require gemination before the past-tense suffix and make up fewer than 1% of the data, accounts for a disproportionate 30-43% share of residual errors and contributes roughly 34-48x its prevalence to total errors. We then move from diagnosis to causal isolation: controlled ablation experiments show that removing this subtype alone produces larger accuracy gains than removing all irregular verbs combined. These findings indicate that error concentration in neural morphological learning is not driven by irregularity per se, but by the interaction between extreme low-frequency morphological patterns and specific orthographic processes. We argue that morphological evaluation should incorporate fine-grained subclass analysis, and discuss implications for data-efficient, developmentally plausible language model pretraining.

[NLP-47] CoLearn: An Agent ic Tutor that Learns its Learner in a Human–AI Co-Learning Loop EMNLP2026

【速读】: 该论文旨在解决现有辅导工具缺乏个性化适应能力的问题,即大多数部署的辅导系统仅提供固定的题库,将错误答案视为单一信号,无法动态追踪学习者的知识掌握情况与常见误解。其核心解决方案是提出CoLearn——一种交互式、自主代理型智能导师,通过构建迭代式辅导循环实现精准个性化教学。关键创新在于:(i)采用基于大语言模型(Large Language Model, LLM)作为连续观测函数的软证据贝叶斯知识追踪(soft-evidence variant of Bayesian Knowledge Tracing),构建持久的学习者状态记忆,持续更新各知识点的掌握程度与误解模式;(ii)基于该记忆自适应生成针对性问题,聚焦学习者最薄弱的知识点及反复出现的认知偏差;(iii)提供可解释的证据视图,通过实时进度可视化与盲态A/B对比测试,使个性化过程透明可验证。实验表明,在盲态A/B评估中,基于学习者记忆生成的问题被偏好68%-69%的时间,且在模拟真实学习者场景下,系统对学习者掌握程度的信念估计趋于收敛于真实水平,验证了其有效性。

链接: https://arxiv.org/abs/2609.21154
作者: Kailai He,Zhihao Wu,Linhai Zhang,Runcong Zhao,Yulan He,Jiazheng Li
机构: King’s College London(伦敦国王学院)
类目: Computation and Language (cs.CL)
备注: Accepted to EMNLP 2026

点击查看摘要

Abstract:Good tutoring adapts to the individual: it tracks what a learner knows, notices why they go wrong, and asks the next question that will help most. Most deployed tutoring tools instead serve fixed item banks and treat a wrong answer as a single bit of signal. We present CoLearn, an interactive, agentic tutor that supports an iterative tutoring loop: the learner practises, and the system builds an evidence-grounded memory of the learner’s mastery and misconceptions. This memory is updated as evidence accumulates and is used to generate the next personalised question. CoLearn has three components: (i) a persistent learner-state memory that updates per-topic mastery with a soft-evidence variant of Bayesian Knowledge Tracing, where a large language model acts as a continuous observation function; (ii) adaptive question generation that targets the learner’s weakest topic and recurring misconceptions; and (iii) an evidence view that makes personalisation visible and testable through live progress visualisation and blind A/B comparison. In blind A/B evaluation, questions conditioned on this memory are preferred over non-personalised ones 68-69% of the time, and in persona simulations with hidden ground-truth mastery the agent’s belief converges toward the learner’s true mastery.

[NLP-48] Clinician-Grounded Quality Assurance for AI-Assisted Psychiatric Intake

【速读】: 该论文旨在解决在临床实践中对基于人工智能的精神病学初诊系统进行常态化质量评估的难题,尤其是在不同临床访谈风格并存的情况下,如何实现高效、可比且具有临床意义的评估。其核心挑战在于:(1)支持跨不同访谈方式的比较;(2)降低临床人员的评估负担;(3)衡量评估结果与临床实践的相关性。为此,研究提出了一种以临床医生为中心的评估平台——InterviewPlayground,其关键创新在于构建了一个基于记忆增强型患者模拟器的开放式AI访谈评估环境。通过专家编写的临床情景(vignettes)生成交互式虚拟患者,并搭建模拟初诊平台及设计符合实际需求的评估模态,在一项包含6名临床医生的试点研究中,对比了基于GPT的大语言模型(LLM)与人工评估的表现。结果显示,尽管LLM在提取临床相关信息方面表现更优(88.0% vs. 38.9%),但其过度推断非访谈依据的临床结论(56.8% vs. 27.8%)且对识别出的安全风险描述不足(33.3% vs. 66.7%),暴露出当前生成式AI在临床推理可靠性方面的局限性,为部署前的质量保障提供了关键评估框架。

链接: https://arxiv.org/abs/2609.21149
作者: King Shi,Amanda Li,Jonathan Ivey,Synthia Qia Wang,Guan Gui,Hyunseo Kim,Peter Zandi,Jason Straub,Jacob Taylor,Ananya Joshi
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 7 pages, 3 figures, submitted to IAAI’27

点击查看摘要

Abstract:Before patients can use AI-assisted psychiatric intake systems, health systems need practical ways to routinely evaluate these tools against their clinical standards for quality assurance. Because clinicians may use different intake styles, evaluation for this task must (1) support comparison across interviewing approaches, (2) minimize clinician burden, and (3) measure clinically relevant performance for health systems deploying these technologies. We present a clinician-grounded evaluation platform built around a memory-augmented patient simulator for open-ended AI interviewing, InterviewPlayground. We created interactive patients using InterviewPlayground with our expert-authored vignettes, constructed a simulated intake platform for the interviews, and designed evaluation modalities relevant to intake. In a pilot of 6 clinicians in a 25-minute assessment compared to a GPT-based LLM intake interviewer, the LLM recovered more of the clinically relevant items embedded in the patient vignettes (88.0% vs. 38.9%), but made more clinical inferences not based on the interview (56.8% vs. 27.8%), and characterized identified safety concerns less often (33.3% vs. 66.7%), setting the stage for deployed quality assurance for this task.

[NLP-49] Scaling Forced Alignment to End-User Devices

【速读】: 该论文旨在解决传统维特比算法(Viterbi algorithm)在语音与文本强制对齐任务中面临的时间与空间复杂度高、难以处理长时序输入的问题。现有实现通常具有二次时间与空间复杂度,导致在处理长达数小时的音频数据时资源消耗巨大且效率低下。其解决方案的关键在于两项优化:一是引入希尔斯伯格算法(Hirschberg algorithm),通过原地计算实现线性内存占用,将三小时音频输入的内存需求从140 GB降至5 MB;二是将语音与文本对齐建模为受限随机游走(constrained random walk),结合可任意置信度的搜索空间剪枝策略,在保留98%以上测试案例对齐准确率的前提下,对超过20分钟的输入实现额外2倍的加速。这两项优化显著提升了算法在大规模语音训练数据挖掘中的可扩展性与实时性。

链接: https://arxiv.org/abs/2609.21145
作者: Lawry Sorenson,Michael Crandall,Eric K. Ringger,Stephen D. Richardson
机构: 未知
类目: Computation and Language (cs.CL); Data Structures and Algorithms (cs.DS)
备注:

点击查看摘要

Abstract:The Viterbi algorithm has been previously used to perform forced alignment of audio to text to mine training data from online resources. However, many existing implementations have quadratic time and space complexity, scaling poorly to long input sequences. We propose two optimizations to address this issue. First, we apply the Hirschberg algorithm to perform the alignment in place using linear memory. Second, we model the alignment between speech and text as a constrained random walk, allowing us to prune the search space with arbitrary confidence while accounting for transcription errors. The Hirschberg optimization reduces memory usage from 140 GB to 5 MB for three-hour inputs while producing identical alignments in one-third the time of torchaudio when both run on a CPU. We achieve an additional 2x speedup with pruning on inputs longer than 20 minutes while preserving alignment accuracy in more than 98% of tested cases.

[NLP-50] Detecting Hallucination in LLM s: Tracing the Topological Signatures of Impaired Context Sharing

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)中幻觉(hallucination)检测的难题,即如何有效区分模型生成内容中的真实信息与虚构内容。其核心问题在于,现有方法难以准确捕捉生成过程中信息流动的内在结构特征,尤其是在复杂注意力机制下的上下文依赖关系。解决方案的关键在于引入形式-里奇曲率(Forman-Ricci curvature)来分析注意力图(attention graph)的拓扑结构,识别出信息流动中的瓶颈区域,并提出一种能够同时捕获注意力头在局部与全局层面信息流特性的单次遍历方法。实验结果表明,该方法在多个主流大模型和基准测试上均显著优于现有的基于注意力或多响应的基线方法,且具备良好的跨架构泛化能力。进一步分析揭示,因果生成过程中词元间上下文共享能力的退化是导致幻觉发生的关键因素,具体表现为对自注意力(self-attention)的过度依赖、早期词元信息检索的弥散性以及最终Transformer层中信息过压缩(information over-squashing)等现象。

链接: https://arxiv.org/abs/2609.21096
作者: Amir Jalilifard,Anderson Rocha,Eric Wong,Marcos Medeiros Raimundo
机构: Institute of Computing, Universidade Estadual de Campinas (UNICAMP), Campinas, SP, Brazil; Department of Computer and Information Science, University of Pennsylvania, Philadelphia, PA, USA
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:In this work, we examine the topology of information flow patterns within attention graphs to effectively distinguish hallucinated from non-hallucinated responses. We analyze the Forman-Ricci curvature to identify structural patterns indicating information bottlenecks in attention graphs. We then introduce a method that captures both semi-local and global information-flow characteristics of attention heads associated with hallucinated responses. We evaluate our approach extensively across several LLMs and established benchmarks. Empirical results demonstrate that our proposed single-pass approach provides consistent improvements over existing attention-based and multi-response baselines across two hallucination-detection benchmarks, while achieving competitive performance across diverse LLM architectures. Further analysis reveals that impaired context sharing among tokens during causal generation is strongly associated with hallucination occurrences in LLMs. In particular, hallucinated responses are consistently characterized by an over-reliance on self-attention, diffused context retrieval from earlier tokens, or information over-squashing, especially in the final transformer layer.

[NLP-51] Geometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models ICML ICML2026

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在处理跨语言道德困境时存在的隐性偏见与指令遵循能力脆弱的问题,尤其关注模型在面对诚实(Honesty)、正义(Justice)与自主性(Autonomy)三者之间冲突时的决策倾向及其跨语言一致性。其解决方案的关键在于构建一个包含12,000个双选项道德困境的多语言数据集,涵盖三种核心价值对的冲突(Honesty vs. Justice、Justice vs. Autonomy、Autonomy vs. Honesty),并将其翻译为印地语、阿拉伯语、西班牙语和中文,以系统评估模型在不同语言环境下的行为表现。研究发现,未经过微调的GPT-5-mini模型在所有语言中均表现出对诚实性的显著偏好;而通过普通微调(plain fine-tuning)或直接偏好优化(Direct Preference Optimization, DPO)可有效消除Llama-3.2-1/3B模型中的首选项偏差,使准确率提升至98%以上。为进一步解耦数据集中学习到的关联性与抽象价值之间的关系,作者提出基于任务向量迁移(task vector transfer)的实验方法:计算特定价值偏好方向的任务向量后,将其与通用指令遵循向量正交化,从而实现对特定价值取向的精确隔离。实验证明该方法能够成功用于任务算术(task arithmetic),生成具有相反价值立场的模型,为可控的价值对齐提供了可解释且可操作的技术路径。

链接: https://arxiv.org/abs/2609.21094
作者: Utkarsh Agarwal,Monojit Choudhury
机构: Mohamed bin Zayed University of Artificial Intelligence (穆罕默德·本·扎耶德人工智能大学); Abu Dhabi (阿布扎比); UAE (阿拉伯联合酋长国)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Accepted at the Pluralistic Alignment Workshop @ ICML 2026, Seoul, South Korea. this https URL

点击查看摘要

Abstract:Large Language Models (LLMs) are increasingly deployed in applications that must weigh clashing moral values, yet even strong models exhibit hidden biases and brittle instruction-following across languages. We introduce a 12,000-instance dataset of two-option dilemmas covering pairwise three value conflicts: Honesty vs. Justice, Justice vs. Autonomy, and Autonomy vs. Honesty, along with their translations into Hindi, Arabic, Spanish, and Chinese, to probe cross-lingual behavior. Benchmarking on GPT-5-mini reveals that it consistently favors Honesty over Autonomy across all five languages when no policy is given. The Llama-3.2-1/3B models exhibit strong first-option bias; however, both plain fine-tuning and Direct Preference Optimization fine-tuning effectively remove this bias, increasing accuracy to greater than 98%. In order to decouple the effect of learning correlations in the dataset from abstract values, we propose a task vector transfer based experiment where after computing the task vectors for a direction of value preference we orthogonalize it with respect to the general instruction following vector. Our experiment shows that this method is effective in isolating the direction of the specific value preference that can successfully be used to conduct task arithmetic to obtain a model with the opposite stance.

[NLP-52] Aligning with Lived Experience: Heterogeneous Benefits of Fine Tuning in Mental Health Support Generation

【速读】: 该论文旨在解决在线心理健康支持平台(如Reddit)中大量求助帖未获回应的现实问题,探索如何利用大语言模型(Large Language Models, LLMs)生成基于社区经验、与用户真实生活情境相契合的同伴支持内容。其核心挑战在于:尽管现有LLMs在临床基准测试中表现优异,但其生成内容是否真正符合特定社群的共情逻辑与支持偏好仍缺乏系统评估。为此,研究提出一个以社区为中心的同伴支持数据集——COmmunity-centered Peer Engaged Support (COPES),并构建三轴评估框架,用于量化模型响应与社区视角在策略一致性(Strategy Alignment)、情绪基调(Emotion Tone)和需求适配性方面的对齐程度。关键解决方案是通过在COPES数据集上进行指令微调(SFT)和直接偏好优化(DPO)等后训练方法,显著提升模型在策略一致性和情绪基调上的社区对齐度(通用模型提升达50%),但同时也揭示出模型改进存在异质性:不同子社区及具体应对策略之间的性能差异显著,且后训练引发分布偏移,过度强化问题聚焦型建议而抑制情感聚焦型支持策略。因此,研究结论表明,虽然基于社区驱动的数据能有效提升模型输出的语境适配性,但其效果仍受限于子群体多样性与特定心理需求的复杂性,提示未来需发展更具细分适应性的对齐机制。

链接: https://arxiv.org/abs/2609.21075
作者: Mohit Chandra,Nabin Kim,Eli Min,Aamogh Sawant,Tanmay Sutar,Munmun De Choudhury
机构: Georgia Institute of Technology(佐治亚理工学院); Magic Hour
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
备注: 25 pages, 6 figures, 17 tables

点击查看摘要

Abstract:As access to professional mental healthcare remains limited, many individuals turn to online platforms such as Reddit to seek peer support situated within human lived experience. However, a significant portion of such queries go unanswered, presenting an opportunity for using Large Language Models (LLMs) to fill this gap. While LLMs have demonstrated strong performance on clinical benchmarks, their ability to generate lived-experience informed and community-aligned peer support is underexplored. Addressing this gap, we introduce the COmmunity-centered Peer Engaged Support (COPES) dataset and a three-axis evaluation framework to assess LLM alignment with community perspectives to mental health support seeking queries. Evaluating zero-shot and post-trained (SFT and DPO) models, we show that post-training on COPES significantly improves Strategy Alignment (50% for general-purpose models) and alignment in Emotion Tone. However, we also observe that such improvements are heterogeneous and alignment improvements vary significantly across subreddits and requested coping strategies. Furthermore, post-training induces distributional shifts, heavily favoring problem-focused recommendations while suppressing emotion-focused strategies. Together, this work shows that while curating community-driven data improves the alignment of LLM responses, model performance remains disparate across distinct sub-communities and specific mental health needs.

[NLP-53] Scaling Discovery through Test-Time Communication

【速读】: 该论文旨在解决现有智能体系统在科学研究中难以有效模拟协作机制的问题,尤其关注在复杂任务中多智能体间测试时通信(test-time communication)是否能显著提升性能。其核心挑战在于:尽管科学进步依赖于协作,但当前的自主智能体系统大多采用孤立运行模式,缺乏动态信息共享与协同决策能力。论文提出的关键解决方案是引入无预设角色的多智能体测试时通信框架,即所有智能体通过一个共享目录进行实时交互,共同探索问题解空间。实验表明,在ARC-AGI-3这一需要创造性求解能力的基准任务上,由 $ k $ 个通信智能体组成的团队(team@ $ k $)可达到相当于 $ 4k $ 个独立智能体的成功率,且优势随 $ k $ 增大而增强,体现规模收益叠加效应。更关键的是,该方法不仅能解决单个智能体无法攻克的任务,还能在真实研究场景中实现超越人类最优解的表现——如在多米诺骨牌拼装(polyomino packing)和MNIST分类器压缩任务中,通信团队分别超越了最佳单智能体结果及已知人类最优方案;其中四智能体团队生成的1,957字节分类器在保持99.4%测试准确率的同时,体积小于此前最优解。然而,这些增益并非无条件成立:当计算资源受限或缺乏明确进展度量时,独立运行可能更具优势。总体而言,在具备充足算力与清晰反馈信号的前提下,多智能体通信能够稳定带来更强的性能表现,揭示了协作式生成式智能体在复杂问题求解中的巨大潜力。

链接: https://arxiv.org/abs/2609.21032
作者: Jongho Park,Vasilis Kontonis,Shivam Garg,Akshay Krishnamurthy,Dimitris Papailiopoulos
机构: 未知
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 34 pages, 12 figures

点击查看摘要

Abstract:Science advances not in isolation but through collaboration, yet existing agentic systems capture little of this. Whether communicating agents help remains an open question with mixed prior results. We show that test-time communication can substantially outperform independent parallel attempts on challenging tasks, where sharing a breakthrough can push the whole group forward. We first study the effect of scaling multi-agent test-time communication, where agents have no predefined roles and communicate via a shared directory, on ARC-AGI-3, a benchmark requiring novel problem solving. We find that a team of k communicating agents, team@ k , matches the success rate of 4k independent agents, and this advantage grows with k , suggesting gains compound with scale. The effect is not merely efficiency: a task that no single agent can solve, a team of agents can solve reliably. Furthermore, these gains transfer to research-oriented tasks, given sufficient compute. On polyomino packing, communicating agents outperform best@ k and exceed the prior best-known score. On MNIST classifier compression, communication surpasses the best-known human solution. A team of four agents produced a 1,957-byte classifier submission achieving 99.4% test accuracy, smaller than both the best-known human solution and the best single-agent result. These gains are not unconditional. Independent agents may outperform communication when compute is limited or when a clear measure of progress is absent. However, under sufficient compute and clear feedback, multi-agent communication consistently yields stronger results.

[NLP-54] Voice-Light: A Full-Duplex Cascaded Voice Agent with Causal Turn-Taking and Speculative Generation

【速读】: 该论文旨在解决全双工语音交互系统在真实对话场景中面临的多重挑战,包括:如何在用户重叠发言时避免误取消(false cutoff)、如何在话语未完全确定前提前生成响应、以及如何确保被取消的音频内容不会进入对话历史。其核心解决方案是提出Voice-Light——一种基于级联架构的全双工语音代理系统,关键在于融合了即时声学起始检测(immediate acoustic onset)、共享流式自动语音识别(ASR)编码器的因果适配器(causal adapter)、可逆回放控制(reversible playback control)以及私有推测性响应生成(private speculative response generation)。通过并行执行结构化工具调用与可听桥接语音(audible bridge speech),并利用浏览器确认机制使渲染音频成为持久对话历史的权威依据,系统实现了低延迟响应与高可靠性。尽管早期学习型完成检查点在假切断率上表现较好(2.70%),但端到端回合召回率仅12.53%,远低于基于Silero时序策略的95.60%,因此系统最终采用混合控制器而非完全依赖学习策略。在三次无脚本的运行录音会话中,36次测量响应回合的中位延迟为758毫秒,其中21次低于800毫秒,验证了系统的实时性。研究结果已通过合成数据、模型权重、评估代码及部署配置等开源支持。

链接: https://arxiv.org/abs/2609.20995
作者: Bertil Braun
机构: 未知
类目: ound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
备注: 9 pages, 4 figures, 6 tables. Code, datasets, and model artifacts: this https URL ; live demo: this https URL

点击查看摘要

Abstract:Natural spoken interaction requires more than streaming ASR, language generation, and speech synthesis: a system must react to overlap without canceling on every acknowledgment, prepare a response before a turn is certain, and ensure canceled audio cannot enter conversation history. We present Voice-Light, a full-duplex cascaded voice agent that combines immediate acoustic onset, a causal adapter sharing a streaming ASR encoder, reversible playback control, and private speculative response generation. Structured tool calls execute concurrently with audible bridge speech, while browser acknowledgments make rendered audio authoritative for durable history. Locked evaluation on 1,673 real-conversation silence candidates found that an earlier learned completion checkpoint preserved a 2.70% false-cutoff rate but reached only 12.53% end-of-turn recall, compared with 95.60% for a Silero timing policy. The deployed system therefore retains a hybrid controller rather than claiming a learned-policy replacement. Across three unscripted operator-run microphone sessions, 36 measured response turns had a 758 ms median from final VAD endpoint to first server audio; 21 turns were below 800 ms. These sessions are an instrumented case study, not a controlled user evaluation. We release the synthetic data, model artifacts, evaluation code and summaries, source code, and deployment configuration supporting the result.

[NLP-55] μ2-Bench: A Multilingual Machine Unlearning Benchmark

【速读】: 该论文旨在解决多语言大模型(Multilingual Large Language Models, LLMs)中潜在有害内容与敏感隐私数据在直接训练及跨语言间接传播过程中难以有效清除的问题。现有研究中的多语言机器遗忘(Multilingual Machine Unlearning, MMU)方法缺乏系统性评估框架,导致其是否真正实现目标知识在所有语言上的彻底消除尚不明确。为此,本文提出μ²-Bench,一个全面模拟记忆化、遗忘与评估全流程的多语言机器遗忘基准测试体系,其关键创新在于:1)覆盖广泛的语言种类,实现跨语言泛化;2)同时在训练语言和未见语言上进行评估,检验遗忘效果的迁移能力;3)衡量知识在多语言间的分散分布特性。实验表明,有效的MMU方法必须充分考虑多语言特征,该工作通过深入分析揭示了多语言遗忘机制的本质要求,为后续研究提供了可复现的评估标准与理论指导。

链接: https://arxiv.org/abs/2609.20945
作者: Kyomin Hwang,Hyeonjin Kim,Hyunho Lee,Yearim Kim,Yeji Song,Nojun Kwak
机构: GSCST, Seoul National University(首尔国立大学高级科学与技术研究所); AIIS, Seoul National University(首尔国立大学人工智能研究所)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Undesired information such as harmful content and private data propagates through Multilingual Large Language Models (LLMs) via direct training and indirect cross-linguistic spread. Multilingual Machine Unlearning (MMU) aims to remove such information, yet its evaluation remains underexplored, leaving unclear whether unlearning truly eliminates target knowledge across all languages. To bridge this gap, we introduce \mu^2 -Bench, an MMU benchmark that simulates the full pipeline of memorization, unlearning, and evaluation across diverse languages. It 1) spans a broad set of languages, 2) evaluates on both training and hold-out languages, and 3) assesses knowledge as dispersed across multiple languages. We show that successful MMU requires methods that reflect multilingual characteristics, and conduct analysis to provide deeper insights into MMU.

[NLP-56] Generative Artificial Intelligence Chatbots for Motivational Interviewing: A Scoping Review From System Design to Intervention Outcomes

【速读】: 该论文旨在解决生成式人工智能(Generative AI, GenAI)在动机访谈(Motivational Interviewing, MI)对话系统中的应用中存在的设计、评估与干预转化证据碎片化问题。其核心解决方案在于通过系统性文献综述(PRISMA-ScR框架),对基于GenAI的MI聊天机器人在系统设计、安全性、MI质量、用户感知及干预效果等方面的现有证据进行整合与描述性综合。研究发现,尽管多数GenAI-MI聊天机器人能实现符合动机访谈原则的交互,并获得用户对共情、可用性与帮助性的积极评价,但其在长期行为或功能改变方面的证据仍十分有限,且存在安全监测不统一、评估方法异质性强等问题。因此,关键突破点在于强化运行时安全监控、标准化MI质量评估体系,并采用更长周期的对照研究设计以验证可持续的健康行为改变效果。

链接: https://arxiv.org/abs/2609.20902
作者: Runze Hu,Jingqi Kong,Yang Yang,Yihang Yang,Jingyao Liu,Haizhou Tang,Shanghang Zhang,Zheng Liu
机构: 未知
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Motivational interviewing (MI) is a collaborative approach to elicit autonomous motivation for health behavior change. Generative AI (GenAI) offers new ways to deliver MI via conversational systems, but evidence on their design, assessment, and translation into interventions remains fragmented. This scoping review characterized evidence on GenAI-MI chatbots across system design, safety, MI quality, user perceptions, and intervention outcomes. We conducted a PRISMA-ScR scoping review. Nine datasets were searched for studies published or publicly available from January 1, 2015 to June 2, 2026 that used GenAI to generate MI chatbot responses or counselor utterances. Data were extracted using a predefined framework and synthesized descriptively. Forty-seven reports (48 studies) were included. Twenty (41.7%) focused on system design without direct participant use; 28 (58.3%) involved direct interaction. Most systems were text based and disembodied; 23 (47.9%) incorporated dynamic adaptation. Safety measures were unevenly reported. Among studies with direct use, 21/28 (75.0%) reported informed consent or user education. Thirty (62.5%) assessed MI quality, generally suggesting MI-consistent interactions. User perceptions were favorable, especially empathy, usability, helpfulness, and intention to use, though measures were heterogeneous. Eighteen (37.5%) reported intervention outcomes, mostly after a single session. Positive findings were more consistent for short-term motivation than sustained behavioral or functional change. GenAI-MI chatbots can deliver MI-consistent interactions perceived favorably, but evidence for sustained behavioral or functional change is limited. Future research should strengthen runtime safety monitoring, standardize MI quality assessment, and use longer-term comparative designs with behavioral and functional outcomes.

[NLP-57] BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence

【速读】: 该论文旨在解决传统商业智能(Business Intelligence, BI)工作流中用户需手动完成复杂数据准备步骤(如表识别、数据转换、关联构建等)所带来的效率低下与使用门槛高的问题。其核心挑战在于如何实现从自然语言业务问题到精准答案的端到端(end-to-end)自动化推理,而无需用户干预数据预处理过程。解决方案的关键在于提出一种工具增强型的BI-Agent框架:该框架将完整的BI任务分解为结构化数据操作子任务(如搜索、连接、转换),并协调使用专用的数据管理方法进行多阶段流程控制;同时设计了一种基于真实BI项目轨迹的后训练(post-training)框架,通过监督微调(SFT)与强化学习(RL)相结合的方式,利用真实工作流中的训练轨迹对模型进行领域适配。实验表明,该方法在自定义基准BI-Bench上显著提升了大语言模型(LLM)的准确率,相较于原始模型最高提升达40个百分点,经后训练后的系统进一步获得最高30个百分点的增益,验证了工具增强推理与领域特定后训练相结合的有效性,为生成式AI在复杂企业级数据分析场景中的应用提供了重要方向。

链接: https://arxiv.org/abs/2609.20886
作者: Chuxuan Hu,Yeye He,Penny Zhou,Wee Hyong Tok,Daniel Kang,Surajit Chaudhuri
机构: University of Illinois Urbana-Champaign (UIUC)(伊利诺伊大学厄本那-香槟分校); Microsoft Research (微软研究院); Microsoft(微软)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Databases (cs.DB)
备注: code and data are available at \url{ this https URL }

点击查看摘要

Abstract:Business intelligence (BI) is a cornerstone of enterprise decision-making and is widely used by enterprise users in software such as Power BI and Tableau. In traditional BI workflows, users need to prepare data by (1) identifying relevant tables, (2) performing data transformations, and (3) building join relationships, before they can (4) answer their business questions. These steps can be complex and time-consuming, making BI challenging. Given the strong capabilities of large language models (LLMs) in working with data, we study their ability to answer BI questions end-to-end, without requiring users to manually perform the tedious preparation steps. To do this, we harvest a large collection of real-world BI projects from public sources, and manually extract pairs of (questions, ground-truth answers) from real user dashboards. The resulting benchmark, BI-Bench, is the first benchmark to systematically study LLMs’ ability on end-to-end BI. We find that even frontier LLMs perform poorly on BI-Bench, with less than 50% accuracy. To address their limitations, we design a tool-augmented BI-Agent that decomposes BI workflows into subtasks on structured data, such as search, join, and transform, and orchestrates specialized data management methods across BI stages. Furthermore, we develop a post-training framework that synthesizes training trajectories from real BI projects, enabling BI-Agent to be further post-trained using both supervised fine-tuning (SFT) and reinforcement learning (RL). BI-Agent achieves substantial accuracy gains of up to 40 percentage points with vanilla LLMs, and post-trained BI-Agent yields gains of up to 30 points. Our results highlight the importance of combining tool-augmented reasoning with domain-specific post-training in complex BI workflows, and point to promising directions for future research. Comments: code and data are available at \urlthis https URL Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Databases (cs.DB) Cite as: arXiv:2609.20886 [cs.LG] (or arXiv:2609.20886v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2609.20886 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[NLP-58] MME-Safety: A Fine-grained Benchmark for Safety Evaluation of MLLM s

【速读】: 该论文旨在解决多模态大语言模型(Multimodal Large Language Models, MLLMs)在跨模态交互中暴露的安全脆弱性问题,尤其针对现有评估基准缺乏细粒度意图相关标注、依赖单一维度指标导致安全鲁棒性评估不全面的缺陷。其解决方案的关键在于提出一个经过严格验证的基准测试框架MME-Safety,采用独特的四维标注体系,对风险场景、危害严重程度及模态特异性隐蔽性水平进行系统分类;同时引入分层评估框架,从基础响应可靠性、实际风险暴露程度到防御行为结构完整性三个层面综合评估模型安全性。通过在17个前沿MLLM上开展零样本评估,揭示了跨模态输入配置与链式思维(Chain-of-Thought, CoT)推理机制带来的潜在安全风险,凸显了在多模态场景下实现具备推理感知能力的安全对齐的紧迫性。

链接: https://arxiv.org/abs/2609.20850
作者: Yueming Lyu,Yilian Shi,Haoxiang Tan,Linzhuang Zou,Qihao Wang,Guihua Yu,Jie Qin,Xin Gao,Chenyang Si,Jing Dong,Caifeng Shan
机构: Nanjing University(南京大学); Meituan(美团); Yale University(耶鲁大学); Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
类目: Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:While Multimodal Large Language Models (MLLMs) show remarkable advancements, their cross-modal capabilities introduce complex vulnerabilities that easily bypass unimodal filters. Existing benchmarks lack fine-grained intent-related annotations and rely on unidimensional metrics, hindering comprehensive robustness evaluation. To address this, we propose MME-Safety, a rigorously verified benchmark featuring a unique four-dimensional annotation schema that categorizes risk scenarios, harm severity, and modality-specific stealth levels. Furthermore, we introduce a hierarchical evaluation framework to assess fundamental response reliability, actual risk exposure, and the structural integrity of defensive behaviors. Extensive zero-shot evaluations across 17 state-of-the-art MLLMs provide a comprehensive safety profile of current multimodal systems. Our analysis systematically investigates cross-modal input configurations and uncovers safety implications associated with Chain-of-Thought (CoT) reasoning. These multifaceted findings underscore the urgent need for robust, reasoning-aware safety alignment in the multimodal landscape.

[NLP-59] Enhancing Audio Reasoning via Semantic Summary Prediction INTERSPEECH2026

【速读】: 该论文旨在解决大型音频语言模型(Large Audio Language Models, LALMs)在复杂问答任务中存在推理鸿沟(reasoning gap)的问题,即当采用显式思维链(Chain-of-Thought, CoT)推理时,模型准确率反而低于直接生成答案的情况。其核心假设是:过长的推理序列会分散模型对音频输入的关注,导致关键音频信息被忽略。为此,作者提出SPARE(Semantic Prediction for Audio REasoning),其关键创新在于引入一个与最终结论对齐的寄存器标记(register token),并通过与Sentence-BERT嵌入之间的余弦相似度损失进行优化,从而在推理开始前将目标语义目标注入模型的隐空间。该方法有效引导模型在推理初期更早地关注音频内容,在不增加额外推理开销的前提下,显著提升了零样本推理性能,并增强了对音频输入的早期注意力。

链接: https://arxiv.org/abs/2609.20849
作者: Francesco Bonzi,Pooneh Mousavi,Cem Subakan,Mirco Ravanelli
机构: Concordia University (康考迪亚大学); Mila - Quebec AI Institute (魁北克人工智能研究所); Université Laval (拉瓦尔大学)
类目: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
备注: Accepted at Interspeech 2026

点击查看摘要

Abstract:Large Audio Language Models (LALMs) perform well on complex question answering but often show a reasoning gap, where explicit Chain-of-Thought (CoT) reduces accuracy compared to direct answers. We hypothesize that long reasoning sequences shift attention away from the audio input. To address this, we propose SPARE (Semantic Prediction for Audio REasoning), which introduces a register token aligned with the final conclusion using a cosine similarity loss with a Sentence-BERT embedding. This conditions the model’s latent space with the target semantic goal before reasoning begins. Experiments on MMAU and MMAR with SALMONN show improved zero-shot reasoning and stronger early attention to audio without additional inference cost.

[NLP-60] Reading Anxiety or Reading the Label? Comparing Fine-Tuned and Frontier Models for Anxiety Detection on Social Media

【速读】: 该论文旨在解决在线文本中焦虑症状的早期检测问题,尤其关注不同模型在真实场景下的有效性与可靠性。其核心挑战在于现有评估方法存在显著偏差:在所使用的Reddit语料库中,69.3%的标注为焦虑的帖子包含“anxiety”或其变体,远高于其他心理状况的出现频率,导致分类器可通过关键词匹配而非真正理解语言语义而获得高分。因此,论文提出的关键解决方案是引入“词汇依赖性(lexical dependence)”作为评估指标——即在同一模型上分别对原始文本和删除关键词后的文本进行测试,并计算性能下降幅度,以此量化模型对显式词汇的依赖程度。研究发现,尽管前沿的生成式模型(如zero-shot模型)在F1分数上表现最优(0.846),但一个仅110M参数、经过领域适应训练的编码器也能达到0.831的性能,且无需外部API;同时,领域预训练仅贡献了约0.7分的提升。更重要的是,词汇依赖性差异巨大(8.6至25.4分),且不反映模型真实能力,例如经过LoRA微调的3B模型表现出最强的词汇依赖性,甚至超过传统的TF-IDF分类器。这表明,以往基于该数据集报告的性能结果存在明显上界偏误,尤其是对微调模型而言,其性能被严重夸大。

链接: https://arxiv.org/abs/2609.20847
作者: Cris Huynh,Arlene Pham
机构: 未知
类目: Computation and Language (cs.CL); Social and Information Networks (cs.SI)
备注: 11 pages, 5 figures, 2 tables

点击查看摘要

Abstract:Anxiety is among the most common mental health conditions, and people often write about it online well before seeking clinical help. Practitioners building detection tools face a concrete choice: call a frontier commercial model, fine-tune a smaller model in-house, or deploy a conventional classifier. We compare six conditions spanning all three on a held-out Reddit test set under a single controlled protocol. We also identify a confound in how this task is evaluated. In the corpus used here, 69.3% of anxiety-labelled posts contain the word “anxiety” or a variant, roughly twice the rate of comparable conditions, so a classifier can score well by keyword matching rather than by modelling the language of the condition. We therefore evaluate every model twice, on original text and with those terms deleted, and report the difference as lexical dependence. A frontier model leads on anxiety F1 (0.846), but a 110M-parameter domain-adapted encoder reaches 0.831 with no external API dependency, and mental-health domain pretraining accounts for only 0.7 of those points. Lexical dependence spans 8.6 to 25.4 points and does not track model capability: the LoRA fine-tuned 3B model is the most keyword-dependent condition tested, above even a TF-IDF classifier, while the frontier zero-shot model is the least. Published figures on this corpus are therefore upper bounds, and the inflation is largest for the fine-tuned models such figures typically report.

[NLP-61] Rewarding Efficient Reasoning Improves Abstention on Underspecified Tasks in Reasoning Models

【速读】: 该论文旨在解决现代大型推理模型(Large Reasoning Models, LRMs)在面对无法回答的任务时,缺乏有效判断能力的问题,即难以识别何时应放弃回答。研究表明,人类在处理无解任务时的推理努力存在上限,而LRMs却倾向于在无解任务上生成更长的思维链(Chain of Thought, CoT),造成计算资源浪费。其解决方案的关键在于引入一种受资源理性认知理论启发的新型广义奖励策略(Generalized Reward for Prompt Optimization, GRPO),该策略通过鼓励模型高效评估任务是否具备解答所需信息,从而实现对“不回答”的合理决策。基于此奖励进行微调后,多个40亿参数规模的LRM在保持原有回答能力的同时,显著提升了类似人类的拒答表现(平均提升12.8%),并使思维链长度平均缩短44%,大幅增强了推理效率。

链接: https://arxiv.org/abs/2609.20846
作者: Polina Tsvilodub,Max Höth,Michael Franke,Björn Deiseroth,Carina Kauf
机构: University of Tübingen(图宾根大学); Aleph Alpha Research(阿尔法阿尔法研究); Lab1141
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 19 pages, 9 figures

点击查看摘要

Abstract:While modern large reasoning models (LRMs) excel at providing correct answers in many tasks, we provide additional evidence for the observation that they often struggle with a critical capability: knowing when to abstain from answering. We analyze this gap by comparing LRM behavior to results from a human study, revealing that human reasoning effort on unanswerable tasks is upper-bounded by answerable tasks, whereas LRMs waste computational resources by generating longer Chains of Thought (CoTs) on unanswerable than on answerable prompts. To overcome this inefficiency, we take inspiration from a resource-rational perspective on human cognition and introduce a novel GRPO reward that encourages efficient reasoning about whether the task contains all the information needed to solve it. Fine-tuning several 4B LRMs with this reward leads to human-like abstention performance gains (+12.8% on average) while retaining answering capabilities and boosting the models’ efficiency (44% shorter CoTs on average).

[NLP-62] Reading Less While Writing: A Closed-Form Bandwidth Dial for Streaming Multimodal Decoders

【速读】: 该论文旨在解决实时语音或视频转文本任务中,传统解码器因需等待完整输入才能生成输出而导致的延迟问题,尤其是在流式(streaming)场景下,固定延迟策略(如wait-k)无法适应输入长度与节奏变化,造成性能瓶颈。其核心解决方案是提出ZENDAYA调度机制,通过一个连续可调参数γ动态控制解码进度与可见源前缀的关系,使解码过程在生成进展中以闭式函数形式依赖于输入的预测长度,从而将普通离线解码器与实时流式解码器统一为同一模型家族的两个端点。该方法的关键在于:利用单一参数γ同时实现对平均每输出词所消耗源内容比例Eˉ(γ)1/(1+γ)\bar{E}(\gamma) \approx 1/(1+\gamma)的精确闭式控制,使γ兼具延迟调节与可解释性预算的功能。论文进一步证明了结构依赖性定理——在任何预先设定且非递减的调度策略下,任意输出词均不会依赖尚未到达的输入,该性质适用于训练与未训练权重,并可扩展至任意异步到达的无界流数据。实验结果表明,在三个公开数据集(Charades-STA、ActivityNet Captions、LibriHeavy)上,仅2900万参数的小型解码器在读取更少源信息的前提下,仍能匹配甚至超越固定调度方案,尤其在低延迟场景下优势显著,且流式METEOR得分在所有数据集上均具有统计显著性,验证了“看得越少反而越好”的反直觉现象——过多源信息会稀释注意力,而此时模型自身输出尚不足以为注意力提供锚点。

链接: https://arxiv.org/abs/2609.20845
作者: Yasir Mehmood,Kashif Javed
机构: University of Engineering and Technology (UET)(巴基斯坦工程与技术大学)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:A decoder that turns video or audio into text conventionally consumes the entire input before emitting a word. Offline this is merely more than the task requires; live it is impossible, since a caption cannot wait for a match to end. Streaming systems bolt on a fixed rule such as wait- k , which waits for the same number of input tokens before every word, regardless of the input’s length or pace. We replace the fixed offset with ZENDAYA, a schedule governed by a single continuous parameter \gamma . It makes the visible source prefix a closed-form function of generation progress, scaled by the input’s own predicted length, so an ordinary offline decoder and a real-time streaming decoder become two endpoints of one family rather than separate models. The same scalar fixes, in closed form, the mean fraction of source consumed per emitted word, \barE(\gamma) \approx 1/(1+\gamma) , making it at once a latency dial and an interpretable budget. We prove a structural dependency theorem: under any schedule fixed in advance and non-decreasing, no emitted token can depend on input that has not yet arrived. The guarantee holds for trained and untrained weights alike, and extends to unbounded streams under arbitrary asynchronous arrival. The empirical result is counterintuitive: seeing less can produce better text, because a flood of source dilutes attention exactly when the model has the least of its own output to anchor on. Trained from scratch across two modalities and three public corpora (Charades-STA, ActivityNet Captions, LibriHeavy), a compact 29M-parameter decoder matches or beats the fixed schedule while reading less of the source, with the sharpest gains at the lowest latencies, where a fixed offset collapses. Streaming METEOR gains are statistically significant on all three corpora. Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG) Cite as: arXiv:2609.20845 [cs.CL] (or arXiv:2609.20845v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2609.20845 Focus to learn more arXiv-issued DOI via DataCite

[NLP-63] Boosting Deepresearch and LongContext Ability with Self-Generated Deepresearch Rollouts Traces

【速读】: 该论文旨在解决深度研究(Deep Research, DR)智能体在多轮搜索与网页访问过程中因上下文持续膨胀而导致的长上下文理解能力不足问题。尽管已有基于强化学习的DR-RL方法,但仍有61.6%的预测误差源于长上下文幻觉及跨文档证据整合失败,表明当前模型在处理长序列信息时存在显著瓶颈。其解决方案的关键在于提出一种名为“DR Rollouts to LongContext-QA (DR-to-Long)”的数据构建方法,通过将原有DR-RL轨迹中的压缩片段和网页摘要替换为对应完整URL内容,生成具有丰富多文档上下文的长文本问答(LongQA)样本,从而在无需额外标注成本的情况下有效扩展上下文长度并保留原始证据关系。在此基础上,论文进一步设计了DLD(DR - LongQA - DR)-RL框架:先执行短周期的DR-RL以收集初始轨迹,再将其转化为LongQA实例,利用长上下文强化学习(LongQA-RL)阶段提升模型对长上下文的理解与推理能力,最后回归全量DR-RL以持续优化任务表现。实验表明,DLD-RL在三个深度研究基准上相比标准DR-RL提升7.3%,在三个长上下文基准上提升13.5%,验证了增强长上下文能力对提升复杂推理性能的核心作用。

链接: https://arxiv.org/abs/2609.20844
作者: Zihan Wang,Hao Wang,Boyuan Jiang,Yiqun Zhang,Shi Feng,Xiaocui Yang,Yiwen Ye,Jianghang Lin,Xiaozhong Ji,Jinghao Lin,Kai Wu
机构: 未知
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Deepresearch (DR) agents interact with real-world web environments through multi-turn search and visit, causing their contexts to grow rapidly over time. We observe that, even after DR Agentic Reinforcement Learning (DR-RL), 61.6% of the model’s remaining prediction errors can still be attributed to insufficient long-context understanding, including longcontext hallucination and failures in cross-document evidence integration. It motivates us to further break the bottleneck of DR-RL by strengthening the model’s long-context ability. However, effective LongContext training requires more than simply increasing context length. To bridge the data gap, we propose `DR Rollouts to LongContext-QA (DR-to-Long)'. The method repurposes DR-RL trajectories, which naturally contain search histories, visited webpages, evidence snippets, and final-answer supervision. It then replaces the compact snippets and webpage summaries in each trajectory with the full contents of their corresponding URLs, producing substantially longer multi-document contexts while preserving the original evidence relationships. Building on DR-to-Long, we introduce DLD (DR - LongQA - DR)-RL. DLD-RL first performs a short DR-RL stage to collect rollout trajectories, which are then converted into LongQA instances at zero annotation cost. The model is subsequently optimized with LongQA-RL to strengthen LongContext ability, followed by full DR-RL to continue improving its DR capability. Experiments show that DLD-RL outperforms standard DR-RL by 7.3% on three Deepresearch benchmarks and improves performance by 13.5% on three long-context benchmarks.

[NLP-64] VISPATH: Visual-Intent-Guided Path Reasoning for Multimodal Knowledge Graph Question Answering

【速读】: 该论文旨在解决多模态知识图谱问答(Multimodal KGQA, MM-KGQA)中模型在多跳推理过程中无法充分利用中间步骤的多模态线索的问题。现有方法通常仅将多模态信息用于初始实体定位或证据检索,随后的多跳推理退化为纯文本图搜索,导致在后续推理阶段丢失重要的视觉或跨模态语义信息。为此,论文提出VISPATH——一种基于视觉意图引导的路径推理框架。其核心创新在于:首先通过融合多模态感知与图结构线索实现可靠的起始实体定位;随后在每一步推理中动态重构与当前路径状态相关的特定跳跃意图(hop-specific multimodal intent),以指导路径扩展,确保每一步推理均受当前上下文和多模态信息的约束;进一步通过推理链剪枝机制,基于问题一致性、推理草图及各跳意图对候选路径进行筛选与优化;最终验证所选证据是否足以生成答案。该方案显著提升了多模态多跳推理能力。研究还构建了VISPATH-Bench,一个涵盖2至4跳推理任务的基准数据集,实验表明,相较于主流基线,VISPATH在多个基准上表现更优,尤其当采用GPT-4o作为主干模型时,在VISPATH-Bench上超越GPT-5.4,平均准确率相对提升10.6%,2跳推理任务提升达13.1%。

链接: https://arxiv.org/abs/2609.20843
作者: Jinke Wu,Zhengpin Li,Mengzhe Jia,Yang Li,Wentao Zhang
机构: 未知
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: Preprint

点击查看摘要

Abstract:Knowledge graph question answering (KGQA) enables models to answer natural-language questions through structured graph reasoning and has achieved substantial progress across many benchmarks and applications. Recently, multimodal KGQA (MM-KGQA) has attracted increasing attention because many questions require jointly using multimodal inputs and KG evidence. However, existing MM-KGQA methods typically use multimodal information only for starting entity grounding or evidence retrieval, after which multi-hop reasoning degenerates into text-only graph search. As a result, they cannot exploit multimodal cues that become important at intermediate hops. To address this limitation, we propose VISPATH, a visual-intent-guided path reasoning framework for MM-KGQA. VISPATH first identifies a reliable starting entity by combining multimodal grounding with graph-structural cues. It then performs intent-guided path discovery by recomputing hop-specific multimodal intent from the input, question, and current partial paths, so that each expansion is guided by the current reasoning state. The discovered paths are further refined through reasoning-chain pruning, which evaluates candidate paths as complete evidence chains based on their consistency with the question, reasoning sketch, and hop-specific intent. Finally, VISPATH checks whether the selected evidence is sufficient for answer generation. We further construct VISPATH-Bench, a benchmark for evaluating multimodal multi-hop reasoning over KGs, covering questions that require two to four hops over KG paths. Extensive experiments on VISPATH-Bench and three additional multimodal QA benchmarks show that VISPATH consistently outperforms strong baselines. Notably, with GPT-4o as the backbone, VISPATH surpasses GPT-5.4 on VISPATH-Bench, achieving a 10.6% relative improvement in average accuracy and a 13.1% improvement at 2-hop reasoning.

[NLP-65] COAL-SQL: Coverag e-Guided Augmentation and Failure-Driven Learning for Text-to-SQL Post-Training

【速读】: 该论文旨在解决开源大语言模型在复杂真实场景下进行文本到SQL(Text-to-SQL)生成时,因缺乏充分覆盖各类SQL结构的高质量训练数据及动态适应模型缺陷的学习策略而导致性能受限的问题。其核心挑战在于现有数据集对SQL语法结构的覆盖不全,而传统数据增强方法无法精准识别并填补结构性缺失;同时,单纯依赖监督微调(Supervised Fine-Tuning, SFT)或强化学习(Reinforcement Learning, RL)难以在训练过程中动态发现并修复模型暴露的薄弱环节。为此,论文提出COAL-SQL统一框架,其关键创新在于融合基于覆盖率引导的数据增强(Coverage-Guided Augmentation, CGA)基于失败驱动的学习(Failure-Driven Learning, FDL):CGA通过贪心选择机制识别原始数据集中缺失的SQL结构,并生成补充样本以提升结构覆盖度;FDL则在步骤层面利用强语言模型生成的可验证推理轨迹对累积失败案例进行针对性微调,在周期层面根据失败模式检索结构相关示例,构建靶向练习,从而系统性增强模型的SQL生成能力。实验表明,仅使用12,600个独特后训练样本,COAL-SQL在BIRD开发集上达到64.9%的执行准确率,显著优于同规模训练基线。

链接: https://arxiv.org/abs/2609.20842
作者: Qifeng Cai,Xuanguang Pan,Hao Liang,Chang Xu,Wentao Zhang
机构: Peking University (北京大学); Beihang University (北京航空航天大学)
类目: Computation and Language (cs.CL); Databases (cs.DB)
备注:

点击查看摘要

Abstract:Text-to-SQL translates natural-language questions into executable SQL queries, but open-source large language models still require task-specific post-training for complex, real-world SQL generation. Effective post-training requires both training data that cover the capabilities demanded by the target task and a learning strategy that enables the model to acquire them. Existing datasets provide valuable supervision but incompletely cover SQL structures, while augmentation methods typically expand data without identifying structural gaps. Moreover, supervised fine-tuning (SFT) or reinforcement learning (RL) alone cannot dynamically address weaknesses exposed during training. We propose COAL-SQL, a unified framework combining Coverage-Guided Augmentation (CGA) and Failure-Driven Learning (FDL). CGA uses greedy selection to identify SQL structures missing from the original dataset and constructs complementary examples, improving structural coverage. FDL retains GRPO as the main optimization objective while supplying targeted supervision for unsolved examples. At the step level, it applies SFT to verified reasoning traces generated by a strong LLM for accumulated failures. At the epoch level, it retrieves structurally related examples based on accumulated failures to create targeted practice, helping the model acquire the corresponding SQL capabilities. With only 12,600 distinct post-training examples, COAL-SQL achieves 64.9% execution accuracy on the BIRD development set and outperforms baselines trained at comparable scale. The code is available at this https URL.

[NLP-66] Curriculum-Based Noise Adaptation for Phoneme-to-Text Reconstruction in Visual Speech Recognition

【速读】: 该论文旨在解决音素中心的视觉语音识别(phoneme-centric visual speech recognition)中,音素到文本重构模型在实际推理时因与训练阶段的音素预测误差分布不匹配而导致性能下降的问题。现有方法通常在干净音素序列或合成噪声输入上进行训练,难以模拟真实场景中的复杂预测错误,从而限制了重构模型的鲁棒性。为此,论文提出渐进式错误课程学习(Progressive Error Curriculum Training, PECT),其核心在于通过逐步引入由视觉语音识别器生成的多域伪标签和目标域伪标签,结合合成音素扰动,使基于No Language Left Behind(NLLB)架构的音素到文本重构模型逐步适应越来越真实的音素预测误差。该方法在保持句子重构准确率的同时显著提升了模型对实际语音识别错误的鲁棒性。在LRS2和LRS3基准上的实验表明,PECT在多种前端模型(包括V-ASR、PV-ASR和HP-VSR)上均能稳定提升性能,例如将HP-VSR-FiLMFuse (L4) 在LRS2上的词错误率(WER)从23.3%降至22.2%,将HP-VSR-ResFiLM在LRS3上的WER从30.3%降至29.7%。消融实验与定性分析进一步验证了渐进式适应真实音素错误的有效性,证明PECT是一种高效且可泛化的音素到文本重构课程学习策略。

链接: https://arxiv.org/abs/2609.20839
作者: Matthew Kit Khinn Teng,Haibo Zhang,Takeshi Saitoh
机构: Kyushu Institute of Technology (九州工业大学)
类目: Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD)
备注: Submitted for journal publication and currently under consideration

点击查看摘要

Abstract:Phoneme-centric visual speech recognition reconstructs sentences from intermediate phoneme predictions, making overall recognition performance highly dependent on the robustness of the phoneme-to-text reconstruction model. Existing reconstruction approaches are commonly trained on clean phoneme sequences or synthetically corrupted inputs, leading to a mismatch between training conditions and the realistic phoneme prediction errors encountered during inference. To address this limitation, this paper proposes progressive error curriculum training (PECT). This curriculum learning framework progressively adapts a No Language Left Behind (NLLB)-based phoneme-to-text reconstruction model using synthetic phoneme perturbations, multi-domain pseudo-labels, and target-domain pseudo-labels generated by a visual speech recognizer. By gradually exposing the reconstruction model to increasingly realistic phoneme prediction errors, the proposed framework improves robustness while preserving sentence-reconstruction accuracy. Experiments on the LRS2 and LRS3 benchmarks demonstrate that PECT consistently improves reconstruction performance across multiple phoneme-based visual speech recognition frontends, including visual automatic speech recognition (V-ASR), point visual automatic speech recognition (PV-ASR), and head-pose-aware visual speech recognition (HP-VSR) variants. In particular, PECT reduces the word error rate (WER) of HP-VSR-FiLMFuse (L4) from 23.3% to 22.2% on LRS2 and reduces the WER of HP-VSR-ResFiLM from 30.3% to 29.7% on LRS3. Comprehensive ablation studies and qualitative analyses further demonstrate the effectiveness of progressively adapting the reconstruction model to realistic phoneme prediction errors. These results show that PECT provides an effective and generalizable curriculum learning strategy for phoneme-to-text reconstruction in phoneme-centric visual speech recognition.

[NLP-67] From Generation to Detection: Exploration of Discourse Driven Scenario based LLM Generated Fake News

【速读】: 该论文旨在系统评估大型语言模型(Large Language Models, LLMs)在典型生成场景下生成与检测虚假新闻的能力,尤其关注不同操纵策略对虚假信息可检测性的影响。研究聚焦于四种基于新闻话语框架的操控场景:开放式生成、重写、基于提示的操纵以及基于属性的提示。其解决方案的关键在于构建了一个包含14,000篇生成文章的合成虚假新闻语料库,并通过分析其语言学特征来评估生成内容在结构与语义上与真实新闻的相似程度;进一步地,采用迭代优化的提示工程方法,从真实-虚假新闻对中提取误导性模式以改进检测提示,进而评估各模型在不同提示下的检测性能。研究发现,不同模型在生成与检测能力上存在显著差异,且生成策略对可检测性具有决定性影响;值得注意的是,经优化的提示反而常导致检测性能下降,揭示了当前基于提示的检测方法在面对生成式虚假信息时的局限性。

链接: https://arxiv.org/abs/2609.20838
作者: Zeynep Özdemir,Murat Osmanoğlu,Sevgi Yiğit-Sert,Ömer Özgür Tanrıöver,Yılmaz Ar
机构: 未知
类目: Computation and Language (cs.CL)
备注: 35 pages, 8 figures, 5 tables

点击查看摘要

Abstract:In this study, we examine how modern LLMs generate and detect fake news under controlled settings across four manipulation scenarios. These are open-ended generation, rewriting, manipulation prompts and attribute based prompts grounded in the journalistic discourse framework. Firstly, using seven widely adapted models, we created a synthetic fake news corpus with 14000 generated articles across these four scenarios. Then we analyzed its linguistic properties to assess how closely model-generated news resembles real news structurally and semantically. Finally, to evaluate detection performance, we conducted experiments where each model judges generated fake news, starting with a basic detection prompt and improved prompts developed through an iterative refinement process that extracts misleading patterns from real-fake pairs. Our results revealed substantial variation across models in both generating and detecting misinformation, demonstrated that the generation strategy strongly influences detectability, and show that the refined prompt does not improve and often harms detection performance. Therefore, the study provides a systematic assessment of LLMs detection capability of LLMs generated fake news across typical generation scenarios.

[NLP-68] PhysioBench: A Unified Benchmark for Physiological Signal Question Answering

【速读】: 该论文旨在解决现有生理信号基础模型在面对多样化临床与监测任务时,普遍依赖特定任务的适配、缺乏统一泛化能力的问题。其核心挑战在于如何实现跨模态(如心电、脑电、呼吸等)和跨任务的通用预测能力,而当前模型对自然语言指令的解析与执行能力仍不充分。解决方案的关键在于提出首个统一的生理信号问答基准——PhysioBench,该基准将22个公开数据集的标注信息标准化为6140万条问题-答案对,覆盖30项任务,每一对均基于具体的信号片段并可追溯来源。通过在三种互补设置下评估21种代表性模型(包括大语言模型、视觉-语言模型、时间序列语言模型及生理信号基础模型),研究发现现有模型在不同信号模态与任务间表现不稳定,尽管引入自然语言可促进跨任务的统一预测,但性能仍高度依赖问题表述方式。因此,该工作不仅揭示了当前模型在多任务、跨模态生理信号理解上的局限性,更提供了一个可扩展的平台,推动未来对生理信号理解的细粒度分析与模型优化。

链接: https://arxiv.org/abs/2609.20836
作者: Mengxuan Li,Junfa Chen,Jinze Xia,Yundan Chen,Lixin Fan,Ke Liu,Keyue Shi,Haishuai Wang
机构: Zhejiang University(浙江大学); Zhejiang University of Technology(浙江工业大学)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Physiological signals support diverse clinical and monitoring tasks, yet existing physiological signal foundation models typically require task-specific adaptation for each task. Natural language provides a common interface for specifying different prediction objectives, but the ability of current models to follow such instructions across physiological signal modalities remains insufficiently evaluated. To address this gap, we introduce PhysioBench, a unified benchmark for physiological signal question answering. PhysioBench harmonizes annotations from 22 public datasets into 61.4 million questions across 30 tasks. Each question-answer pair is grounded in a signal segment and traceable to its source annotation. We evaluate 21 representative models, including large language models, vision-language models, time-series language models, and physiological signal foundation models under three complementary settings. The results show that none of the evaluated models achieves consistently strong performance across physiological signal modalities and tasks. The incorporation of natural language supports unified prediction across tasks, although performance remains sensitive to question formulation. Beyond these findings, PhysioBench offers an extensible platform for fine-grained analysis and future research on physiological signal understanding. Our codes are available at this https URL.

[NLP-69] A Generative Grammar Underlying the Voynich Manuscript the Pastiche Hypothesis: Evidence from Large Language Models

【速读】: 该论文旨在解决《伏尼契手稿》这一15世纪神秘文献中未知文字系统的语言性质与内容含义问题,特别是其符号是否构成真实语言、以及图文之间是否存在可解释的对应关系。核心挑战在于如何在缺乏已知语料和解码线索的情况下,通过跨学科方法揭示其潜在的生成机制与语义结构。解决方案的关键在于构建一个整合多模态分析框架:首先基于新转写语料,采用位置依赖的概率语法模型对词级与字母级分布进行建模,揭示符号行为更接近字母而非音节;其次通过语音模式对比分析发现其与希伯来语、阿拉伯语等辅音主导语言具有更高相似性;再者利用大语言模型驱动的图文对齐技术,识别出手稿植物插图与地中海传统伪阿普勒乌斯草药志之间的强对应关系;最后,概率建模重现了类似齐夫定律(Zipf-like)的分布特征,并揭示重复起始字母序列的极低概率,表明其为模仿自然语言的结构性生成系统。整体表明,《伏尼契手稿》并非无意义乱码,而是基于特定语言规则与中世纪医药知识体系构建的复杂仿拟文本,验证了将概率建模与生成式 AI 辅助分析相结合在历史文献研究中的有效性。

链接: https://arxiv.org/abs/2609.20835
作者: Nicolas Turenne
机构: 未知
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Background: The Voynich Manuscript is a fifteenth-century codex written in an unknown script whose content remains undeciphered. Previous studies suggest that its statistical properties resemble those of natural languages, while its illustrations - primarily plants - recall medieval herbals. Methods: We present a multidisciplinary analysis combining probabilistic modeling, phonetic decomposition, rare-event detection, and multimodal image analysis, based on a newly transliterated corpus. Word- and letter-level distributions are modeled using position-dependent probabilistic grammars, while phonetic patterns are compared across Indo-European, Semitic, and Asian languages. Image-text alignment methods based on large language models are applied to identify potential botanical correspondences. Results: The results indicate that Voynich symbols behave as letters rather than syllabic units, while word-length distributions resemble syllabic structures. Phonetic analyses show closer alignment with consonant-heavy languages such as Hebrew or Arabic than with Indo-European languages. Probabilistic modeling reproduces Zipf-like distributions and reveals extremely low probabilities for repeated initial-letter sequences, indicating a structured imitation of natural language. Image analysis suggests strong correspondences between Voynich plant illustrations and those found in Pseudo-Apuleius herbals from the Mediterranean tradition, consistent with an imitation of medieval medicinal books. Perspectives: These findings support the hypothesis that the Voynich Manuscript follows a structured generative system combining linguistic regularities and herbal knowledge, and demonstrate the value of integrating probabilistic and AI-assisted approaches in the analysis of historical manuscripts. Subjects: Computation and Language (cs.CL) Cite as: arXiv:2609.20835 [cs.CL] (or arXiv:2609.20835v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2609.20835 Focus to learn more arXiv-issued DOI via DataCite Submission history From: Nicolas Turenne [view email] [v1] Tue, 28 Jul 2026 16:51:57 UTC (3,859 KB)

[NLP-70] owards Secure Cloud-Native Computing: Unveiling Kubernetes Misconfigurations with Large Language Models

【速读】: 该论文旨在解决云原生环境下Kubernetes集群中存在的配置错误(misconfigurations)问题,这类错误会严重威胁系统的安全性和性能。随着容器化应用在Kubernetes上的广泛应用,其复杂性导致了大量潜在的配置缺陷,而传统检测手段难以高效、准确地识别这些问题。论文的关键解决方案在于引入大语言模型(Large Language Models, LLMs)并结合系统化的分类框架,构建一种基于机器学习的新型误配置检测方法。其核心创新点包括:提出一个全面的常见误配置类型分类体系(taxonomy),通过实证评估现有主流检测工具的表现以进行基准对比,并深入分析最易发生误配置的Kubernetes对象及其风险等级,从而为提升自动化检测能力提供可量化的技术路径和理论支持。

链接: https://arxiv.org/abs/2609.20834
作者: Mostafa Anouar Ghorab,Mohamed Aymen Saied
机构: 未知
类目: Computation and Language (cs.CL); Software Engineering (cs.SE)
备注:

点击查看摘要

Abstract:In the rapidly evolving landscape of cloud-native computing, Organizations are increasingly adopting infrastructure models that emphasize scalability, flexibility, and efficiency. Kubernetes has become the de facto standard for orchestrating containerized applications in these environments. However, the inherent complexity of cloud-native ecosystems introduces significant challenges, particularly in the form of misconfigurations that can compromise both security and performance. This study explores the potential of Large Language Models (LLMs) in identifying Kubernetes misconfigurations. We introduce a comprehensive taxonomy of common misconfiguration types, offering a structured framework to better understand and categorize these issues. Additionally, we conduct an empirical evaluation of state-of-the-art detection tools to benchmark their effectiveness. Furthermore, we analyze the Kubernetes objects most prone to misconfiguration and evaluate the severity of the identified issues. By leveraging advanced machine learning techniques, including LLMs, we provide novel insights into enhancing misconfiguration detection methodologies.

[NLP-71] ranssions Speaker-Attributed Multilingual ASR System for the MLC-SLM 2026 Challenge

【速读】: 该论文旨在解决多语言对话语音中的说话人归属转录(speaker-attributed transcription)问题,即在多说话人、多语言环境下准确识别每个语音片段的说话人身份并生成带说话人标注的转录文本。其核心挑战在于如何在复杂对话场景中实现高精度的说话人分离、跨语言语音识别以及精确的时间对齐与说话人-文本融合。解决方案的关键在于提出一个级联式框架,包含三个核心模块:基于DiariZen的说话人聚类(speaker diarization)模块,通过局部说话人活动检测与全局聚类生成说话人同质段;基于Qwen3-Omni的长序列多语言自动语音识别(ASR)模块,支持多语言转录,并结合外部基于CTC的对齐模型提供词级与字符级时间戳;最后通过说话人-转录融合模块,将说话人标签与带时间戳的转录结果进行精确融合,生成带说话人归属的STM(Speaker-Transcribed Markdown)格式输出。该框架在MLC-SLM 2026挑战赛官方评测集上表现优异,以15.41%的tcpMER(transcription and speaker attribution error rate)成绩位列第二,验证了其在多语言、多说话人场景下的有效性与鲁棒性。

链接: https://arxiv.org/abs/2609.20833
作者: Zhecheng Ren,Xuanji He,Xiaoxiao Li,Zhichen Han,Gaoyang Dong,Gaosheng Zhang,Minchuan Chen,Fengjie Zhu
机构: 未知
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:This paper presents the Transsion Speech Team submission to Task 1 of the MLC-SLM 2026 Challenge, which focuses on speaker-attributed transcription for multilingual conversational speech. We propose a cascaded framework consisting of three components: a speaker diarization module, a long-form multilingual ASR module, and a speaker-transcription fusion module. The diarization module is built upon DiariZen and produces speaker-homogeneous segments through local speaker activity estimation and global speaker clustering. The ASR module is based on Qwen3-Omni and generates multilingual transcriptions, while an external CTC-based alignment model provides precise word- and character-level timestamps. Finally, the fusion module combines diarization outputs with timestamped transcriptions to generate speaker-attributed STM outputs. Experimental results on the official evaluation set demonstrate the effectiveness of the proposed framework. The submitted system achieves a tcpMER of 15.41% and ranks second among all participating teams.

[NLP-72] atBLiMP: A Benchmark of Linguistic Minimal Pairs for Tatar

【速读】: 该论文旨在解决图尔克语族语言塔塔尔语(Tatar, tt)缺乏系统性语法正确性评估基准的问题,尤其针对当前主流多语言模型在塔塔尔语上的表现评估缺失。现有基准如101语言的MultiBLiMP未包含塔塔尔语,而生成式模型对低资源语言的适配仍缺乏可量化、精细化的评测工具。为此,作者提出TatBLiMP——首个面向塔塔尔语的句对最小对(linguistic minimal pairs)基准,涵盖16类形态句法现象,共1248组句子对,每组仅通过单个词素差异区分语法正确与错误。其核心解决方案在于:采用基于Apertium-TAT转换器的确定性单词素扰动生成非语法项,并确保非语法项符合母语者实际可能犯的错误模式(遵循可解释性原则),同时所有句子均经母语者验证,保证真实语料来源。该基准不依赖文本生成或句法解析,仅比较模型对句对的概率输出,适用于基础模型及训练中检查点,实现对塔塔尔语特定训练效果的精准追踪。实验表明,尽管参数规模越大模型表现越强的趋势存在,但真正提升关键在于针对塔塔尔语的专门化训练,而非单纯扩大参数量。然而,该基准当前仍受限于继承自TurBLiMP的分类体系,未能覆盖塔塔尔语中对母语者尤为显著的音系学特征,如元音和谐(vowel harmony)和辅音同化(consonant assimilation),作者建议后续引入以母语者认知为基础的第二层评估框架以弥补此缺陷。

链接: https://arxiv.org/abs/2609.20832
作者: Ilshat Saetov,Dmitry Gaynullin
机构: 未知
类目: Computation and Language (cs.CL)
备注: 11 pages. Dataset: this https URL

点击查看摘要

Abstract:We introduce TatBLiMP, the first benchmark of linguistic minimal pairs for Tatar (tt, ISO 639-3 tat), a Qypchaq Turkic language written in Cyrillic. To our knowledge it is the first grammaticality evaluation for Tatar language models of any kind, since even the 101-language MultiBLiMP does not include Tatar. TatBLiMP covers 16 morphosyntactic phenomena in 1248 sentence pairs. Each pair differs by a single morpheme, one grammatical and one ungrammatical. A model passes a pair when it assigns higher probability to the grammatical member. Scoring compares probabilities the model already assigns, so the benchmark needs no text generation and no parser, and it runs on base models and on mid-training checkpoints. TatBLiMP adapts the phenomenon inventory and single-morpheme breaking operations of TurBLiMP to Tatar and adds one phenomenon specific to Tatar, bare-noun number after numerals and quantifiers. The grammatical member of every pair is an attested sentence from Tatar literary prose. The ungrammatical member is produced by a deterministic single-morpheme perturbation with the apertium-tat transducer. Every pair is ratified by a native speaker. A plausibility principle governs construction, so the ungrammatical member is a plausible real-world error rather than an arbitrary corruption. Across from-scratch Tatar models, cross-lingual adaptations, and frontier multilingual LLMs, the benchmark tracks focused Tatar training rather than parameter scale. A 478M from-scratch model and a 125M monolingual model lead near 0.97, a 7B adaptation trails, frontier LLMs of 30-120B parameters fall to 0.80-0.92, and a lightly tuned multilingual model is weakest. We close with the benchmark’s main limitation. Its inherited taxonomy omits the morphophonology, vowel harmony and consonant assimilation, that is most salient to native speakers, and we sketch a native second layer that would add it.

[NLP-73] Recursive Language Models Generalize Out of Domain

【速读】: 该论文旨在解决大语言模型在推理过程中因依赖训练数据中的特定上下文线索(contextual shortcut)而导致泛化能力下降的问题,尤其是在分布外(out-of-distribution)场景下。其核心问题在于:当标准链式思维(Chain-of-Thought, CoT)能够通过观察完整推理轨迹而“模拟”出递归规则时,它可能倾向于利用与当前子任务无关的外部上下文信息作为捷径进行拟合,这种捷径在分布外情形下失效,从而导致推理失败。为此,论文提出以递归语言模型(recursive language models)为解决方案,其关键在于对每个子任务采用上下文隔离(context isolation)机制,强制模型仅基于局部上下文完成推理,从而排除了依赖全局上下文的捷径学习路径。尽管从理论上讲,CoT 的表达能力仍包含递归规则,但受模型自身“简洁性偏好”(simplicity bias)影响,会优先选择更简单的捷径而非正确的推理规则。因此,真正的推理能力不仅要求覆盖正确的规则形式,还需通过结构设计(如上下文隔离)抑制错误捷径,这挑战了经典学习理论中“覆盖正确假设即能泛化”的基本前提。

链接: https://arxiv.org/abs/2609.20831
作者: Chenxiao Yang,Zhiyuan Li,David McAllester,Nathan Srebro
机构: 未知
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:We study when limiting what a language model can see improves learning. We compare standard CoT, the more general learner that reads the full trace, with recursive language models, which restricts itself by solving each subtask in an isolated context. In-distribution, this generality comes for free: CoT can efficiently simulate the recursive rule, so the IID generalization guarantee changes only by a constant factor, and recursion does not offer much. But out of domain, CoT can fit training by relying on context outside the current subtask, i.e. a shortcut that breaks once those tokens change; recursive context isolation rules out this failure mode. Even though CoT’s class still covers the recursive rule, simplicity bias picks the shortcut over the truth. Thus, to go beyond distributional accuracy and truly reason, covering the right rule is not enough; this contrasts with classical learning theory.

[NLP-74] Reviser: Revision-Capable Text Generation via Autoregressive Cursor Actions

【速读】: 该论文旨在解决生成式模型在文本生成过程中缺乏灵活编辑能力的问题,尤其是如何实现真正非单调(non-monotonic)的生成机制,以支持在生成过程中插入或修改已有内容。传统非自回归或基于编辑的方法虽具备修订能力,但通常依赖多次序列级计算,导致效率低下。本文提出Reviser,一种仅包含解码器结构的Transformer模型,其通过在可变画布(mutable canvas)上执行相对于光标位置的动作序列来生成响应。每个生成步骤仅预测一个动作标记:INSERT(token)、MOVE(Δ) 或 STOP,且模型在编辑历史动作序列上进行自回归建模,而非最终文本顺序。这一设计的关键在于将生成过程解耦为一系列可逆的编辑动作,从而实现真正的非单调生成,同时保持简单直观的下一步动作接口。实验结果表明,Reviser在续写任务中显著优于SEDD和MDLM,在对弈评估中更受青睐;轨迹分析显示其频繁执行回退移动与画布中段插入操作,而非简单模拟尾部追加的自回归模式。在相同规模下,其性能可与自回归基线模型相媲美,并在共享浮点运算量(FLOPs)标准下,显著降低推理计算开销,优于多轮精炼及扩散类方法。

链接: https://arxiv.org/abs/2609.20830
作者: Sean Diab
机构: 未知
类目: Computation and Language (cs.CL)
备注: 45 pages, 2 figures. Code: this https URL Checkpoints: this https URL

点击查看摘要

Abstract:Revision-capable generation is appealing because it can insert or revise earlier content, but many non-autoregressive and edit-based approaches obtain this flexibility through repeated sequence-level computation. We propose Reviser, a decoder-only Transformer that generates a response as a sequence of cursor-relative actions on a mutable canvas. At each step, Reviser predicts exactly one action token: INSERT(token), MOVE( \Delta ), or STOP, and is autoregressive over edit-history actions rather than final text order. This design enables genuinely non-monotonic generation while preserving a simple next-action interface. On a continuation benchmark, Reviser is strongly preferred to SEDD and MDLM in our arena evaluations, and trajectory statistics confirm that the model performs frequent backward moves and mid-canvas insertions rather than merely emulating end-append decoding. Against size-matched autoregressive baselines, Reviser is competitive at both the 100M and 300M scales. Under our shared FLOPs convention, Reviser also requires substantially less inference compute than representative multi-pass refinement and diffusion-style baselines.

[NLP-75] SAGE: Schema-Guided LLM s for Grant Review

【速读】: 该论文旨在解决科研资助申请评审过程中评审标准主观性强、评估过程缺乏透明度与可审计性的问题。传统评审依赖评审人对申请材料的定性判断,难以确保一致性与可复现性,且评审意见往往缺乏明确的证据支撑。为此,论文提出SAGE(Schema-Guided Aspect-Based Grant Evaluation)系统,其核心解决方案是将资助评审标准(grant rubric)转化为结构化的评估条目,并将每项评判结果与申请材料中的具体证据进行显式关联。通过这一机制,SAGE生成可追溯、可审计的结构化评审草案,提升了评审过程的透明度与可验证性。实验表明,在辅助评审阶段,SAGE在准则层面达到kappa = 0.58的一致性,显著优于基线方法(kappa = 0.33),并展现出更高的等级相关性和更低的错误率。此外,基于条款级别的审计功能可识别出已确认、存在争议及未覆盖的评审内容,进一步增强了评审过程的可控性与可靠性。因此,SAGE的关键创新在于实现了评审流程的结构化、证据驱动化和可审计化,为提升资助评审的客观性与效率提供了可落地的技术范式。

链接: https://arxiv.org/abs/2609.20829
作者: Erik Varapaev,Andrei Chetvergov,Stepan Ukolov,Timofei Sivoraksha,Alexander Evseev,Sergey Bolovtsov
机构: 未知
类目: Computation and Language (cs.CL)
备注: 14 pages, 2 figures, 10 tables

点击查看摘要

Abstract:Grant reviewers must apply detailed criteria to application forms, budgets, and supporting documents while producing assessments that colleagues can inspect. We present SAGE, Schema-Guided Aspect-Based Grant Evaluation, a system that translates a grant rubric into structured checks and links its judgements to evidence from the application package. We evaluate SAGE in two stages on 35 nonprofit grant applications. A post-factum comparison with 105 reviews from the original competition shows fair ordinal agreement (kappa = 0.29). The foundation then conducted a criterion-level re-review after inspecting SAGE, producing 202 assessments. In this assisted round, SAGE reached kappa = 0.58 and outperformed a one-prompt-per-criterion baseline (kappa = 0.33 on the common subset), with higher rank correlation and lower error. A claim-level audit further identifies confirmed, disputed, and unaddressed parts of the structured draft. SAGE operationalizes the review methodology by producing a detailed, evidence-linked, and auditable draft for expert correction.

[NLP-76] Beyond WER: Entity and Disfluency Recall in Accented Conversational ASR INTERSPEECH2026

【速读】: 该论文旨在解决自动语音识别(ASR)系统在优化词错误率(WER)时,对带有口音的对话英语中命名实体和填充停顿(filled pauses)识别能力不足的问题,这对语言学习反馈至关重要。其解决方案的关键在于提出一个三阶段流水线:首先,通过启发式SQL过滤器从训练数据中筛选出高实体密度的数据,使实体密度达到随机采样的2.8倍;其次,基于Qwen2.5-Omni-3B模型微调区域化低秩适配器(LoRA),实现单次前向传播同时输出原始转录与修正转录;最后,构建一个六类错误分类体系,并通过大语言模型(LLM)裁判验证,达成83.8%的一致性(基于210个手工标注样本)。实验结果表明,该方法在6,000条测试语句上实现了80–85%的实体召回率(较原有53–55%显著提升)、76–86%的填充停顿召回率(从5%大幅提升),同时保持6–10%的词错误率,优于Whisper及商用ASR系统在实体召回上的表现,并以10倍更少参数量媲美零样本30B模型。配对自举检验进一步证实,仅数据筛选环节即可带来2.8–4.2个百分点的实体召回率提升(p < 0.0001)。

链接: https://arxiv.org/abs/2609.20828
作者: Fiza Husain,Ankit Pandey,Yash Singh
机构: Stimuler(斯蒂姆勒); Bengaluru(班加罗尔); India(印度)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 5 pages, 1 figure, 1 table, accepted at Interspeech 2026

点击查看摘要

Abstract:ASR systems optimised for Word Error Rate (WER) often miss named entities and filled pauses in accented conversational English, both critical for language-learning feedback. We present a three-stage pipeline for speakers from India, Indonesia, and Latin America: (1) heuristic SQL filters curating entity-rich training data at 2.8x the entity density of random sampling, (2) regional LoRA adapters fine-tuned on Qwen2.5-Omni-3B producing both verbatim and corrected transcripts in a single forward pass, and (3) a six-category error taxonomy validated by an LLM-based judge (83.8% agreement, 210 human-labelled samples). The pipeline achieves 80-85% entity recall (up from 53-55%), 76-86% filler recall (up from 5%), and 6-10% WER across 6k test utterances, outperforming Whisper and a commercial ASR on entity recall while matching a zero-shot 30B model with 10x fewer parameters. Paired bootstrap tests confirm that curation alone accounts for 2.8-4.2 pp of entity recall gain (p0.0001).

[NLP-77] From Discharge Notes to Patient Understanding: Persona-Grounded Open-Ended Simulation of LLM s as Discharge Educators

【速读】: 该论文旨在解决现有大语言模型(LLM)评估体系在医院出院教育场景中无法有效衡量患者理解程度的问题。当前的评估方法多聚焦于静态文本生成或特定格式的任务,难以捕捉真实对话情境下患者对医疗信息的理解与反馈。为此,研究提出DischargeBench——一个基于角色设定(persona-grounded)的模拟系统,通过让候选LLM教育者与虚拟患者(Virtual Patient)进行多轮互动,由教育监控代理(Education Monitor Agent)动态调控患者行为的真实性,从而在不修改教育者模型的前提下保持评估信号的可靠性。研究构建了基于MIMIC-IV-Ext数据集的DischargeBench数据集,涵盖24个ICD章节、477例病例,并引入人格、教育水平、健康素养及既往病史回忆等多维度角色轴,支持分层分析。每个模拟会话依据对话质量、主题覆盖、理解度和事实一致性四个维度,由经医师标注对齐的“以LLM为裁判”(LLM-as-a-Judge)系统评分。实验结果表明,尽管整体平均分相近,但不同疾病章节与患者角色间存在显著临床差异,高难度角色暴露了模型在内容覆盖、理解能力及答案一致性方面的系统性缺陷。因此,该研究强调:针对出院教育的LLM评估应以患者理解为核心,而非仅依赖文本质量或答案准确性。

链接: https://arxiv.org/abs/2609.20827
作者: Won Seok Jang,Zonghai Yao,Hong Yu
机构: University of Massachusetts Lowell(马萨诸塞大学洛厄尔分校); University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Hospital discharge education is an interactive teaching task: a clinician adapts a discharge plan to a patient’s literacy, recall, and personality. Existing LLM evaluations target static or artifact-generation tasks and do not measure patient understanding under open-ended dialogue. We introduce DischargeBench, a persona-grounded simulation in which a candidate LLM educator conducts a multi-turn session with a Virtual Patient, while an Education Monitor Agent regulates patient realism without modifying the educator, protecting the evaluation signal. We curate MIMIC-IV-Ext-DischargeBench, 477 cases over 24 ICD chapters with persona axes (personality, education level, health literacy, past-medical-history recall) for stratified analysis. Each simulation is scored on four axes – Conversation Quality, Topic Checklist, Comprehension, and Factual Consistency – by an LLM-as-a-Judge aligned against physician annotations. Across closed- and open-source LLMs, aggregate scores conceal clinically relevant variation across ICD chapters and patient personas; difficult personas expose coverage failures, comprehension gaps, and reduced source-answer agreement. LLM evaluation for discharge education should center patient understanding, not text quality or answer accuracy alone.

[NLP-78] ALON: A Temporally Aware Longitudinal Framework for Radiology Report Generation

【速读】: 该论文旨在解决当前放射科报告生成(Radiology Report Generation, RRG)模型在纵向分析能力上的局限性,即多数模型仅基于单次检查或最近一次既往检查生成报告,难以实现精准且有意义的纵向对比与细微时间间隔变化的检测。其核心解决方案在于提出一种时序感知的纵向报告生成框架TALON(Temporally Aware LONgitudinal RRG),其关键创新在于双通道时序融合模块(Dual-Channel Temporal Fusion Module, DCTFM)。DCTFM通过互补的相似性通道与变化通道,分别捕捉持续存在的病灶特征与时间间隔内的变化信息,并结合通道特异性注意力机制动态评估各既往检查的相关性,同时利用可学习的先验特定门控机制自适应融合有价值的纵向证据并抑制冗余信息。该设计使TALON能够有效处理变长患者历史数据,在MIMIC-CXR数据集上显著优于现有最先进方法,且随着可用既往检查数量增加,性能持续提升,验证了其在更长、更复杂的患者病史建模中的优越性。

链接: https://arxiv.org/abs/2609.20826
作者: Nien-Tsyr Sun,Min-Chen Chen,Hui Nien Hung,Vincent S. Tseng
机构: National Taiwan University(国立台湾大学); National Cheng Kung University(成功大学); National Chiao Tung University(交通大学); Academia Sinica(中央研究院)
类目: Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Current radiology report generation (RRG) models usually produce descriptive reports based on a single examination or only the most recent prior examination, limiting their ability to perform accurate and meaningful longitudinal comparisons and detect subtle interval changes. Although recent approaches have begun to incorporate multiple prior examinations, they usually aggregate a fixed-length history without explicitly modeling the role-dependent relevance of each prior examination before fusion. To address this, we propose TALON, a Temporally Aware LONgitudinal RRG framework that adaptively integrates variable-length patient histories. The underlying Dual-Channel Temporal Fusion Module (DCTFM) compares the current examination with each prior examination through complementary similarity and change channels to capture persistent findings and interval changes, respectively. The specially designed channel-specific attention estimates the relevance of each prior examination, while a learned prior-specific gate adaptively integrates informative longitudinal evidence and suppresses redundancy. Experiments on MIMIC-CXR show that TALON outperforms the current state-of-the-art method on various clinical efficacy and graph-based metrics. When more prior examinations become available, TALON’s performance on these metrics improves even further, emphasizing the strength of TALON’s DCTFM in modeling longitudinal RRG across longer and more complex patient histories than existing approaches.

[NLP-79] HERMES: Contrast-Aware Knowledge Graph Reasoning from Clinical Notes for Patient Outcome Prediction

【速读】: 该论文旨在解决临床预测模型在利用非结构化临床文本时,因将文本编码为扁平序列而丢失临床叙述中隐含的语义关系与时间动态性的问题。其核心解决方案是提出一种基于图结构的框架HERMES,关键在于通过大语言模型(Large Language Model, LLM)引导从临床笔记中提取信息,并结合对比逻辑建模(Contrastive Logic Modeling)构建个性化知识图谱(Personalized Knowledge Graph, KG),显式捕捉治疗失败、病情变化等时间动态特征;随后采用图注意力网络(Graph Attention Network)在知识图谱上进行图学习,以生成更具表达力的患者表征。实验结果表明,该方法在MIMIC-III和MIMIC-IV数据集上的院内死亡率与30天再入院预测任务中均显著优于现有纯文本基线模型,验证了显式关系建模与对比逻辑建模对提升预测性能的关键作用。

链接: https://arxiv.org/abs/2609.20825
作者: Gia-Bach Nguyen,Hoang-Ha Nguyen,Tuan-Cuong Vuong,Trang Mai Xuan,Duy Quoc Ngo,Tien-Cuong Nguyen,Huan Vu,Thien Van Luong
机构: 未知
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 12 pages, 4 figures, The 15th Conference on Information Technology and its Applications

点击查看摘要

Abstract:Clinical predictive models often rely on structured Electronic Health Record data, such as time-series and procedure codes. While recent approaches have begun leveraging unstructured clinical notes, they typically encode them as flat sequences, which may lose explicit relational and temporal structure present in clinical narratives. In response, we propose HERMES, a graph-based framework that operates exclusively on clinical text while preserving clinical relationships. This approach builds on two key ideas. First, personalized Knowledge Graphs (KGs) are constructed through Large-Language-Model-guided extraction from clinical notes with Contrastive Logic Modeling that explicitly captures temporal dynamics and treatment failures and changes in outcomes. Second, a Graph Attention Network synthesizes patient representations through graph-based learning over the KGs. Experiments on MIMIC-III and MIMIC-IV for in-hospital mortality and 30-day readmission prediction show that HERMES consistently outperforms strong text-only baselines. Our findings demonstrate that explicit relational modeling with Contrastive Logic Modeling significantly advances predictive performance.

[NLP-80] Do small language models know what they dont know?

【速读】: 该论文旨在解决小型语言模型(Small Language Models, SLMs)在参数量少于30亿且仅运行于消费级硬件上的情况下,如何提升其准确率的问题。其核心挑战在于,传统基于熵的置信度信号在小模型中失效——研究发现,在91%的数据集-模型组合中,词元级熵(token-level entropy)均接近零,无法有效区分正确与错误预测,因而无法作为可靠的置信度指标。解决方案的关键在于引入语义熵(semantic entropy),通过生成多个回答样本、基于语义聚类并测量分布不确定性来重建有效的置信信号。利用该语义熵实现对不确定查询的智能路由,将其传递至更大的专家模型(expert model),可带来最高达+50个百分点的准确率提升。尤为关键的是,跨架构路由(如SmolLM 360M路由至Phi-3.5-mini)平均带来+22.0%的性能增益,显著优于同架构路由(+6.8%),表明专家模型的质量比架构兼容性更为重要。因此,该研究揭示:基于熵的方法在SLMs中的价值并非降低计算开销,而在于实现智能算力分配,即在关键位置投入更多计算资源以最大化性能收益。

链接: https://arxiv.org/abs/2609.20824
作者: Prashant Mudgal
机构: Nagarro(纳格罗)
类目: Computation and Language (cs.CL)
备注: 9 pages, 8 figures

点击查看摘要

Abstract:We explore whether entropy-based confidence signals can be leveraged to improve the accuracy of Small Language Models (SLMs) with fewer than 3 billion parameters, running entirely on consumer hardware. We evaluate seven distinct approaches, including token-level entropy early stopping, semantic entropy estimation, and uncertainty-aware routing to larger expert models, across 7 model pairs and 5 standard NLU benchmarks. Our key finding is that token-level entropy is effectively blind in SLMs: in 91% of dataset-model combinations, mean token entropy is near zero regardless of answer correctness, rendering token-based confidence signals unusable at this scale. We demonstrate that semantic entropy, computed by generating multiple samples, clustering answers by meaning, and measuring distributional uncertainty, recovers a viable confidence signal. Using semantic entropy to selectively route uncertain queries to a larger expert model yields accuracy improvements of up to +50 percentage points. Notably, cross-family routing (e.g., SmolLM 360M to Phi-3.5-mini) averages +22.0% improvement compared to +6.8% for same-family routing, revealing that expert model quality matters more than architectural compatibility. Our results suggest that the value proposition for entropy-based methods in SLMs is not computational savings but intelligent compute allocation: spending more tokens where they matter most.

[NLP-81] he Spoken Wikipedia Presentation Corpus

【速读】: 该论文旨在解决多模态自动语音识别(ASR)中缺乏高质量、结构化对齐数据的问题,尤其针对语音与视觉内容协同提升识别性能的挑战。其核心解决方案是构建一个新型语料库——“口语维基百科演示语料库”(Spoken Wikipedia Presentation Corpus),通过大语言模型(LLM)生成的幻灯片内容实现音频、文本与视觉信息的精准对齐。关键在于采用混合式流水线:首先利用LLM对维基百科内容进行分段并生成标题、要点、核心观点及视觉描述;随后通过基于规则的匹配策略选择布局、主题与风格以完成幻灯片设计;最终由视觉大语言模型(Vision LLM)提取幻灯片中的文本为Markdown格式。该多模态对齐数据显著提升了跨模态上下文的利用能力,实验表明在仅使用音频输入时,最优模型达到平均微平均词错误率(micro-WER)10.23%和字符错误率(CER)6.48%,且英语表现最佳,低资源语言性能下降明显,验证了跨模态上下文在提升语音识别中的潜力。

链接: https://arxiv.org/abs/2609.21676
作者: Thomas Ranzenberger,Steffen Freisinger,Tobias Bocklet,Korbinian Riedhammer
机构: Technische Hochschule Nürnberg(纽伦堡应用技术大学)
类目: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Multimedia (cs.MM); Sound (cs.SD)
备注: Accepted at SLT 2026

点击查看摘要

Abstract:We present the Spoken Wikipedia Presentation Corpus, an extension of the Spoken Wikipedia Corpora featuring LLM-generated slide decks for multimodal ASR. Slides are created from LLM-segmented sections using a hybrid pipeline that combines LLM-based content planning with rule-based design decisions. For each section, an LLM generates a slide title, bullet points, a takeaway message, and a visual description that is used to create an illustration. Rule-based matching then selects layouts, themes, and styles to produce the final slides. A vision LLM extracts slide text as Markdown. We evaluate multiple ASR and spoken language models (SLMs). The best model achieves an average micro-WER of 10.23% and an average micro-CER of 6.48% on audio-only inputs. English yields the lowest error rates, followed by German and Dutch, while performance declines across lower-resource languages. Although audio-only baselines are strong, multimodal zero-shot prompting of omni models remains challenging. The aligned slide, text, and audio data show a strong potential to improve recognition through cross-modal context.

[NLP-82] he Hidden Cost of Digits: Number Normalization and WER in ASR Systems ICASSP2027

【速读】: 该论文旨在解决多语言自动语音识别(ASR)系统评估中因数字表达未进行适当文本归一化而导致的不公平比较问题。现有主流方法通常仅将文本转为小写并移除标点,对非英语语言缺乏针对性的归一化处理,尤其忽视了数字表达的标准化。研究以波兰语这一高度屈折的语言为例,分析了不同文本归一化策略对ASR评估性能的影响,基于VoxPopuli和波兰议会演讲数据集进行实验,量化了词错误率(WER)在不同归一化方式下的差异。研究表明,若不进行数字归一化,导致的WER偏差可能超过2个百分点,甚至高于主流多语言基准中不同模型间的性能差距,凸显了在多语言ASR评估中引入数字归一化的必要性与关键作用。

链接: https://arxiv.org/abs/2609.21084
作者: Stanisław Kacprzak,Mieszko Fraś
机构: 未知
类目: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL)
备注: Submitted to ICASSP 2027

点击查看摘要

Abstract:Modern automatic speech recognition (ASR) systems trained on extremely large datasets can produce transcripts with numbers written in Arabic numerals. This creates a need for fair comparison with models that output verbatim texts and proper processing of reference transcripts. Popular approaches often reduce text normalization to lowercase and remove punctuation, with no additional normalization applied to languages other than English. In this work, we analyze the impact of normalization of numerical expressions in the evaluation of ASR systems in various languages, using Polish as an example of a highly inflective language. We perform experiments on VoxPopuli and The Polish Parliamentary speech datasets and estimate word error rate (WER) differences for different text normalization approaches. We show that the difference due to the lack of number normalization in WER may be substantial - more than 2 percentage points, and often higher than the differences between systems in popular multilingual benchmarks.

[NLP-83] Cross-Lingual Parkinsons Disease Severity Assessment Using Pre-trained Speech Embeddings: A Multi-Class Evaluation

【速读】: 该论文旨在解决帕金森病(Parkinson’s disease, PD)严重程度多分类评估中跨语言泛化能力不足的问题,尤其针对标注数据稀缺、缺乏可解释性方法以及不同语言和数据集间差异显著等挑战。其核心解决方案在于评估四种先进的开源语音基础模型(Speech Foundation Models, SFMs)在零样本(zero-shot)与少样本(k-shot)跨语言设置下,利用预训练语音嵌入进行多类别PD严重程度分类的性能。研究发现,尽管预训练语音嵌入能够实现有意义的跨语言迁移,但模型表现对数据集特性、预处理方式及适配策略高度敏感;误分类现象主要源于说话人之间的个体差异及非典型语音模式,凸显了在PD严重程度评估中构建更鲁棒的特征提取与建模机制的重要性,同时强调了可解释性对于提供可靠临床洞察的关键作用。

链接: https://arxiv.org/abs/2609.20875
作者: Simon Pals,Cristian Tejedor-Garcia
机构: Radboud University (奈梅亨大学)
类目: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
备注: Accepted and published at IEEE SLT 2026 - IEEE Spoken Language Technology 2026. OneVoice-MSD 2026: Multilingual Speech Technologies for Motor Speech Disorders. this https URL Please cite the conference version

点击查看摘要

Abstract:Parkinson’s disease (PD) often manifests through speech impairments, facilitating accessible, non-invasive, and cost-effective severity assessment for early diagnosis and progression tracking. Despite advances in speech foundation models (SFMs), their cross-lingual generalization for PD severity multi-class classification remains underexplored due to limited labeled data, a lack of explainable methods and variability across languages and datasets. In this work, we evaluate pre-trained embeddings from four state-of-the-art open-source SFMs across three datasets in zero-shot and k-shot cross-lingual settings for multi-class PD severity assessment. Our results show that pre-trained speech embeddings enable meaningful cross-lingual transfer, although performance is sensitive to dataset properties, preprocessing, and adaptation strategy. Misclassifications under these conditions related to inter-speaker variability and atypical speech patterns highlight the need for more robust feature extraction and modeling for PD severity assessment while emphasizing the importance of explainability for reliable clinical insights.

信息检索

[IR-0] Predictable Failure in Multi-Hop Retrieval: Score-Distributional Confidence Scoring and Abstention

链接: https://arxiv.org/abs/2609.22056
作者: Andre Bacellar
类目: Information Retrieval (cs.IR); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 8 pages, 2 figures, 4 tables

点击查看摘要

Abstract:Multi-hop retrieval failures are not uniformly distributed across queries: they cluster in structurally predictable subpopulations. We prove two results formalizing this structure. First (CWAR Reducibility): confident-failure reduction is achievable if and only if retrieval features carry mutual information about success, a condition satisfied by LLM-judge pipelines but substantially weaker in dense-only settings, explaining the AUC-AC gap between regimes. Second (Feature Regime Complementarity): no single ANN score feature achieves best predictive performance across all failure regimes; the dominant feature differs between datasets (query length on MuSiQue, hop-1 concentration on HoVer), and a constructive witness pair shows each is necessary in one regime and non-contributory in the other. We instantiate these principles in RegimeAbstain, which computes a Retrieval Confidence Score (RCS), a logistic function of up to nine query-ANN structural features, all available without any additional LLM call, and uses it to implement a calibrated abstention policy. We define the Confident-Wrong-Answer Rate (CWAR) metric and evaluate across three multi-hop benchmarks (MuSiQue, 2WikiMultiHopQA, HoVer) and two retrieval architectures (LLM-judge and dense-only), covering five failure regimes with CWAR from 14.5% to 62.1%. RCS achieves best or co-best AUC-AC in all five conditions against eight confidence baselines. On MuSiQue (LLM-judge), RCS reduces CWAR from 39.5% to 20.6% at 50% coverage (47.8% relative reduction), with ECE=0.035. A model trained on MuSiQue transfers to 2WikiMultiHopQA with only -0.5pp AUC loss, confirming the domain-agnostic structure of regime features.

[IR-1] AutoRecLab: Describe the Experiment Get the Code! RECSYS’26

链接: https://arxiv.org/abs/2609.21863
作者: Moritz Baumgart,Philipp Meister,Justus Krell,Michael Schmidt,Bela Gipp,Joeran Beel
类目: Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注: Accepted at the 20th ACM Conference on Recommender Systems (RecSys '26), Demo Track. 4 pages, 2 figures

点击查看摘要

Abstract:Empirical evaluation is central to recommender-systems (RecSys) research, but turning experimental designs into executable code remains a manual and error-prone task. We present AutoRecLab, a Python-based autonomous RecSys lab that automates RecSys experiments from natural-language prompts. Given a research idea, AutoRecLab derives explicit experiment requirements, builds and validates a prototype, and iteratively expands it into the requested full experiment. The workflow combines retrieval-augmented generation (RAG) for documentation lookup, static type verification, and execution-steered tree search. In our demonstration, AutoRecLab autonomously implements an explicit-to-implicit feedback conversion study. In a baseline comparison across six algorithms and three datasets, 8 of 9 runs succeed at an average cost of approx- imately 1 per run with GPT-5.4-mini.

[IR-2] Do We Care About Personalization and Explainability? An Interview Study with News Recommendation Engineers RECSYS2026

链接: https://arxiv.org/abs/2609.21547
作者: Jasmin Kareem,Siddharth Mehrotra,Martijn C. Willemsen,Maarten de Rijke
类目: Human-Computer Interaction (cs.HC); Information Retrieval (cs.IR)
备注: 10 pages, Accepted at ACM RecSys 2026 Main Track

点击查看摘要

Abstract:Research on explainability in recommender systems largely centers on end users, overlooking the perspectives of those who build and maintain these systems and their potential use cases such as model debugging. In this study, we examine how news engineers and related technical stakeholders perceive and implement personalization and explainability in practice. We conducted 15 semi-structured interviews across nine news organizations, spanning diverse regions in both public and private sectors, to investigate the challenges and motivations shaping their approaches. Our findings reveal that personalization is not always a straightforward or desirable choice for news organizations, as concerns around user tracking, editorial control, and resource constraints often limit its adoption. Even among organizations implementing personalized news recommender systems in production, explainability is rarely prioritized, with day-to-day operational demands frequently taking precedence over longer-term transparency goals. Definitions of explainability vary widely across organizations, though some demonstrate promising internal practices and visualization tools that facilitate communication between engineering teams and newsrooms. Based on our analysis, we provide actionable and practical guidelines for news engineers and researchers on how to adopt explainability methods within a news personalization pipeline.

[IR-3] Adaptive Preference Modeling via Explicit Indirect Relational Learning for Personalized Fashion Matching

链接: https://arxiv.org/abs/2609.21475
作者: Shuiying Liao,Li Li,P. Y. Mok
类目: Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:Personalized fashion complementary recommendation requires jointly modeling user preferences and item compatibility under sparse and multimodal data conditions. Existing approaches often capture higher-order relational signals implicitly through graph propagation or rely on direct interaction data, limiting their ability to explicitly model indirect preference and compatibility relationships. To address this limitation, we propose an Adaptive Preference with Contrastive Learning framework (APCL) that explicitly models both direct and indirect relational signals within a unified recommendation architecture. Specifically, APCL constructs indirect user-item and item-item relationships through a correlation-guided adaptive aggregation mechanism and represents them as dedicated personalization and compatibility views. To improve representation learning, we further introduce a functional view contrastive learning strategy that aligns direct and indirect preference representations and direct and indirect compatibility representations, encouraging consistency across relational contexts. By integrating multimodal visual and textual information with explicit indirect relational modeling, APCL captures richer semantic characteristics while improving robustness in sparse-interaction settings. Experiments on two benchmark fashion recommendation datasets demonstrate that APCL consistently outperforms representative baseline methods.

[IR-4] Auto-Bidding with Disentangled Advertiser Profiles and Train-Free Adaptation

链接: https://arxiv.org/abs/2609.21308
作者: Songyue Cai,Shan Gu,Wei Chen,Ziru Xu,Lianyu Wang,Jian Xu,Xiaofeng Zhu
类目: Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:Auto-bidding is a key component of modern advertising systems that provides a personalized bidding strategy for each advertiser. By characterizing each individual, profile-based methods achieve personalization and have proven effective in domains such as recommendation. However, despite the diverse bidding behavior of advertisers, their application to auto-bidding remains limited. A primary reason is that constructing and leveraging advertiser profiles face several challenges: extracting pure profiles is non-trivial, modeling common and private information simultaneously is difficult, and profile updating and cold-start adaptation remain challenging. To tackle these issues, we propose \textbfADAPT, an \underline\textbfAuto-bidding framework with \underline\textbfDisentangled \underline\textbfAdvertiser \underline\textbfProfiles and \underline\textbfTraining-free adaptation. ADAPT introduces a two-stage training paradigm and supports training-free adaptation. Specifically, (i) the stage 1 extracts pure static and dynamic profiles via contrastive learning over the advertiser memory bank; (ii) the stage 2 disentangles the dynamic profile into a common profile and a private profile, and combines them with the static profile to jointly condition the bidding strategy; (iii) once trained, ADAPT constructs profiles for new advertisers and updates profiles of existing advertisers without retraining. Our experiments on a large-scale auto-bidding benchmark demonstrate that ADAPT consistently achieves superior performance, and ablation studies further validate the effectiveness of each module. The source code will be released at this https URL.

[IR-5] Hybrid GPU-CPU Retrieval for Personalized Search at Ultra-Large Scale KDD2027

链接: https://arxiv.org/abs/2609.21281
作者: Hao Fu,Jichao Sun,Baiting Zhu,Qiaoling Liu,Yan Shi,Cheng Lu,Liu Liu,Yubo Wang,Xin Yao,Xiangyu Niu,Xu Dong,Wenhan Lyu,Chiyao Shen,Yinjie Huang,Minglei Chen,Shuai Ding,Li Fan,Xiao Kong
类目: Information Retrieval (cs.IR); Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG); Performance (cs.PF)
备注: 10 pages, 5 figures, 9 tables. ACM sigconf format; submitted to the KDD 2027 Applied Data Science Track

点击查看摘要

Abstract:Embedding-based retrieval on user-generated content at the trillion-document scale exposes a sharp conflict between two production demands: deep, expressive personalization for queries with rich user intent, and broad coverage of a massive inventory under fixed latency and resource budgets. We characterize this as the personalization-scale paradox: hosting the full serving inventory in GPU memory is too resource intensive, while CPU compute cannot execute the same interaction-heavy model on the latency-critical path. We present a hybrid GPU-CPU co-serving system that resolves the paradox through orchestration rather than a new model class. A high-depth GPU pathway fuses retrieval and interaction pre-ranking over a curated online pool on the order of a billion documents, while a high-breadth CPU pathway searches an independently selected online inventory roughly twenty times larger with lightweight personalized scoring. Either or both pathways can run per request; candidates are deduplicated before shared downstream ranking. The system is deployed in production. A full-system A/B test against the legacy CPU-only configuration improves model-scored relevance and substantive engagement, while separate pathway experiments show positive value at their own deployment scopes. Retrieval logs show that the pathways contribute structurally distinct candidates, production serving measurements characterize their latency, and a matched capacity plan quantifies the economic rationale for assigning modeling depth to GPUs and inventory breadth to CPUs. Together, these results validate a practical, independently evolvable depth-breadth architecture for ultra-large-scale personalized search. Comments: 10 pages, 5 figures, 9 tables. ACM sigconf format; submitted to the KDD 2027 Applied Data Science Track Subjects: Information Retrieval (cs.IR); Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG); Performance (cs.PF) Cite as: arXiv:2609.21281 [cs.IR] (or arXiv:2609.21281v1 [cs.IR] for this version) https://doi.org/10.48550/arXiv.2609.21281 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[IR-6] Verify Dont Trust: Agent ic Model Development for Video Discovery Retrieval at Scale KDD2027

链接: https://arxiv.org/abs/2609.21257
作者: Hao Fu,Baiting Zhu,Minglei Chen,Yinjie Huang,Shuai Ding
类目: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 9 pages, 1 figure, 8 tables. ACM sigconf format; submitted to the KDD 2027 Applied Data Science Track

点击查看摘要

Abstract:Large language model (LLM) agents can propose, implement, and evaluate model changes. Autoresearch loops demonstrate this capability through minutes-scale iterations on a self-contained program. Online autoresearch instead spans asynchronous systems, hours-long variants, and weeks-long campaigns that can influence a product. A completed run can still support an invalid conclusion when a code change is a no-op, data windows leak, evaluator semantics drift, or the two arms traverse different serving funnels. We present EvoPilot, a human-gated method for long-horizon online autoresearch. Role-specific agents execute each round through a versioned domain skill and typed adapter. Durable records preserve experiments and failures; deterministic checks enforce recorded lessons. We study a 37-day campaign for the retrieval system that powers Video Deep Dive (VDD), an online experience for discovering follow-on videos after a user opens a seed video. The campaign covered seven directions and used an hourly refreshed index of hundreds of millions of videos. Earlier manual experiments had not established a benefit from an interaction head. A primitive autoresearch attempt revisited the direction but incorrectly attributed an offline hit-rate decline of 22 percentage points to the head. We then introduced EvoPilot. Its human-gated verification traced the drop to a pre-existing evaluation defect that produced output depths of 3,000 and 600. After repair, a matched comparison measured an offline improvement of 3.20 percentage points. Post-study replay and mutation tests rejected invalid comparisons while admitting valid counterparts. Durable state recovered an interrupted round, and artifact reuse avoided approximately five GPU-hours. Separately, a seven-day randomized online evaluation estimated a 0.66% relative increase in the VDD slice of Good Search Result Rate for Retention (GSRR). Comments: 9 pages, 1 figure, 8 tables. ACM sigconf format; submitted to the KDD 2027 Applied Data Science Track Subjects: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) Cite as: arXiv:2609.21257 [cs.IR] (or arXiv:2609.21257v1 [cs.IR] for this version) https://doi.org/10.48550/arXiv.2609.21257 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[IR-7] MAGIC: Marginal-Guided Compression with Optimal Transport for Efficient Visual Document Retrieval

链接: https://arxiv.org/abs/2609.21018
作者: Xu Yuan,Hua Liu,Wenqi Fan,Qing Li
类目: Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:Recent visual document retrieval (VDR) systems such as ColPali use multi-vector page embeddings, in which patch-level vectors enable fine-grained evidence matching but incur substantial index storage and MaxSim scoring overhead. Post-hoc merging offers a practical route to efficient VDR by reducing this cost without retraining the retriever, but its uniform reconstruction objectives are poorly aligned with the sparse, non-uniform patch usage induced by late-interaction retrieval. Under aggressive compression, this misalignment can preserve rarely used patches while concentrating retrieval activity on too few retained representatives. To address this misalignment, we propose Marginal-Guided Compression with Optimal Transport (MAGIC), a training-free post-hoc compressor for efficient retrieval with frozen multi-vector embeddings. MAGIC derives a MaxSim-induced compression surrogate and optimizes it through a two-marginal entropic optimal-transport formulation, where a retrieval-demand source marginal prioritizes high-use patches and a balanced target marginal regularizes retained-facet usage. Across ViDoRe benchmarks, keep ratios, and retrieval backbones, MAGIC consistently outperforms strong post-hoc compressors, with particularly large gains in the aggressive-compression regime; component ablations verify the complementary effects of its two marginals. We release the code at: this https URL.

人机交互

[HC-0] A Sociotechnical Review of Algorithms in Health Systems: Technical Cost and Human-Centered Considerations

链接: https://arxiv.org/abs/2609.22070
作者: Victoria Chui,Kelly McConvey,Shion Guha
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Artificial intelligence (AI) applications in healthcare are becoming increasingly prevalent, to assist health systems, providers, and patients with tasks such as decision-making, risk prediction, and diagnosis. This increasing computational potential brings AI applications to the forefront of workplace decision making, often without full consideration of subsequent computational, organizational, and social costs. These applications are leveraged to reduce healthcare costs and increase efficiency of daily tasks, with model-related costs being considered at varying levels of granularity. To understand these trends, we critically analyze 114 papers to examine how cost-aware AI models have been developed for health systems. We explore the data, method, and outcome choices of these models, as well as their intersection with cost and human-centered concerns, highlighting the gaps in rigorous sociotechnical model design. From these trends, we define model costs and subsequent dimensions, presenting insight into those studies reporting financial, computational, organizational and/or social measures. Further, we critique the benefits and challenges of evaluating model-related costs and sustainability concerns when developing AI models for health systems.

[HC-1] Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw

链接: https://arxiv.org/abs/2609.22067
作者: Renkai Ma,Ruyuan Wan,Xuan Lu,Fan Yang,Chen Chen,Lingyao Li
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Users increasingly delegate work to autonomous AI agents, yet evaluations typically measure task completion rather than the values users prioritize. Using Value Sensitive Design, we analyzed, with LLM assistance, 73,093 first-person Reddit posts about using OpenClaw, each for its human value, agent aspect, value fulfillment, and user outcome. The 21 values form six value groups, including Autonomous, Dependable, and Affordable Operation, Bounded Reach, Reviewability, and Equitable Access. Relative to each aspect’s corpus share, values clustered not at the agent’s outputs but at the operating conditions users set around a run. Values were usually met where users described what the agent delivered, in five of six groups, and mostly unmet where users described supervising it, in all six groups. We conceptualize this pattern as value-sensitive delegation. Supporting human values requires attention not only to what an agent accomplishes, but to the conditions users set around delegation, including cost, access, and oversight.

[HC-2] How Researchers Use and Verify AI Coding Assistants: Tasks and Validation Practices in Scientific Programming

链接: https://arxiv.org/abs/2609.22049
作者: Gabrielle O’Brien,Reed Milewicz,Nasir Eisty
类目: oftware Engineering (cs.SE); Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Generative AI has entered research programming, yet there is little evidence about which tasks researchers hand to it or how they decide whether its code is correct. We draw on 527 free-text responses to a 2025 survey of researchers who write code, most of them at U.S. universities. In each response, a researcher recounts a single task from their own work, the way they used an AI tool for it, and what they did to assess the result. We coded the task and the evaluation strategies reported, and related both to programming experience, research area, and confidence ratings. Use was concentrated in five tasks: data handling, visualization, debugging, mathematical/scientific computing, and statistical analysis. Evaluation was informal and individual. Over half of accounts described running the generated code, while automated tests and review by another person were rare. Use cases and evaluation strategies varied little with programming experience, but confidence did: Less experienced programmers trusted the AI more than themselves, and experienced programmers the reverse. Evaluation confidence was not associated with the strategies reported. Its strongest correlates were confidence in the tool and in oneself. Validating AI contributions to scientific code rested largely on individual judgment, outside shared infrastructure for testing or review. Interfaces could support task-appropriate evaluation rather than leave it to the user.

[HC-3] Gricea: An Open Science Platform for Conversational AI Research

链接: https://arxiv.org/abs/2609.22039
作者: Nikhil Sharma,Yunlin Gong,Xinyang Cheng,Ziang Xiao
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI)
备注: 19 pages, 3 figures, 4 tables. Pre-print

点击查看摘要

Abstract:We need studies on conversational AI (CAI) at scale to understand human behavior and shape CAI design. However, fragmented reporting of systems and study configurations hinders replication, extension, and knowledge accumulation. We present Gricea, an open-science platform representing studies as configurable, deployable research artifacts that researchers can run, inspect, share, and reuse. Informed by a formative analysis of prior CAI research, Gricea couples study procedures, participant-facing systems, and conversational task behavior in. In a replication study using Gricea, we replicated configurations 93% of eligible CUI 2026 papers; while also flagging missing information in 96% of papers that hinder faithful replication — further motivating Gricea’s need. In a user study, researchers and practitioners from diverse backgrounds successfully constructed runnable studies addressing various open-ended research questions. Together, these findings demonstrate Gricea’s support for constructing, reproducing, and extending CAI studies through shared research artifacts, enabling cumulative knowledge building through open science.

[HC-4] Beyond Reactive Assistance: PV-Care Using Low-Density EEG and AI to Provide Proactive Context-Aware Help for MCI

链接: https://arxiv.org/abs/2609.22024
作者: Simon L Liu,Manish Kumar Krishne Gowda
类目: Human-Computer Interaction (cs.HC)
备注: Includes supplementary materials

点击查看摘要

Abstract:The growing elderly population gives rise to an urgent need for intelligent support systems, particularly for individuals with Mild Cognitive Impairment (MCI). This paper presents PV-Care, a proactive AI-driven assistance scheme that integrates wearable electroencephalogram (EEG) sensing with visual environmental perception to provide real-time, context-aware voice assistance for MCI users. Unlike traditional assistant systems that passively wait for user commands, PV-Care actively initiates helpful interactions based on the user’s detected brain states, including Learning, Memory Recall, and Resting, using a novel deep neural architecture named Spatial and Frequency Refinement Network (SFR-Net). By combining EEG-based cognitive-state recognition with AI-based visual analysis, PV-Care generates structured “4W-UT” prompts to guide the output of large language models (LLMs). Simulation results and user studies validate the high accuracy of the proposed SFR-Net and the effectiveness of PV-Care’s context-aware assistance. These results indicate that PV-Care is a feasible and promising solution for MCI caring.

[HC-5] When Should a Failing Robot Ask? Initiating Corrective Human-Robot Dialogue from Audited Sensor Evidence IROS2026

链接: https://arxiv.org/abs/2609.21942
作者: Eshika Pathak,Leela Krishna
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
备注: Accepted at the IROS 2026 Workshop on Human-Robot Dialogue

点击查看摘要

Abstract:A robot that fails at a task faces the first decision in corrective dialogue: act on its own diagnosis, consult another onboard sensor, or interrupt a person. Choosing well requires knowing how much the robot’s sensors reveal about the cause and how reliable the robot’s own diagnosis is. We build a simulated benchmark in which every failure’s true cause is known, because we injected it, and measure what each sensor reveals, with explicit checks against data leakage. Some failures are diagnosable from camera images; others only from the robot’s force data (0.99 from force data, no image method above 0.55). We then test six open vision-language models. Their behavior tracks the surface of the prompt, not the evidence: moving the refusal option from last to first in the answer list collapses refusal rates from 78-100% to 0-6% in three of the six swept model-and-family pairs. Accuracy from frames stays at or below a majority-class baseline under every prompt variant, with or without worked examples, and stated confidence carries no information about correctness. Handing the same models the force data as ten lines of text produces the first above-baseline diagnoses, in four of the six models: much of the failure reflects missing sensor data, not missing ability. We pose the choice as a three-action decision problem, act, consult your own sensors, or ask a human, whose optimal policy follows from measured accuracy. The models do not follow it, and their ask rates ignore a fourfold change in question cost. One question to a human still lifts them from that baseline to roughly the answerer’s own reliability (0.70-0.81 when they ask). The decision to ask should be tied to measured accuracy and stated costs, not to the model’s confidence.

[HC-6] Can I Trust My Body? A Three-Year Autoethnography of ChatGPT s Place in My Support System for Panic Attacks

链接: https://arxiv.org/abs/2609.21925
作者: Dongyijie Primo Pan,Pan Hui,Mirjana Prpa
类目: Human-Computer Interaction (cs.HC)
备注: 19 pages, 4 figures, 5 tables

点击查看摘要

Abstract:People increasingly seek mental health support from large language models, yet little is known about their use across years of recurrent panic. We present a three-year analytic autoethnography of the first author’s ChatGPT use while living with panic disorder, drawing on conversations, personal records, and accounts from friends or family members and professionals. Narrative analysis traces how my questions shaped ChatGPT’s roles and how earlier experiences influenced later responses to symptoms. Familiar explanations could make sensations less frightening, while changed symptoms renewed fears of serious illness. During sudden panic, advice could be difficult to follow, and some replies prompted further checking. Conversations could end while symptoms, checking, or help-seeking continued. We propose trajectory-level safety during and after panic: usable advice (Fit), a stopping point for repeated checking and reassurance seeking (Closure), and useful understanding and human support that remain available over time (Continuity).

[HC-7] Depressive symptoms are reflected differently across digital contexts

链接: https://arxiv.org/abs/2609.21919
作者: Yajing Wang,Emilia Marchese,Talayeh Aledavood,Juhi Kulshrestha
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:As more of everyday life takes place online, digital behavior may provide a potential window into how depressive symptoms are reflected in daily life. Yet digital mental health studies have produced mixed findings. These inconsistencies may partly reflect how digital behavior is measured: self reported use, single device studies, and aggregate screen time can obscure differences across devices, activities, and patterns of engagement. We combined monthly assessments of depressive symptoms with passively recorded mobile and desktop web traces from 1,146 adults in Germany over six months. We examined how general, cognitive-affective, and somatic depressive symptoms are reflected across digital contexts defined by device and activity type. Associations varied markedly across these contexts. On mobile, more severe symptoms were associated with more nighttime activity, greater use of social media, messaging, and entertainment, and fewer but longer sessions. On desktop, associations were fewer and largely involved reduced engagement with news, shopping, and adult content. Mobile associations primarily arose for general and cognitive-affective symptoms, whereas desktop associations were concentrated in somatic symptoms. Our findings suggest that characterizing how depressive symptoms are reflected in digital behavior requires attending to what people do online and where, not only how much screens are used.

[HC-8] Comparing Haptic Feedback Across Hand Tracking and Controllers in VR Object Interaction Tasks

链接: https://arxiv.org/abs/2609.21869
作者: Natalia Ocampo,J. Felipe Gonzalez,Robert J. Teather,Kiyoshi Kiyokawa
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Hand tracking offers natural VR interaction but lacks controllers’ inherent tactile feedback, while haptic feedback across input methods remains understudied. We conducted a participant study comparing vibration, impulse, and no feedback in hand- and controller-based VR interaction across grasp-, pinch-, and tap-like gestures assessing performance, workload, and preference. A custom glove and controller-mounted device provided both feedback types, aiming for consistency across input modalities. Controllers were faster for grasp and pinch, whereas hand tracking was more accurate for pinch and preferred overall. Haptics had limited performance effects, although impulse reduced grasp accuracy with controllers. Participants preferred vibration and impulse over no haptic feedback, favouring vibration overall. Our findings reveal a more nuanced relationship between haptic feedback, input device, and interaction context than suggested by previous work.

[HC-9] Beyond Counting Blessings: Tracing the Evolution of Gratitude Practices and Technology Needs

链接: https://arxiv.org/abs/2609.21853
作者: Qiuyue(Joy)Zhong,Jeongah Lee,Drishti Goel,Violeta J. Rodríguez,Dong Whi Yoo,Koustuv Saha,Ravi Karkar
类目: Human-Computer Interaction (cs.HC)
备注: 30 pages, 8 figures, 3 tables, including references and appendices

点击查看摘要

Abstract:Gratitude technologies support well-being by prompting reflection on what people appreciate. But gratitude does not serve the same purpose in every circumstance: as life situations change, so does what people seek from it, and whether it feels appropriate at all. To understand how technology can adapt to and support such shifts, we conducted retrospective, artifact-elicitation interviews with 17 adults who had practiced gratitude for one to fifteen years. Participants’ appraisals of their situations shaped what they needed, yielding six recurring practice patterns, including a boundary where gratitude felt forced. We contribute the Adaptive Gratitude Practice Model, which explains how appraisals shifted even within the same life situation, how participants adapted activities, modalities, and rhythms, lapsed under competing demands or emotional unreadiness, and resumed when gratitude again felt useful. Additionally, we derive design implications for situated support, self-understanding through past records, and relational care with changing life situations.

[HC-10] ouvigation: Embodied Adaptive Object Acquisition for Blind and Low-Vision Users in Unfamiliar Indoor Environments

链接: https://arxiv.org/abs/2609.21828
作者: George Xi Wang,Xiangyu Li,Shaoyue Wen,Jiaqian Hu,Junan Xie,Yupeng Wang,Ziyue Shi,Qijun Chen,Maaike Bouwmeester,Yuhua Jin,Jing Qian
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI)
备注: 12 pages, including figures and references

点击查看摘要

Abstract:Blind and low-vision users often face challenges when locating and physically acquiring objects in unfamiliar indoor environments. Existing vision-language-model-based assistants can provide semantic descriptions but may introduce latency, hallucinations, and guidance that is poorly aligned with embodied action. We present Touvigation, a hands-free object acquisition system that combines vision-language understanding with persistent local spatial modeling to provide low-latency, body-relative guidance. Drawing on formative interviews with eight blind and low-vision participants, we design a multi-stage guidance framework that adapts spatial references as users transition from orienting, to walking, to reaching and tactile verification. We evaluated Touvigation with 12 blind and low-vision participants against a multimodal large-language-model assistant and unassisted search. Touvigation achieved 100% task success, compared with 58% for the multimodal assistant and 85% for unassisted search, while reducing completion time and cognitive workload. Our findings demonstrate how persistent spatial grounding and adaptive embodied guidance can improve object acquisition for blind and low-vision users.

[HC-11] An Agent ic Just-in-Time Adaptive Intervention System for Personalized Sleep Support: Proof-of-Concept Study with N of 1 Data

链接: https://arxiv.org/abs/2609.21805
作者: Nick Rezaee,Chelsea Boccagno
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI)
备注: 7 pages, Submitted to ACM CHI

点击查看摘要

Abstract:Background: Just-in-time adaptive interventions (JITAIs) can use behavioral data to adapt support to changing contexts, but many rely on predefined rules and manual configuration. Objective: We developed a proof-of-concept sleep JITAI using an AI agent to review personal data, evaluate reminders, adapt interventions, and record decisions for human review. Methods: Running in Home Assistant on a configurable schedule, the agent follows a reusable skill file to review 30 days of sleep and behavioral data, including physical activity, smartphone use, and bedtime routines, to identify patterns and create or update automated reminders. Results: Initial runs demonstrated technical feasibility, successfully completing data review and intervention decisions while limiting reminders to three per day and saving decision records. Conclusions: Agentic AI may enable flexible, adaptive sleep JITAIs. The architecture supports future comparison with fixed or rulebased interventions, requires human oversight, and could extend to other health behaviors. Comments: 7 pages, Submitted to ACM CHI Subjects: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI) Cite as: arXiv:2609.21805 [cs.HC] (or arXiv:2609.21805v1 [cs.HC] for this version) https://doi.org/10.48550/arXiv.2609.21805 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Nick Rezaee [view email] [v1] Fri, 18 Sep 2026 14:16:57 UTC (338 KB) Full-text links: Access Paper: View a PDF of the paper titled An Agentic Just-in-Time Adaptive Intervention System for Personalized Sleep Support: Proof-of-Concept Study with N of 1 Data, by Nick Rezaee and 1 other authorsView PDF view license Current browse context: cs.HC prev | next new | recent | 2026-09 Change to browse by: cs cs.AI References Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading… BibTeX formatted citation loading… Data provided by: Bookmark checked="checked"class=“labs-tab-input”> Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs. Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?) mathjaxToggle(); We gratefully acknowledge support from our major funders, member institutions, , and all contributors. About Help Contact Subscribe Copyright Privacy Accessibility Operational Status (opens in new tab) Major funding support from

[HC-12] Comparing Hand and Controller Avatars with Hand Tracking and Controller-Based Interaction

链接: https://arxiv.org/abs/2609.21799
作者: Natalia Ocampo,J. Felipe Gonzalez,Robert J. Teather
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Previous research suggests that the congruency between common VR input devices - such as controllers or hand tracking - and their visual representations (e.g., hand or controller avatars) influences user experience and performance. However, the specific effects of input-avatar combinations remain underexplored. We study the effects of common input devices (hand tracking and controllers) and visual representations (hand and controller avatars) on performance and perceived success in target acquisition tasks. We included both grasping and pinching gestures across 16 combinations of input, avatar, and target size. Results indicate that hand tracking benefits from any form of visual representation - even when mismatched - achieving up to 5.8% greater accuracy compared to having no avatar, likely due to its reliance on visual feedback in the absence of a physical prop. Controllers were generally preferred and offered faster task completion. However, mismatched avatars had a stronger negative effect with controllers, particularly when the virtual gesture did not align with the physical action, leading to a 5.6% drop in accuracy compared to the matched condition - suggesting that inaccurate feedback can be more disruptive than having no avatar feedback at all.

[HC-13] Notrix: Understanding Machine Learning Solutions Across Computational Notebooks at Scale

链接: https://arxiv.org/abs/2609.21775
作者: Xiaotian Su,Hongxin Fu,Xiaoyu Zhang,April Yi Wang
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Computational notebooks make problem-solving visible, but typically only one notebook at a time. Meanwhile, in data science platforms like Kaggle, one competition can accumulate hundreds of notebooks. Effective collection-level analysis requires characterizing recurring solution patterns across all notebooks, as well as isolating specific notebooks for closer examination and learning. However, standard notebooks provide no common basis for this. Their workflows are nonlinear, cells declare no intent, and identical code can serve different ends, leaving hundreds of notebooks as separate documents. In this paper, we present Notrix, an interactive visual analytics tool for profiling hundreds of notebooks as one collection. Inspired by a formative study (N = 11), Notrix classifies every cell into one of thirteen machine learning (ML) stages, turning each notebook into a stage sequence, and clusters those sequences by structure rather than by code. To keep the representation constant as the scope narrows from the whole collection to a single cell, Notrix features three coordinated views—Workflow, Structural Matrix, and Detail—that appear at all four levels of granularity. In a within-subject study (N = 17) using two Kaggle collections of over 400 notebooks each, we observed participants answered questions about all notebooks more accurately with Notrix (median 88% vs. 50%) while opening 80% fewer notebooks per minute. Notably, four of the fourteen answered it without opening a single notebook (interaction logs, N = 14). Participants also reported significantly lower mental demand, temporal demand, and stress with Notrix (Holm-Bonferroni adjusted).

[HC-14] Learned Parametric Emotion Editing: Real-Time Affective Filtering for On-Device Social Media Video

链接: https://arxiv.org/abs/2609.21624
作者: Musa Rochi,Marcel Schubert,Christoph Gebhardt
类目: Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)
备注: 18 pages, 13 figures

点击查看摘要

Abstract:Problematic internet use affects a growing share of the population, yet common interventions, e.g., time limits, blocking, forced breaks, are coercive and easily circumvented. We explore a less restrictive alternative: adapting the emotional intensity of visual content. Prior work has shown that optimization can steer an image’s affective content, but its per-image optimization cost makes it impractical for real-time deployment. We instead learn a model that predicts this transformation in a single forward pass: a MobileNetV4 backbone with FiLM-based emotion conditioning outputs parameters for differentiable global transformations. This replaces prior iterative optimization (80 s per image) with a single 3.7 ms forward pass. In a user study (N = 54), the model reduced viewer-reported arousal relative to unedited images, comparably to the grayscale well-being filter, while being rated higher in perceived quality. We integrate the model into an Android app that adapts Instagram video in real time, sustaining 60 fps on a Samsung Galaxy S23.

[HC-15] Open Platform Field Experiments: Expanding the Design Space of Experimental Research on Social Media

链接: https://arxiv.org/abs/2609.21608
作者: Jordi Guillem Condom-Tibau,Giovanni Puccetti,Clara Bacciu,Matteo Abrate,Stefano Cresci
类目: Computers and Society (cs.CY); Human-Computer Interaction (cs.HC); Social and Information Networks (cs.SI)
备注:

点击查看摘要

Abstract:Despite a growing demand for causal evidence about social media, independent researchers remain severely constrained in their ability to conduct experiments directly on online platforms. To cope, multiple methodological workarounds have emerged - from controlled surveys and simulations to client-side overlays and platform partnerships - each requiring distinct trade-offs between desirable experimental properties. The recent emergence of open social media platforms offers a qualitatively different methodological opportunity. Here we propose a design space of social media experimentation and discuss Open Platform Field Experiments (OPFEs). OPFEs represent a distinct class of experimental approaches that enable independent researchers to directly intervene on functional platform components - such as clients, recommendation systems, and moderation services - within live social media environments. Through a comparative analysis of experimental archetypes, we show that OPFEs occupy a previously unexplored region of the design space. We then bridge theory and practice by characterizing the architectural and governance elements that enable OPFEs, mapping them onto Bluesky and the AT Protocol, and illustrating the end-to-end lifecycle of a complete OPFE design. Overall, this work establishes OPFEs as a practical methodological paradigm for independent, transparent, and ecologically grounded experimentation on open social media.

[HC-16] Reducing Barriers to Academic Support: Evaluating a Course-Specific RAG System for Addressing Help-Seeking Disparities in Higher Education

链接: https://arxiv.org/abs/2609.21600
作者: Andy Gray,Jake Hobbs
类目: Artificial Intelligence (cs.AI); Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Access to academic support is a key determinant of student success, yet students experience it unequally: some readily seek help from lecturers or tutors, while others hesitate due to anxiety, fear of judgement, uncertainty about expectations, or low confidence in their understanding. This may be especially evident in computing education, where programming tasks are cumulative and cognitively demanding. Although students increasingly turn to general-purpose generative AI tools, these can produce responses that are inaccurate, insufficiently contextualised, or misaligned with module expectations. This study presents and evaluates Beacon, a course-specific Retrieval-Augmented Generation (RAG) system providing private, immediate, module-aligned academic support. Grounding responses in approved teaching materials, Beacon was designed to lower barriers to help-seeking while encouraging independent learning. Using a design-based research approach, Beacon was developed iteratively and evaluated via mixed methods, combining questionnaires and semi-structured interviews with students and staff at a Higher Education institution. Students described Beacon’s responses as closely aligned with module content and more trustworthy than unrestricted generative AI tools, valuing its use of pseudocode and scaffolded explanations over direct solutions. Although participants remained cautious about trusting AI-generated responses without verification, they viewed the system as a valuable first point of support before consulting lecturers or official resources. The findings suggest that carefully designed course-specific AI systems may reduce barriers to academic support by occupying an intermediary space between independent study and formal support. Rather than replacing educators, educational AI may be most valuable when it broadens access to guidance while preserving the pedagogical role of lecturers.

[HC-17] GestureFAR: Streaming Co-Speech Gesture Generation with Flow Autoregression

链接: https://arxiv.org/abs/2609.21576
作者: Pinxin Liu,Haiyang Liu,Jiahao Luo,Junhua Huang,Chunhao Zou,Luchuan Song
类目: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Generating natural co-speech gestures from streaming speech is essential for embodied conversational agents, where motion must be produced while a user is still speaking. Recent streaming gesture systems make online generation possible by autoregressing over discrete motion tokens, but this design compresses high-dimensional continuous motion into finite codebooks and can limit the realism and diversity of generated gestures. To preserve both causality and continuous expressiveness, we propose \textbfGestureFAR, a flow-autoregressive framework for streaming co-speech gesture generation. First, GestureFAR autoregresses over causal continuous motion latents, using a transformer to model streaming audio-motion context and a per-token flow-matching head to sample the next latent from a continuous distribution. Second, we introduce a head-only flow distillation strategy that freezes the causal backbone and distills the multi-step per-token flow head into a single network evaluation using consistency and distribution-matching objectives. This keeps the model token-causal while removing the main latency bottleneck for live interaction. Experiments on BEAT2 show that GestureFAR significantly improves the quality–latency trade-off among streaming-capable methods, preserving strong gesture quality while enabling real-time token-causal generation. Project Page: this https URL

[HC-18] People escalate against a competitor labelled human and hold back against one labelled an optimising machine

链接: https://arxiv.org/abs/2609.21439
作者: Vinicius Ferraz,Leon Houf
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:People increasingly compete against AI agents rather than other human opponents. We distinguish two channels: an opponent effect and an information effect. These are different elements with different consequences: the opponent effect is specific to a given computational system, the information effect a property of the information environment that an organisation or policymaker can control. We separate them in a preregistered experiment (N = 1,395) using a dynamic all-pay auction, a repeated contest in which escalation of commitment arises from the incentives. What participants are told about the opponent (human, an AI trained to imitate people, or an AI trained to compete well) is varied and crossed with who they actually face, in a deception-free design. What people are told influences escalation: the median price rises by 6.7 points when a human might be the opponent and falls by 8.8 when an optimising machine might be, a spread of about 15% of the prize value of the competition, produced by information alone. Competing against the AI agents lowers prices, yet reduces the chance that both sides finish with positive earnings, showing distinct effects of the opponent channel. The information effect is not explained by articulated strategy, or individual differences, and is consistent with a competitive response engaged when a human is a live possibility. This shows that describing an AI competitor is not behaviourally neutral.

[HC-19] AESSI: An Around-Ear Silent Speech Interface for Cross-Day Online Reuse without Test-Day Calibration

链接: https://arxiv.org/abs/2609.21436
作者: Xiran Xu,Mochu Dong,Yujie Yan,Chenxi Wang,Yu Jiao,Jing Chen
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Silent speech interfaces (SSIs) enable private communication without audible speech and may support people with post-stroke dysarthria. Everyday reuse requires articulation-related representations that generalize across days despite sensor repositioning and physiological changes. We present AESSI, an around-ear SSI using masked-context representation pretraining (MCRP): a student predicts teacher representations from masked time-frequency inputs to encourage robustness to recording variability. We collected 44 electrophysiological recordings from 24 participants for 25 everyday Mandarin sentences. With six participants’ complete final recordings held out, AESSI achieved 92.24 percent mean accuracy without test-day calibration. AESSI exceeded the best adapted baseline by 50.57 percentage points. At least 21 days after each participant’s last recording, five participants each completed 50 independently randomized online tests without calibration, achieving 98.0 percent overall accuracy. Median preprocessing and inference time in CPU replay was 31.70 ms. These results demonstrate end-to-end system operation and support online reuse without test-day calibration. Demo is included with the paper.

[HC-20] MarineCraft: Enabling Rapid Prototyping of Underwater Robots via Modular Construction

链接: https://arxiv.org/abs/2609.21396
作者: Yuta Sugiura
类目: Robotics (cs.RO); Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Underwater robot development is often hindered by the complexities of waterproofing and wiring, which significantly delay the rapid prototyping process. This paper presents MarineCraft, a modular toolkit designed to accelerate the development cycle through structural reconfiguration. The system features self-contained, waterproof propulsion modules that integrate power, wireless communication, and actuation. By eliminating centralized wiring and the need for repeated sealing, MarineCraft allows diverse robot geometries to be assembled and tested in minutes rather than days. Experimental results demonstrate that this reconfigurable architecture enables fast, iterative design cycles while maintaining reliable operation and leak-free performance at depths of up to 2.5 meters. Our toolkit effectively lowers the barrier to underwater robotics by transforming modularity into a vehicle for rapid physical prototyping.

[HC-21] RobotEQ-Video: A Video-Centric Benchmark for Social Proactive Intelligence with World-State Taxonomy

链接: https://arxiv.org/abs/2609.21371
作者: Xinyi Che,Zheng Lian,Kuofei Fang,Xuehao Wang,Xinghai Gao,Junqing Wu,Chuyu Wu,Liyi Liu,Yanhan Huang,Keyi Xie,Haomin Ouyang,Jinyang Wu,Fan Zhang,Runhao Zeng,Xun Yang,Bin He
类目: Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Social Proactive Intelligence (SPI) extends proactive assistance beyond task completeness to consider social appropriateness in diverse embodied scenarios. However, prior SPI research faces two key limitations. First, existing work focuses on static images, whereas dynamic videos provide crucial cues for inferring human states and needs, offering richer information than isolated images. Second, prior work often relies on free-form data collection pipelines, which fail to guarantee comprehensive coverage of diverse scenarios. To address these gaps, we introduce RobotEQ-Video, shifting the focus from image-centric to video-centric analysis. To ensure comprehensive video coverage, we construct a hierarchical world-state taxonomy organized into a four-level coarse-to-fine structure, comprising 6 domains, 20 dimensions, 142 level-1 attributes, and 816 level-2 attributes. The resulting benchmark comprises 2K+ videos with 100K+ human annotations and 16K+ labels for assessing behavior properness. Benchmark evaluation reveals that current systems remain unreliable and fall short of human performance. We further explore how world models can help tackle this task. This work advances SPI research from static images to dynamic videos and ensures more comprehensive scenario coverage during benchmarking.

[HC-22] GUIDE: Designer-in-the-loop Authoring of Conformant Generative User Interfaces

链接: https://arxiv.org/abs/2609.21285
作者: Hyewon Lee,Ziying Wang,Aiden Moy,Saran Nagubandi,Jason Wu
类目: Human-Computer Interaction (cs.HC)
备注: 24 pages, 13 figures, 5 tables

点击查看摘要

Abstract:Generative User Interfaces (GenUIs) enable applications to generate interfaces on demand from user needs and context. Like conventional UIs, they must still reflect designers’ intent and conform to requirements such as brand identity. Unlike conventional UIs, designers cannot directly specify or see every interface a GenUI may produce, making design intent harder to enforce. We introduce GUIDE (GenUI Development Environment), a system that lets designers continuously inspect and refine GenUI behavior as they create and edit interfaces. GUIDE uses designers’ modifications and interactions to adapt GenUIs through prompt optimization and a novel adaptive conformance scoring model. We validate GUIDE’s scoring model and system. The scoring model matched or outperformed proprietary LLM baselines on held-out comparisons of real and synthetic application screens. In a study with 12 UI/UX practitioners, participants found GUIDE effective and usable and significantly preferred aligned outputs over a strong baseline using exemplars and a model-generated this http URL.

[HC-23] wos a Crowd: Human and AI-Based Copresence for Developers with ADHD

链接: https://arxiv.org/abs/2609.21254
作者: Veronica Pimenova,Seth Bernstein,Shalini Madan,Dhruv Jain,Venkatesh Potluri
类目: Human-Computer Interaction (cs.HC); Software Engineering (cs.SE)
备注: 28 pages, 3 figures

点击查看摘要

Abstract:Effective collaboration and communication are vital to developer productivity and well-being, yet remain constrained by human factors such as attention, intrinsic motivation, and interpersonal accountability. These constraints are particularly vital for developers identifying with Attention Deficit Hyperactivity Disorder (ADHD), who navigate persistent environmental barriers in modern hybrid workplace settings. While developers with ADHD frequently rely on collaborative copresence practices (such as body doubling or pair programming) to support executive function, the recent emergence of agentic AI coding assistants has begun reshaping these collaborative dynamics. To investigate how developers with ADHD engage in human and AI-based copresence practices, we conducted semi-structured interviews with 14 software engineers with ADHD. Our findings reveal that while traditional human-human copresence provides critical social support and onboarding structure, it forces developers to constantly manage professional reputation and sacrifice personal privacy. Conversely, developers leverage emerging human-AI copresence to maintain accountability and cognitive flow without the social anxiety, performance judgment, or surveillance associated with human observation. Based on these empirical insights, we map developer copresence practices onto core dimensions of Goffman’s copresence theory and Forsgren et al.'s SPACE framework of developer productivity, and provide design recommendations for AI-based tools that promote inclusive collaboration for developers with ADHD.

[HC-24] Self-Care and Mental Health: Mapping Over A Decade of HCI Interventions

链接: https://arxiv.org/abs/2609.21239
作者: Anna Fang,Tony Wang,Jenny Fu
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Technology increasingly supports self-care for understanding and improving one’s own mental health. HCI is at the center of the turn towards self-care technology, yet we lack an account of who these interventions serve, what practices they support, how technology mediates those practices, and assumptions underlying design for self-care. In order to characterize the current landscape and inform future research, we analyzed 91 SIGCHI papers that contribute HCI interventions for mental health self-care from the ACM Digital Library from 2015 through June 2026. Then, we conducted an interpretive synthesis to surface six orientations of self-care, which describe how HCI self-care interventions constitute care through shared assumptions regarding self, care, and technology. Overall, our work provides an interconnected vocabulary for positioning HCI mental health self-care, highlights changing responsibilities of care towards users, and discusses implications for providing a more situated account of HCI self-care technology in addressing the ‘general’ user.

[HC-25] Human Driver Temperament and the Safety Impact of a C-V2X Denial-of-Service Flooding Attack in Mixed-Autonomy Traffic

链接: https://arxiv.org/abs/2609.21206
作者: Rasheed Bello,Gurcan Comert,Varghese Vaidyan,Akinbobola Jegede,Vijay Bendigeri,Judith Mwakalonge
类目: Human-Computer Interaction (cs.HC)
备注: 12 pages, 6 figures, Submitted to AHFE Hawaii International Conference

点击查看摘要

Abstract:Cooperative and connected automated vehicles (CAVs) rely on Signal Phase and Timing (SPaT) messages to cross signalized intersections; a denial-of-service (DoS) flood that blocks SPaT forces CAVs into a fail-safe mode. Because human-driven vehicles share the intersection, the safety consequence depends not only on the attack and the CAV fail-safe policy, but on how the surrounding human drivers behave. We investigate this human-factors dimension with a coupled OMNeT++/INET (5G NR-V2X) and SUMO microsimulation of a signalized corridor, sweeping CAV market penetration (10-90%), four calibrated driver temperaments (cautious to aggressive) and two standards-based fail-safe policies, a minimal-risk maneuver (MRM) and an adaptive cruise control (ACC) keep-driving fallback, with each attack arm differenced against its policy-matched no-attack baseline. Temperament’s effect on the attack is specific and modest rather than a blanket amplification. Aggressive surroundings worsen one metric, the hard-braking a keep-driving fail-safe forces on nearby drivers (p = 0.03), rising from near zero to +8 episodes/1000 veh-s. They appear to dampen rear-end conflicts, but only because the flood clears the queues aggressive drivers build, so the gain is in flow, not safety. On the attack’s primary signatures, CAV red-light running and crossing conflicts, temperament has no detectable effect. It instead dominates baseline risk, producing a 13- to 18-fold cautious-to-aggressive gradient far larger than the attack itself, which acts through a channel already congested by CAV adoption. Human driver populations determine how dangerous the intersection is but do not systematically amplify this attack, so fail-safe design cannot assume a cautious test population bounds the risk.

[HC-26] OpenRoIS: A Community-Driven Open-Source Middleware Implementing the Robotic Interaction Service (RoIS) Framework for Physical Robots and Virtual Agents

链接: https://arxiv.org/abs/2609.21178
作者: Sebastian Carrera Villalobos,Christopher Nolan Arellano,Arne Hitzmann,Edilson Morais Brito,Akira Utsumi,Yukiko Horikawa,Takahiro Miyashita,Lotfi El Hafi
类目: Robotics (cs.RO); Human-Computer Interaction (cs.HC)
备注: Submitted for presentation at the 2027 IEEE/SICE International Symposium on System Integration (SII), Kobe, Japan

点击查看摘要

Abstract:Service applications for human-robot interaction are commonly written against the hardware-specific interfaces of one platform, so a change of hardware forces a rewrite of the application. The Robotic Interaction Service (RoIS) Framework 2.0, standardized by the Object Management Group (OMG), addresses this fragmentation by defining a platform-independent model in which Service Applications interact with Human-Robot Interaction (HRI) Engines through standardized interfaces and hardware-independent symbolic messages. A specification alone, however, does not provide the maintained implementation, Software Development Kits (SDKs), and adapters needed for practical adoption. This paper presents OpenRoIS, a community-driven open-source middleware providing a concrete implementation of the RoIS Framework 2.0. It takes the position that an openly developed, paradigm-neutral implementation is what carries the standard from specification to practice. OpenRoIS contributes a recursive engine architecture in which a single engine class realizes the main and sub HRI Engine roles, an internal five-method component contract distinct from the five external RoIS interfaces, a mapping of those interfaces onto JSON-RPC 2.0 over WebSocket, a single-source-of-truth type pipeline that generates three consistent language stacks, TypeScript and C# client SDKs that include web and Unity support, and a Python adapter SDK that includes ROS 2 support. Through the common RoIS interfaces, a Service Application can address physical robots and virtual agents over the internet. All source code, interface types, and documentation are released under the Apache-2.0 license and openly developed at this https URL.

[HC-27] Exploring Text Classification Models with Sparse Autoencoders IEEE-VIS2026

链接: https://arxiv.org/abs/2609.21142
作者: Daniel Kerrigan,Brian Barr,Enrico Bertini
类目: Human-Computer Interaction (cs.HC)
备注: 11 pages, 13 figures. Accepted as a short paper at IEEE VIS 2026

点击查看摘要

Abstract:As language models (LMs) rise in prominence, there is interest in making them more transparent in order to better understand their internal behavior. Recent interpretability work has focused on using sparse autoencoders (SAEs) to break down neuron activations at a given layer in the LM into human-understandable features, where each feature represents a concept that the model has learned. In this paper, we share work on using SAEs to analyze the behavior of text classification LMs. We present techniques for exploring the relationships between the SAE’s features and the model’s predictions and errors. We integrate these techniques into SAEfarer, a tool for analyzing concepts learned by text classification LMs. We assess SAEfarer in an expert pilot evaluation with five Ph.D. students.

[HC-28] Decoding the Dashboard: Data Comics to Support Students Understanding of Learning Analytics Visualisations

链接: https://arxiv.org/abs/2609.21141
作者: Mikaela Elizabeth Milesi,Vanessa Echeverria,Lixiang Yan,Yueqiao Jin,Riordan Alfredo,Jie Xiang Fan,Linxuan Zhao,Dragan Gašević,Yi-Shan Tsai,Roberto Martinez-Maldonado
类目: Human-Computer Interaction (cs.HC)
备注: Accepted to ECTEL 2026

点击查看摘要

Abstract:Learning analytics dashboards (LADs) are intended to help students make sense of their learning data to support reflection and decision-making. However, their visualisations can be complex, particularly for students with low visualisation literacy. Narrative techniques, such as annotated charts and data comics, have been used to communicate insights directly, but not as supplementary materials to empower students to explore their visualisations themselves. In response, we conducted a qualitative study examining how data comics can complement LADs. We interviewed 18 nursing students and 4 of their teachers about a multimodal LAD containing visualisations with data comics explaining them. Analysis showed that data comics were clear, engaging, and helped make complex visualisations more accessible, though they must be carefully designed to avoid overwhelming students with information. The findings suggest that both students and teachers are receptive to data comics as a means of supporting the interpretability of LADs.

[HC-29] From Task Success to Productive Success: Evaluating Human-AI Collaboration by Quality and Cost EMNLP2026

链接: https://arxiv.org/abs/2609.21117
作者: Saki Imai,Mert İnan,Malihe Alikhani
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
备注: EMNLP 2026

点击查看摘要

Abstract:AI productivity is often measured by task completion time, economic value, or improvements in outcome quality. However, these measures usually treat collaboration as a black box where they capture what output was produced, but not the interaction cost required to produce it. Motivated by economics literature, we introduce a productivity-oriented framework for evaluating human-AI collaboration as outcome quality relative to interaction cost. Across two datasets spanning four tasks, we show that: (1) sessions with identical quality ratings can differ by up to 70 times in interaction cost; (2) quality-cost relationships vary by task, with some tasks rewarding extended interaction and others favoring fast convergence; (3) subjective user ratings are not reliable substitutes for productivity; and (4) productive sessions are characterized by agents probing earlier and users spending less effort repairing the interaction. By distinguishing productive success from costly success, our framework makes interactional cost visible and shows how dialogue analysis can inform the evaluation and design of AI systems.

[HC-30] Understanding How Educators Configure GenAI Support for Open-Ended Learning – An Exploratory Study of K-12 Career Exploration

链接: https://arxiv.org/abs/2609.21019
作者: Si Chen,Xinyue Chen,Artur Mullagaliyev,Alexander Nwanganga,Shifu Hou,Deng Pan,Ronald Metoyer,Sugana Vijay Chawla
类目: Human-Computer Interaction (cs.HC)
备注: 32 pages

点击查看摘要

Abstract:Generative AI (GenAI) can support open-ended learning through generation, personalization, and learner modeling, yet educators need ways to shape these capabilities around educational goals. Through interviews and design activities with 15 U.S. educators, we examined educator configuration of GenAI using K-12 career exploration as an exploratory context. Educators configured not only AI-generated experiences, but also when student activity became an inference, whether learner information persisted, who could access it, and how it informed subsequent human action. They also faced challenges translating teaching needs into configurations: recognizing possibilities for control beyond familiar uses of GenAI, decomposing general-purpose AI into understandable functions and responsibilities, and identifying useful information through intended teaching actions. We discuss how GenAI systems can support educators in expressing and testing configurations, while establishing boundaries around personalization, inference, persistence, disclosure, and action to keep AI-supported learning aligned with evolving learner needs.

[HC-31] rustworthy FinAInce: Unpacking How AI-Mediated Financial Advice is Judged

链接: https://arxiv.org/abs/2609.20989
作者: Aryan Ramchandra Kapadia,Eshwar Chandrasekharan,Koustuv Saha
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY)
备注:

点击查看摘要

Abstract:As generative AI is increasingly used as a source of personal financial guidance, understanding how people appraise such advice is important for supporting appropriate reliance. We conducted a randomized vignette experiment with 285 U.S. adults across eight financial decisions, independently varying three advice styles—AI, expert, and online community—and displayed source labels while holding the underlying recommendation consistent. Advice style most strongly shaped message and safety appraisals, Expert labels selectively increased perceived source knowledge, and decision context primarily shaped risk and safety appraisals. These appraisals were associated with downstream judgments, with models explaining 69.2% of overall quality, 75.9% of trust, and 82.9% of intended reliance. Expert-style advice also remained most preferred when shown without source labels. Our findings have implications for understanding financial advice evaluation, distinguishing the roles of advice style and source labels, and designing financial AI that supports grounded evaluation rather than simply maximizing trust.

计算机视觉

[CV-0] Designer-RSI: Evolving Procedural Memory from User Traffic for Agent ic Graphic Design

链接: https://arxiv.org/abs/2609.22086
作者: Hongyang Du,Lan Yan,Christian Flores,Asim Kadav
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注: 9 pages, 7 figures

点击查看摘要

Abstract:Professional graphic design is a long-horizon agentic task in which structured, editable artifacts emerge from many interdependent actions, yet outcomes admit no reliable programmatic oracle. We introduce a continual adaptation framework in which a frozen frontier model operates professional design software through more than 230 tools, while an external procedural memory of natural-language skills accumulates and refines reusable design procedures from experience. The memory widens by acquiring procedures for recurring uncovered subtasks and deepens by revising existing procedures against their own successful and failed executions, while a matched replay gate admits only changes that repair failures without regressing observed successes. Five rounds over 1,406 real user briefs and 1,869 automatically graded trajectories, with no weight updates and no human labels, grow the bank from 76 documentation-derived skills to 139 and raise GenEval2 execution success on Claude-Sonnet-4 from 72.7% to 99.3% (+11.99 points in generation quality), with 61.8% and 67.6% win rates against the no-skill agent across four specialized design benchmarks on Claude-Sonnet-4 and Claude-Opus-4.6. We further show the two mechanisms are effective in combination: on 200 held-out briefs from user-traffic benchmark, widening or deepening alone reaches a 49.4% / 48.6% win rate over the no-skill agent, while their combination reaches 58.5% (p = 0.025). Procedural memory offers a practical route to continual adaptation of agents under noisy, unverifiable feedback.

[CV-1] MintAct: A Unified Visual Agent for Digital Environments

链接: https://arxiv.org/abs/2609.22083
作者: Mingfei Gao,Rui Tian,Haiming Gang,Bohan Zhai,Le Zhang,Yuanzheng Gong,Di Feng,Ege Özsoy,Kaixin Ma,Vishwesh Kirthivasan,Oğuzhan Fatih Kar,Roman Bachmann,Anders Boesen Lindbo Larsen,Afshin Dehghan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:We present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, trained at 2B, 4B, and 8B scales. Through careful design of our environments, data, and training recipes, MintAct models match the performance of per-domain specialists across all of these capabilities. To enable this, we develop a scalable environment and reinforcement learning (RL) infrastructure. On the environment side, we host hundreds of concurrent instances across heterogeneous per-domain backends, serving both trajectory data collection and online RL. To enable efficient and scalable RL training, an asynchronous framework keeps explicit control over the cross-domain training distribution and remains stable under noisy environment feedback and off-policy drift. Experimental results show that MintAct achieves state-of-the-art performance (48.9 on OSWorld-Verified) across a wide range of benchmarks at comparable model sizes.

[CV-2] OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation

链接: https://arxiv.org/abs/2609.22069
作者: Wenxue Li,Peiyan Guan,Haoyang Jiang,Junxian Cai,Hualuo Liu,Chunjie Zhang,Chong Guan,Songlian Li,Taiyi Wu,Yongjian Yu,Xiaotong Zhao,Alan Zhao,Eric Liu,Xi Chen,Yu Liu,Lei Zhu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Reference-to-video (R2V) generation is evolving toward increasingly general and versatile reference control, giving rise to the emerging paradigm of omni R2V generation. However, existing benchmarks fall short of these emerging capabilities: their test cases cover limited reference types and compositions, and their evaluation protocols largely assess holistic reference consistency, overlooking whether reference factors are properly preserved, disentangled, and routed. Meanwhile, the high cost of constructing omni R2V training data makes suitable training resources scarce. To address these gaps, we introduce OmniVBench and the Omni-R2V Dataset for evaluating and training omni R2V models. OmniVBench expands R2V evaluation across broader reference types, fine-grained control tasks, and richer reference compositions, covering 7 task families and 18 fine-grained tasks spanning content, motion, style, structure, narrative, and multi-reference settings. We introduce factor-grounded evaluation with 12,172 case-specific checklist items, assessing whether intended reference factors are faithfully preserved, correctly disentangled and bound to their targets, and properly realized according to the instruction. We further introduce the Omni-R2V Dataset, bringing industrial-grade training resources for diverse R2V tasks to the broader research community. Drawing primarily on a large-scale corpus of professional video footage, it comprises 340K processed training samples spanning diverse reference types and multi-reference compositions. We develop task-specific pipelines for reference-target pair construction, offering a practical and scalable recipe for omni R2V data construction. Extensive evaluation of advanced open- and closed-source R2V models reveals clear performance gaps across task families and evaluation dimensions on OmniVBench, highlighting remaining limitations of current R2V models.

[CV-3] raffic Sign Recognition for Autonomous Driving Using Branched YOLOv2 and Geometric Features

链接: https://arxiv.org/abs/2609.22060
作者: Arefeh Rezaei
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Traffic sign recognition (TSR) is an important perception task for autonomous driving and advanced driver-assistance systems, where a system must both localize traffic signs and determine their semantic classes efficiently. This work presents a TSR system based on YOLOv2 for simultaneous detection and classification. Two complementary modifications are studied. First, YOLOv2 is extended with intermediate prediction layers, forming a branched architecture that can terminate inference early for easy cases and reduce computation time. Both whole-image and cell-wise branching strategies are investigated. Second, geometric information is introduced to reduce classification errors between visually similar signs. An unsupervised Bayesian image-segmentation method produces binary representations that are compared with class-specific geometric templates inside YOLOv2 bounding boxes. This information is used either during inference or as an additional signal during training. A dedicated dataset is constructed by combining GTSDB and GTSRB samples using seamless cloning and controlled image transformations. Experiments cover ten traffic-sign classes, with 3,000 training and 300 test samples. The selected branched architecture reports 0.647 s runtime and 0.680 mAP, compared with 0.6607 s and 0.680 mAP for baseline YOLOv2. Geometric verification during inference increases mAP to 0.713, while the geometric-feature training variant achieves 0.697 mAP with a reported runtime of 0.6608 s.

[CV-4] PRIME: Perception Feedback with Situational Memory Embeddings in VLA Models

链接: https://arxiv.org/abs/2609.22040
作者: Erik Deinzer,Naya Baslan,Luca Paparusso,Narunas Vaskevicius,Peter Knott,Luigi Palmieri
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Current Vision-Language-Action (VLA) models for autonomous driving operate primarily through feedforward inference across the perception–reasoning–planning hierarchy. While modern architectures maintain temporal recurrence within the perceptual module, early perception remains blind to downstream reasoning and navigation goals, processing visual inputs agnostically without prioritizing cues informed by prior decisions. To bridge this gap, this paper introduces PRIME, a learned feedback mechanism that conditions the VLA perceptual queries on a novel Situational Memory. By aggregating latent representations of past perception, reasoning, navigation goals, and predicted behaviors across an L-step window via cross-attention, PRIME enables intent-driven perceptual attention at minimal computational cost, adding only a maximum of 29.7M parameters (0.41% of the 7.3B-parameter base model). Evaluated on the Bench2Drive closed-loop benchmark, PRIME achieves a state-of-the-art Driving Score of 82.47 (+4.73 over ORION) and a Success Rate of 60.00% (+5.38 percentage points), the highest reported Driving Score among published VLAs trained on Think2Drive demonstrations.

[CV-5] GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments

链接: https://arxiv.org/abs/2609.21948
作者: Yichen Liu,Puzhen Yuan,Xiang Zhu,Yanjiang Guo,Jianyu Chen
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Learning large-scale vision-language-action (VLA) models from multi-embodiment datasets remains challenging due to heterogeneous action spaces across end effectors. Although latent action models (LAMs) can learn embodiment-agnostic action representations from diverse video data, existing image-based LAMs often fail to capture fine-grained end-effector articulation, particularly finger-level geometric changes in human and dexterous robot hands. To address this limitation, we propose GALA, a Geometry-Aware Latent-Action modeling framework that augments image-based latent actions with 3D end-effector geometric motion. However, naively incorporating point clouds yields fine-grained action representations with limited shared semantics, hindering cross-embodiment pretraining. To address this issue, we introduce the Unified End-effector Motion Representation (UEMR), which preserves fine-grained motion information while improving the cross-embodiment generalizability of latent actions. Building upon UEMR, GALA combines visual latent actions that capture scene-level dynamics with geometric latent actions that capture shared fine-grained end-effector articulation, providing effective supervision for VLA pretraining from multi-embodiment data, including action-free ego-centric human videos. Experiments on fine-grained motion probing, cross-embodiment retrieval, and downstream VLA evaluation demonstrate GALA’s effectiveness in modeling generalizable fine-grained motions across embodiments, achieving 68.3% RoboCasa-GR1 success rate and 75.5% real-world success rate. Code, appendix, and demos are available at this https URL.

[CV-6] Info3R: Information-Adaptive Test-Time Training for 3D Reconstruction

链接: https://arxiv.org/abs/2609.21938
作者: Sunghyun Baek,Hanna Bae,Minchan Kwon,Junmo Kim
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Transformer-based models have recently achieved strong performance on 3D reconstruction from images, and recent works extend them to process video streams in an online manner for real-world deployment. However, existing methods overlook two key signals when handling long image streams: the importance of each incoming frame and the information saturation of the model’s internal state. In this paper, we propose Info3R, a novel information-adaptive test-time training method for the online 3D reconstruction. We introduce an information-aware state update that modulates the state update strength based on the redundancy and informativeness of each incoming frame. To restore the state’s plasticity – its capacity to incorporate new observations – we propose a dynamic state reset, triggered by the cumulative magnitude of state updates and the model’s prediction confidence and accompanied by an anchor-to-world alignment. Our method achieves consistent improvements on camera pose estimation, video depth estimation, and 3D reconstruction, while substantially mitigating the performance degradation in the long sequence evaluation. Notably, on KITTI Odometry, our method achieves on average 1.68x lower ATE than LongStream, demonstrating its robustness on extended outdoor sequences.

[CV-7] he Role of Radiometric Features in Cross-Site Leaf-Wood Segmentation of LiDAR Point Clouds

链接: https://arxiv.org/abs/2609.21903
作者: Roman Kaharlytskyi,Derek T. Robinson,Roberto Guglielmi
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 5 pages, 6 figures. Accepted for presentation at IGARSS 2026 (IEEE International Geoscience and Remote Sensing Symposium)

点击查看摘要

Abstract:Leaf-wood segmentation of individual trees from LiDAR point clouds is essential for quantitative structure models (QSMs) used in non-destructive biomass estimation. Existing segmentation methods typically exclude radiometric features (e.g., intensity, return number) to maximize cross-sensor compatibility. We challenge this design choice by evaluating cross-site and cross-platform generalization: training on the public Heidelberg dataset (terrestrial TLS, 1550nm) and testing on a novel dataset from Ontario, Canada (RPA-LS, 905nm). Results show that geometry-only methods - including state-of-the-art deep learning models trained on high-density LiDAR datasets - fail to generalize to the sparse, top-down geometry of aerial scans, achieving F1 scores = 0.56. Incorporating radiometric features (intensity, return number, number of returns) improves F1 to 0.61, but more critically, increases wood recall by 119% from 0.16 to 0.35. Furthermore, geometry-only approaches often result in fragmented stem and branch components. We find that leveraging radiometric features preserves greater structural connectivity, resulting in more coherent architectures that are better suited for QSM reconstruction. We demonstrate that while geometric patterns are view-dependent and prone to overfitting scan patterns, radiometric features encode physical material properties that generalize across disparate sensors and environments.

[CV-8] Catena: A Comprehensive Software Suite for Large-Scale Connectomics

链接: https://arxiv.org/abs/2609.21887
作者: Samia Mohinta,Pedro Gómez-Gálvez,Shi Yan Lee,Daniel Franco-Barranco,Michael Clayton,Stephan Preibisch,Jan Funke,Albert Cardona
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:The gold standard datasets for mapping connectomes are electron microscopy volumes of densely labeled neural tissue at nanometer resolution. Yet reconstructing and proofreading neuronal arbors and annotating all synapses requires pipelining multiple software tools that are often fragmented, inconsistently maintained, or proprietary, hindering reproducibility and automation. Here, we introduce Catena, an open-source, comprehensive, developer-centric software suite for connectomics that integrates modules for 3D neuron and organelle segmentation, synapse detection, microtubule tracking, and neurotransmitter inference. Catena organizes its modules in composable, chunk-wise processing pipelines in a completely documented, extensible, and adaptable design. We further reduce compute and ground-truth data requirements with pretrained machine learning models, facilitating fine-tuning. Catena ships fully containerized modules that encapsulate evolving dependencies for consistent execution across workstations and clusters. By consolidating open components, shareable models, and containerized runtimes, Catena delivers a reproducible and scalable approach to mapping cellular connectomes from electron microscopy volumes. Code and documentation: this https URL

[CV-9] Benchmarking the Explanatory Quality of Open-Weight Vision-Language Models in Face Recognition

链接: https://arxiv.org/abs/2609.21879
作者: Laurent Colbois,Sébastien Marcel
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 11 pages

点击查看摘要

Abstract:Vision-Language Models (VLMs) have recently been proposed as promising tools for face recognition, as they can produce natural language explanations alongside similarity scores. This capability is considered appealing for face comparisons in forensic contexts, which require decisions to be transparent and auditable. However, existing evaluations of VLMs for that use case focus mostly on recognition accuracy, while the validity of generated explanations remains unquantified. In this work, we introduce a benchmarking framework for VLM-based face recognition that treats explanation quality as a core evaluation axis. We propose two criteria that explanations should satisfy: relevance, i.e., reliance on identity-stable facial features; and faithfulness, i.e., alignment with the visible image content without hallucinated features. We jointly develop a methodology enabling the quantification of relevance and faithfulness of evaluated models, based on constraining model outputs to a structured explanation format that supports automated querying and auditing. Using this framework, we benchmark several families of open-weight VLMs, jointly evaluating face verification accuracy and explanation quality. Our results highlight remaining shortcomings of produced explanations, and emphasize the need for such explanation quality metrics to get a complete picture of model performance. The proposed benchmark and open-source evaluation harness provide a foundation for proper benchmarking and future fine-tuning of explainable face recognition systems.

[CV-10] Chronosphere: Space-Time Tessellation of Local Climate Experts

链接: https://arxiv.org/abs/2609.21872
作者: Daniel Cher,Eric Xing,Kexing Li,Brian Wei,Isaac Corley,Nathan Jacobs
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:We introduce Chronosphere, a spatio-temporal neural field that learns representations of climate. A central challenge in geographic representation learning is modeling environmental processes whose spatial and temporal complexity varies widely. Yet existing location encoders typically fix a single level of detail everywhere. Global bases such as spherical harmonics spread capacity uniformly across space and time. Localized bases resolve only predefined regions. Learned tessellations adapt, but are inefficient at representing higher frequencies. Chronosphere unifies these approaches, pairing an adaptive tessellation of learnable sites on the spacetime torus S^2\times S^1 with a shared bank of local basis functions. Both where capacity is placed and how much detail each region carries adapt to the data, across space and time. Trained to reconstruct climatology, Chronosphere matches or leads state-of-the-art location encoders across spatial and temporal tasks, with the largest gains under spatial and temporal transfer.

[CV-11] Morphology-Aware Ambiguity Learning for Wafer Defect Decision Support

链接: https://arxiv.org/abs/2609.21866
作者: Seungjun Chu,Seokhyun Chung
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Wafer map defect recognition is commonly formulated as a fixed-taxonomy classification problem that assigns each wafer to a single defect class. However, some wafers exhibit morphologies near class boundaries, for which forcing a single prediction may be less informative than providing plausible diagnostic alternatives. This paper proposes a morphology-aware ambiguity learning framework that supports three diagnostic actions: automatic single-class diagnosis, assisted diagnosis with two plausible defect classes, and full review. Using the radial, angular, and geometric characteristics of training wafer maps, the framework constructs a class-level ambiguity matrix representing defect-class pairs with similar morphology and plausible diagnostic alternatives. It guides the model to learn plausible alternative classes rather than treating all incorrect classes equally. During inference, the matrix determines whether an uncertain prediction can be represented by a meaningful two-class diagnostic set or should be escalated for full review. Experiments on WM-811K show that the proposed framework outperforms conventional approaches in defect recognition and diagnostic decision support, providing meaningful two-class alternatives while reserving full review for cases with unresolved ambiguity. Illustrative cost analyses further show the potential cost advantage of the proposed routing strategy. The diagnostic behavior of the framework remains consistent across different backbone architectures.

[CV-12] he Weight Is Over - Interactive Diffusion on Consumer GPUs

链接: https://arxiv.org/abs/2609.21849
作者: Frieder Ganz,Maximilian Müller
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Performance (cs.PF)
备注:

点击查看摘要

Abstract:On-device inference is booming, but the momentum is almost all in language models. Diffusion pipelines are memory hungry, latency-sensitive, and require orchestrating an embedder, a transformer, a decoder, and often further postprocessing that is not as standardized as LLM inference loops are. We navigate the trade-off between performance, quality, and model footprint to reach as many client devices in the wild as possible. We make three contributions: an embedding translator that maps a small text encoder into a large encoder space to cut weight and latency; a reproducible sweep recipe for navigating the speed/quality/memory triangle in diffusion pipelines; and an interactive on-device image generation editor achieving sub-second TTFI on recent GPUs.

[CV-13] Object Detection Benchmarks are Incomplete: The Role of Label Errors and Annotation Uncertainty

链接: https://arxiv.org/abs/2609.21822
作者: Sarina Penquitt,Jonathan Klees,Antonia van Betteray,Parssa Jashnieh,Peter Stehr,Matthias Rottmann,Lars Schmarje
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:While object detection has advanced through improved architectures and open-vocabulary models, we provide strong evidence that benchmark quality is limited by annotation incompleteness. Across four widely used datasets (COCO, Pascal VOC, Cityscapes, KITTI), re-annotation reveals substantial increases in annotated objects (e.g., up to +60% on KITTI and +40% on COCO), driven primarily by previously unlabeled small, occluded, or densely packed instances. While some differences arise from dataset-specific annotation conventions, we consistently find that missing annotations are the main source of label errors across all datasets. To achieve high data quality, we introduce a scalable annotation pipeline that emphasizes high recall and captures ambiguity through soft labels aggregated from at least 11 annotators per object. The resulting annotations improve coverage and align well with human calibration. We show that benchmark performance is highly sensitive to annotation quality, although model rankings remain largely stable. We introduce two large-scale benchmarks: (i) an uncertainty-aware object detection benchmark, and (ii) a label error detection benchmark grounded in real label errors. We show that current detectors are strongly depended on annotation quality and are misaligned with human perception. Current label error detection methods, which have been shown to perform well on synthetic noise, struggle to achieve high recall and precision on real label errors. Our results highlight the need for future object detection benchmarks to move beyond deterministic annotations toward high-recall, uncertainty-aware evaluation that maximizes valid instances and better reflects real-world ambiguity.

[CV-14] MIST: Multimodal Survival Prediction with Genomic-Guided Histology Attention MICCAI2026

链接: https://arxiv.org/abs/2609.21811
作者: Muhammet Sami Yavuz,Sabri Mustafa Kahya,Richard R. Chen,Jana Lipkova,Benedikt Wiestler
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at the COMPAYL 2026 Workshop on Computational Pathology and Multimodal Data at MICCAI 2026. 11 pages, 2 figures, 4 tables

点击查看摘要

Abstract:Multimodal survival models can combine complementary prognostic information from whole-slide images and genomic profiles, but effective fusion remains challenging amid external cohort shift and computational complexity. To address these challenges, we propose MIST, multimodal survival prediction with genomic-guided histology attention. MIST represents genomic features as tokens and allows them to query compact foundation-model-derived histology context tokens before survival prediction. This design enriches molecular information with histology context rather than merging separately encoded modalities only at the final stage. Training combines discrete-time survival prediction with genomic feature masking, WSI dropout, and paired WSI-genomics contrastive alignment. Across four external evaluations in colon, renal, lung, and glioblastoma cohorts, MIST improves external C-index over standard fusion baselines in the primary comparisons. These results support genomic-guided histology attention as a compact and effective strategy for multimodal oncology outcome prediction. Our code is available at this https URL .

[CV-15] VideoReloc: Long-Term Indoor Video Relocalization against a Kilobyte-Scale Semantic Scene Graph

链接: https://arxiv.org/abs/2609.21804
作者: Qianru Li,Xuyang Chen,Xuqin Wang,Zhenghao Zhang,Hongyi Luo,Tao Wu,Daniel Cremers,Lu Liu,Yanfeng Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 8 pages, 3 figures, 4 tables. Project page: this https URL

点击查看摘要

Abstract:Given a compact semantic scene graph, long-term indoor video relocalization estimates a map-frame trajectory after lighting and furniture changes. Visual methods rely on appearance and become unreliable under these changes; localizing one frame at a time from object classes and geometry instead leaves sparse, ambiguous evidence. We introduce VideoReloc, whose adaptive clips use odometry to gather spatial evidence until object and motion criteria are met, adapting query length to the observed scene. Its run-level decision rechecks conflicting placements using evidence accumulated across connected clips, stabilizing the trajectory beyond adjacent-clip tracking. Hypothesis-first registration proposes poses from object triplets and verifies each using clip-wide object centers and box surfaces. Orientation-aware refinement uses box faces, gravity and wall directions to resolve ambiguity in camera orientation and refine the full pose. This reframes sparse-map relocalization as verification of spatially extended video queries, moving discriminative support from stored appearance to temporal context and permitting a 100 kB map of class-labelled boxes. On RIO10 and ReplicaCAD, the all-frame localization success rate at 1 m/10 ^\circ is 73.5% and 61.1% under causal evaluation, rising to 90.6% and 74.8% with clip closure. The evaluated per-frame scene coordinate regressors reach up to 47.6% and 49.8%, respectively, with maps of 12.6-42 MB. Project page: this https URL

[CV-16] A Principled Approach to Unsupervised Anomaly Detection

链接: https://arxiv.org/abs/2609.21800
作者: James Myles,Matthew Baugh,Johanna P. Müller,Bernhard Kainz,Yingzhen Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 14 pages, 2 figures, 3 tables

点击查看摘要

Abstract:Traditional unsupervised anomaly detection (UAD) methods are designed to flag or localise deviations from a normative distribution, ignoring the underlying generative mechanisms of the anomalies. Yet the nature of an anomaly is often as important as its presence. We reformulate UAD as a Bayesian inverse problem, in which the objective is to infer the most probable corruption responsible for each observation. Our framework yields a probabilistic anomaly score as the energy of the inferred corruption parameters, and serves as a principled recipe for developing new UAD algorithms. We derive several existing methods as instances of the general framework, each corresponding to the same energy score under different modelling choices. Experimentally, we study the framework’s components in a controlled setting, and improve object-class AUROC on the MVTec AD dataset by 2.3% by adapting the underlying corruption model. Finally, we validate the framework on a brain MRI benchmark, achieving strong detection performance while producing estimates of pathology intensity, bias, and geometry. Code is available at this https URL.

[CV-17] PointLAM: Local Attentive Mamba for Efficient Point-based 3D Object Detection ECCV2026

链接: https://arxiv.org/abs/2609.21780
作者: Xuanming Shang,Weijia Zhang,Chao Ma
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to ECCV 2026

点击查看摘要

Abstract:3D object detection from LiDAR point clouds faces a fundamental dilemma: voxel-based methods achieve efficiency at the cost of geometric quantization, while point-based methods preserve fidelity but suffer from prohibitive computational bottlenecks. Specifically, point-based architectures are crippled by slow downsampling strategies (e.g., FPS) and expensive dynamic neighbor queries (e.g., k-NN) coupled with costly continuous interactions. To tackle these systemic inefficiencies, we propose PointLAM, a highly efficient and powerful point-based architecture driven by two synergistic innovations. First, to resolve the downsampling bottleneck, we develop the Laplacian Point Sampler (LPS). LPS employs an implicit discrete Laplacian high-pass filter and Doubly Sorted Sampling to achieve fast, structure-aware foreground preservation. Second, to overcome local modeling latency, we design the Local Hadamard Aggregator (LHA). LHA decouples spatial indexing from feature representation using transient grids, and replaces complex continuous interactions with a Hadamard Gating mechanism for topology-aware, attentive modulation. By coupling this local gating with Bi-Directional Mamba (BDM) layers for global sequence modeling, we formulate the Local Attentive Mamba (LAM) block. Powered by this architecture, PointLAM achieves competitive performance on nuScenes and Waymo for point-based detectors. It rivals highly optimized voxel competitors while requiring a fraction of the computational footprint, demonstrating marked superiority in detecting small instances and handling extreme sparsity. Project page: this https URL.

[CV-18] XCalib Depth-Guided Geometric Optimization for Dense Thermal-Visible Video Registration

链接: https://arxiv.org/abs/2609.21770
作者: Aurelien Godet,Gabriel Jobert,Mauro Dalla Mura
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 10 pages, 6 figures

点击查看摘要

Abstract:Image registration is a vital preprocessing step in multimodal perception tasks, including image fusion, object detection, and semantic segmentation. In Advanced Driver- Assistance Systems (ADAS), spatial misalignment between visible (RGB) and infrared (IR) cameras -caused by non-coincident optical axes and field-of-view differences- introduces non-uniform parallax and visual ghosting. Classical keypoint-based methods are restricted to global homographies that fail under dynamic depth, while unconstrained dense flow algorithms lack structural regularization and suffer from temporal instability. In this paper, we propose XCalib, an unsupervised dense thermal-visible registration framework that bridges this gap. Rather than serving as an absolute metric calibration tool, XCalib leverages virtual pinhole camera parameterization strictly as a geometric constraint space. By optimizing effective relative pose and intrinsics alongside predicted monocular metric depth, XCalib restricts the search space of spatial displacements to physically valid projection geometries. Our key contributions are: (1) a novel registration paradigm that uses camera parameterization as an implicit regularizer for dense cross-modal warping; (2) Normalized Edges Correlation (NEC), a robust structural similarity metric tailored to cross- spectral alignment; and (3) extensive quantitative and qualitative evaluations across public ADAS datasets, demonstrating superior temporal stability and alignment accuracy over unconstrained dense flow baselines.

[CV-19] Beyond Benchmark Scores: Auditing Medical Vision-Language Models for Chest X-Ray Tuberculosis Screening

链接: https://arxiv.org/abs/2609.21763
作者: Mushir Akhtar,M. Tanveer,Mohd. Arshad
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 27 pages, 7 figures, and 21 tables; includes extended methods, statistical analyses, and robustness evaluations

点击查看摘要

Abstract:A medical model’s benchmark score does not establish that the same conclusion holds under a different evaluation. This study tests whether claims about model ranking, score reliability and screening performance survive changes in cohort, prompt, negative spectrum, specified prevalence and operating threshold. We audit three medical vision-language models (BioMedCLIP, CheXficient, and MedSigLIP) and a general-domain OpenCLIP comparator on 12,200 chest radiograph records from four datasets (Montgomery, Shenzhen, TBX11K, and VinDr-CXR). Five fixed prompt families yield 244,000 model–image–prompt scores. No model leads every cohort and reliability criterion. Prompt-family changes alter AUROC in 21 of 48 multiplicity-controlled comparisons. Replacing healthy controls with sick non-tuberculosis controls reduces AUROC by 0.075–0.306 across all four models. On VinDr-CXR, the three medical models distinguish tuberculosis from no-finding controls substantially better than from pneumonia or lung tumor; their AUROC point estimates for both named diseases fall below 0.5. CheXficient has documented VinDr-CXR pretraining exposure, which limits the interpretation of its results. Thresholds chosen for 95% sensitivity on TBX11K training retain that constraint by point estimate in only four of sixteen target evaluations. A five-seed supervised source model reaches 0.999 AUROC on TBX11K validation but 0.629 on each of two external cohorts. Conservative exclusion of perceptual-overlap candidates narrows this gap without closing it. These retrospective, single-task results show that discrimination, score reliability and threshold retention support different portability claims. Evidence for chest X-ray tuberculosis screening should identify the complete evaluation specification rather than attribute clinical portability to a checkpoint alone.

[CV-20] SFVO: Decoupled Confidence-Guided Stereo-Flow Visual Odometry with Bidirectional PnP

链接: https://arxiv.org/abs/2609.21754
作者: Kai Zhang,Guoyang Zhao,Jun Ma
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Deep learning-based visual odometry (VO) has achieved significant progress, yet most existing methods focus on a monocular approach, which suffers from scale ambiguity. Stereo VO provides real metric by its nature, but remains less studied in deep learning VO due to its high computational cost and modeling complexity. Recent advances in stereo matching and optical flow estimation have made dense visual correspondence increasingly accurate and reliable, but their complementary geometric information has not been fully exploited for VO. In this paper, we present SFVO, a correspondence-driven stereo VO framework that directly builds upon pretrained stereo matching and optical flow models. SFVO exploits pretrained stereo matching and optical flow models to estimate stereo and temporal correspondences. Instead of learning pose directly from images, SFVO maps learned correspondences into geometric constraints and predicts which points are trustworthy. To improve the reliability of visual correspondence-based geometric constraints, we introduce decoupled confidence maps for rotation and translation. This design better aligns the characteristics of visual correspondence and 6-DoF transformations. Extensive experiments on outdoor and indoor datasets demonstrate that SFVO achieves robust and accurate pose estimation with strong generalization capability. The code will be released.

[CV-21] Balanced Prompt Adaptation against Entropy-Induced Collapse for Test-Time Binary Segmentation

链接: https://arxiv.org/abs/2609.21743
作者: Zhengshan Wang,Joshua Charles Webster-Ford,Yifei Tian,Xinxin Wang,Long Chen,Weiping Ding
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Entropy minimization is a standard objective for test-time adaptation (TTA), but it can fail in imbalanced binary segmentation. Unlike image classification, dense segmentation aggregates thousands of pixel predictions, allowing the larger predicted class to dominate the update, pull minority predictions toward itself, and produce a degenerate mask as predictions saturate and their entropy gradients vanish. We theoretically establish this collapse in a shared-shift model. This analysis motivates Balanced-Anchor Prompt Adaptation (BAPA), which combines two complementary modules. The Class-Balanced Anchors (CBA) module selects high-confidence anchors separately from each predicted class and gives foreground and background equal total loss weight, preventing the larger region from dominating the update. Dynamic Prompt Adaptation (DPA) refreshes these anchors after each prediction update and optimizes only text-side prompt residuals while keeping the vision-language encoders frozen. This prompt-only update refines the foreground-background decision boundary without altering the pretrained dense visual representation. Across experiments from four domains, BAPA achieves the highest mean Dice among the evaluated methods. Factorized ablations further validate the complementary roles of CBA and DPA, supporting balanced prompt adaptation as an effective alternative to entropy minimization for test-time binary segmentation.

[CV-22] ZYT-World: A Real-Time Controllable World Model for Closed-Loop Autonomous-Driving Simulation

链接: https://arxiv.org/abs/2609.21712
作者: Boni Hu,Xiong Wei,Haoming Huang,Yong Huang,Chenbo Wang,Yi Yang,Jiancheng Wang,Ruicheng Zhu,Zhimin Yang,Guanglai Liu,Qiaowan Jin,Dongzhuo Wang,Haiwei Kuang,Jiajun Fan,Yue Wu,Jiaxin Wei,Hao Sun,Feihong Yan,Wei Bi,Kaixuan Wang,Zichao Guo,Xiaozhi Chen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Technical Report. Videos and additional results are available at this http URL

点击查看摘要

Abstract:Generative world models offer controllable and repeatable closed-loop simulation for end-to-end and vision-language-action driving policies, but production deployment exposes three unresolved requirements: faithfully reproducing a mixed fisheye-pinhole rig at native resolutions; reconciling causal, per-timestep interaction with long-horizon stability and low latency; and preserving scene identity when a location is revisited. We present ZYT-World, a single architecture that natively generates four fisheye views with field of view 180° and three pinhole views. Projection-specific Plucker adapters encode camera geometry, ego-motion adaptive layer normalization provides global motion control, and a lightweight pixel-aligned layout conditions traffic participants and signals through instance-level boxes, headings and colors. Heterogeneous training combines full-rig geometric coverage with high-resolution detail. Teacher forcing, causal consistency distillation, self-rollout distribution matching distillation, and RigCritic transform a 40-step bidirectional teacher into a one-step, per-latent streaming generator, with RigCritic evaluating the seven-view rig jointly. A 19M-parameter variational autoencoder decoder (TinyVAE), W8A8 quantization, and our inference engine reduce decoding, backbone, and incremental-execution costs, respectively. Finally, cross-trajectory pairs derived from real captures train a plug-in implicit-memory module that preserves place-specific evidence. On the internal multi-view test set, the one-step model retains more than 90% of the teacher’s PSNR and SSIM, while FID, FVD, and LPIPS stay within 11% of the teacher. Under the generator-only timing in Figure 2, it is 107.7 times faster than the 40-step bidirectional teacher. TinyVAE decodes 59.8 times faster than Wan. 30s rollouts and cross-trajectory revisits show the intended long-horizon and memory behavior.

[CV-23] SignGPT : Toward LLM -Mediated Sign Language Interaction through Gloss-Free Translation and Generation

链接: https://arxiv.org/abs/2609.21709
作者: Ronghui Li,Jun Dong,Zhongyuan Hu,Zunnan Xu,Jun Zhou,Liyuan Chen,Shuoling Liu,Jiangpeng Yan,Jie Guo,Xiu Li,Linchao Bao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Large language models (LLMs) provide limited support for sign language interaction. Unifying sign language translation (SLT) and generation (SLG) to enable sign language as both input and output can reduce switching between separate models during sign-text interaction. We present SignGPT, a unified, pose-based framework for gloss-free SLT and SLG. SignGPT integrates part-aware hierarchical representations of body, hand, and facial motion into a shared language model and employs asymmetric multi-token prediction and progressive training for bidirectional modeling. We evaluate SignGPT on How2Sign (ASL) and Phoenix-2014T (DGS) through benchmark comparisons, qualitative analyses, and component ablations. An exploratory study with 12 Deaf ASL signers assesses an LLM-mediated sign-to-sign response pipeline, highlighting the potential of unified modeling to support sign language conversation (SLC).

[CV-24] Diffusion-Based Tumor Inpainting for Renal Segmentation under Clinical Data Scarcity

链接: https://arxiv.org/abs/2609.21698
作者: Ekaterina Sedykh,Salme Ussanov,Dmytro Fedorenko,Dmytro Fishman
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Deep learning segmentation of renal tumors requires large annotated datasets, yet clinical deployments typically offer only a handful of tumor-positive cases from the target site. We propose a diffusion-based inpainting framework that synthesizes anatomically plausible renal tumors within healthy CT scans, requiring no additional annotation, and provide the first systematic comparison of 2D, 2.5D, and full 3D (MAISI) synthesis strategies for this task. Training the diffusion model on public data (KiTS23, KIRC) and evaluating nnU-Net segmentation on a internal cohort across three low-data regimes, we find that 2.5D and 3D augmentation substantially reduce false positives (from \sim 18–20% to \sim 3–6%) while maintaining Dice, whereas 2D provides no consistent benefit. Crucially, the proposed 2.5D method matches full 3D synthesis on every metric at substantially lower computational cost, indicating that local volumetric consistency alone is sufficient for effective augmentation in data- and resource-scarce clinical settings.

[CV-25] DRT: Dense Reasoning Trace for Efficient and Grounded Multimodal Reasoning

链接: https://arxiv.org/abs/2609.21675
作者: Wan Xu,Yuanfan Guo,Kevin Han,LaLa Chen,Wangmeng Zuo
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Despite the remarkable progress in Multimodal Large Language Models (MLLMs), prevailing Chain-of-Thought (CoT) paradigms remain confined to the natural-language expression space. Consequently, they inherently incur excessive linguistic overhead, leading to information dilution and weak visual grounding. To address this challenge, we propose Dense Reasoning Trace (DRT), a paradigm that departs from natural-language-centered CoT by expressing reasoning as compact structured traces, which include concise intermediate states with symbolic connectors and disentangle visual observations from logical deductions. First, we introduce the Dense Trace Initialization to internalize the DRT reasoning mode into the model, substantially improving token efficiency while preserving visual evidence. To further enable the model to faithfully capture the logical relations within traces, we propose the Trace-Grounded Reinforcement Learning framework, which builds reference traces through a tri-perspective verification pipeline and employs Trace-Grounded GRPO with structured rewards, encouraging the model to generate concise DRT-style traces with reduced hallucination and stronger logical grounding. Extensive experiments on challenging reasoning benchmarks show that DRT achieves 5.5 \times token efficiency improvement while improving 1.3 accuracy points over the Qwen3-VL baseline. These findings suggest that complex multimodal reasoning may not require verbose natural-language traces, opening a more efficient path for next-generation MLLMs. Our code and data are available at: this https URL

[CV-26] Extending Decoupled Attention to Dense Prediction and Masked Training for Multi-Channel Images

链接: https://arxiv.org/abs/2609.21629
作者: Umar Marikkar,Sameed Husain,Muhammad Awais,Sara Atito
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Multi-Channel imaging (MCI) data differs fundamentally from natural images, as each channel records a semantically distinct signal rather than a colour band. To adapt vision encoders to MCI data, Multi-Channel Vision Transformers (MC-ViTs) tokenize each channel independently and concatenate the resulting tokens into one sequence, and the channel count is no longer fixed by the architecture. Self-attention is then computed across all channel-patch tokens with no restriction on which channels attend to which, which dilutes the features of individual channels. The Decoupled Vision Transformer (DC-ViT) regulates this by separating updates computed within a channel from updates computed across channels, and by forming a representation per channel before the channels are combined. Its formulation, however, pairs tokens by spatial position, and thus requires the same visible tokens in every channel. Correspondence under independent per-channel masking is recovered by solving a linear assignment between the retained patches of each channel, which allows decoupled attention to be combined with current masked multi-channel training in its standard configuration rather than a restricted one. Across three classification and three segmentation benchmarks spanning fluorescence microscopy, imaging mass cytometry and satellite imaging, including dense prediction at high channel counts, the resulting formulation outperforms the strongest MC-ViT baseline.

[CV-27] Detection is solved delineation is not: what governs tooth segmentation on panoramic radiographs

链接: https://arxiv.org/abs/2609.21628
作者: Muhammad Rehan,Moaz Amjad,Syed Danial Ahmed,Mariam Adnan,Haider Ali
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 15 pages, 6 figures, 5 tables. Code: this https URL

点击查看摘要

Abstract:Automatic tooth segmentation and FDI numbering on panoramic radiographs underpins computer-assisted dental diagnosis, yet which factors govern performance remains unclear. We assemble a corpus of 1,422 panoramic radiographs containing 42,142 expert-delineated tooth polygons across the 32-class FDI taxonomy, annotated by 30 dental practitioners and independently reviewed by two others, and use it to isolate input resolution, architecture and anatomical priors under a single evaluation protocol. First, resolution dominates: across a controlled 640/1024/1280 ablation, mask mAP50-95 rises 0.656 - 0.710 - 0.717 while mAP50 stays flat at ~0.982. Both gains are significant under a paired bootstrap over images (p 0.001, p = 0.024); neither mAP50 change is distinguishable from zero. Added resolution buys boundary precision, not detection. Second, architecture is nearly irrelevant in-domain: a query-based transformer with 2.1x the parameters is statistically equivalent to a one-stage detector (95% CI [-0.0064, +0.0064]), only marginally better under domain shift, 5.5x slower on CPU and not executable under standard ONNX runtimes. Third, three targeted interventions fail: a LoRA-adapted self-supervised encoder underperforms, a promptable foundation segmenter degrades masks by 39%, and globally optimal anatomical label assignment yields +0.0007 despite correcting a constraint violated in 40% of out-of-domain predictions. Zero-shot transfer to an independent multi-centre cohort, verified overlap-free, costs 62% of mask mAP50-95 but only 18% of mAP50, reproducing the dissociation. Decomposing masks along the tooth axis localises the residual error to the apical third. Boundary precision is therefore the binding constraint, and effort is better directed at resolution and acquisition diversity than at architectural novelty. Comments: 15 pages, 6 figures, 5 tables. Code: this https URL Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) Cite as: arXiv:2609.21628 [cs.CV] (or arXiv:2609.21628v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2609.21628 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[CV-28] HAT: Hypothesis-Anchored Tracking for Video Monocular Spacecraft Pose Estimation

链接: https://arxiv.org/abs/2609.21597
作者: André Lopo,Atabak Dehban,Rodrigo Ventura
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 8 pages, 3 figures, 4 tables

点击查看摘要

Abstract:Monocular 6-DoF pose estimation of non-cooperative targets is important for on-orbit servicing and debris removal. A single-image estimator can confuse near-symmetric spacecraft orientations, and tracking can preserve an incorrect pose. We present Hypothesis-Anchored Tracking (HAT), a causal framework that uses inter-frame motion to select among competing CAD-based pose hypotheses before alignment and fusion. Rather than independently choosing the highest-scoring hypothesis in each image, HAT retains competing orientation histories and selects a pose to anchor the relative trajectory estimated by monocular SLAM. Sparse anchors and pose fusion provide per-frame estimates after initialization without revising past outputs. The method requires only a calibrated RGB sequence, a metric CAD model, and target image regions, which can be supplied by detection or segmentation. The pretrained pose and SLAM networks require no target-specific training or fine-tuning. We evaluate two versions, Mega-HAT and Pico-HAT, using MegaPose and PicoPose, on SPARK-2024, SwissCube and SHIRT, with YCB-Video assessing performance outside the space domain. Using one temporal configuration per method, the arithmetic means of the four dataset-wise comparisons show 9.4% lower mean pose error and 3.76 times the sustained input FPS for Mega-HAT relative to independent MegaPose, and 23.9% lower mean pose error and 2.42 times the FPS for Pico-HAT relative to independent PicoPose. Mega-HAT ablations on SPARK and an offline reference examine component contributions and the effect of revising past estimates.

[CV-29] A benchmark dataset and baseline methods for four-dimensional STEM diffraction patterns

链接: https://arxiv.org/abs/2609.21593
作者: Yuyan Guan,Haoran Zhang,Zian Mao,Antong Yang,Caifei Li,Jialong Wang,Chuying Ouyang,Hong Wang,Xiaoqin Zeng,Yujun Xie
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 16 pages, 5 figures. Data and trained model weights: this https URL . Code: this https URL

点击查看摘要

Abstract:Four-dimensional scanning transmission electron microscopy (4D-STEM) records a two-dimensional diffraction pattern at each electron-probe position, yielding spatially resolved reciprocal-space information but large, heterogeneous data volumes. Here we describe 4D-ImageNet, a collection of 174,000 diffraction patterns comprising 145,000 experimental patterns selected from 29 acquisitions and 29,000 multislice simulations. The experimental data cover acquisition-level labels for Ag, Au, mixed Au-Ag, CoO, Pd and ZnO specimens across multiple fields of view, scan dimensions, camera lengths and exposure times. Each acquisition contributes 5,000 quality-ranked patterns with source scan coordinates and acquisition metadata. A set-prediction detector provides model-derived pseudo-labels for the direct-beam position and Bragg-disk centres, with a confidence score for each disk. The simulation data cover 13 crystal structures and include Euler rotations, reciprocal-space sampling and approximate low-index beam directions. A grouped mixed-domain masked-reconstruction benchmark is provided to assess leakage-resistant loading and evaluation across experimental and simulated data. The dataset is intended for representation learning, disk detection, diffraction-pattern retrieval, orientation analysis and simulation-to-experiment studies.

[CV-30] From Retrieval to Recognition:How Vision–Language Models Become OCR Specialists

链接: https://arxiv.org/abs/2609.21543
作者: Yuanxiang Huangfu,Hanmeng Zhong,Linqing Chen,Jeffrey Tiong Jee Hui
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Does a general vision–language model acquire specialized OCR ability by developing a new reading circuit or by reusing an existing mechanism? We address this question in the setting of full-sequence OCR, rather than local-answer retrieval. Using an evidence-grounded protocol with held-out causal interventions, we identify sparse and stable OCR-head sets in GLM-OCR, MinerU2.5, and PaddleOCR-VL-1.6. We then investigate the mechanistic origin of these OCR heads by comparing them with independently identified textual retrieval/copy heads in general VLMs. Across two general VLMs, visual OCR heads strongly overlap independently identified textual retrieval/copy heads, yielding untuned top-20 intersections of 73.3% and all-head Spearman correlations of 0.677-0.886. The overlap and causal interventions suggest that full-sequence OCR operates as dense sequential multimodal copy-and-paste, repeatedly retrieving visual evidence and routing it to the current output position. Finally, we examine how this shared circuit changes as a general VLM becomes an OCR specialist. Matched base-to-specialized comparisons show that OCR specialization largely preserves head identity, retaining 17-20 of the top 20 heads per task with all-head rank correlations of 0.874-0.942, while redistributing their functional and causal strengths.

[CV-31] Purification and Regulation: Comorbidity-Aware Multi-Label Few-Shot Learning for Medical Image Classification

链接: https://arxiv.org/abs/2609.21541
作者: Ying-Chih Lin,Po-Chih Kuo,Yong-Sheng Chen
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Multi-label few-shot learning (MLFSL) remains a significant challenge in medical image analysis (MIA). Current metric-based meta-learning methods face two critical limitations in MIA. First, conventional prototype generation often entangles irrelevant disease information, leading to contaminated prototypes and degraded performance. Second, prior studies typically enforce inter-class separability in embedding space, largely neglecting the inherent correlations among diseases. To overcome these challenges, we propose Prototype Purification and Regulation (PPR), a novel MLFSL framework for MIA. PPR first performs prototype purification by leveraging sample-level comorbidity scores to emphasize disease-specific features, producing purified prototypes that better characterize each disease. Building upon these purified prototypes, PPR further addresses the underexplored problem of inter-class prototype distance in MIA by incorporating disease-level comorbidity statistics to adaptively regulate inter-class similarity, forming a comorbidity-aware embedding space. Overall, PPR sequentially enables the model to capture pure disease features and inter-class relationships for reliable MLFSL in MIA. Extensive experiments across four chest X-ray benchmark datasets, including cross-domain evaluation, show that PPR consistently outperforms state-of-the-art methods, significantly improving disease detection while demonstrating robust generalization and clinical applicability.

[CV-32] Refine Then Fusion: Training-Free 3D Point Cloud Adaptation with Priority Refinement and Multi-Modal Knowledge Fusion

链接: https://arxiv.org/abs/2609.21522
作者: Hang Cheng,Yan Chen,Mingyu Fan,Long Zeng
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Recent pre-trained foundation models provide rich multi-modal priors for downstream 3D vision tasks. However, the effectiveness of these representations in few-shot scenarios is limited by two fundamental challenges: High-dimensional features often contain substantial channel redundancy and task-irrelevant noise, while the reliability of different modalities varies across samples. Consequently, direct aggregation of heterogeneous representations overlooks sample-dependent modality reliability and may obscure the discriminative cues essential. To address these limitations, we propose Refine Then Fusion(RTF), a training-free framework for few-shot 3D recognition. RTF first identifies discriminative feature channels by jointly modeling inter-class similarity and intra-class stability, thereby decoupling domain-specific knowledge refinement from the cached representations of pre-trained models. It then introduces a reliability-aware fusion mechanism that estimates sample-wise modality reliability from the distribution shifts induced by feature refinement, enabling adaptive aggregation of multi-modal representations. Furthermore, RTF constructs a memory cache that integrates instance-level support features with class-level prototypes to infer query labels. Extensive experiments on five benchmarks demonstrate that RTF consistently outperforms single-modal baselines, partial-fusion variants, and existing lightweight adaptation methods, achieving state-of-the-art few-shot 3D recognition performance without gradient optimization, additional training data, auxiliary training, or parameter updates.

[CV-33] VidOmni-Bench: A Benchmark for Fine-Grained Video Understanding via Spatio-Temporal Event Verification across Complexity and Duration

链接: https://arxiv.org/abs/2609.21521
作者: Changbeen Kim,Junwon Chang,Kipyo Kim,Risa Shinoda,Kuniaki Saito,Donghyun Kim
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:While Video Large Language Models (Video-LLMs) have recently demonstrated strong performance, reliably evaluating their fine-grained video understanding remains challenging. Existing benchmarks often rely on question answering or ground-truth caption matching, where models may succeed through superficial cues and incomplete annotations. To this end, we introduce VidOmni-Bench, a benchmark that requires models to verify whether each event in dense video captions is supported by the video. VidOmni-Bench consists of 500 videos spanning five complexity types and diverse durations from 4 seconds to 90 minutes. After collecting videos along these axes, we use diverse Video-LLMs to generate dense captions and obtain human-verified sentence-level labels, where sentences containing incorrect events serve as hard negatives for evaluation. Our experiments on VidOmni-Bench reveal three key findings: (i) Video-LLMs frequently generate hallucinated descriptions in dense video captioning; (ii) they also struggle as verifiers, failing to reliably detect plausible but incorrect event descriptions; and (iii) model weaknesses vary across video complexity and duration, revealing diverse, model-specific bottlenecks in current Video-LLMs.

[CV-34] 2D GauSS-MI: Efficient Active Scene Reconstruction with Balanced Visual and Geometric Quality

链接: https://arxiv.org/abs/2609.21516
作者: Yuhan Xie,Jia Pan
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Active reconstruction requires efficient active view selection to achieve high-quality reconstruction within limited onboard computational resources. Existing methods face challenges in adequately balancing visual and geometric quality with the computational efficiency required for real-time operation. In this work, we present an active reconstruction framework based on 2D Gaussian Splatting (2DGS). We develop an efficient online 2DGS mapping pipeline for incremental RGB-D observations and introduce a probabilistic reliability model that characterizes the view-dependent reconstruction quality of individual 2D Gaussian splats. Building on this model, we formulate 2D Gaussian Splatting Shannon Mutual Information (2D GauSS-MI), a mutual-information-based metric that exploits the explicit surface orientation of 2DGS to evaluate the expected information gain of candidate views. The proposed metric enables active view selection to account for both visual and geometric reconstruction quality. We evaluate the proposed system against three state-of-the-art baselines on eight Replica scenes. Experimental results demonstrate that our method achieves a favorable balance between visual and geometric reconstruction quality with substantially lower computational cost and competitive model storage.

[CV-35] 2nd Place Solution to the HANDS 2026 Workshop Challenge-Dexterous Grasp Motion Track: Single-Shot Trajectory Warping for Grasp Motion Generation

链接: https://arxiv.org/abs/2609.21511
作者: Muneeb A. Khan,Woojin Kim,Shinwoo Kim,Muhammad Munsif,Binod Bhattarai,Seungryul Baek
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:This report describes our 2nd place solution to the HANDS 2026 workshop challenge (Dexterous Grasp Motion track) in conjunction with ECCV 2026. In this challenge, we address grasp motion generation for the 12-DoF LinkerHand O6, aiming to produce physically plausible reach-and-lift trajectories for unseen objects from randomized initial hand poses in simulation. This task is particularly challenging because each grasp requires a per-step policy to make approximately 70 twelve-dimensional decisions, with errors accumulating over time, while test objects and physical dynamics may differ from those encountered during training. To address these challenges, we propose editing a single successful GraspM3 demonstration instead of generating the motion step by step: a policy observes the object once and outputs a 12-D warp of the demonstration, which is then replayed open-loop. Moreover, we train the warp policy with one-step PPO over all 4,824 training objects in parallel. As a result, our method achieved success rates of 94.61% on the easy track, the highest of all submissions, and 57.18% on the hard track of the private test set.

[CV-36] Adaptive World Memory 3D Foundation Model for Scalable 3D Mapping Localization and Rendering

链接: https://arxiv.org/abs/2609.21502
作者: Tianchen Deng,Guole Shen,Yilin Shen,Wenhua Wu,Yilin Fang,Ziqi Ma,Tianjun Zhang,Shenghai Yuan,Wolfram Burgard,Hesheng Wang
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Recent 3D foundation models enable generalizable geometric reasoning from RGB images but remain limited in persistent memory, scalability, and renderable scene modeling. We present a memory-centric 3D foundation model for scalable robotic localization, reconstruction, and Gaussian rendering. Its core is an adaptive world memory mechanism that combines transformer-based gated updates with test-time temporal-spatial regulation. Learned gates control recurrent memory propagation, while temporal state evolution and spatial observation-state consistency regulate token-wise updates and forgetting over long image sequences. To support large-scale mapping, we organize memory into local submaps and integrate progressive mapping and tracking, loop closure, and SL(4)-based global refinement to maintain local accuracy and global consistency. A Gaussian reconstruction head decodes memory-enhanced features into renderable primitives, unifying camera pose estimation, dense point-cloud reconstruction, and photorealistic rendering within a single model. Experiments on public benchmarks and self-collected datasets from diverse robotic platforms demonstrate improved trajectory accuracy, reconstruction completeness, and rendering quality over existing 3D foundation reconstruction and SLAM baselines. These results support adaptive memory as a foundation for persistent robotic world modeling. The dataset and code will be made publicly available at \hrefthis https URLthis https URL.

[CV-37] VoxelTTO: Voxel-Aligned Feed-Forward 3D Gaussian Splatting with Test-Time Optimization

链接: https://arxiv.org/abs/2609.21498
作者: Yibin Zhao,Yihan Pan,Yangwen Li,Jun Nan,Jianjun Yi
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Recent feed-forward 3D Gaussian Splatting (3DGS) methods typically regress pixel-aligned Gaussian primitives, often causing excessive overlap and artifacts, while inaccuracies in predicted camera poses can lead to misalignment in novel-view synthesis (NVS). We present VoxelTTO, a feed-forward framework for reconstructing geometrically accurate 3DGS scenes from an arbitrary number of images and optional camera parameters. VoxelTTO aggregates dense image features into a global voxel representation and decodes Gaussians from voxel features, breaking the pixel-to-Gaussian correspondence. To exploit known camera parameters while keeping the pretrained visual foundation model (VFM) parameters frozen, we introduce test-time optimization (TTO) that adapts lightweight LoRA modules using pose supervision. We further replace vanilla 3DGS rasterization with stochastic solid volume rendering during training and inference, improving geometric fidelity. Training updates only the voxel-aligned Gaussian reconstruction modules, requiring 80 GPU hours. Experiments on Replica, Tanks and Temples, and DTU demonstrate improved RGB-D NVS and camera-pose estimation relative to prior methods.

[CV-38] OpenSAL360: Open-Source Crowdsourcing Platform for Omnidirectional Video Saliency Collection ACM-MM2026

链接: https://arxiv.org/abs/2609.21480
作者: Alexey Bryncev,Andrey Moskalenko,Kira Shilovskaya,Ivan Kosmynin,Dmitriy Vatolin
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by ACM MM 2026

点击查看摘要

Abstract:Omnidirectional video saliency prediction plays an important role in many immersive multimedia applications, including viewport-adaptive streaming and compression, foveated rendering, mesh simplification, perceptual quality assessment. Yet progress in this area remains constrained by the cost and complexity of collecting eye-tracking data with VR headsets, which makes large-scale dataset creation difficult to extend. We present OpenSAL360, the first open-source platform for scalable, low-cost 360° video saliency collection. Unlike conventional VR-based protocols, it requires only a standard screen, mouse, and internet connection, enabling parallel saliency data collection from common crowdsourcing assessors without specialized hardware. We validate our collection protocol against seven well-established VR eye-tracking datasets and conduct ablation studies on key interface, pre-, and post-processing parameters. To demonstrate the effectiveness and scalability of the proposed methodology, we collect and publicly release a saliency dataset covering 500 omnidirectional videos annotated by 2,000+ crowdsourcing assessors, making it, to the best of our knowledge, the largest dataset in this field. We make OpenSAL360 publicly available at this https URL.

[CV-39] MT-WAM: Reorienting the One-Pass Predictive Representation Toward Action Generation

链接: https://arxiv.org/abs/2609.21474
作者: Yiguang Yang,Jiankun Peng,Xiaoming Wang,Yiran Zhang,Zhibo Fang
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Fast-WAM shows that video-action co-training improves control without generating future video at inference, making the representation from a single video diffusion Transformer forward central to action generation. However, future-observation prediction does not explicitly prioritize the future dynamics and visual structure needed for control. We present MT-WAM, which retains the original training objectives and adds complementary supervision for future two-dimensional point trajectories and visual features. A lightweight dual-stream branch copied from the video backbone’s final blocks provides target-specific processing, while a structured attention mask prevents cross-stream attention. Motion-stream tokens supply additional dynamics conditions to the action expert. Future visual-feature prediction provides supervision in a feature space that captures object and spatial structure. This supervision trains the video backbone to provide more informative visual context for action generation under changing visual conditions, without adding visual-feature-stream tokens to action conditioning. At inference, MT-WAM uses video and motion caches computed once per replan and skips future-video prediction. Without additional embodied policy pretraining, MT-WAM achieves 98.2% success on LIBERO and 73.7% on LIBERO-Plus, exceeding Fast-WAM by 23.8 percentage points on the latter. On RoboTwin 2.0 Clean2Rand, Random success increases from 6.30% to 19.40%; across four real-world tasks, average success increases from 67.0% to 77.8%.

[CV-40] SkillIR: Evolving Scene-Aware Skills for Agent ic Image Restoration

链接: https://arxiv.org/abs/2609.21468
作者: Jie Shao,Shengkai Hu,Xu Zhang,Beihang Song,Yongcheng Jing,Xu Wu,Jun Wan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:This paper studies agentic image restoration, in which multimodal agents coordinate specialized restoration tools to recover images affected by complex degradations. Existing restoration agents often derive complete tool-use plans from the original degraded image or retrieve previously successful trajectories, providing limited support for adapting individual actions to evolving intermediate restoration states. We find that accepted tool executions can change the residual degradation state and, consequently, the applicability of subsequent tools. To address this issue, we propose SkillIR, a skill-guided framework that represents restoration experience as degradation-centered action evidence rather than complete tool-use trajectories. SkillIR consolidates context-dependent action outcomes into scene-aware restoration skills that characterize applicable conditions, expected effects, and attributable failure cases. Instead of prescribing a complete restoration plan, the retrieved skills guide one bounded action at a time within a verified residual-state loop: each tool output is treated as a candidate, committed only after transition verification, and followed by reassessment of the active residual degradations. After each rollout, the resulting evidence is used to create, refine, or patch dynamic skills, enabling accumulated restoration experience to improve decision-making for subsequent inputs. Experiments on synthetic and real-world multi-degradation datasets demonstrate that SkillIR improves restoration quality and enables more reliable and effective tool use.

[CV-41] PSEE: Progressive Sensor Event Expansion for Point-Supervised Temporal Action Localization

链接: https://arxiv.org/abs/2609.21462
作者: Jiaxi Yin,Ge Wang,Han Ding,Fei Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Temporal action localization (TAL) in wearable sensor streams identifies action classes and temporal boundaries, enabling finer-grained activity understanding than conventional action recognition. However, training typically requires costly start–end annotations for every action instance. To reduce this burden, we study point-supervised TAL, where each instance is labeled with only one timestamp and its class. We propose Progressive Sensor Event Expansion (PSEE), which combines semantic activations, sensor-specific transition evidence, and adaptive temporal ownership to recover point-supervised pseudo segments. These segments supervise standard TAL detectors without modifying their inference procedures. Cross-subject experiments on four inertial-sensing benchmarks demonstrate improved pseudo-boundary quality over adapted point-supervised baselines, compatibility with different TAL detectors, and robustness to point sampling. Code is available at this https URL.

[CV-42] CompAdapt: Adaptable Composite Motion Modeling for Physics-Consistent Text-to-Video Generation NEURIPS2026

链接: https://arxiv.org/abs/2609.21455
作者: Haoran Qin(1),Renlong Wu(1),Tianyu Huang(1),Yukang Ding(2),Hui Li(1),Wangmeng Zuo(1) ((1) Harbin Institute of Technology, China, (2) Taobao, Alibaba Group, China)
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 23 pages, 4 figures. Submitted to the 40th Conference on Neural Information Processing Systems (NeurIPS 2026). Project page: this https URL

点击查看摘要

Abstract:While diffusion-based text-to-video (T2V) models have demonstrated impressive capability in generating realistic and temporally coherent videos, they often fail to respect fundamental physical dynamics. Although recent physics-constrained methods incorporate explicit dynamics priors to improve physical plausibility, they remain limited to simple single-type motions, depend on manually specified parameters, and struggle to generalize to unseen physical laws. In this work, we propose CompAdapt, a physics-consistent T2V framework for adaptable generation across complex real-world scenarios. It extends neural dynamics modeling beyond single-type motions to encompass composite physical behaviors, including coupled motions, multi-stage transitions, and multi-object collisions. Furthermore, CompAdapt translates natural language prompts into structured physical semantics, enabling end-to-end specification of motion types, temporal relations, and initial physical parameters. To generalize to novel physical environments, CompAdapt introduces dynamics-aware prior matching, achieving one-shot adaptation without retraining the core dynamics module. In addition, a physics-aware latent feature fusion module improves visual fidelity under fast and complex motion. Experiments on physics-focused T2V benchmarks demonstrate that CompAdapt improves physical consistency over both general T2V models and physics-constrained baselines, while preserving high visual quality and adaptability to unseen dynamics. The project page is available at this https URL .

[CV-43] ME-Dex 1.0: Bringing Heterogeneous Tactile Sensing into World Action Modeling

链接: https://arxiv.org/abs/2609.21449
作者: Xuancheng Zhang,Xuetao Liu,Qianying Tang,Jizhe Wang,Zhijing Cheng,Bochen Lin,Haoran Wen,Ming Li,Kun Zhan,Yu Liu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:World Action Models bring the predictive capabilities of video models into robot action generation, providing a rich foundation for modeling future visual states. Tactile sensing complements this foundation with direct measurements of physical interaction. Some existing methods use tactile features as conditioning inputs without jointly predicting future tactile states, visual observations, and actions. Our key insight is that tactile signals, like video, provide observations of the evolving world state and should be modeled as future observations alongside video. We present ME-Dex-1.0 (MachEmbodied-Dex-1.0), a unified World Action Tactile Model for joint visual, tactile, and action learning. ME-Dex-1.0 adopts a Mixture-of-Transformers architecture comprising a Video Expert, a Tactile Expert, and an Action Expert, all trained with flow matching. We use shared attention connects the experts in intermediate layers, allowing action generation to draw on learned representations of visual and tactile dynamics during joint denoising. To support multi-source heterogeneous tactile inputs, a Canonical Hand Model and a Unified Tactile Autoencoder map tactile observations from different embodiments and sensing layouts into shared spatial and latent spaces. To address the limited availability of paired visual, tactile, and action data, we develop the Agentic Tactile Data Engine, an agent-based data production platform. It supplements RoboTwin and DexJoCo with tactile data recorded directly from force sensors during trajectory replay in simulation. Experiments on the RoboTwin, DexJoCo, and ManiFeel simulation platforms, together with real robot evaluations, demonstrate improved manipulation performance using both grippers and dexterous hands equipped with tactile sensing.

[CV-44] hink Locally Refine Globally for Memory-Efficient 3D Reconstruction

链接: https://arxiv.org/abs/2609.21437
作者: Jingke Zhou,Chenhang Ma,Zhizhou Zhong,Mingkai Liu,Zhuang Zhou,Yicheng ji,Binghua Su,Bo Cai,Xianliang Huang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 9 pages,4 figures

点击查看摘要

Abstract:We propose LoG-VGGT, a memory-efficient framework for long-sequence 3D reconstruction that balances local temporal modeling with global camera consistency. Instead of relying on full global attention, our method introduces cross-window attention at a small subset of transformer blocks, enabling effective information propagation across adjacent temporal windows while keeping memory usage bounded. To mitigate long-term pose drift, we further design a global camera consistency refinement module, where camera tokens interact with compact register tokens via cross-attention to enforce scene-level constraints across the entire sequence. This design enables joint optimization of camera representations and significantly improves long-horizon pose stability without incurring the high cost of sequence-wide attention. Extensive experiments demonstrate that LoG-VGGT achieves improved depth accuracy and robust camera pose estimation across multiple long-sequence benchmarks, while delivering competitive streaming reconstruction performance.

[CV-45] P3-SAM: SAM with Perceptual Parallel Prompt for Few-Shot Strip Steel Surface Defect Segmentation ICME2026

链接: https://arxiv.org/abs/2609.21424
作者: Qian Xu,Hang Xiong,Anpeng Wang,Sam Kwong,Cong Zhang,Runmin Cong
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by ICME 2026, 6 pages, 3 figures. Corresponding authors: Anpeng Wang and Runmin Cong

点击查看摘要

Abstract:Few-shot semantic segmentation (FSS) of strip steel surface defects (S ^3 D) has posed significant challenges distinct from natural scenes. Unlike natural images, S ^3 D task exhibits unique characteristics including low local contrast, uneven illumination, and complex fine-grained texture patterns. Although recent methods based on Segment Anything Model (SAM) have shown promise in FSS on natural images by leveraging SAM’s powerful pre-trained representations, these unique industrial characteristics of S ^3 D images lead to performance drop when directly applying SAM to industrial defect scenarios. In this paper, we propose a novel Perceptual Parallel Prompt (P ^3 ) framework that empowers SAM, creating the P ^3 -SAM model to address these challenges through two core strategies. First, we develop a Perceptual-Optimized Encoding (POE) strategy that enhances local contrast and preserves critical texture details for S ^3 D segmentation. Second, we introduce the Parallel Prompt Generator (PPG) strategy that simultaneously generates both semantic and spatial prompts, enabling comprehensive guidance for SAM’s decoder across varying images. Extensive experiments on three few-shot S ^3 D benchmarks demonstrate that P ^3 -SAM achieves state-of-the-art performance, with particularly notable improvements of 12.00% in mIoU on Surface Defects-4i dataset.

[CV-46] When Online Adaptation Hurts: Parameter-Frozen Test-Time Ensembling for Continual Medical Image Segmentation

链接: https://arxiv.org/abs/2609.21412
作者: Ruijie Huang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 7 pages, 2 figures

点击查看摘要

Abstract:Medical image segmenters often get worse when sites, scanner vendors, or protocols change. Continual test-time adaptation (CTTA) addresses this problem without target labels, but it can be impossible to update a model on a non-stationary stream and can lead to a lot of errors. We examine a more reasonable and meaningful alternative: parameter-frozen inference enhancement(PIE). We use a source-trained segmenter that learns about anatomy-preserving scale and flip views, maps their predictions back to the native location, and averages the probabilities. We do not modify the weights of the model or the normalization statistics. On a cardiac MRI stream from M\Ms, which is trained on vendor A and evaluated sequentially on vendors B, C, and D, PIE has 0.7786 mean Dice, compared to 0.7680 for source-only inference and 0.7388–0.7416 for five other online-adaptation baselines. The controlled ablations show that performance saturates at 28 views, and confidence weighting, class-prior correction, connected-component filtering, morphological refinement, and inter-slice smoothing have no effect or cause negative transfer. Qualitative results on cardiac MRI and fundus images are also consistent with the frozen ensemble keeping thinner and nested anatomical structures. These results provide a strong, stable baseline for medical CTTA and expose an important failure mode: adaptation and handcrafted refinement can be less reliable than carefully designed inference.

[CV-47] Quantization-Aware Kalman Estimation for Diffusion Sampling

链接: https://arxiv.org/abs/2609.21407
作者: Qitan Shi,Cheng Jin,Jiawei Zhang,Yuantao Gu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 20 pages, 6 figures, 2 tables

点击查看摘要

Abstract:Quantization offers a practical path to deploying diffusion models with reduced memory and computation, but aggressive compression can cause quantized outputs to deviate substantially from their full-precision counterparts. Sampling-stage correction methods seek to compensate for such deviations during sampling, but existing approaches rely primarily on local information and underexploit trajectory history, limiting their ability to correct errors that propagate across timesteps. In this work, we formulate sampling with a quantized denoiser as an online estimation problem, using the history of quantized denoiser outputs to recover the underlying full-precision outputs required by the sampler. We propose QuAKE, a Quantization-Aware Kalman Estimator that combines a smooth trajectory prior with a conditional Gaussian observation model. At each sampling step, QuAKE recursively updates the posterior over the output window in closed form and feeds its posterior mean to the sampler. QuAKE is a lightweight plug-and-play corrector that requires no modification to the quantized network and naturally supports arbitrary high-order multistep ODE samplers. Experiments across W4A4-quantized text-to-image diffusion models show that QuAKE consistently outperforms existing methods in reducing the distributional discrepancy from full-precision sampling.

[CV-48] SIRA: Reasoning -Aware Surgical Instrument Segmentation via Query-Anchored Alignment

链接: https://arxiv.org/abs/2609.21402
作者: Zhibo Zhang,Qijie Wang,Zengqiang Yan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Surgical instrument segmentation (SIS) plays a critical role in robotic assistance and surgical workflow analysis. However, most existing SIS methods formulate segmentation as a category-driven localization problem, limiting their ability to capture procedural context and task-dependent semantics in surgical workflows. We introduce Reasoning-Aware Surgical Instrument Segmentation (RA-SIS), a task formulation that frames segmentation as query-conditioned inference under surgical context. To benchmark this setting, we construct SurgRS, a surgical reasoning segmentation dataset consisting of 41,000 image-text pairs, which aligns instance-level masks with structured query-answer supervision to enable semantic grounding at the pixel level. Based on SurgRS, we propose Surgical Instrument Reasoning and Segmentation Assistant (SIRA), a multimodal framework that disentangles target-level and query-level semantics and integrates them with visual features through query-anchored dual alignment. By aligning query semantics with spatial features and segmentation prompts, SIRA enhances semantic-visual consistency in mask prediction. Extensive experiments on SurgRS demonstrate improvements over existing reasoning-aware baselines. Code is available at this https URL.

[CV-49] A Scene Language Model for Open-Vocabulary Scene Mapping

链接: https://arxiv.org/abs/2609.21400
作者: Adam Lilja,Fabio Hübel,Siming He,Junsheng Fu,Claire Tomlin,Lars Hammarstrand,Jitendra Malik,Jonas Frey,Marco Pavone
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Open-vocabulary 3D scene mapping aims to build a persistent representation of the objects in an environment. Existing systems typically rely on engineered mapping pipelines to associate observations, merge information across views, and maintain a consistent scene representation over time. Many additionally store feature-rich object representations, such as embeddings or image crops, increasing the size and complexity of the persistent memory. We introduce SceneLM, a Scene-Language Model that directly maintains a textual scene map. The full scene is represented as a structured text list of objects, which serves as the model’s only persistent memory. For each input image, the model reads the current scene state and updates the map by adding, editing, and removing objects. To learn this behavior, we introduce supervision tasks for iterative scene map maintenance together with an automatic annotation pipeline that generates training data from images without human labels. We evaluate SceneLM on both a language-grounded retrieval benchmark and a localization benchmark. Across both benchmarks, the model produces a scene map that achieves competitive performance with complete mapping systems built from dedicated perception and geometric modules while producing a scene representation that is 6-12x more compact. We further show that SceneLM can be run online on an edge device through experiments on a quadruped. These results show that a persistent open-vocabulary 3D scene map can be maintained directly by a single vision-language model using only a lightweight text representation. Training and inference code is available on this https URL.

[CV-50] Agent VidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents KR

链接: https://arxiv.org/abs/2609.21386
作者: Seoyeon An,Hyeonseo Jang,Minsu Kim,Chanho Lee,Younghan Park,Kangwook Lee
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 36 pages, 8 figures. Code: this https URL Dataset: this https URL

点击查看摘要

Abstract:Comprehensive video understanding is crucial for advancing artificial intelligence toward the intricate dynamics of the physical world. While recent advances in Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in video understanding, existing benchmarks remain confined to simple scene-level queries or global summaries that require only single-step inference. Real-world video understanding involves more challenging tasks that require multi-hop multimodal reasoning, and there is a critical absence of video benchmarks equipped to rigorously evaluate these agentic capabilities. To bridge this gap, we introduce AgentVidBench, a multi-hop video question answering benchmark focused on evaluating the spatial, temporal, and causal reasoning capabilities of MLLM agents. Beyond standard question-answer pairs, AgentVidBench provides step-by-step solution traces to support trajectory evaluation that assesses whether agents explicitly acquire the evidence needed to justify their answers. Experiments with 12 proprietary and open-source MLLMs show that single-turn performance remains limited on AgentVidBench, while integrating these models into state-of-the-art agentic workflows generally improves performance with respect to both accuracy and trajectory scores. We further present a simple yet effective agentic strategy that serves as a competitive baseline on AgentVidBench, establishing our benchmark as a holistic testbed for future research on agentic video understanding. Code and datasets are available at this https URL and this https URL.

[CV-51] JEPA Guided Diffusion: Predictive Vision-Language Conditioning for Generative Traffic Forecasting ECCV

链接: https://arxiv.org/abs/2609.21379
作者: Trinh Tra Giang Nguyen,Thanh Nguyen Vo,Nguyen Hoai Thuong Bui,Ha Duc Bui
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: ECCV Workshop 2026, AI City Challenge 2026 Track 5

点击查看摘要

Abstract:Accurate traffic forecasting requires both understanding scene dynamics and synthesizing realistic future observations. Recent diffusion-based video generation models produce visually plausible predictions but require expensive end-to-end training and often entangle scene understanding with image synthesis. In this work, we propose a decoupled forecasting framework that separates future representation learning from video generation. A frozen V-JEPA encoder first extracts predictive latent representations from the observed traffic videos, capturing the underlying scene dynamics in a semantic latent space. A lightweight latent alignment module then projects these representations into the conditioning space of a frozen Cosmos diffusion module, enabling future video synthesis without retraining the large generative model. By freezing all foundation models and training only the lightweight alignment module, the proposed framework substantially reduces optimization complexity while preserving forecasting capability. Experimental results on the AI City Challenge 2026 Track 5 benchmark demonstrate that the proposed method achieved a score of 75.1297, ranking third in the competition. These results suggest that predictive world representations learned by V-JEPA can effectively guide downstream video generation, providing a practical and efficient alternative to end-to-end diffusion-based forecasting.

[CV-52] ProTracer: Proprioception-Guided Failure Diagnosis in Robot Manipulation

链接: https://arxiv.org/abs/2609.21369
作者: Chang Dong,Mehdi Hosseinzadeh,King Hang Wong,Lingqiao Liu,Francois Fraysse,Feras Dayoub,Minh Hoai Nguyen
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: 9pages, 5 figures, 5 tables

点击查看摘要

Abstract:This paper presents a comprehensive framework for robot manipulation failure analysis that includes binary failure detection, failure categorization, explanation generation, and the additional capability of failure onset localization, which aims to identify the earliest moment at which a robot execution deviates from a valid task-completion trajectory and is ultimately followed by task failure. To address these tasks, we propose ProTracer, a training-free framework that leverages existing Vision-Language Models (VLMs) together with proprioceptive signals for failure analysis. Our method uses proprioceptive dynamics to identify temporally informative action boundaries and converts richer robot-state signals into structured natural-language descriptions that can be jointly analyzed together with visual observations by the VLM. This design combines the temporal precision of proprioceptive signals with the multimodal reasoning capabilities of modern VLMs without requiring additional model training. We further introduce FailTime, a benchmark with synchronized visual and proprioceptive observations for evaluating conventional failure diagnosis tasks as well as failure onset localization. Experiments demonstrate that ProTracer achieves strong performance across both conventional failure diagnosis tasks and the newly introduced failure onset localization task, highlighting the importance of proprioceptive reasoning for fine-grained temporal failure analysis.

[CV-53] Hiding in Plain Sight: A Diffusion-based Mitigation of Geolocation Privacy Leakage in Vision-Language Models NDSS2027

链接: https://arxiv.org/abs/2609.21363
作者: Yining Wang,Xi Li,Mi Zhang,Xiaohan Zhang,Xiaoyu You,Zhenxing Qian,Mi Wen
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: NDSS 2027

点击查看摘要

Abstract:Multimodal large reasoning models (MLRMs) have demonstrated remarkable capabilities in complex visual understanding. However, this very power introduces a critical yet underexplored privacy threat: adversaries can exploit MLRMs to precisely infer users’ geographic locations from casually shared photographs, by performing structured reasoning over subtle visual cues such as architectural styles, vegetation, and lighting conditions. In this work, we present a systematic study of MLRM-driven geolocation privacy leakage. We first reveal that refusal-based safeguards are critically insufficient, as carefully crafted jailbreak prompts can raise model response rates to 100%. We further identify that existing defenses, which inject imperceptible perturbations into shared images, suffer from structural limitations intrinsic to their pixel-space optimization, resulting in degraded black-box transferability and pronounced visual artifacts. Motivated by these findings, we propose a diffusion-based framework that provides targeted, proactive defense against geolocation privacy leakage. By injecting perturbations into the latent space of a diffusion model during reverse sampling, our method operates directly on high-level semantic representations, thereby resolving the effectiveness-utility bottlenecks by construction. We further ground our optimization with GeoCLIP, a model explicitly aligned with GPS coordinates, as a surrogate to pinpoint and disrupt the geographic signals that MLRMs exploit for location inference. This targeted semantic disruption yields significantly stronger black-box transferability while preserving perceptual image quality, offering a seamless integration on social media platforms.

[CV-54] Field Tracking of Insects Using a Stereoscopic Event-Based Camera Setup

链接: https://arxiv.org/abs/2609.21354
作者: Pratham G. Shenwai,Martin J. Lankheet,John T. Hrynuk,Mandiyam Y. Mahadeeswara,Mandyam V. Srinivasan,Sridhar Ravi
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:High-speed tracking of small, fast-moving organisms in their natural environments is important to better understand their behavior and ecology. Traditional frame-based imaging suffers from motion blur due to low temporal resolution, and data storage limitations, propelling a search for more adaptive solutions. Event cameras, which capture changes in brightness at pixel level instead of entire frames, have emerged as a promising solution by increasing temporal resolution and data efficiency. Here, we demonstrate the use of event-based imaging with standard video-based processing methods by converting the asynchronous events into conventional video formats, allowing us to leverage the event camera’s enhanced temporal detail to capture intricate insect flight movements and apply established image analysis techniques. Coupling this conversion process with a stereoscopic configuration provides continuous, low-latency, three-dimensional tracking of fast-moving subjects in field conditions. As a result, we substantially mitigate motion artifacts and achieve more accurate representations of animal movements. By making event-based imaging more readily applicable in natural field settings, our method support broader applications across animal behavior and ecological research, agricultural management, and other fields requiring high-fidelity object tracking in the wild.

[CV-55] PrismAlign: Prior-Steered Multi-View VLM Alignment for Hallucination-Robust Table OCR EMNLP

链接: https://arxiv.org/abs/2609.21351
作者: Guangyi Liu,Qianjun Huang,Boyu Hou
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by EMNLP industry track

点击查看摘要

Abstract:Table extraction suffers from frequent structural errors and semantic hallucinations. We propose PrismAlign, a multi-VLM framework aligning diverse visual perspectives to resolve ambiguity. It integrates priors of table logic to assess output plausibility, decoupling structural alignment from cell content alignment. A Bayesian decision strategy maximizes alignment accuracy by exploiting the correlation between extraction errors and computable rule violations. Evaluated on open-source and custom VLMs, PrismAlign reduces hallucinations and achieves state-of-the-art performance on OmniDocBench 1.5, as well as on the table category of CC-OCR and PureDocBench.

[CV-56] Cube-Splat: High-Fidelity 360° Gaussian Splatting SLAM via Cubemap Factorization and Adjoint-Consistent Optimization ECCV2026

链接: https://arxiv.org/abs/2609.21347
作者: Xiangfei Guo,Hao Shi,Yufan Zhang,Zhonghua Yi,Yongqi Mao,Xiaoting Yin,Kaiwei Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to ECCV 2026. Source code : this https URL

点击查看摘要

Abstract:Recent progress in 3D Gaussian Splatting (3DGS) has enabled dense visual SLAM with pinhole cameras, yet most pipelines are not designed for panoramic imagery. We present Cube-Splat, the first panoramic GS-SLAM framework that factorizes each 360° frame into a cubemap of four fixed-orientation virtual pinhole views sharing a single optical center. By designating the front face as the primary pose state, we accumulate gradients from all faces via an adjoint mapping, thereby enabling multi-face observations to coherently update a single state while strictly preserving cross-view geometric consistency. Concurrently, our mapping module densifies and optimizes anisotropic Gaussians using aggregated cubemap rays for high-fidelity, dense reconstruction. Furthermore, to rigorously evaluate panoramic SLAM under diverse and challenging conditions, we introduce SynPano, a highly scalable, photorealistic synthetic dataset featuring parameterized complex trajectories and multi-modal ground truth. Extensive evaluations on two public benchmarks (PALVIO and OmniBlender) and our SynPano dataset, collectively encompassing both indoor and outdoor scenes, demonstrate that Cube-Splat achieves state-of-the-art (SOTA) performance in tracking accuracy and reconstruction fidelity. Both the source code and the SynPano dataset are available at this https URL.

[CV-57] VeriFuse: Bounded Vision-Language Arbitration and Reason -Guided Refinement for Cooperative 3D Perception

链接: https://arxiv.org/abs/2609.21323
作者: Hongyi Lin,Yiyao Liu,Qi Kang,Heye Huang,Yang Liu,Haris Koutsopoulos,Jinhua Zhao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 8 pages, 4 figures

点击查看摘要

Abstract:Vision-language models (VLMs) have demonstrated strong scene understanding and semantic judgment across diverse tasks, but their appropriate role in cooperative perception remains unclear. Directly asking a VLM to regress 3D detections is unreliable and computationally expensive, whereas using it to select the output of a single source discards useful information from other agents. We introduce VeriFuse, a bounded arbitration framework for vehicle-infrastructure cooperative 3D detection. Each agent first produces detections independently. Around each vehicle and roadside proposal, VeriFuse generates source-conditioned geometric candidates and combines the original detections, their perturbations, and cross-source hypotheses into a unified candidate pool. A frozen VLM then chooses among three admissible actions: SELECT an adequate candidate; REFINE an existing anchor when an object is supported but all candidates are geometrically inadequate; or REJECT an unsupported infrastructure-only proposal. Experiments on the DAIR-V2X dataset show that VeriFuse achieves 0.494/0.357 cooperative 3D AP50/AP70 and limits the relative vehicle-side BEV AP50 drop under a 300 ms delay to 1.7%. Overall, VeriFuse assigns the VLM a clear and constrained role in cooperative perception: semantic reasoning resolves ambiguity among cross-agent hypotheses, while deterministic constraints determine the final 3D geometry.

[CV-58] S3VD: Semantic-Guidance Spatio-Temporal Scanning for Video Deraining

链接: https://arxiv.org/abs/2609.21322
作者: Kui Jiang,Yiang Chen,Yan Luo,Zhaocheng Yu,Junjun Jiang,Xianming Liu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Heavy rainfall severely degrades outdoor videos by corrupting high-frequency details and introducing motion blur, critically undermining the reliability of visual tasks. Recently, State Space Models (SSMs), particularly Mamba, have emerged as efficient alternatives for vision tasks with their linear complexity and ability to model long-range dependencies. However, when confronted with the poor visual representations in rainy videos, Mamba still faces difficulties in preserving the integrity of 2D spatial semantics and modeling 3D spatio-temporal correlations. To break these limitations, we introduce S3VD, a Semantic-Guidance Spatio-Temporal Scanning framework for video deraining, featuring two key innovations: Multi-Scale Semantic Fusion (MSSF) Module and Spatio-Temporal Scanning Fusion (STSF) Module. The former integrates temporal semantic priors from DINOv2 to guide precise feature representation and counteract the loss of local semantic context inherent to Mamba’s 1D flatten operation, enhancing robustness against extreme degradation. The latter introduces a spatio-temporal scanning mechanism and devises a Decoupled-Gating Mamba (DG-Mamba) layer, which employs two independent gates to adaptively control preceding and subsequent contextual information within the input clip, optimizing intra-frame and inter-frame correlation modeling. Experiments on video deraining benchmarks demonstrate the superiority of S3VD, achieving state-of-the-art performance with an average 0.84 dB PSNR improvement over Mamba-based baselines.

[CV-59] Combining Object Detection with Geometry-Aware Clustering to Distinguish Overlapping Plants in UAV Imagery

链接: https://arxiv.org/abs/2609.21304
作者: Ik Jae Lee,Hieu D. Nguyen,Mahbubur Meenar,Carlos Morrison Martinez,Cameron Connelly
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 34 pages

点击查看摘要

Abstract:Reliable plant-level information from unmanned aerial vehicle (UAV) imagery is important for automated crop monitoring. However, in dense crop canopies, adjacent plants frequently overlap and are detected as a single object, reducing the reliability of plant-level measurements. This study presents a geometry-aware post-detection framework for resolving overlapping plant instances using standard RGB UAV imagery. The framework combines object detection with geometric clustering of plant components. Leaves or branches detected within each bush-level region are represented using two complementary geometric features: component centroids and radial intersection points (RIPs) derived from detected plant structures. K-means and Gaussian mixture models determine whether a detected region contains a single plant or two overlapping plants. Density filtering suppresses spurious radial intersections, and a post-pipeline ensemble combines spatial and directional geometric information. The framework was evaluated using UAV imagery of eggplant and tomato crops under field conditions. Centroid-based clustering achieved an F1-score of 0.89 for eggplant, while the combined centroid-RIP approach achieved the best tomato performance, with an accuracy of 0.80, precision of 1.00, and F1-score of 0.75 using K-means. Density filtering substantially improved RIP-based clustering for tomato. The proposed approach provides a lightweight, modular engineering solution that can be integrated with existing RGB UAV monitoring pipelines without additional depth sensors, pixel-level segmentation, three-dimensional reconstruction, or retraining of the primary bush detector. The results demonstrate that geometric reasoning applied to existing detector outputs can complement deep-learning-based object detection and improve plant-level interpretation in dense agricultural canopies. Comments: 34 pages Subjects: Computer Vision and Pattern Recognition (cs.CV) ACMclasses: I.4.9 Cite as: arXiv:2609.21304 [cs.CV] (or arXiv:2609.21304v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2609.21304 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Hieu Nguyen [view email] [v1] Fri, 18 Sep 2026 04:25:07 UTC (34,223 KB)

[CV-60] Beyond Exact Match: Task-Aware GRPO for Cross-Domain PCBA Visual Question Answering ACM-MM2026

链接: https://arxiv.org/abs/2609.21276
作者: Jia Li,Li Dai,Peng Jia,Zhenzhen Hu,Chee Seng Chan,Bingkun Bao,Richang Hong
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 8 pages, 2 figures. Accepted to the 34th ACM International Conference on Multimedia (ACM MM 2026)

点击查看摘要

Abstract:In automated Printed Circuit Board Assembly (PCBA) inspection, standards-guided decisions require systems to jointly reason over fine-grained visual cues, component semantics, and manufacturing knowledge. Although large vision-language models (VLMs) provide a promising foundation, their deployment is hindered by the domain shift between standards-derived samples and real-world production-line imagery, together with heterogeneous output spaces spanning choice-based and numerical counting tasks. To address these challenges, we propose a multimodal reasoning framework for cross-domain PCBA visual question answering. The framework converts standards-derived, real-world, and auxiliary PCB-domain data into a unified instruction format and constructs verified reasoning traces aligned with visual evidence, question semantics, candidate options, and ground-truth answers. We further introduce Task-Aware Group Relative Policy Optimization (GRPO), which moves beyond exact-match supervision by integrating multi-component semantic rewards for choice-based questions, distance-aware rewards for counting questions, and an auxiliary format reward for valid outputs. During inference, answer-option semantic consistency correction, self-consistency voting, and multi-model arbitration are combined to improve prediction robustness. The proposed system achieves an Overall Score of 83.24 on the official PCBA Standard-to-Real Grand Challenge leaderboard, demonstrating the effectiveness of task-aware reward design and robust inference for cross-domain PCBA visual question answering.

[CV-61] Edit-VAR: Taming Visual Autoregressive Model for Precise Video Editing

链接: https://arxiv.org/abs/2609.21268
作者: Chongbo Zhao,Jiangming Wang,Xilai Wang,Xinyu Wang,Jingyi Tang,Chunjie Hao,Pengjie Song,Yue Ma
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL . Code: this https URL

点击查看摘要

Abstract:Text-guided video editing modifies target content while preserving the appearance and temporal coherence of unedited regions. Training-based approaches provide strong control but demand substantial data and computation. Training-free methods fall into inversion-free and inversion-based paradigms. Inversion-free approaches avoid trajectory recovery, but their source-preserving guidance can limit editing strength and leave semantic changes incomplete. Inversion-based approaches recover a latent trajectory before regeneration, where approximation errors can accumulate and cause source-content drift and temporal inconsistency. We introduce Edit-VAR, the first training-free and inversion-free framework for text-guided video editing with a pretrained visual autoregressive video model. Edit-VAR directly encodes the source video into multi-scale discrete tokens and performs probability-guided conditional token replacement for source preservation. Attention-guided token-wise and scale-aware modulation selectively relaxes source constraints over edit-relevant positions and generation stages. Scale-Decoupled Generation, implemented as late-scale constraint release, regenerates motion-consistent details and reduces texture fragmentation. Residual-guided token pruning further exploits redundancy at the final two high-resolution scales to reduce inference cost. Extensive experiments and a blind user study demonstrate that Edit-VAR outperforms existing training-free video editing methods overall in editing fidelity, source preservation, temporal coherence, and inference efficiency.

[CV-62] Geometry-Aware Diffusion Guidance via Curvature-Adaptive Tubular Correction

链接: https://arxiv.org/abs/2609.21251
作者: Enze Jiang,Jinwei He,Zheng Ma
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Gradient-guided diffusion samplers provide flexible priors for inverse problems and conditional generation, but strong guidance can move the sampling trajectory into regions where the learned score is poorly supported. Existing tangent-projection strategies limit first-order departure from an iso-density surface, yet discard potentially useful normal motion and overlook the second-order departure induced by tangent motion on a curved surface. We introduce curvature-adaptive tubular correction (CAT), a training-free plugin that regulates both effects within a shared, noise-dependent geometric budget. CAT decomposes the guidance gradient into normal and tangent components, charges normal displacement at first order and tangent displacement according to directional curvature, and obtains their jointly optimal magnitudes from a one-dimensional dual equation. Armijo backtracking calibrates the resulting finite step against the actual guidance objective, while matrix-free directional derivatives avoid constructing the full score Jacobian. We establish local guarantees for the tubular approximation, uniqueness of the correction, and sufficient objective decrease. Across seven inverse problems on FFHQ and ImageNet, CAT improves the evaluated pixel- and latent-space host samplers, with particularly consistent gains in perceptual metrics. It also improves black hole reconstruction on InverseBench and yields the lowest FID among the compared methods at every tested classifier-free guidance scale, while maintaining stable saturation and contrast. These results support curvature-aware tubular control as a reusable mechanism for stabilizing diffusion guidance.

[CV-63] VLA-Scope: Shift-Aware Failure Prediction for Vision-Language-Action Models

链接: https://arxiv.org/abs/2609.21246
作者: Kaiwen Zhu,Dongfang Liu,Liangkai Liu
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Systems and Control (eess.SY)
备注: 9 pages, 3 figures

点击查看摘要

Abstract:Vision-language-action (VLA) models map visual observations and natural-language instructions to robotic actions, but distribution shifts can compromise their reliability. Because these models may still succeed under out-of-distribution (OOD) conditions, detecting OOD inputs alone is insufficient to predict execution failure. In this paper, we introduce VLA-Scope, a two-stage framework that combines input-shift characterization with execution history to predict failure during OOD rollouts. The first stage uses pooled image and language representations to detect OOD inputs and classify their shift categories. For inputs flagged as OOD, the second stage combines the predicted category, action-prefix features, and execution progress features. A logistic regression model shared across shift categories updates failure risk as execution proceeds. We evaluate the framework with OpenVLA on ten LIBERO-Spatial tasks using leave-one-group-out cross-validation. OOD detection achieves a ROC-AUC of 0.9454, and shift classification achieves 91% accuracy. Evaluated independently of the OOD gate on all 1,400 OOD rollouts, the failure predictor achieves a ROC-AUC of 0.8497 after 60 executed actions, compared with 0.7906 without execution progress features. It also achieves a higher ROC-AUC than the evaluated ActProbe and SAFE-MLP baselines. These results suggest that combining action features with temporally aggregated execution step representations improves failure prediction under input shifts.

[CV-64] SafeStyle: Calibrated Style Residual Injection for Controllable Style-Leakage Trade-off in Diffusion Stylization

链接: https://arxiv.org/abs/2609.21242
作者: Zhangping Yang,Min Li,Song Yan,Rong Gao,Xinliang Bi,Guanye Xiong,Yujie He
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 5pages, 6figures

点击查看摘要

Abstract:Reference-guided diffusion stylization aims to transfer visual style from a reference image while preserving the semantics specified by a text prompt. However, image conditioning often entangles transferable style cues with reference-specific content, leading to an inherent trade-off: stronger conditioning improves style fidelity but increases content leakage, whereas aggressive suppression reduces leakage at the cost of style expression. This challenge is further complicated by the distinct spatial organization of texture- and geometry-dominant styles. To address these issues, we propose SafeStyle, a training-free framework for calibrated style residual injection in frozen diffusion models. SafeStyle first estimates style-supported and content-associated subspaces from compact calibration sets, preserving their informative overlap while suppressing useless content variations. It then transports the purified style evidence over adaptive spatial granularity and constrains its effective influence through an explicit residual-norm budget. Experiments across texture- and geometry-dominant styles show that SafeStyle achieves a DINO style similarity of 0.432 while maintaining competitive text alignment. On a semantically disjoint leakage-stress benchmark, it further achieves a DINO style similarity of 0.474 with only 0.8% semantic leakage, demonstrating an effective balance between style fidelity and reference-content suppression.

[CV-65] Multiclass Semantic Segmentation of Wildland Fire Images Using Context-Aware Centralized Copy-Paste Data Augmentation

链接: https://arxiv.org/abs/2609.21241
作者: Joon Tai Kim,Nishanth Kunchala,Vishv Patel,Tianle Chen,Ziyu Dong,Daniel Ospina Acero,Roger Williams,Mrinal Kumar
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 14 pages, 9 figures

点击查看摘要

Abstract:Producing accurate annotations for deep learning based image segmentation is both costly and labor intensive. This challenge is especially evident in wildland fire applications, where accurately labeled datasets are scarce due to the difficulty of collecting and annotating dynamic fire scenes. To address this problem, our previous work introduced the Centralized Copy-Paste Data Augmentation (CCPDA) method for semantic segmentation of wildland fire imagery, which generates artificial training samples by randomly pasting fire clusters from source images onto target images. However, random placement can produce contextually unrealistic scenes, such as fire burning on asphalt. In this paper, we present a context-aware strategy designed specifically to improve data quality and realism in small multiclass wildland fire datasets, ensuring that augmented samples remain contextually meaningful. The proposed method restricts fire placement to semantically valid target regions and selects the location whose Ash-Vegetation composition most closely matches the source context. This approach preserves existing fire regions in the target image, prevents unrealistic placements, and maintains contextual accuracy by generating images that resemble real wildland fire scenes. We evaluate the Context-Aware CCPDA strategy through numerical analysis and comparisons with other augmentation methods by a weighted sum-based multi-objective optimization (MOO) approach. The results confirm that the context-aware data augmentation strategy leads to improved segmentation performance and contextual realism, outperforming other augmentation procedures.

[CV-66] FOCAL-VLA: Subtask-Guided Geometry Distillation and Implicit World Modeling for Vision-Language-Action Models

链接: https://arxiv.org/abs/2609.21228
作者: Zhiyuan Gao,Di Wen,Yanxiang Zhan,Mohammad Khoshnazar,Jeroen Schäfer,Kunyu Peng,Michael Beetz
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Vision-language-action (VLA) models built on pretrained vision-language models have demonstrated strong performance across diverse robotic manipulation tasks. However, VLA models that directly map current 2D observations to actions often lack sufficient spatial and temporal understanding, limiting their performance in precise and long-horizon manipulation. Recent methods enhance VLA models through geometric supervision and future-state prediction across the entire scene. However, these methods can suffer from redundant scene information, distracting the model from learning the geometry and dynamics relevant to the current interaction. To address this issue, we propose FOCAL-VLA, a framework that combines subtask-guided geometry distillation with implicit world modeling to learn representations of current spatial structure and future interaction dynamics. To focus geometric learning on the current subtask, we transfer geometric knowledge from VGGT to the VLA model by aligning geometry latents with features from subtask-relevant image regions. To capture the future 3D evolution of the current interaction, we incorporate implicit world modeling using Track4World features from current and future demonstration frames. The two complementary representations jointly guide action generation without running VGGT or Track4World at inference time. Experiments show that FOCAL-VLA outperforms baselines on both simulation benchmarks and real-world manipulation tasks. Project website: this https URL.

[CV-67] VGGT-CAD: Reconstructing Parametric CAD 3D Model with Geometric Grounding

链接: https://arxiv.org/abs/2609.21225
作者: Chunan Yu,Tianrun Chen,Fu Shen,Cheng Chen,Lanyun Zhu,Yang Yang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Parametric CAD reconstruction requires recovering both precise geometry and editable modeling operations from visual observations, making it challenging under limited and ambiguous views. Existing methods mainly rely on 2D appearance cues and lack strong multi-view geometric priors. In this work, we present VGGT-CAD, a geometry-aware framework for parametric CAD reconstruction from single- and multi-view observations. We transfer pretrained 3D geometric priors into CAD reconstruction by encoding camera parameters as condition tokens and jointly modeling them with image tokens. To handle varying numbers of viewpoints, we introduce a variable-view cross-view context aggregation module that adaptively fuses multi-view features. We further develop a training-free geometry-aware view selection strategy to select complementary and reliable frames during inference. The resulting representation is decoded into CAD command sequences using a non-autoregressive decoder. We also develop VideoCAD, a large-scale multi-view video benchmark derived from existing CAD data through multi-view re-rendering. Extensive experiments demonstrate the effectiveness of VGGT-CAD for visual CAD reconstruction under different observation configurations.

[CV-68] Multi-viewpoint Geo-localization with Event Cameras

链接: https://arxiv.org/abs/2609.21219
作者: Adam D. Hines,Michael Milford,Tobias Fischer
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: 8 pages, 4 figures, 4 tables, under review

点击查看摘要

Abstract:Robot localization is an ongoing challenge that demands mapping and positioning systems that are tolerant to viewpoint change. Event cameras are attracting increasing interest and adoption in robotics; however, dealing with viewpoint variance is an under-investigated problem in existing event-based localizers. In addition, event-based datasets that emphasize viewpoint variance for challenging localization situations are scarce. Here, we introduce an event-based visual place recognition (VPR) system that performs robustly under viewpoint changes. We converted five large-scale geo-tagged datasets, conventionally used to train frame-based localization systems, into synthetic event streams using Image-to-Event (I2E) conversion, and used them to fine-tune a pre-trained event-based vision transformer backbone with a multi-loss function, yielding a system we call MegaEvent that learns viewpoint-robust features for place recognition. We achieved an average Recall@1 of 82% across three existing event-based localization datasets, leading the next best event-based method by 20 recall points, and frame-based VPR models applied directly to event frames by 8 to 26 recall points. We introduce a new, challenging dataset - Springfield-Event-VPR - which features a 3.7km walking route recorded in three camera orientations for a total of 11.1km, which MegaEvent outperforms the strongest baseline by 9 recall points. The code for MegaEvent is available at this https URL.

[CV-69] Hand-Aware Transition Modeling for Bimanual Procedural Anomaly Detection KR

链接: https://arxiv.org/abs/2609.21207
作者: Di Wen,Jimmy Weissert,Luc Maria Scherrer,Cedric Zöllner,Kailun Yang,Ruiping Liu,Yufan Chen,Jiale Wei,Junwei Zheng,Kunyu Peng
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 6 pages, 1 figure, 3 tables. Code: this https URL

点击查看摘要

Abstract:Procedural anomaly detection in bimanual assembly requires judging each hand action against the execution so far. A corrective action may look unusual in isolation, while a visually plausible action can violate the order of the procedure. We present HACT, a transition model over predicted per-hand events. A role-preserving history keeps the concurrent responsibilities of both hands, and a marked temporal point process assigns each observed transition a semantic and temporal surprisal. A supervised evidence head and a two-state filter convert these surprisals into per-hand anomaly posteriors. A recovery-aware protocol on predicted events and participant-disjoint folds reports the recovery false-positive rate at an operating point selected on validation participants. On two bimanual power-tool procedures HACT has the highest AUPRC and F1 among the compared methods and the fewest recovery alarms. Applied without retraining to a different assembly order of the same product, it retains the highest AUPRC and F1. The source code is available at this https URL.

[CV-70] OnomatoBridge: Onomatopoeia Translation and Rendering Pipeline in Manga

链接: https://arxiv.org/abs/2609.21199
作者: Takara Taniguchi,Wataru Shimoda,Kota Yamaguchi,Hideki Nakayama
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Manga is a comic drawn by black and white paints gaining popularity around the world. Onomatopoeia in Manga specifically appeals to the audience with its unique visual styles, which convey sound, motion, and emotion. Visual onomatopoeia translation requires the clean replacement of Japanese onomatopoeia with onomatopoeia in the other language while preserving their visual style. Existing approaches often produce residual artifacts or style inconsistency when removing the Japanese onomatopoeia and rendering stylized English onomatopoeia. To approach these problems, we present OnomatoBridge, a filtering pipeline for visual onomatopoeia translation. We evaluate OnomatoBridge from Japanese to English on the Manga109 onomatopoeia dataset and compare it with baseline image editing models. Experimental results show that the filtered outputs by the proposed method outperform those of conventional methods. OnomatoBridge improves English text correctness by roughly 10 to 25 points and reduces residual Japanese text by about 20 to 50% in relative terms.

[CV-71] Robust Structureless Monocular Visual Inertial Initialization Exploiting Line Features and Vanishing Points IROS2026

链接: https://arxiv.org/abs/2609.21186
作者: Junwan Choi,Woongrae Jo,Dong-Uk Seo,Jinwoo Jeon,Hyun Myung
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: 8 pages, 5 figures, Accepted to IROS 2026

点击查看摘要

Abstract:Accurate initialization is essential for reliable visual-inertial odometry (VIO), but it is often ill-conditioned under degenerate motions. Existing methods typically require restrictive excitation motions to ensure sufficient observability or rely on computationally expensive 3D structure reconstruction, limiting efficient and practical deployment. To address these limitations, we propose SLIM-init, a structureless monocular VIO initializer that directly exploits geometric constraints from tracked 2D line features without explicit 3D landmark reconstruction. Specifically, SLIM-init leverages line-derived vanishing points (VPs) as translation-invariant orientation cues to provide robust rotation-only constraints under degenerate scenarios such as low-parallax or translation-dominant motions. It further incorporates a line epipolar residual to constrain translation and a line-normal projection residual to improve the conditioning of linear alignment, enhancing the accuracy and robustness of initial state estimation. Extensive experiments on a public benchmark and challenging custom degenerate-motion sequences demonstrate improved accuracy and robustness over state-of-the-art initialization methods. The source code is available at: this https URL.

[CV-72] 4DGS-Fixer: Generative Sparse-View 4D Gaussian Splatting with Iterative Refinement Guided by Video Diffusion Priors SIGGRAPH

链接: https://arxiv.org/abs/2609.21176
作者: Haitao Huang,Shenghao Zhao,Boyuan Tian,Shin-Fang Chng,Songlin Yang,Sheila Lim,Huangying Zhan,Yi Xu,Anyi Rao,Frank Guan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to SIGGRAPH Asia TC

点击查看摘要

Abstract:This paper addresses the challenges of dynamic scene synthesis from sparse-view videos. Existing methods employ geometric priors, adaptive optimization, or density-control strategies to improve 4D Gaussian modeling under sparse observations. However, they cannot fundamentally resolve the ill-posed problem caused by insufficient observations and missing scene information. Moreover, sparse-view 4D Gaussian Splatting (4DGS) often suffers from poor geometric initialization: with only a few input views, COLMAP typically reconstructs sparse and incomplete point clouds, leaving large scene regions without sufficient Gaussian support and making them difficult to recover through subsequent optimization. To address these limitations, we propose a novel iterative refinement framework based on a video diffusion model to improve the completeness and consistency of dynamic 4D scenes. Specifically, we first estimate multi-view depth maps and fuse them into dense point clouds to provide more complete geometric initialization for a dynamic 4DGS representation. We then employ a pretrained video restoration model to refine sequences rendered along novel camera trajectories at different time steps. The restored sequences serve as pseudo-supervision to regularize and iteratively refine the 4DGS representation. Experiments on a widely used benchmark dataset demonstrate that our method substantially outperforms existing baselines, achieving nearly a 2 dB PSNR improvement over the previous best-performing method.

[CV-73] MarsFM: Shading-Regularized Flow Matching for Martian Relief Estimation

链接: https://arxiv.org/abs/2609.21095
作者: Marius F. R. Juston
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 73 pages, 62 figures

点击查看摘要

Abstract:We present MarsFM, an image-conditioned latent flow-matching model for local Martian relief estimation from single-band HiRISE RED orthoimagery. The method combines a pretrained generative prior with stereo-derived geometric supervision and a differentiable Lunar–Lambert shading objective. Relief, normal, gradient, curvature, and ordinal terms constrain complementary aspects of terrain structure, while a positive-affine-invariant image comparison constrains rendered appearance. An evaluation comprising 2024 gathered patch records per integration-step count yields mean affine-aligned RMSE between 0.0935 and 0.0957 in normalized signed-log relief space for one to twenty Euler steps. These scores measure agreement with VAE-reconstructed references on positive-reference support. Their narrow range supports low-step inference under this protocol. Spatial, differential, and spectral diagnostics show broad terrain correspondence alongside smoothing, amplitude compression, and boundary mismatch. MarsFM provides a framework for combining learned terrain priors with image-based constraints; establishing improved physical terrain resolution requires matched baselines and independent high-resolution reference data. Data: this https URL code: this https URL.

[CV-74] Frag ment-Aware Vision Transformers for Fresco-Frag ment Style Classification ECCV2026

链接: https://arxiv.org/abs/2609.21012
作者: Sara Miketek,Biagio Barchielli,Nadeem Iqbal Kajla,Sinem Aslan
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: VISART Workshop, ECCV 2026 (Oral)

点击查看摘要

Abstract:Artistic style classification is usually studied on complete artworks, where models can exploit global composition, spatial organisation, and iconographic structure. In archaeological settings, however, artworks often survive only as fragmented remains, forcing recognition from incomplete, irregular, and context-limited visual evidence. We study fresco-fragment style classification using a progressive transformer-based framework. Starting from a ViT-B/16 baseline, we introduce foreground-guided masking to suppress background-only tokens, inpainting-based geometric regularisation to align irregular fragment supports with the ViT patch grid, and a supervised contrastive objective that operates on predictive distributions through a Kullback-Leibler similarity and consistently improves every branch. We combine the branches with a deliberately simple learnable logit ensemble. Experiments on CLEOPATRA and POMPAAF show that fragment-aware modelling improves over the standard ViT baseline, with the ensemble increasing accuracy from 0.604 to 0.656 and macro-F1 from 0.596 to 0.648 on CLEOPATRA, and outperforming the best single branch in four of six fragmentation settings on POMPAAF. We additionally evaluate a more complex graph-fusion variant and find that it matches the simple ensemble on POMPAAF while offering only a small, dataset-specific gain on CLEOPATRA, which does not justify its added complexity. Beyond these empirical gains, our contribution is twofold: a distribution-level contrastive objective that consistently sharpens single-branch recognition, and an interpretability analysis that verifies the models exploit genuine painted evidence, while quantifying that the inpainting-based branch draws part of its attribution from the synthesised surround.

[CV-75] Do Spinning Radar Doppler Velocity Measurements Improve Vehicle Detection and Tracking?

链接: https://arxiv.org/abs/2609.21000
作者: Eric Xie,Daniil Lisus,Timothy D. Barfoot
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Signal Processing (eess.SP)
备注: 8 pages, 8 figures

点击查看摘要

Abstract:Spinning frequency-modulated continuous-wave (FMCW) radars have been gaining popularity in autonomous vehicle perception on account of their robustness to adverse weather conditions and 360° field of view. Recently, scanning radars have also been shown capable of generating per-azimuth Doppler velocity. In this paper, we investigate whether these Doppler velocity measurements improve spinning radar vehicle detection and tracking performance. For detection, we estimate the ego motion and use it to undo the Doppler range distortion of the radar image before passing it to a network. For tracking, we propose a new way to estimate a per-vehicle velocity and use it as a prior for the tracker’s motion model. Since Doppler-enabled spinning radar data is not available in any dataset with ground-truth dynamic object labels, our first contribution is an automatic labelling pipeline that uses an ensemble of fine-tuned off-the-shelf lidar detectors to label all 643 km of the Boreas Road Trip dataset. We then transfer detections to radar, and use over 250 km of vehicle-dense sequences as ground-truth training data. By training and evaluating two state-of-the-art detectors, we show that Doppler undistortion can improve detection accuracy by up to 2.37 points on mean average precision. Furthermore, we show that the Doppler velocity prior can improve tracking accuracy by 13.68 points on multi-object tracking accuracy (MOTA) versus the zero-velocity initialization baseline, while achieving 99.7% of the MOTA obtained using ground-truth velocities as the prior.

[CV-76] Image-Derived PM10 Estimation in Cattle Feedlot Using Machine Learning: Addressing Concentration Ranges Beyond Existing Digital Imaging Methods

链接: https://arxiv.org/abs/2609.20975
作者: Sirapoom Peanusaha,Greg B. Ferguson,K. Jack Bush,Peiyang Li,Brent W. Auvermann
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Affordable dust monitoring remains a pressing need for the cattle feedlot industry, yet camera-based PM estimation, despite its growing body of research in urban air quality settings, has not been evaluated under the extended concentration ranges characteristic of intensive livestock operations. This study developed an image-based approach using contrast panel features and machine learning to estimate PM10 concentrations in a commercial cattle feedlot, where hourly average PM10 ranged from 250 to 1,000 ug/m^-3 and instantaneous concentrations reached 5,000 to 20,000 ug/m^-3. Grayscale images were captured during the evening dust peak period, and features including panel contrast, black and white panel pixel values, and overall image brightness were extracted. The model also incorporated recent past values from preceding images and solar zenith angle as predictors. Among the candidate models evaluated, XGBoost achieved the highest predictive performance, with an R^2 of 0.792 and a median absolute error of 103 ug/m^-3. Feature importance analysis revealed that (a) panels positioned farthest from the camera contributed most strongly to predictions and (b) that black panel pixel values were more sensitive than white panel values to changes in PM10 concentration. Prediction accuracy during the sunset transition, which coincides with the onset of the feedlot evening dust peak, remains an area for further refinement. These findings demonstrate the feasibility of image-based PM10 estimation across PM concentration ranges substantially exceeding those reported in prior urban studies and provide practical guidelines for future deployment in feedlot environments.

[CV-77] MemeTAG: Keyword-Driven Meme Classification through Tag Embedding Reconstruction WACV

链接: https://arxiv.org/abs/2609.20962
作者: Akshit Sharma,Prashant W. Patil
类目: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
备注: 10 pages, 3 figures; published in the Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2026

点击查看摘要

Abstract:The proliferation of harmful internet memes poses a significant societal threat, yet their automated classification remains a formidable algorithmic challenge due to the nuanced, multimodal nature of their content. To address this, we introduce MemeTAG, a novel dual-objective framework that pioneers a keyword-aware approach to meme classification. Our core innovation is a two-part semantic guidance mechanism: first, we leverage a pretrained Vision-Language Model to generate a set of descriptive keywords, that capture the high-level semantics. Second, we introduce the Aggregated Tag Inference Network (ATIN), an attention-based module that distills these keywords into a single, rich semantic embedding. This embedding serves as a target for a novel auxiliary reconstruction loss, which compels the model to learn deeply aligned visual and textual features. This approach, combined with an efficient three-stage training strategy, establishes a new state-of-the-art on the HarMeme, Hateful Memes Challenge (HMC), and PrideMM datasets, decisively outperforming existing state-of-the-art methods.

[CV-78] WM-VS: Progress-Aligned World Models for Closed-Loop Visual Servoing

链接: https://arxiv.org/abs/2609.20892
作者: Guanzhong Sun,Junyi Ma,Yixuan Zhou,Yuxuan Wu,Yanzi Miao,Hesheng Wang
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Closed-loop visual servoing requires predictions that indicate whether an action reduces task error, not only whether the action is plausible. We call this gap the prediction-control mismatch and introduce WM-VS, a target-centric progress-aligned world-model framework for closed-loop visual servoing. Offline target-region DINOv2 correspondences define a signed four-dimensional servo coordinate for translation, scale, and in-plane rotation. Stage 1 aligns action-conditioned latent transitions with this coordinate; Stage 2 freezes the world model and trains a reactive joint-velocity policy with action imitation, consequence supervision, and short imagined rollouts that favor error contraction. Deployment is RGB-only and reactive, without online trajectory optimization. On a real 7-DoF eye-to-hand system, WM-VS reaches a corner RMSE no larger than 10 percent of its initial value in 30/30 trials and retains this criterion at the final valid frame in 25/30 (83.33 percent). Removing future-error alignment reduces retention to 26.67 percent. The learned progress signal agrees with an external AprilTag corner error not used for training or control (mean Spearman rho = 0.8778). Without retraining, two unseen 3D targets achieve translation-error reductions of 86.48 percent and 90.27 percent and rotation-error reductions of 70.01 percent and 65.70 percent. These results link progress-aligned action consequences to repeated closed-loop correction and transfer. Code and data will be released as open source.

[CV-79] APeML: A Compact Structured Representation for Multi-Task Computer Vision

链接: https://arxiv.org/abs/2609.20869
作者: Sergey Kurinov(1),Alexey Upatov(1) ((1) Comexp Research Lab, TAPe + ML Project, Nizhniy Novgorod, Russia)
类目: Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)
备注: 39 pages, 4 figures, 11 tables. Project page: this https URL

点击查看摘要

Abstract:We present TAPe+ML v3, a compact computer vision system based on TAPe (Theory of Active Perception), a structured representation that encodes relations among perceptual elements before recognition. Instead of operating directly on pixel tensors, the system uses a shared TAPe representation and a modular recognition architecture for image classification, object detection, and instance segmentation. TAPe+ML v3 combines background and contour processing, local object localization, prototype-based classification, and a coordinator for specialized submodels. Across the reported experiments, it uses fewer than 100,000 parameters. On COCO object detection, it obtains 84.7 mAP50 and 65.3 mAP50-95. On COCO instance segmentation, it obtains 80.7 mask mAP50 and 58.4 mask mAP50-95. In classification experiments, it reaches 92 percent validation accuracy on Imagenette under an identical-training comparison with a raw-pixel baseline, and 89.9 percent Top-1 accuracy on ImageNet-Real. We also evaluate compactness in video scene detection and adaptation under distribution shift in an industrial pilot. The results suggest that shifting part of the modeling burden from network parameters to a structured input representation can support compact multi-task vision systems with reduced data, memory, and compute requirements. Comments: 39 pages, 4 figures, 11 tables. Project page: this https URL Subjects: Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV) Cite as: arXiv:2609.20869 [cs.CV] (or arXiv:2609.20869v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2609.20869 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[CV-80] AnyviewMeter: Adapting Robotic Reward Models with Camera Geometry and Multi-View Attention

链接: https://arxiv.org/abs/2609.20106
作者: Yuang Tu,Runjia Tan,Yujie Yan,Jinghan Hu,Chen Lv
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: 8 pages, 3 figures, 5 tables

点击查看摘要

Abstract:Robotic reward models evaluate task execution from visual observations, but their predictions can change with camera viewpoint and occlusion even when the underlying task state is unchanged. Adapting a pretrained reward model to a local task therefore requires accounting for how that task is observed. We introduce AnyviewMeter, a geometry-conditioned adaptation framework for robotic reward models that represent task progress as a scalar reward signal. It combines low-rank fine-tuning with token-aligned Plucker rays and synchronous block attention: ray conditioning incorporates camera geometry into visual features and attention queries and keys, while block attention fuses synchronized views inside the pretrained decoder. The framework supports both single-view reward prediction and joint multi-view evaluation through parameter-efficient adaptation of a pretrained Robometer model. On PickCube, single-view adaptation improves progress prediction in every camera group and reduces mean absolute error under a changed field of view by approximately 21% relative to RGB fine-tuning. Across simulated manipulation tasks, joint multi-view prediction reduces progress error by 41-69% compared with averaging single-view RGB predictions and improves temporal ordering in approximately 88% of task-camera groups. On real tasks with fixed and wrist-mounted cameras, mean absolute error decreases by approximately 21% relative to averaged RGB fine-tuning. These results support camera geometry and joint visual evidence as useful components of task-specific robotic reward adaptation.

[CV-81] raining-Adaptive Convolutional Sparse Coding via Information Bottleneck for Robust Visual Representation

链接: https://arxiv.org/abs/2609.19122
作者: Meng’en Qin,Yinchen Liu,Mingxuan Cui,Youlu Xing
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Machine Learning (stat.ML)
备注:

点击查看摘要

Abstract:Visual signals require compact yet sufficient representations for robust downstream prediction. Convolutional sparse coding (CSC) provides an explicit mechanism for suppressing redundant components while preserving signal content, but its sparsity coefficient is typically fixed and manually selected. We propose a training-adaptive convolutional sparse coding framework for robust visual signal representation. Specifically, we unfold the CSC optimization with the Fast Iterative Shrinkage-Thresholding Algorithm (FISTA) and treat the sparsity coefficient as a differentiable variable jointly learned with the network parameters. From the information bottleneck perspective, this coefficient controls the trade-off between information retention and compression: the sparsity term promotes compact representations, while the reconstruction term together with task loss preserves task-relevant signal content. We further introduce a label-free post-training strategy that adjusts the compression strength for corrupted inputs with the main network parameters fixed. Experiments on CIFAR and ImageNet demonstrate competitive clean-data recognition and greatly improved robustness under different input perturbations.

[CV-82] How Many Posterior Samples? Calibrated Stopping for Adaptive Sensing

链接: https://arxiv.org/abs/2609.21813
作者: Vincent Corlay,Andriy Enttsel
类目: Machine Learning (stat.ML); Computer Vision and Pattern Recognition (cs.CV); Signal Processing (eess.SP)
备注:

点击查看摘要

Abstract:In classification-oriented adaptive sensing, posterior samples characterize uncertainty at the current measurement state and can serve two roles: they may guide the next sensing direction, while their class labels provide votes for the candidate classes and determine whether sensing should continue. We focus on the stopping layer that turns these votes into a declaration, without modifying the posterior sampler or sensing directions. A natural plug-in rule declares when the observed vote share exceeds a threshold. We show that this threshold is not itself a confidence guarantee: when the underlying vote mass equals the threshold, the plug-in rule declares about half the time. As alternatives, we calibrate a fixed-sample rule and a finite-horizon sequential rule to a prescribed false-declaration probability, and study exact curtailment, which stops a fixed-pool rule once its final verdict is forced. We then derive how one-round declaration probabilities determine posterior-sample cost and classification accuracy along a sensing path. On MNIST with DDRM and a fixed PCA-guided probe sequence, curtailment saves up to 62% of posterior samples. Among the evaluated rules at matched operating points, sequential stopping reduces the cost the most. At a high accuracy, that same sequential rule can trade more posterior samples for fewer measurements.

[CV-83] Classification-oriented adaptive sensing via posterior sampling

链接: https://arxiv.org/abs/2609.21812
作者: Andriy Enttsel,Maxime Rousselot,Vincent Corlay
类目: ignal Processing (eess.SP); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Recent advances in diffusion models have enabled high-performance, instance-adaptive compressed sensing through posterior sampling, without task-specific policy training. Existing methods select sensing probes by maximizing total posterior signal variance and are therefore primarily reconstruction-driven. We introduce a classification-driven extension motivated by the closed-form posterior covariance of a class-conditional Gaussian mixture model, which decomposes into within-class and between-class uncertainty. Using calibrated soft classifier outputs, we estimate these uncertainty terms from diffusion posterior samples and propose a classification-oriented criterion for selecting the dominant sensing direction in the unmeasured subspace. Experiments on MNIST and CIFAR-10 compare the resulting classification accuracy, measurement cost, and reconstruction quality with those of reconstruction-oriented counterparts. The results identify regimes in which semantic posterior uncertainty yields a more favorable classification–measurement trade-off and quantify the associated reconstruction cost.

[CV-84] WS-NeRF: A Mamba-Driven World-State-Aware Adaptive Deblurring Neural Radiance Field

链接: https://arxiv.org/abs/2609.21391
作者: Hang Jiang,Jinghao Wang,Yiming Zhang,Xinhong Wang,Luwei Ran,Yinfeng Yu
类目: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注: Main paper (6 pages). Accepted for publication by IEEE International Conference on Systems, Man, and Cybernetics 2026 (IEEE SMC 2026)

点击查看摘要

Abstract:Neural Radiance Fields (NeRF) have attracted extensive attention in recent years due to their strong capability for high-quality 3D reconstruction and novel view synthesis from multi-view images. Existing methods usually rely on high-quality sharp inputs, while real-world image acquisition is highly susceptible to blur degradation, which severely affects the reconstruction quality of NeRF. In this paper, we propose a novel Mamba-driven world-state-aware adaptive deblurring neural radiance field, termed WS-NeRF, to address image degradation and 3D inconsistency. We formulate the alternating optimization of radiance fields as a dynamic evolution process with temporal memory, and jointly exploit comprehensive multi-dimensional world states and a mixture-of-experts mechanism to dynamically adjust the confidence of deblurring priors. Experimental results show that WS-NeRF significantly improves blurry radiance field reconstruction quality, achieving better performance on PSNR, SSIM, and LPIPS, while exhibiting more stable iterative recovery behavior.

[CV-85] Adaptive Color Grading

链接: https://arxiv.org/abs/2609.21169
作者: Trevor D. Canham,Abhijith Punnappurath,Michael S. Brown
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted @ 34th Color Imaging Conference

点击查看摘要

Abstract:Independent control of tonescale regions (e.g., shadows, highlights) is essential for painters, photographers and cinematographers to bring 2D images to life. In image manipulation software this is most directly addressed by color grading modules, which use intensity thresholds to segment distinct illumination regions for local manipulation. In this work we develop an open source color grading tool and use it to annotate a large dataset of video frames with tonescale region thresholds. Using these thresholds we conduct modeling experiments with strategies based on both practitioners’ conventional wisdom and machine learning. Results show that K-nearest neighbors is an effective prediction strategy, outperforming state-of-the-art end-to-end methods for image enhancement. This outcome demonstrates the benefit of focusing on a compact set of core parameters when modeling creative stylization processes. Our adaptive color grading interface and data are available at this https URL.

[CV-86] Uncertainty-driven training for three-dimensional calibrated lung nodule classification

链接: https://arxiv.org/abs/2609.20905
作者: Giuseppe Tripodi,Alessandro De Rosis,Saleh Rezaeiravesh
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:In this work, we present an uncertainty-driven training framework for three-dimensional computed tomography (CT) lung nodule classification, where validation-based uncertainty estimates guide loss reweighting to enhance predictive performance and probability calibration. Two Uncertainty Quantification (UQ) methods are considered: Monte Carlo Dropout (MCD) and Evidential Deep Learning (EDL). Both provide per-class uncertainty estimates that modulate the loss and encourage focus on hard or unreliable classes. The framework is evaluated with ResNet, DenseNet, EfficientNet, Vision Transformer (ViT), and Swin Transformer backbones on two datasets: the clinical LIDC-IDRI cohort and the NoduleMNIST3D benchmark. Uncertainty-driven training achieves classification performance similar to conventional training while substantially improving calibration, with an expected calibration error (ECE) reduced by up to 65% on LIDC-IDRI. EDL attains competitive performance on shallower architectures with single-pass inference, whereas MCD is more robust on deeper networks. Analysis across architectural families reveals that uncertainty-driven training benefits convolutional backbones more consistently than transformer-based architectures: EDL in particular degrades on ViT, suggesting that the Dirichlet evidence parameterisation may interact unfavourably with attention-based architectures at lower input resolutions. A posteriori temperature scaling proves highly effective across all configurations, indicating that a simple scalar calibration can be competitive even without explicit uncertainty-aware training. Our results indicate that integrating UQ into the training loop can significantly improve probabilistic calibration and support more trustworthy deployment of three-dimensional medical imaging models.

[CV-87] A Differentiable Ray-Wave Framework for Hybrid Refractive-Diffractive System Modeling and Optimization

链接: https://arxiv.org/abs/2605.15418
作者: Jiazhou Cheng,Margaret Gao,Yixuan Shao,Chenkai Mao,Tom D. Milster,Jonathan A. Fan
类目: Optics (physics.optics); Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Signal Processing (eess.SP); Computational Physics (physics.comp-ph)
备注: 9 pages, 7 figures

点击查看摘要

Abstract:Hybrid optical systems combining refractive and diffractive optical responses have the potential to support new types of optical behavior, but they are difficult to model and optimize due to the disparate spatial scales and physics exhibited by ray and wave phenomena. In this work, we present a differentiable ray-wave framework for modeling hybrid refractive-diffractive optical systems that operates as a plug-and-play module within standard ray tracing pipelines. Our model uniquely applies to both planar and curvilinear diffractive surfaces and accommodates arbitrary scalar holographic profiles with high spatial frequency responses, with each simulation evaluated at a single wavelength. We analyze ray-wave modeling regimes that optimally account for the spatial frequency properties and spatial curvature of the diffractive surfaces, and we demonstrate the gradient-based end-to-end optimization of hybrid refractive-diffractive systems featuring planar and conformal diffractive surfaces. We anticipate that these modeling capabilities will enable new classes of hybrid optical systems relevant to computational imaging and display applications.

[CV-88] MultiHU-TD: Multifeature Hyperspectral Unmixing Based on Tensor Decomposition

链接: https://arxiv.org/abs/2310.03860
作者: Mohamad Jouni,Mauro Dalla Mura,Lucas Drumetz,Pierre Comon
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Signal Processing (eess.SP)
备注:

点击查看摘要

Abstract:Hyperspectral unmixing allows representing mixed pixels as a set of pure materials weighted by their abundances. Spectral features alone are often insufficient, so it is common to rely on other features of the scene. Matrix models become insufficient when the hyperspectral image (HSI) is represented as a high-order tensor with additional features in a multimodal, multifeature framework. Tensor models such as canonical polyadic decomposition allow for this kind of unmixing but lack a general framework and interpretability of the results. In this article, we propose an interpretable methodological framework for low-rank multifeature hyperspectral unmixing based on tensor decomposition (MultiHU-TD) that incorporates the abundance sum-to-one constraint in the alternating optimization alternating direction method of multipliers (ADMM) algorithm and provide in-depth mathematical, physical, and graphical interpretation and connections with the extended linear mixing model. As additional features, we propose to incorporate mathematical morphology and reframe a previous work on neighborhood patches within MultiHU-TD. Experiments on real HSIs showcase the interpretability of the model and the analysis of the results. Python and MATLAB implementations are made available on GitHub.

人工智能

[AI-0] CodeMidas: Scaling Agent ic Coding RL Environments from Code Itself

链接: https://arxiv.org/abs/2609.22068
作者: Bowen Ye,Lei Li,Shicheng Li,Zihao Yue,Linghao Zhang,Hanglong Lv,Yuanxin Liu,Wenhan Ma,Hao Tian,Rang Li,Jinhao Dong,Yikai Zhao,Xiangwei Deng,Hailin Zhang,Liang Zhao,Qi Liu,Lingpeng Kong,Tong Yang,Fuli Luo
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers. Open-source codebases offer a rich source of such tasks, while existing methods typically rely on development artifacts such as issues and commits, limiting the range of tasks that can be extracted. To better scale RL environments, we present CodeMidas, an agentic pipeline that turns implemented functionality in existing codebases into executable RL environments using source code as its only task-specific input. CodeMidas allocates agentic compute to every stage of environment construction: agents explore implemented functionality to formulate behavioral specifications, construct tests grounded in execution of the original code, and validate and filter candidate tasks through execution checks and repeated solution rollouts. The resulting dataset has 5,545 training tasks from 3,185 open-source codebases spanning 23 programming languages and 15 technical domains. Training MiMo-V2.5 on these tasks with GRPO improves performance on all five diverse benchmarks, covering issue repair (DeepSWE + 11.7%), whole-program construction (ProgramBench +17%), and terminal work (Terminal-Bench v2.1 +8.5%). Ablations show that increasing the number of high-quality training tasks improves performance. Trajectory analysis shows the RL-trained agent demonstrates better behaviors like increasing codebase exploration and more diverse self-verification. These results establish source code as a scalable foundation for constructing RL environments that improve coding agents across diverse software tasks.

[AI-1] A Lie Detector Test for Language Models: Reading Knowledge a Model Wont Reveal

链接: https://arxiv.org/abs/2609.21996
作者: Hiskias Dingeto
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models can hold knowledge they do not report. A model may sandbag on a capability evaluation, or answer against what it internally knows, and its outputs alone cannot tell whether it is hiding an answer or simply does not have one. We borrow the Concealed Information Test, a forensic method that identifies guilty knowledge by presenting a suspect with the true detail among plausible decoys and measuring a stronger response to the item they recognize. Our method, Probe of Internal Recognition (PIR), does the same inside a model. It presents a question with its candidate answers and reads, from the model’s internal states, which candidate the model recognizes as correct. PIR is reference-free, needing no honest reference model and no labeled truth corpus. Across eight models from five families (Gemma, Qwen, Llama, Mistral, and Phi), PIR recovers the recognized answer at 0.70 to 0.87 balanced accuracy, well above the 0.28 to 0.40 unknown-item baseline and the 0.25 chance rate. It stays readable across every form of concealment we test, from prompted deception and trained sandbagging to external password-locked and circuit-broken checkpoints, with recognition between 0.85 and 0.93. When the model hides a known answer, recognition stays high. When unlearning removes the knowledge, recognition drops to the level of a question the model never knew. PIR therefore separates a model that will not answer from one that cannot, which supports sandbagging audits and unlearning verification. The signal is causal, adds information beyond black-box behavioral cues, and extends from multiple-choice questions to free-form generation.

[AI-2] Learning Cardiac Features: ECG Biometrics Across Time and~Exercise

链接: https://arxiv.org/abs/2609.21962
作者: Luca Thiebaud(AMU, AMU SCI, DIAPRO, LIS),Paul Chauchat(AMU SCI, AMU, LIS, DIAPRO),Mustapha Ouladsine(AMU SCI, AMU, LIS, DIAPRO),Stéphane Delliaux(AMU, APHM, C2VN)
类目: Artificial Intelligence (cs.AI); Tissues and Organs (q-bio.TO)
备注:

点击查看摘要

Abstract:Electrocardiograms (ECGs) carry subject-specific patterns enabling reliable individual discrimination, forming the basis of ECG biometrics. Beyond authentication, this paradigm holds significant potential to secure sensitive cardiac data and to serve as a pretext task in self-supervised learning. Yet, most studies remain confined to singlesession, resting data, leaving robustness to temporal and physiological variations largely untested. We address this gap by evaluating ECG biometrics under realistic conditions involving exercise-induced stress and cross-session variability. A Siamese ResNet with late multi-lead fusion strategy is trained on a large ECG dataset extracted from cardiopulmonary exercise tests and evaluated with a exercise-and time-aware protocol, as well as on public benchmarks. This first extensive assessment of ECG biometrics under combined physiological and temporal variability achieves an intra-session rest-to-peak EER of 1.7% and stateof-the-art 3.9% on the CYBHi dataset. Findings support the presence of an intrinsic cardiac signature resilient to physiological and temporal drift.

[AI-3] AutoViewMem: Self-Configuring Orthogonal Views for Conversational Long-Term Memory

链接: https://arxiv.org/abs/2609.21940
作者: Zijie Cao,Xijun Qu,Zhicheng Gu,Xiaoshu Chen,Duanyang Yuan,Yanning Hou,Sihang Zhou,Jianxing Gong,Jian Huang,Yang Mei
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Long-term memory is essential for large language model (LLM) agents to maintain consistency and personalization over extended interactions. Existing memory systems typically rely on fixed granularities or static schemas, but these designs struggle when heterogeneous information, such as preferences, events, constraints, and temporal updates, is embedded in a single mixed representation. The resulting semantic interference makes top-K retrieval sensitive to noise and often leaves relevant evidence poorly ranked. We present AutoViewMem, a data-driven framework that organizes long-term conversational memory into self-configuring, low-overlap semantic views before indexing. AutoViewMem discovers candidate views from interaction traces, selects a compact complementary view set, and uses these views to guide write-time structured extraction of provenance-grounded memories. This representation-first design moves semantic disentanglement from retrieval time to write time, allowing standard top-K similarity search to retrieve focused evidence without explicit routing or iterative retrieval. We further apply offline consolidation to improve memory compactness and consistency. Experiments on the LoCoMo and PersonaMem benchmarks, under both Qwen3-8B and Qwen3-14B backbones, show that AutoViewMem improves long-horizon question answering and personalization over strong memory baselines while preserving a simple inference pipeline.

[AI-4] What Should We Ask Next? Retrieval-Aware Question Learning under Partial Evidence

链接: https://arxiv.org/abs/2609.21924
作者: Lyucheng Qian,John Yuehan Zhang,Pingyu Wang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Interactive retrieval under partial evidence is a sequential information-acquisition problem: an agent must decide which question will create the most useful evidence for the next retrieval update. Existing systems train this decision by imitating an offline ordering of candidate QA pairs, although question value is determined by the response it elicits and its downstream effect on retrieval. We establish that candidate discriminativeness and perceived usefulness provide weak supervision for this objective, then introduce RAVEL, a retrieval-aware online reinforcement learning framework for interactive person re-identification. RAVEL initializes from supervised question generation, observes the current Top-4 candidates directly, and optimizes the question policy with rank feedback from the full question-answer-retrieval loop. Experiments on Interactive-PEDES show that RAVEL delivers progressively stronger retrieval performance across five interaction rounds. Further analysis shows that RAVEL reallocates the questioning budget toward localized open-ended attributes, which provide more useful retrieval evidence and yield the largest gains on initially difficult queries.

[AI-5] Neural Cellular Automata Learn General Features in their Hidden Channels

链接: https://arxiv.org/abs/2609.21870
作者: Etienne Guichard,Stefano Nichele
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Modern deep learning models achieve impressive generalization through over-parameterization, but this paradigm often struggles with overfitting and memorization in few-shot regimes. Neural Cellular Automata (NCAs) offer a highly parameter-efficient alternative, yet research has focused primarily on their output, leaving the role of their internal hidden channels largely unexplored. In this paper, we investigate the internal dynamics of NCA hidden channels and introduce a novel transfer-learning mechanism that injects a pretrained teacher’s hidden states into a student model to guide early optimization. Evaluated on few-shot and scale-variant MNIST benchmarks, NCAs outperform comparable recurrent and feed-forward architectures, demonstrating superior generalization with a minimal parameter budget (~9,800 parameters). Mechanistic analysis reveals that the hidden channels decouple feature extraction from uniform classification consensus by absorbing morphological complexity and converging to mutually orthogonal states. Furthermore, we demonstrate that these hidden channels capture general, scale-invariant topological primitives rather than class-specific templates. This allows a student model to achieve strong few-shot performance on unseen classes using features transferred from a teacher trained only on a subset of digits (0-5). Our results highlight the potential of utilizing hidden-state dynamics as a robust, decentralized computational substrate for parameter-efficient transfer learning

[AI-6] EnterpriseVal: Quantifying the Efficacy Reliability and Value of Generative AI in the Enterprise

链接: https://arxiv.org/abs/2609.21841
作者: Abbas Raza Ali,Muhammad Ajmal Siddiqui,Moona Zahid
类目: Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
备注:

点击查看摘要

Abstract:Frontier language models now produce professional deliverables that expert graders judge to match human work on a substantial share of economically valuable tasks, yet most enterprise GenAI initiatives fail to show a measurable business effect and a large fraction of agentic projects are expected to be cancelled. We argue that this is substantially a measurement problem: public benchmarks answer “what can the model do?”, whereas a deployment decision requires “is this workflow fit, reliable, safe and worth scaling - here, on our data, under our controls?”. We present EnterpriseVal, a use-case-level evaluation system that closes this gap. It comprises (i) a formal specification of the use case and of the frozen socio-technical configuration under test, model, prompts, retrieval, tools, guardrails and human oversight, with an autonomy level and consequence tier that jointly set the required evaluation intensity; (ii) a metric catalogue spanning fidelity, utility, efficiency, reliability, assurance and oversight; (iii) a grading protocol that scales blinded expert judgement with calibrated LLM-as-judge scoring through prediction-powered inference; (iv) a two-tier threshold gate, stated as an executable algorithm, that maps metric vectors with confidence bounds to REJECT/CONDITIONAL/SCALE decisions; and (v) a value-and-risk model in which the reviewer catch rate is a measured parameter. We report a pilot across three workflows in a global bank. In credit-memo drafting, human-graded citation precision reached 88% and hallucination rate 1.6% for the best model against gates of 70% and 5%; in procedure transformation, analyst refinement effort fell from an estimated 27.4 to 2.9 hours per document. We separate established results, documented pilot evidence, the proposed system and open hypotheses, and specify the experiments required for full validation

[AI-7] Federated Deep Clustering Networks for High-Dimensional and Heterogeneous Data

链接: https://arxiv.org/abs/2609.21829
作者: Morris Stallmann,Charalampos S. Kouzinopoulos,Marcin Pietrasik,Anna Wilbik
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Accepted to the 4th International Conference on Federated Learning Technologies and Applications (FLTA 2026)

点击查看摘要

Abstract:Clustering high-dimensional data is a fundamental task in unsupervised machine learning with applications to a variety of domains. In the centralized data scenario, this task is commonly solved using deep clustering methods that utilize deep neural network architectures to learn clustering-friendly latent space representations. In Federated Learning, where data is distributed between clients and is private, deep clustering methods are less explored. In particular, recently introduced federated deep clustering methods, despite showing very promising performance, still fall short in reliably providing good performance if data across clients are non-identically-independently distributed. In this work, we introduce a generalization of Deep Clustering Networks to the federated scenario, named FedDCN, that simultaneously optimizes a reconstruction loss and a clustering loss. To ensure robustness and latent space alignment in non-identically-independently distributed data scenarios, FedDCN generates synthetic data augmentations, and its learning objective includes a geometric regularization for latent space alignment. Through experimental evaluation, the effectiveness of the approach under IID and non-IID assumptions is demonstrated, and future research directions are identified.

[AI-8] Matrix AdaGrad: Row-wise and Column-wise Adaptive Subgradient Methods

链接: https://arxiv.org/abs/2609.21815
作者: Wenpeng Zhang,Runsheng Yu,Peilin Zhao
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Adaptive optimization methods such as AdaGrad and Adam are widely used in modern neural-network training, but their adaptive scaling is primarily designed for vector-valued parameters and does not explicitly exploit matrix structure. Recent matrix-aware optimizers demonstrate the benefits of structured optimization, yet a general theoretical framework for deriving matrix-aware adaptivity comparable to that of AdaGrad remains lacking. In this work, we develop a general Online Mirror Descent framework with adaptive proximal functions for matrix-valued parameters, providing a principled approach to deriving matrix-aware adaptive optimization through online regret minimization. By introducing row-wise and column-wise matrix proximal functions and analyzing the resulting regret trade-off, we derive Row-wise Matrix AdaGrad (Row-AdaGrad) and Column-wise Matrix AdaGrad (Column-AdaGrad), with adaptive scaling determined by the accumulated row-wise or column-wise gradient norms. We establish regret guarantees and show that these matrix-aware bounds can be strictly tighter than those of entry-wise AdaGrad under structured gradients. Experiments on matrix factorization and deep neural-network training further demonstrate the benefits of aligning adaptive scaling with matrix structure, including improved optimization stability and trainability at larger learning rates and greater network depths.

[AI-9] LLM -Generated Feature Pools for Time Series Anomaly Detection

链接: https://arxiv.org/abs/2609.21801
作者: Youssef Attia El Hili,Malik Tiomoko,Corinne Ancourt
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:We study how far a simple statistical pipeline can go on univariate time series anomaly detection under a strict selection protocol. The method extracts a small pool of statistics over sliding windows, scores each window with a transductive robust (MAD) model, and selects a feature subset per domain on a held-out tuning split. On TSB-AD-U it reaches 0.529 per-series VUS-PR, above the best neural ( 0.45 ) and statistical ( 0.44 ) entries on the public leaderboard and within 0.06 of the strongest pretrained foundation model, several of which use more supervision than ours. Ablations locate the cause: across three selection strategies and a hindsight oracle the score moves by 0.031 , and across the aggregation grid by 0.096 , while changing the candidate pool moves it by 0.226 . The candidate pool sets the ceiling; the search over it is second-order. We therefore generate a pool per domain by prompting a multimodal LLM with in-context example windows from that domain. The generated pools match the hand-crafted one under matched selection, and the two cover different domains: selecting over their union improves on the generated pool in all twelve generator-seed pairs and lifts the pipeline to 0.588 , matching the performance of the best entry on the leaderboard.

[AI-10] ECG Mirag e: Revealing and Mitigating the Underutilisation of ECGs in Vision-Language Models for Clinical Prediction

链接: https://arxiv.org/abs/2609.21755
作者: Jinning Liang,Mingcheng Zhu,Tingting Zhu
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Emergency department (ED) decision-making relies on heterogeneous clinical information, including patient history, vital signs, laboratory results, and electrocardiograms (ECGs). Vision–language models (VLMs) can jointly process these modalities, but strong predictive performance does not necessarily imply meaningful use of the correct patient’s ECG. We term this failure mode ECG Mirage: apparent multimodal capability without useful dependence on patient-specific ECG information. We distinguish two forms: ECG neglect, where ECGs provide little predictive benefit, and ECG confusion, where matched ECGs outperform no-image inputs but not mismatched ECGs. To evaluate these behaviours, we compare predictions obtained with matched ECGs, outcome-discordant mismatched ECGs, and no-image inputs while holding the clinical text and prediction targets fixed. Across four VLMs on MDS-ED, matched ECGs provide no consistent advantage for either ICU admission or clinical deterioration prediction. We then train four restricted visual prompts using supervised learning followed by conditional direct preference optimisation, while keeping the VLM backbone frozen. The resulting models achieve balanced accuracies of 70.6% for ICU admission and 67.5% for deterioration and increase the matched-versus-mismatched performance gap to approximately 16.5 and 5.5 percentage points, respectively. Overall, our study identifies ECG Mirage in multimodal clinical prediction and introduces visual prompt tuning as an efficient mitigation strategy.

[AI-11] ForceTwin: Physics-informed Digital Twins for Robotic Manipulation from Instrumented Human Interaction

链接: https://arxiv.org/abs/2609.21751
作者: Tim Engelbracht,René Zurbrügg,Mayank Mittal,Marco Hutter,Marc Pollefeys,Hermann Blum,Zuria Bauer
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Manipulating objects requires understanding not only their motion, but also the physical properties that determine it. For articulated objects, these include inertia, friction, and mechanisms such as springs or door closers, whose effects can vary with configuration and velocity. Such properties are not directly observable from appearance: visually identical doors may require very different effort to manipulate. Existing digital-twin pipelines recover primarily kinematics or assign static physical parameters from visual and language priors, which can yield physically implausible estimates. As a result, state-dependent mechanism dynamics remain unidentified and are not represented in standard asset formats. We present ForceTwin, a system for identifying physics-informed digital twins of articulated objects from instrumented human interaction. A person probes an object using a handheld force-sensing gripper, providing synchronized poses and interaction forces from which we estimate the articulation, parametric dynamics including inertia, Coulomb friction, viscous damping, and a structured neural residual capturing state-dependent mechanism forces. ForceTwin nearly halves the inertial-parameter error of a VLM prior. As a feedforward dynamics model for impedance control on a Spot and a Franka FR3, ForceTwin achieves 87% goal completion across nine object-embodiment pairs, compared with 60% using VLM-prior and 57% using kinematics-only twins, with the largest gains on objects whose strong mechanisms cause both baselines to stall. We further use the identified twins to train whole-body door-traversal policies and deploy them in the real world. Project Page: this https URL

[AI-12] ERMon: Detecting Persistent Behavioral Threats in Edge AI via Hardware-Native Ternary Runtime Monitor

链接: https://arxiv.org/abs/2609.21713
作者: Arish Sateesan,Edlira Dushku
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Edge AI accelerators are increasingly deployed in safety-critical environments, where model outputs may control physical actuators, make access-control decisions, or trigger alarms. In these settings, runtime failures often remain undetected because model corruption, distribution shift, and adversarial inputs can still produce well-formed, confident predictions. This paper presents TERMon, a lightweight hardware runtime monitor that detects such anomalies by observing inference behavior rather than re-executing or formally verifying the model. TERMon represents class-conditional trusted behavior as hardware-efficient ternary patterns that are matched in parallel against a thermometer-encoded fingerprint. The ternary encoding reproduces the corresponding unquantized range decision exactly. TERMon detects harmful weight corruptions in proportion to their behavioral impact, while out-of-distribution and adversarial inputs are largely not separable using the monitored features at a strict false-positive operating point. We implemented TERMon on a PYNQ-Z2 FPGA, and the pipelined design requires no on-chip block RAM or DSPs and has a two-cycle decision latency.

[AI-13] CIPL: A Channel-Aware Framework for Recoverable Privacy Leakage in LLM Agents

链接: https://arxiv.org/abs/2609.21686
作者: Tao Huang,Guosen Wu,Guolong Zheng,Jiayang Meng,Chen Hou,Xu Yang,Xuechao Yang,Feng Xia
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 58 pages, 4 figures; includes appendix

点击查看摘要

Abstract:Privacy leakage in LLM agents is commonly evaluated within individual components such as memory, retrieval, or tool-use pipelines, which makes it difficult to distinguish internal exposure from information that an external observer can actually recover. We present CIPL (Channel Inversion for Privacy Leakage), a channel-aware evaluation framework for black-box privacy leakage in LLM agents. CIPL represents a target through sensitive source, selection, assembly, execution, observation, and extraction stages and evaluates the transition from selected sensitive units to attacker-recoverable output under a shared protocol. Experiments across memory-based, retrieval-mediated, and tool-mediated targets, together with a BrowserUse live-agent case study, show that storage labels alone do not determine recoverability. Memory targets form a near-saturated reference case, retrieval-mediated leakage is frequently partial, and tool-mediated and live-agent leakage varies strongly with observation surface, prompt-to-channel alignment, retrieval depth, and provider behavior. A stratified semantic audit further identifies attacker-useful disclosures that canonical exact matching misses. CIPL therefore provides a common framework for comparing how internal sensitive dependence is realized as externally recoverable leakage across heterogeneous agent pipelines.

[AI-14] GUARD: Natural Forgetting in Large Reasoning Models via Guided Answer-Reasoning Distillation EMNLP2026

链接: https://arxiv.org/abs/2609.21677
作者: Zeyu Yan,Guanghao Zhou,Minghui Qiu,Ming Gao,Cen Chen
类目: Artificial Intelligence (cs.AI)
备注: Accepted to EMNLP 2026 main conference

点击查看摘要

Abstract:Recent advances in large reasoning models (LRMs) have made machine unlearning more challenging, as protected facts or unsafe rationales may surface in intermediate chain-of-thought (CoT) traces before the final answer is produced. Existing unlearning objectives typically suppress the target content or redirect internal representations, but they never specify how the post-forgetting trajectory should continue, which can lead to hallucinated substitutes, malformed boundaries, or repetitive outputs. We argue that LRM unlearning should instead learn a natural forgetting trajectory: a coherent non-disclosing CoT followed by a stable refusal-style answer that replace the original disclosure. To this end, we propose Guided Answer-Reasoning Distillation (GUARD), which converts model-generated unsafe disclosures into safe-exit trajectories, aligns a frozen LRM via guidance tokens, and distills the guided behavior into model this http URL address the lack of metrics for replacement quality beyond leakage, we further introduce Natural Forgetting Reasoning Score (NFRS), which captures structural stability, fluency, and unsupported substitutes in forgotten outputs. Extensive experiments on R-TOFU and a STAR-1-derived harmful-intent setting show that GUARD substantially reduces unsafe and privacy disclosures across two widely adopted distilled LRMs while preserving reasoning utility. Codes are available at this https URL

[AI-15] From Code Archival to Knowledge Graph: Bridging Software Heritage COAR Notify and Wikidata ISWC2026

链接: https://arxiv.org/abs/2609.21667
作者: Camillo Carlo Pellizzari di San Girolamo,Francesco Tosoni
类目: Digital Libraries (cs.DL); Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注: 15 pages, 4 figures. Accepted at the 7th Wikidata Workshop (Wikidata 2026), co-located with ISWC 2026. Open-source pipeline and code available at this https URL

点击查看摘要

Abstract:Software is a first-class scientific object, yet validated links between source code and the scholarly record remain largely absent from the Linked Open Data (LOD) cloud, isolating archived artefacts from semantic discovery. This paper presents an end-to-end reconciliation pipeline that harvests, validates, and models publication-to-repository pairs from sources where the link between a paper and its source code is explicit and editorially verified: the software-centric journals JOSS, SoftwareX, and IPOL, together with the reproducibility reports of the SIGMOD Availability and Reproducibility Initiative (ARI). This yields a curated corpus of 4,397 \langle DOI, repository-URL \rangle pairs. We design two distinct application profiles grounded in Wikidata classes (one for scholarly articles, one for software instances) aligned with the this http URL and CodeMeta vocabularies. This architectural separation enables rule-based reconciliation at two granularities: lightweight, inline publication references or standalone, first-class Wikidata software nodes equipped with SWHIDs, Software Heritage’s content-addressed identifiers. A read-only lookup against Wikidata shows that only 82 of the harvested repositories were already modelled there; human-reviewed batches have since created 4,182 new software items cross-linked to their articles. We further show that payloads of the emerging COAR Notify protocol, an external effort we do not develop, map natively onto our input format, so the same backend could later serve a live enrichment stream. Our core contribution is a pair of application profiles that turn Wikidata into a connector between the scholarly record and archived source code; we openly release all code, application profiles, and harvested datasets.

[AI-16] Outcome-Conditioned End-Effector Geometry Across Vision-Language-Action Policies ICRA2027

链接: https://arxiv.org/abs/2609.21659
作者: Xingyu Lin,Zhuang Li,Zhongrun Wu,Shouquan Zhou,Dehui Du
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: 8 pages, 3 figures, 7 tables, 23 references. Submitted to ICRA 2027

点击查看摘要

Abstract:Vision-language-action (VLA) policies solve the same manipulation task through different action interfaces, but task success alone does not establish whether their physical executions agree. We study cross-policy end-effector geometry in 15,000 closed-loop LIBERO rollouts from four policies. The primary clean-condition analysis forms 3,600 configuration-matched, and therefore dependent, policy pairs. Both-success pairs have a median normalized dynamic time warping distance of 0.0120 m versus 0.0380 m when exactly one policy succeeds. This ordering holds in every task, every policy pair, and nine sampling and band-limited representations; however, the ratio varies severalfold across representations, so we report the direction rather than a fixed multiple. Both-failure pairs are more separated again but rest on thin, uneven support, so we report them as exploratory. Within successful executions, partner replacements separate more across tasks than across initial states. A matched baseline still reveals measurable, heterogeneous residual policy differences, so a low cross-policy distance does not imply interchangeability. Successful executions sit about as far from same-task demonstrations as those demonstrations sit from each other, compatible with task-associated geometry without separating training-data overlap from task constraints. A common 72-action window preserves the ordering but reduces its magnitude; endpoint and duration adjustment likewise leaves a positive mixed-outcome coefficient relative to both-success pairs, though its magnitude is specification-dependent. Under composite visual stress, policy rankings and pair composition change together.

[AI-17] SynthDemo-RL: Breaking the Zero-Reward Barrier in VLA Adaptation with LLM -Guided Synthetic Demonstrations

链接: https://arxiv.org/abs/2609.21650
作者: Hiroaki Kingetsu,Hiroaki Kurihara,Kaoru Yokoo,Kenji Fukumizu,Manohar Kaul
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: Under review

点击查看摘要

Abstract:Fine-tuning Vision-Language-Action (VLA) models commonly relies on human teleoperation demonstrations, while reinforcement learning (RL) with sparse binary rewards faces an exploration challenge when successful trajectories are rarely sampled. We propose SynthDemo-RL, a teacher-student framework in which an automated teacher converts simulator-privileged state into successful manipulation trajectories, a VLA student is distilled from them by supervised fine-tuning (SFT), and PPO with binary task-success rewards refines the student. We study reward coverage, the fraction of tasks for which at least one success is observed under the fixed evaluation protocol, as a complement to the average success rate. On LIBERO-PRO, a public benchmark of perturbed LIBERO tasks for which no demonstrations exist, 27 of 57 scored tasks are at exactly 0% success for a pi_0.5 policy fine-tuned on the original LIBERO tasks. Direct PPO from this policy, under the same PPO recipe and the same RL compute as SynthDemo-RL’s refinement stage, rescues 10 of these 27 tasks and leaves 17 at 0%. SynthDemo-RL, with 50 synthesized trajectories per task and no new human demonstrations, rescues all 27 and reaches average success rates of 97.8% and 97.1% on the Position and Task axes of LIBERO-PRO, respectively. On standard LIBERO, the same pipeline reaches 96.0% with no human demonstrations, within 1.7 points of pi_0.5 trained on 50 human demonstrations per task. We further validate the pipeline on RoboTwin 2.0 and verify that trajectories from a policy trained in a MuJoCo twin execute open-loop on a physical robot.

[AI-18] One Prompt Does Not Fit All: Self-Meta-Evolve for Personalized Information Extraction AACL

链接: https://arxiv.org/abs/2609.21626
作者: Hongliang Li,Lu Wang,Yong Xu,Hanyang Chen,Zhitao Hou,Xiaoting Qin,Song Ge,Qingwei Lin,Dongmei Zhang
类目: Artificial Intelligence (cs.AI)
备注: 16 pages, 4 figures, Findings of AACL-IJCNLP 2026

点击查看摘要

Abstract:Large language models (LLMs) are increasingly deployed for enterprise information extraction (IE), where the same document must be reorganized differently for each user. Existing prompt optimization methods, however, rely on a single prompt optimized against a global objective, which is misaligned with the inherent user heterogeneity of real workplaces. We formulate enterprise IE as per-user prompt adaptation under interaction feedback and propose Self-Meta-Evolve, a hierarchical framework that maintains a dedicated prompt for each user and continuously refines it through a dual-loop process: an inner loop that edits structured prompts based on persona-conditioned feedback, and an outer loop that evolves the meta-prompt itself by distilling successful editing patterns. To enable scalable training and evaluation, we release a persona-driven IE benchmark of 292 simulated enterprise users, paired with a reproducible persona-generation pipeline grounded in O*NET occupational taxonomies. On this benchmark, Self-Meta-Evolve achieves a 74.58% success rate, outperforming the strongest prompt-optimization baseline by 13.56 absolute points, and reaches 52.54% within only two iterations. A double-blind human study with twenty real professionals further confirms that prompts adapted by our framework win against static baselines in 71% of pairwise comparisons.

[AI-19] Calibrating Teacher–Student Discrepancy for On-Policy Distillation

链接: https://arxiv.org/abs/2609.21619
作者: Qiangqiang He,Jin Li,MingCai Chen
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:On-policy distillation (OPD) improves reasoning models by learning the token-level discrepancy between a stronger teacher and an on-policy student. However, this discrepancy does not purely reflect the capability gap between the teacher and the student: it also contains deviations arising from the teacher itself, which are consequently mixed into the observed teacher–student discrepancy and indiscriminately learned by standard OPD during training. This issue is further exacerbated by privileged OPD, where privileged information induces larger teacher-side likelihood shifts, thereby encouraging the student to learn more of the teacher’s own deviation. We introduce \textbfCalibrated On-Policy Distillation (Cal-OPD), which estimates the teacher’s self-deviation region through positive and negative privileged interventions and calibrates the original teacher–student discrepancy by retaining only the component that lies beyond this region. Experiments on mathematical reasoning benchmarks show that, while retaining only about 52–65% of the original teacher–student discrepancy as the optimization signal, Cal-OPD consistently outperforms standard OPD and its variants across model scales.

[AI-20] Potential-Field Action Representation for Reinforcement Learning in Contact-Rich Manipulation

链接: https://arxiv.org/abs/2609.21609
作者: Xinyu Liu,Gökhan Solak,Arash Ajoudani
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Model-free reinforcement learning can acquire contact-rich robotic manipulation skills through trial-and-error interaction, but it often requires the policy to learn both task strategy and low-level motion generation. In this setting, the action representation is critical because it determines how policy outputs are converted into robot motion, shaping both exploration and physical execution. Direct Cartesian command interfaces require the policy to generate motion at every decision step, coupling task-level adaptation with continuous low-level control and increasing the learning burden. We propose PA-RL, a reinforcement-learning framework that uses artificial potential fields as the action representation. Instead of commanding motion directly, the policy adapts the parameters of an energy-like potential field, which generates a state-dependent guidance direction executed through a Cartesian impedance controller. We evaluate PA-RL on peg-in-hole insertion, a representative contact-rich task with nonlinear dynamics and discontinuous contact transitions. In simulation, PA-RL is compared with Cartesian velocity, Cartesian pose, and variable-impedance action spaces using the same RL algorithm. PA-RL is the only method to reach a 100% evaluation success rate within the allotted training time, while the best baseline reaches 92.6%. It also reduces joint-torque variation by 55.4% and Cartesian acceleration variation by 70.8% relative to the best baseline, without explicit motion-quality penalties in the reward. The simulation-trained policy further completes 9/9 real-robot insertions without fine-tuning, demonstrating the deployment feasibility of the learned potential-field interface.

[AI-21] Beyond Accuracy: Centroid-Guided Contrastive Loss for Structured Fraudulent Job Posting Detection

链接: https://arxiv.org/abs/2609.21599
作者: Syed Ali Ahmed(1),Malaika Raza(1),Muhammad Shoaib Siddiqui(2),Muhammad Rafi(1) ((1) National University of Computer and Emerging Sciences, Karachi, Pakistan, (2) Islamic University of Madinah, Madinah, Saudi Arabia)
类目: Artificial Intelligence (cs.AI)
备注: 12 pages, 6 figures, 8 tables. Submitted to IEEE Open Journal of the Computer Society. Code: this https URL

点击查看摘要

Abstract:Fraudulent job posting detection aims to identify job advertisements that are corrupted either through fake content, misleading information, or negative intent, disrupting the online eco-system of job-seekers and employers. Existing studies in this domain lack effective methods to simultaneously achieve high accuracy and meaningful structure of latent-space representations that capture subtleties among fake posts. To this end, we propose Centroid-Guided Contrastive Loss (CGCL), a loss function which unifies classification with densely formulated clustering to consistently reshape latent-space through a centroid-driven top- k push-and-pull mechanism. The complementary nature of CGCL enables the model to enforce accurate decision boundaries and maintain high clustering compactness, effectively capturing both class separability and latent structure. Extensive experiments demonstrate the state-of-the-art (SOTA) performance of our method on EMSCAD, a public benchmark dataset. The code associated with this work is available at: this https URL

[AI-22] Micro-Collaborative Poisoning: A Distributed Attack on RAG Systems ESORICS

链接: https://arxiv.org/abs/2609.21573
作者: Pedro Pereira,Eva Maia,Isabel Praça
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 19 pages, 3 images, 4 tables, conference: 31st European Symposium on Research in Computer Security (ESORICS) 2026, Workshop: 2nd Workshop on the Use of Large Language Models for Cybersecurity

点击查看摘要

Abstract:Retrieval-Augmented Generation (RAG) improves large language models by grounding outputs in external knowledge sources, but this dependency also creates a surface for poisoning attacks. This paper introduces Micro-Collaborative Poisoning, a distributed attack in which a false target claim is divided across multiple locally plausible documents instead of being concentrated in a single malicious passage. We evaluate the attack across 108 RAG configurations by varying dataset, retriever architecture, retrieval depth, database composition, number of poisoned databases, and generator model. The results indicate that Micro-Collaborative Poisoning is not driven by a single dominant poisoned passage, but by the accumulation of weak adversarial signals across retrieved sources. Increasing top- k and poisoning multiple databases make it more likely that these signals will appear together in the retrieved context, while clean database diversity and stronger retrievers can reduce their influence. The document-level poisoning visibility analysis further shows that this threat is difficult to expose through isolated document inspection, since Micro-Collaborative Poisoning achieves downstream influence while leaving a weaker explicit poisoning signature than direct poisoning.

[AI-23] On Repulsive and Attractive Teachers: Separating Correctness from Behavior in Self-Distillation

链接: https://arxiv.org/abs/2609.21561
作者: Anton Baumann,Akmal Ashirmatov,Leo Schmidt-Traub,Frederike Lübeck,Jonas Hübotter,Thomas Kleine Buening,Andreas Krause
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:On-policy self-distillation provides dense, token-level supervision by conditioning a model on privileged information and distilling the resulting teacher distribution back into the model. However, privileged information can change not only what the teacher knows, but also how it behaves, entangling correctness-relevant learning signals with unintended behavioral shifts. We study this effect in reasoning tasks by contrasting attractive self-distillation, which moves the model toward a privileged teacher, with repulsive self-distillation, which moves it away from a privileged teacher. We find that both objectives can induce strong and opposing behavioral shifts: attraction suppresses exploratory reasoning and promotes shorter, more confident responses, whereas repulsion increases response length, can trigger unintended switches into a model’s latent thinking mode, and ultimately becomes unstable. Motivated by these observations, we study contrastive self-distillation, which combines attraction toward a correct-solution-conditioned teacher with repulsion from an incorrect-solution-conditioned teacher. In contrast to prior work that combines such distillation signals with a GRPO objective, we isolate the self-distillation objective and study its behavior on its own. We find that the shared behavioral shifts of the two teachers largely cancel, leaving a token-level signal that more directly reflects correctness. Across non-thinking, instruct-only, and already-thinking models, this contrastive objective improves reasoning performance while maintaining stable response lengths.

[AI-24] OneBid: A Unified Auto-Bidding Foundation Model for Diverse oCPX Advertising Scenarios

链接: https://arxiv.org/abs/2609.21550
作者: Yewen Li,Peng Jiang,Yitian Li,Pengfei Lv,Xialong Liu,Peng Jiang,Qingpeng Cai
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Auto-bidding is central to computational advertising, where strategies must maximize advertisers’ conversion value under economic constraints. It has evolved from rule-based controllers to reinforcement learning and generative methods such as Decision Transformer (DT). Yet these methods increasingly mismatch the prevailing optimized cost-per-X (oCPX) paradigm, which spans heterogeneous scenarios (e.g., registration, purchase), each served by a separate model, leading to fragmented pipelines and underexploring cross-scenario modeling. Inspired by foundation models like LLMs, unifying these oCPX scenarios into one model raises three challenges: multi-objective control, scalable capacity under strict latency, and safe offline policy improvement. We present OneBid, a unified auto-bidding foundation model that learns a reusable backbone from heterogeneous oCPX logs and adapts it to scenario-specific deployments via offline post-training. Building on DT, OneBid extends single Return-to-Go conditioning to two atomic signals, Return-to-Go for conversion value and Cost-to-Go for cost ratio, plus value-aware regularization on next-action prediction. To absorb distributional heterogeneity, we design a sequence-level Mixture-of-Experts architecture, where shared experts encode cross-scenario knowledge and sparsely-routed experts capture scenario-specific patterns at low latency, yielding consistent scaling with model size and data. During post-training, we align the backbone with scenario preferences via Critic-guided Relative Offline Policy optimization (CROP): a learned critic scores candidate actions group-relatively, avoiding the unsafe online exploration of GRPO-style fine-tuning while constraining policy shift to reduce OOD risk. Validated via online A/B tests and fully deployed at Kuaishou, OneBid delivers an overall +2.2% ADVV gain on oCPX Ads, peaking at +13.1% in the ROAS scenario.

[AI-25] Dual-Interest Sequential Product Recommendation With Multi-Granular SSM

链接: https://arxiv.org/abs/2609.21548
作者: Shuiying Liao,P. Y. Mok
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Sequential recommendation aims to predict the next item a user will interact with based on their historical behavior. Advances in Transformers have significantly improved sequential recommendation but are still limited by cost efficiency. Although State Space Models (SSMs) have recently enabled efficient long-range modeling, most existing methods encode each item with a single static contextual role, overlooking the phenomenon of item polysemy. In fact, the same item often plays different semantic roles depending on user context, and existing methods are limited in capturing dynamic behavior across different temporal granularities. In this work, we propose DSRec, a novel dual-interest cross-SSM model that explicitly disentangles item roles across long-term and short-term semantic context. Sequential items are encoded into long-term interest embeddings that capture stable preferences via historical aggregation, and a short-term interest branch that emphasizes local session intent modulated by inter-click time intervals. These interest embeddings are processed through distinct SSM encoders: a full-sequence Mamba for long-term modeling, and a time-modulated SSM that dynamically adjusts state evolution based on temporal gaps. To enable effective cross-granularity alignment, we adopt a residual cross-fusion mechanism that exchanges contextual information between the two branches while preserving semantic independence. Experiments on public benchmarks demonstrate that DSRec outperforms other state-of-the-art methods.

[AI-26] Learning-to-Optimize as the Missing Architectural Layer of AI-Native Networks

链接: https://arxiv.org/abs/2609.21519
作者: Giambattista Amati,Federica Mangiatordi,Pierpaolo Salvo,Emiliano Pallotti,Simone Angelini
类目: Artificial Intelligence (cs.AI)
备注: 6 pages, conference

点击查看摘要

Abstract:Artificial Intelligence (AI) is becoming a fundamental design principle of future AI-native communication networks, enabling autonomous resource management, adaptive control, and zero-touch network operation. While current AI-native architectures increasingly embed intelligence across network functions, they provide little guidance on how optimisation knowledge should be systematically generated, transferred, and exploited by AI models. This paper argues that the Learning-to-Optimize (L2O) represents the missing architectural layer between optimisation and AI-native intelligence. Rather than viewing optimisation merely as an online decision engine, the proposed paradigm redefines optimisation algorithms as offline knowledge generators that produce high-quality supervisory information for neural surrogate models. The resulting models inherit optimisation expertise while enabling low-latency runtime inference suitable for dynamic network environments. A generic four-stage L2O workflow is introduced, comprising optimisation, knowledge generation, surrogate learning, and runtime inference. Unlike existing Learning-to-Optimize approaches, which primarily focus on algorithm acceleration, the proposed framework establishes L2O as an architectural abstraction applicable across heterogeneous communication and computing systems. The proposed paradigm is illustrated by an NR-V2X relay-selection problem, in which optimisation-generated solutions from a Mixed-Integer Linear Programming (MILP) solver are used to train a Graph Neural Network that can reproduce near-optimal decisions in real time. The presented perspective positions Learning-to-Optimize as a key architectural enabler for future AI-native networks.

[AI-27] he Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models

链接: https://arxiv.org/abs/2609.21509
作者: Xavier Suau,Alex Ferrando de las Morenas,Luca Zappella,Samy Bengio
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:When language models reason in chain-of-thought or exchange free-text intermediates, they serialize structured information into natural language. How much tree-structured compositional content survives this bottleneck? We propose a round-trip protocol that answers this question empirically for tree-structured expressions. A generator converts a procedurally generated arithmetic expression into a word problem, a separate extractor recovers the expression from the word problem alone, and symbolic equivalence provides an exact oracle. Evaluating all pairwise combinations of sixteen models yields a communication matrix whose marginals separate generation quality from extraction quality. Three main findings emerge. First, the channel is lossy and asymmetric: swapping which model generates and which extracts shifts accuracy by up to 60.4 points, and the best pair reaches 92.9% by combining different models on each end rather than the same model on both. Second, at least 73.6% of round-trip failures originate at generation, and difficulty is driven by tree structure (operator count, depth, right-branching) rather than model family. Third, the channel is trainable: ~3600 fine-tuning examples that share the evaluation’s operators and tree shapes lift every open-weight model above untrained Gemini-3.1-Pro, an upper bound under matched semantics. A disjoint-domain regime with new operators and vocabulary also raises every open-weight model, confirming the gain is not an artifact of matched semantics, though a gap to the frontier remains. Together these results identify tree-structured expression serialization as a primary limiting factor when models communicate hierarchical structure through natural language.

[AI-28] PolyBridgeBench: Benchmarking Multimodal LLM s for Physics-Grounded Bridge Design

链接: https://arxiv.org/abs/2609.21493
作者: Zicheng Zhao,Dongyin Chen,Rui Xu,Yinghui Xu
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Multimodal large language models, or MLLMs, perform well at visual understanding and structured generation, yet these capabilities do not establish whether an engineering design will work when executed. Existing benchmarks assess spatial reasoning, structural validity, or physics-grounded construction, but they do not determine whether MLLMs can synthesize complete load-bearing structures and repair them after simulator execution exposes a failure. We introduce PolyBridgeBench, an executable benchmark for multimodal bridge design. A model receives a visual scene and structured engineering constraints and generates a complete node–member–material topology. Deterministic legality checks gate execution in a native dynamic physics simulation. Following an execution failure, the benchmark returns temporal visual evidence from the failed rollout and evaluates repair under a fixed interaction budget. Separate measurements of deterministic validity, dynamic functional success, and post-failure recovery identify the stage at which design fails. Experiments with six representative MLLMs across 189 levels expose a substantial gap between deterministic validity and dynamic success, pronounced sensitivity to material budgets, and limited post-failure recovery under the primary strict-budget setting.

[AI-29] LogicTrack: Auditing Reasoning Trajectories of Large Language Models with Formal Logic Solvers

链接: https://arxiv.org/abs/2609.21492
作者: Jingyu Hu,Shu Yang,Weiru Liu,Di Wang
类目: Artificial Intelligence (cs.AI); Logic in Computer Science (cs.LO); Symbolic Computation (cs.SC)
备注:

点击查看摘要

Abstract:Chain-of-Thought (CoT) reasoning has been shown to improve the performance of large language models (LLMs), yet existing optimization methods largely rely on outcome-based feedback, leaving the logical validity of intermediate reasoning steps largely unverified. To address the gap whereby LLMs arrive at correct final answers through logically flawed intermediate reasoning chains, we propose LogicTrack, a neuro-symbolic framework that audits reasoning trajectories by auto-formalizing each reasoning step into symbolic representations and verifying it with automated theorem provers. LogicTrack introduces Solver-Based Backtracking Reward (SBR), a step-wise scoring mechanism that quantifies logical soundness and guides backtracking tree search at inference time. We further extend LogicTrack to construct supervised fine-tuning (SFT) data with backtracking traces from its trajectories, enabling fine-tuned models to internalize step-wise auditing as an intrinsic capability. Extensive experiments across 8 reasoning benchmarks and 7 LLMs demonstrate that LogicTrack effectively improves both the verifiability of reasoning chains and final answer pass rate, thereby enhancing overall CoT quality and trustworthiness in high-stakes domains.

[AI-30] Driving on Registers Reasoning on Risk: Risk-Aware Occupancy for Register-Based End-to-End Autonomous Driving

链接: https://arxiv.org/abs/2609.21486
作者: Jiaxing Chen,Hengduo Zou,YuKai Qin,Yiren Zhao,Lidong Yu,Bolin Gao
类目: Artificial Intelligence (cs.AI)
备注: This version of this research was completed in early 2026

点击查看摘要

Abstract:Multimodal trajectory prediction improves behavioral coverage in end-to-end autonomous driving, but existing methods remain limited by sparse scene representations. Incomplete evidence leads to low-quality candidate generation and unreliable ranking among geometrically similar trajectories. On a register-based baseline, bad and poor candidates constitute 19.74% of the candidate set, while the oracle-best candidate ranks only 33.9th on average. We propose RRDrive, which introduces risk-aware occupancy as a dense, temporally aligned, and trajectory-queryable representation. Its global structure guides high-quality multimodal generation, while candidate-conditioned risk queries support fine-grained selection. We further construct RiskOcc4D-NAVSIM with automatic risk annotations. RRDrive achieves a selected-trajectory PDMS of 0.951, representing a 1.5% relative improvement over the baseline (0.937), and improves the average candidate PDMS by 7.7%. In challenging scenes, it improves candidate PDMS by 30.2% and increases the Spearman correlation among good candidates by 0.41, from 0.26 to 0.67. To move beyond this oracle setting, we further develop an external RiskOcc predictor, a perception module that estimates risk-aware occupancy directly from sensor inputs. The competitive performance validates the representation’s feasibility.

[AI-31] HE-Guardrail: A Homomorphic Guardrail Against Jailbreak Attacks for Encrypted Large Language Model Inference

链接: https://arxiv.org/abs/2609.21484
作者: Byeongseo Min,Yongwoo Lee,Young-Sik Kim,Yongjune Kim
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Homomorphic encryption (HE) has emerged as a promising approach to privacy-preserving machine learning (PPML), enabling computation directly over encrypted data. In HE-based PPML, a client submits an encrypted input to the server, which evaluates models such as large language models (LLMs) without access to the underlying plaintext. However, we identify a critical security vulnerability in this setting: HE-LLM inference is vulnerable to malicious clients that submit adversarial prompts, such as jailbreak attacks. The same confidentiality that protects benign clients also prevents the server from inspecting incoming prompts or generated responses, making adversarial attempts difficult to detect or block and potentially allowing successful attacks to remain entirely invisible to the server. To address this vulnerability, we propose HE-Guardrail, a framework that evaluates guardrail mechanisms entirely over encrypted data and homomorphically controls whether the target-model response is returned to the client. We instantiate HE-Guardrail with three representative guardrails - Llama Guard, JBShield, and GradSafe. Our results show that HE-Guardrail closely reproduces the decisions of the corresponding plaintext guardrails in the encrypted domain, with distinct security-efficiency-utility trade-offs.

[AI-32] Risk-Aware Occupancy for Safety-Oriented End-to-End Autonomous Driving

链接: https://arxiv.org/abs/2609.21470
作者: Jiaxing Chen,Hengduo Zou,Yiren Zhao,Bolin Gao
类目: Artificial Intelligence (cs.AI)
备注: The first version of this research was completed in early 2025

点击查看摘要

Abstract:Sparse representation formulates the environment perception for the end-to-end driving system as a set of discrete elements like objects and lane lines. This formulation meets safety risks in crowded, occluded scenes dealing with unstructured obstacles, uncertain regions, and intricate interactions. In this paper, we propose a dense representation, risk-aware occupancy, to characterize planning-relevant risks in an explicit and uniform manner. It jointly encodes global scene occupancy, map-derived traffic constraints, and future dynamic agent occupancy into a unified BEV map. The unified BEV map captures the risk evidence for trajectory planning in both spatial and temporal dimensions. We design an E2E network, ROIDrive, to realize risk-aware occupancy. It predicts risk-aware occupancy with an independent branch and injects it into planning queries for safety-oriented trajectory generation. In addition, to quantify the safety problem, we introduce RiskOcc4D-nuScenes built upon nuscenes and occ3d-nuscenes. Our risk-aware occupancy yields relative open-loop collision reductions of 52.9% under the UniAD metric and 35.0% under the ST-P3 metric on nuScenes.

[AI-33] AtomEgo: Exploring Ego-Robot Integration for Embodied Foundation Model Pretraining

链接: https://arxiv.org/abs/2609.21461
作者: Di Wu,Dongchen Zheng,Junhe Sheng,Zhongxing Wei,Songxin Zhang,Zejian Xie,Xiaoquan Sun,Junyang Zheng,Zhuoyang Song,Jiaxing Zhang,Jiayu Chen
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Embodied foundation models are constrained by the limited scale and diversity of robot demonstrations, motivating the use of large-scale egocentric human interaction data. However, how to effectively incorporate such data into embodied-model pre-training remains unclear because of substantial embodiment and action-space gaps between humans and robots. We present AtomEgo, a systematic study of ego–robot co-training supported by a curated corpus of approximately 2,659 hours and a scalable data processing pipeline. Across vision–language–action and world–action model architectures, we investigate three representative paradigms: joint co-training with domain-specific action heads, progressive ego-to-robot transfer through embodiment alignment, and joint video–action modeling. We evaluate these paradigms through multi-task real-robot experiments and language-conditioned cross-embodiment representation analysis. Our results reveal a simple principle: Data Scale * Alignment Quality – Capability Gain; egocentric data can improve generalization, but their value depends on how effectively they are aligned and utilized. This principle can provide practical guidance for scalable ego–robot pre-training.

[AI-34] Interference-Driven Clustered Optimisation for FM Spectrum Coordination

链接: https://arxiv.org/abs/2609.21441
作者: Federica Mangiatordi,Emiliano Pallotti
类目: Information Theory (cs.IT); Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE)
备注: 6 pages, conference

点击查看摘要

Abstract:Cross-border FM spectrum coordination involves protecting foreign broadcasting services while preserving domestic coverage, amid increasingly large radio-planning datasets containing thousands of transmitters and millions of transmitter-pixel relationships. In such scenarios, conventional optimisation approaches become computationally demanding due to the high dimensionality of the associated power-control problem. This paper proposes an interference-driven clustered optimisation framework for large-scale FM spectrum coordination. The proposed method exploits the observation that violations of foreign-service protection are typically dominated by a limited subset of transmitters. Protected services are therefore analysed to identify dominant interferers and quantify their impact on interference. These relationships are represented through an interference graph from which optimisation-oriented transmitter clusters are extracted. The clusters decompose the global power-control problem into smaller optimisation tasks solved with clustered simulated annealing, followed by a global refinement that captures residual inter-cluster interactions. Coverage and interference are evaluated using frequency-dependent protection criteria and a dynamic strongest-service assignment model. To enable operational-scale planning, the framework uses sparse matrices and GPU-accelerated computations. Tests on realistic cross-border FM coordination scenarios show that the clustering strategy greatly reduces optimisation complexity and runtime while maintaining foreign-service protection and domestic coverage. The method also yields an interpretable ranking of transmitters that contribute most to harmful interference, supporting optimisation and spectrum planning.

[AI-35] GVPO: Group Variance Policy Optimization for LLM Post-Training and On-Policy Distillation NEURIPS2025

链接: https://arxiv.org/abs/2609.21432
作者: Kaichen Zhang,Yuzhong Hong,Junwei Bao,Hongfei Jiang,Yang Song,Dingqian Hong,Hui Xiong
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Extended version of the NeurIPS 2025 paper “GVPO: Group Variance Policy Optimization for Large Language Model Post-Training”

点击查看摘要

Abstract:Post-training plays a pivotal role in enhancing the reasoning capabilities and task-specific expertise of large language models (LLMs). Despite recent advances in post-training methods, such as Group Relative Policy Optimization (GRPO), their practical deployment remains impeded by training instability arising from the reliance on importance sampling. We introduce Group Variance Policy Optimization (GVPO), a novel post-training method that integrates the analytical solution of KL-constrained reward maximization into its gradient weighting scheme. This formulation provides an intuitive interpretation: GVPO’s gradient corresponds to the mean squared error between the central distance of implicit rewards and that of actual rewards. GVPO offers two key advantages: (1) it guarantees a unique optimal solution, exactly to the KL-constrained reward maximization objective, and (2) it enables flexible sampling distributions without requiring importance sampling. Beyond general post-training, we show that GVPO naturally extends to on-policy distillation (OPD). Furthermore, GVPO enables the optimization of a broad family of extended OPD objectives, providing a principled foundation for diverse objective design. By unifying theoretical guarantees with practical adaptability, GVPO establishes a new paradigm for reliable and versatile LLM post-training and on-policy distillation. Comments: Extended version of the NeurIPS 2025 paper “GVPO: Group Variance Policy Optimization for Large Language Model Post-Training” Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG) Cite as: arXiv:2609.21432 [cs.AI] (or arXiv:2609.21432v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2609.21432 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-36] DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement

链接: https://arxiv.org/abs/2609.21423
作者: Siyuan Liu(1 and 2),Fan Yu(1 and 2),Dongyu Ru(2),Yizhu Liu(2),Yifan Yang(2),Xuezhi Cao(2),Xunliang Cai(2),Yixin Cao(1) ((1) Fudan University, (2) Meituan Longcat Team)
类目: Artificial Intelligence (cs.AI)
备注: 42 pages, including appendices

点击查看摘要

Abstract:Online agent deployments produce abundant execution traces, while task-specific verification and expert annotation are costly to scale. We study how to distill these traces into reusable feedback without post-hoc outcome labels, drawing on their evidence of local progress, recovery, and unfinished requirements. We introduce DENSE (Distilling Evidence from Nested Subtask Executions), which organizes this evidence into evidence-grounded nested shortcut trees. DENSE compresses redundant attempts, reconciles issues across levels using recovery evidence, and summarizes completed branches while expanding unresolved ones, linking reusable progress to remaining obligations. We introduce REFIT, a source-paired protocol comparing feedback from shared initial trajectories under post-hoc outcome blindness, with environments and model contexts reset for fresh attempts at the same tasks. On Terminal-Bench 2.1, DENSE achieves the highest strict pass rate among tested non-privileged feedback methods across four recipient models. Relative to initial executions, strict pass rate improves by 7.12-15.64 pp, with 19.0-43.6% fewer observed recipient tokens in reruns. GPT-5.5 ablations support combining nested subtask analysis with shortcut construction and issue reconciliation. These findings point toward agent self-refinement through evidence-grounded trajectory reuse with less reliance on external supervision.

[AI-37] Knowledge-Graph-Augmented Chronos-2 for HEC-RAS Surrogate Forecasting

链接: https://arxiv.org/abs/2609.21381
作者: Edward Holmberg,Elias Ioup,Mahdi Abdelguerfi
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 9 pages, 4 figures, 4 tables

点击查看摘要

Abstract:We investigate whether coupling a time-series foundation model to hydraulic project knowledge improves surrogate forecasting of HEC-RAS water-surface elevation (WSE). We present KG-Chronos-2, which combines a frozen Chronos-2 predictor with exact-state residual decoding, graph-conditioned historical retrieval, and input-aligned correction. We compare the method with persistence, a residual LSTM, project-conditioned recurrent GeoFNO, a hydraulic DCRNN-style model, and frozen Chronos-2. Task-specific fitting uses the 2008 simulation. Evaluation covers 64 fixed 24-hour windows from the 2011 and 2002 simulations at 4,675 cross sections in 71 reaches on a shared geometry. KG-Chronos-2 achieves event-balanced root-mean-square error 0.246970 in native WSE units. It reduces RMSE by 14.13% relative to frozen Chronos-2, 29.38% relative to the hydraulic DCRNN-style model, and 39.54% relative to recurrent GeoFNO. The 95% hierarchical-bootstrap interval for its event-balanced RMSE difference from frozen Chronos-2 is [-0.075177, -0.016317]. KG-Chronos-2 also achieves the lowest active-window and final-lead RMSE among the six completed systems. These results support coupling a frozen temporal predictor to project knowledge for warm-start HEC-RAS forecasting on the fixed benchmark.

[AI-38] CESBench: Benchmarking Large Language Models on Cryptographic Engineering Security for IoT Devices

链接: https://arxiv.org/abs/2609.21344
作者: Wenquan Zhou,An Wang,Jing Liang,Peien Feng,Jingqi Zhang,Yaoling Ding,Liehuang Zhu
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:For Internet of Things (IoT) devices, a secure algorithm alone is not enough: an attacker with physical access can attack the implementation directly, and its flaws are hard to fix once deployed. Large language models (LLMs) are now used to build and analyze such implementations. LLM benchmarks exist for cryptography and general cybersecurity, but none covers cryptographic engineering. In this paper, we present CESBench, 380 expert-written items across six sub-domains of cryptographic engineering security for IoT devices: side-channel, fault injection, implementation, countermeasures, evaluation, and integration. Four task types target different competences: 209 multiple-choice items test recall, 67 judgment items require a security verdict and its justification, 63 scenario items require an engineering diagnosis, and 41 code tasks are graded by 572 test cases. To validate the benchmark, 11 open-weight and proprietary LLMs answer every item. Multiple-choice and code responses are scored automatically, and judgment and scenario responses by an LLM judge, whose scores are checked against a second judge from another model family and human re-scoring. Composite scores range from 54.4% to 83.6%. The top score on each task type is 98.6% for multiple choice, 95.1% for code, and 88.4% for scenario diagnosis, but only 58.8% for judgment. Across models, 88.5% of verdicts are correct, yet their justifications earn only 53.4% of the rubric marks. Multiple choice is near its ceiling for the strongest models and most code tasks are solved, whereas justifying a security verdict remains the weakest competence. The benchmark, prompts, and per-item results are public.

[AI-39] Co-Evolving Zero-Day Jamming: Adaptive Attack Synthesis and Graph Attention-Based Online Detection

链接: https://arxiv.org/abs/2609.21334
作者: Ghilas Aissou,Rémi A. Chou,Taejoon Kim
类目: Information Theory (cs.IT); Artificial Intelligence (cs.AI); Signal Processing (eess.SP)
备注: Accepted for publication in the 2026 IEEE Global Communications Conference (GLOBECOM)

点击查看摘要

Abstract:Effective evaluation of zero-day jamming detectors requires robust adversarial models. However, existing attack models often assume prior knowledge of the target receiver, limiting their utility as evaluation benchmarks. On the detection side, existing detectors fail to capture the global temporal-spectral structure of jamming behavior and cannot differentiate zero-day strategies as they emerge. This paper addresses these limitations through a two-pronged framework. First, an online detection framework is introduced that combines a graph attention network (GAT) for temporal-spectral representation learning with Dirichlet process (DP)-means clustering. This framework jointly classifies known and discovers zero-day strategies within a unified learning objective. Second, an inference-driven reinforcement learning (RL) jammer is proposed as an adversarial benchmark. The jammer treats the target receiver as a black-box, infers the detector state via hypothesis testing, and optimizes the trade-off between attack impact and stealth. Simulation results show that the proposed RL jammer outperforms benchmarks, achieving 33% higher attack efficacy and 67% higher stealth. The proposed detection framework against the proposed RL jammer is shown to achieve 20% higher detection accuracy than the benchmarks.

[AI-40] Deep Reinforcement Learning with Buffered Quantile Objectives

链接: https://arxiv.org/abs/2609.21327
作者: Mohammad Alipour-vaezi,Sajad Khodadadian
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Optimization and Control (math.OC)
备注:

点击查看摘要

Abstract:Quantile-based reinforcement learning provides an interpretable approach to risk-sensitive decision-making by optimizing a prescribed quantile of the cumulative-return distribution. Despite this appeal, learning under a point quantile objective is challenging: quantiles can change abruptly under small perturbations of the return distribution, and exact quantile-sensitive planning requires computationally demanding distributional optimization. Lower-buffered quantiles alleviate the former difficulty by averaging neighboring quantiles immediately below the target level, providing a smoother surrogate while preserving the underlying point-quantile objective. Existing methods based on this principle, however, remain model-based and rely on explicit return-law planning, limiting their applicability beyond small tabular problems. We develop Deep-BQRL, a model-free distributional reinforcement-learning framework that extends buffered-quantile learning to neural function approximation. The method learns conditional return quantiles directly from sampled transitions, constructs buffered action scores from the relevant region of the learned quantile function, and uses ensemble disagreement to guide exploration. An augmented input representation allows the learned policy to respond to trajectory information without explicitly reproducing the quantile-state recursion required by exact planning. Experiments on an asset-selling optimal-stopping problem and slippery FrozenLake compare Deep-BQRL with model-based UCB-BQRL and tabular PPO and TRPO implementations. In asset selling, Deep-BQRL attains smaller mean cumulative point-quantile policy gaps than PPO and TRPO at the reported target levels, while UCB-BQRL retains the smallest gaps. The learned stopping decisions also vary with the target quantile, providing an interpretable illustration of the method’s risk-sensitive behavior.

[AI-41] LEGIT: Credentialing Protocol for Trustworthy AI Agent Marketplaces

链接: https://arxiv.org/abs/2609.21325
作者: Steve Drew,Jiayu Zhou
类目: Artificial Intelligence (cs.AI)
备注: 28 pages, 7 figures, 11 tables

点击查看摘要

Abstract:Agentic marketplaces are emerging where AI agents with varying capabilities autonomously complete specialized tasks for buyers. A major challenge of such marketplaces is that buyers cannot easily determine which agent will perform best on their tasks. Reported benchmark scores may be difficult to verify or compare across tasks, software, and budgets. We introduce LEGIT, a credentialing protocol connecting certification, reputation, and proposed marketplace allocation. Certification binds measured quality and cost per solved task to an agent configuration, task domain, evaluation budget, and evidence through a signed record. Reputation links records of past task outcomes to the same identity, subject to the reliability of the reported feedback. Buyers and agents can verify credential records and inspect optional visual profiles. Evaluations reveal cost differences between agent configurations with similar observed task success, and show that comparisons depend on the evaluation budget. These results support binding performance measurements to the tested configuration and resource limits. A complementary analysis quantifies the deposits and fees required for reputation manipulation under a stated Sybil attack model.

[AI-42] GameASG-Bench: Benchmarking Autonomous Software Generation for Game Development

链接: https://arxiv.org/abs/2609.21293
作者: Xiuhui Zhang,Yi Chen,Shusheng Xu,Fan Li,Huan Wang,Tongkai Yang,Binhang Yuan
类目: Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注: 17 pages. Code: this https URL

点击查看摘要

Abstract:Autonomous software generation (ASG) aims to turn human requirements into executable applications, but delivering these applications does not necessarily establish that their interacting components satisfy the specified behavioral requirements. We introduce GameASG-Bench, a benchmark that makes behavioral testability part of the generation task for game development. Our design declares an evaluation interface specification before generation, fixing legal starting scenarios, player-level actions, stable snapshots, rejection behavior, and invariants while leaving private implementations open. Concretely, we include: (i) static L1 checks that assess source-level compliance; and (ii) browser-executed L2 checks that combine semantic observations with real input and runtime evidence. We implement this protocol as 47 browser-native game-generation tasks spanning 12 primary genres and both 2D and 3D interaction, each with executable checks and an independently verified reference implementation. Our experiments answer four key questions about end-to-end agent performance, tool access and nominal turn budget, reasoning effort, and harness choice. Across nine agent stacks, the highest observed mean L2 check pass rate is 93.2%, yet the highest observed strict task success rate, requiring all L1 and applicable L2 prerequisite and core requirement checks, is only 55.3% (26/47 tasks). For DeepSeek-V4-Flash, full tool access and larger nominal turn budgets yield more strict task successes, while the strict task success rate is not monotonic in reasoning effort. Both tested harnesses achieve 18 strict task successes, but only ten tasks succeed under both. These results expose task-level compliance gaps that high average check pass rates actually obscure.

[AI-43] Authorization Revocation for Long-Running AI Agents : Root-Scoped Quiescence under Delegation and Asynchronous Execution

链接: https://arxiv.org/abs/2609.21284
作者: Genliang Zhu,Chu Wang
类目: Programming Languages (cs.PL); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
备注: 39 pages, 2 figures, 7 tables; includes a complete proof appendix

点击查看摘要

Abstract:Long-running AI agents outlive initiating processes through credentials, delegated tasks, queues, callbacks, reservations, and provider-side operations. Cancellation, process exit, and credential revocation neither close every pre-cut carrier nor distinguish independently authorized shared work. We define root-scoped authorization quiescence: for each manifested sink, a certificate accounts for every cut-relevant acceptance under the retired root-epoch atom that precedes its local fence and excludes protected acceptance under that atom after the fence, while permitting exact rebind to a current, independently sufficient support. The root-scoped quiescence protocol linearizes a root cut, fences old-root expansion and protected sinks, represents alternative and conjunctive authority as antichains of minimal sufficient root sets, and composes provider-frontier certificates into a cutset over registered old-root paths. Exact channel-token accounting reconciles transfers; missing or conflicting evidence remains indeterminate. Under stated assumptions, we prove post-cut issuer non-expansion, support-sound projection, compositional soundness under exact channel conservation, independent-support preservation, merge-order independence, and crash/replay stability. A provider-free late-effect test suite matches 17/17 registered outcomes. Two cancellation-only and one cut-only execution accept the same class of already scheduled late effect; two cut-plus-fence executions, one restart, and one stale-process execution reject it. A separately implemented checker verifies 17/17 traces and rejects 44/44 consistently rehashed semantic regressions. The certificate establishes root-relative authorization quiescence within its bound manifest and configuration, not global idleness, rollback, or business completion. Comments: 39 pages, 2 figures, 7 tables; includes a complete proof appendix Subjects: Programming Languages (cs.PL); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR) Cite as: arXiv:2609.21284 [cs.PL] (or arXiv:2609.21284v1 [cs.PL] for this version) https://doi.org/10.48550/arXiv.2609.21284 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-44] Efficient Benchmarking in Production: A Study of an Evolving LLM Agent

链接: https://arxiv.org/abs/2609.21267
作者: Yining She,Lei Lin
类目: Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注: A study of efficient recurring evaluation of a production LLM agent based on real-world historical data

点击查看摘要

Abstract:Production LLM agents are evaluated repeatedly as they evolve, but full agent benchmarks are costly to rerun. We study efficient recurring evaluation for a production analytics agent serving tens of thousands of monthly active users and report first-hand deployment experience. Using 574 historical runs of the production benchmark, split chronologically into calibration and held-out periods, we compare random sampling, historical caching, fixed representative subsets, and IRT-based adaptive testing. The results show that multidimensional 2PL adaptive testing achieves the best overall score fidelity: executing 200 questions, 38.5% of a full run, yields 1.03 pp of MAE. We nevertheless deployed difficulty-stratified fixed subsets because of their operational simplicity, and show they transfer without recalibration to five other agent families and remain stable across calibration windows as short as one day. Drawing on this deployment experience, we report practical recommendations for recurring production-agent evaluation.

[AI-45] PlaceReason er-Beta: Reasoning -Driven Macro Placement and Benchmarking

链接: https://arxiv.org/abs/2609.21263
作者: Qiufeng Li,Chengxuan Wang,Rongqian Chen,Quan Cheng,Yihui Ren,Chia-Tung Ho,David Z. Pan,Tian Lan,Weidong Cao
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Automated macro placement remains a fundamental challenge in VLSI physical design. Despite decades of research, existing approaches predominantly optimize hand-crafted proxy objectives, such as estimated wirelength, and typically produce placements through one-shot numerical optimization, limiting their ability to incorporate visual layout context, codified design expertise, and downstream physical-design feedback in a unified loop. We present PlaceReasoner-Beta, a verifier-guided multi-agent framework that reformulates macro placement as a closed-loop reasoning problem rather than black-box optimization. A vision-language model (VLM) planner generates candidate placements from the floorplan image, macro specifications, and connectivity structure; a geometric verifier enforces physical legality and expert placement principles; a physical verifier refines candidates using early implementation feedback; and a post-route optimizer further improves promising layouts using final PPA. To enable reproducible evaluation, we introduce PlaceReasoner-Bench, a fully open end-to-end benchmark built from open RTL designs, EDA tools, and technology libraries. It comprises 8 designs at two aspect ratios, yielding 16 tasks with fixed floorplans and I/O assignments, so methods differ only in macro positions and orientations and are evaluated using routed PPA and DRC rather than pre-route proxies. Across the benchmark, PlaceReasoner-Beta achieves the best timing among DRC-clean methods on all square tasks, reducing post-route TNS by 61.2% at 1:1 and 53.0% at 2:1 relative to the classical baseline field. It also shortens routed wirelength on most designs despite never explicitly optimizing it, demonstrating that reasoning over spatial structure under physical-design feedback can improve end-to-end layout quality beyond proxy-objective optimization.

[AI-46] CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition

链接: https://arxiv.org/abs/2609.21259
作者: Lance Ying,Jinzhou Wu,Yingshan Susan Wang,Shivam Aarya,Luca M. Schulze Buschoff,Harry Chen,Katherine M. Collins,Andrea de Varda,Shuhao Fu,Sean Dae Houlihan,Akshay K. Jagadish,Guangyuan Jiang,Samuel Kiegeland,Tetsu Kurumisawa,Rongzhi Liu,Ryan Liu,Ningshan Ma,Kathryn McGregor,Younes Strittmatter,Polina Tsvilodub,Jacob Hoover Vigly,Sarah Wu,Enjie Xu,Yiling Yun,Kelsey Allen,Tyler Brooke-Wilson,Brian Christian,Evelina Fedorenko,Michael C. Frank,Michael Franke,Tao Gao,Samuel J. Gershman,Robert D. Hawkins,Jennifer Hu,Julian Jara-Ettinger,Max Kleiman-Weiner,Sydney Levine,Tal Linzen,Hongjing Lu,Timothy O’Donnell,Desmond C. Ong,Steven T. Piantadosi,Rebecca Saxe,Eric Schulz,Tianmin Shu,Felix A. Sosa,Ilia Sucholutsky,Tan Zhi-Xuan,Tomer Ullman,Fei Xu,Ilker Yildirim,Jian-Qiao Zhu,Thomas L. Griffiths,Tobias Gerstenberg,Kevin Smith,Joshua B. Tenenbaum
类目: Artificial Intelligence (cs.AI)
备注: Project website – this https URL

点击查看摘要

Abstract:Understanding and modeling human intelligence are parallel goals shared by artificial intelligence (AI) and cognitive science. As AI systems grow increasingly capable, in what ways do model responses resemble human responses, and where do they systematically diverge? The sheer breadth and diversity of the tasks humans can perform and think about pose a challenge for scalable and rigorous comparison between humans and models. We introduce CogGym, a scalable, unified framework grounded in cognitive science for systematically comparing model and human behavior on matched experimental trials. CogGym uses a semi-automated, human-in-the-loop pipeline to standardize diverse experimental paradigms into a task-agnostic Experiment Markup Language (EML), enabling reproducible and faithful comparison at scale. For initial release, we curate and standardize 258 cognitive experiments from 100 papers that focuses on human commonsense reasoning, and evaluate 50 large language models against human responses. We find a clear scaling trend where larger and more recent AI models better reproduce human judgments. Yet AI models’ improvement on such common reasoning tasks is considerably slower than the gains observed on formal-reasoning benchmarks like math and coding, and model–human fit remains well below human splithalf reliability ( R^2 = 0.93 on text, 0.95 on image, and 0.92 on video) with the best models achieving R^2 = 0.59 on text, 0.58 on image, and 0.43 on video experiments. We intend for CogGym to provide a living evaluation framework that continually incorporates new cognitive science experiments to characterize where model behavior resembles human behavior, where it systematically diverges, and how those patterns change as models and experiments evolve.

[AI-47] KnowDemo: Knowledge-Guided Robot Demonstration Generation from Human Videos

链接: https://arxiv.org/abs/2609.21229
作者: Zhiyuan Gao,Yanxiang Zhan,Mohammad Khoshnazar,Jeroen Schäfer,Michael Beetz
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Learning robot manipulation policies typically requires substantial demonstration data, which are costly to collect on real robots. Recent methods generate robot demonstrations from human videos by adapting recovered motion and validating the resulting trajectories in simulation. However, methods centered on motion-reference adaptation can limit behavioral diversity by retaining the demonstrated contact strategies and subtask orders, while insufficient understanding of task requirements and scene relations can reduce demonstration generation efficiency by generating invalid candidates. To address these limitations, we propose KnowDemo, a framework that uses structured manipulation knowledge from human videos to generate diverse robot demonstrations for a target workspace. To distinguish task requirements from demonstration-specific choices, we develop a knowledge extraction and reasoning module based on a vision-language model (VLM) that associates object and action descriptions with inferred task conditions, demonstration references, and permissible execution variations. To translate this knowledge into executable demonstrations, we resolve the descriptions against target-scene entities and geometry to guide candidate generation and screening before motion planning and simulation. The resulting demonstrations exhibit multimodal behavior through alternative contact strategies and valid subtask orders, with structured execution labels. Experiments demonstrate additional verified execution modes beyond a reference-only configuration and improved candidate planning success through task-guided grasp sampling. To validate the generated data for policy learning, we fine-tune the pretrained \pi_0.5 model on simulation data, achieving sim-to-real transfer across three tasks. Project page: this https URL

[AI-48] A Fully Differentiable Neuro-Soft-Symbolic Framework for Perceptual Task Planning

链接: https://arxiv.org/abs/2609.21221
作者: Hongyan Wei,Wael AbdAlmageed
类目: Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Perceptual planning tasks require two key capabilities: accurately perceiving uncertain scenes and planning valid action sequences following logical rules. Conventional methods convert perception into discrete symbolic facts and then plan, discarding perceptual uncertainty and severing task-level feedback to perception. We introduce a generic, fully differentiable neuro-soft-symbolic framework that connects visual perception and task planning within a single computational graph. The framework maintains a continuous soft symbolic state, lifts domain rules into a differentiable soft- T_P transition operator, and optimizes action logits over a short planning horizon. Gradients from the planning objective can also update the perception parameters, allowing task-relevant perceptual representations to be refined during planning. On Blocksworld, our method solves 40/40 LatPlan-40 tasks and 596/600 PlanBench-600 tasks, compared with 33/40 for LatPlan and 587/600 for the reasoning-model baseline, while requiring substantially less computation and time. In the perceptual-uncertainty ablation, our method improves the success rate from 59% with frozen perception to 83%. We further conduct task-and-motion simulations on Blocksworld scenes, providing an execution-level validation of the compatibility between decoded task plans and downstream robotic motion execution.

[AI-49] Fewer Steps Better Actions: Rethinking Flow-Matching Inference for VLA Policies

链接: https://arxiv.org/abs/2609.21216
作者: Zhipeng Tang,Xinda Chen,Weining Rao,Xiao Li,Wenting Tan,Yuning Wang,Xiao Shi,Xiaofang Zhao
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: 15 pages, 6 figures, 7 tables

点击查看摘要

Abstract:Vision-language-action (VLA) policies based on flow matching generate action chunks through repeated evaluations of an action expert. Increasing the number of integration steps raises inference cost, but does not necessarily improve closed-loop success. We propose Coda, which reallocates part of this integration budget to a single learned endpoint correction. A frozen policy first completes a few-step noise-to-action trajectory; a lightweight Transformer then predicts a demonstration-supervised residual using the candidate action, source noise, and shared observation-prefix cache. Only the corrector is trained. On 50 RoboTwin Easy tasks, five-step Coda improves success from 71.64% to 74.68% over the matched five-step baseline, while reducing forward latency by 30.2% relative to the default ten-step policy. A two-step configuration achieves 71.88% success with a 2.12 \times speedup. An independent 13-task control shows a 5.69-percentage-point gain at nearly equal latency, supporting correction as an effective alternative to additional integration. The same design also improves frozen official SmolVLA, raising two-step success from 60.8% to 69.4%. These results show that endpoint correction improves the quality-latency trade-off of frozen flow-matching policies.

[AI-50] Ability-Residual Decoupled Modeling for Affective Cognitive Diagnosis

链接: https://arxiv.org/abs/2609.21214
作者: Boyuan Zhao,Meng Ye
类目: Artificial Intelligence (cs.AI)
备注: 33pages

点击查看摘要

Abstract:Cognitive diagnosis infers students’ concept mastery from response logs. However, students’ responses are not determined by mastery alone: non-cognitive factors such as emotion, engagement, and fatigue can also affect performance. Affective cognitive diagnosis therefore extends conventional cognitive diagnosis by incorporating affective states. Existing methods often assume that the cognitive diagnosis backbone has already explained ability, item, and concept effects, so the remaining errors can be attributed mainly to affect. We argue that this assumption can be insufficient in real educational data: item calibration bias, systematic concept bias, personalized student-concept deviations, and latent student-item matching can form stable cognitive residuals. Without an explicit modeling pathway, these residuals may leak into affective representations, producing affect contamination. To address this problem, we propose an ability-residual decoupled framework for affective cognitive diagnosis. The model first captures unmodeled cognitive residuals through student, item, concept, student-concept, and low-rank student-item components, and then uses an affective module to modulate guess/slip effects. A Q-matrix-constrained concept residual attention mechanism adaptively aggregates only item-relevant concept residuals. Experiments on ASSIST2017, ASSIST2012, ASSIST2009, and Junyi with six cognitive diagnosis backbones show response-prediction gains across the reported comparisons and generally improved affect alignment when affect labels are available. Ablation studies, leakage probes, principal component analysis visualization, long-tail analysis, and case studies further indicate that ability residuals absorb stable cognitive bias, reduce cognitive contamination in the affective branch, and enhance the robustness and predictive accuracy of cognitive diagnosis models.

[AI-51] Visual Navigation Transformer with Pose Attention

链接: https://arxiv.org/abs/2609.21212
作者: Beiming Li,Jaime Romero,Jonathan Diller,Vijay Kumar,Alejandro Ribeiro
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Learned navigation policies typically consume observations as a temporally ordered history, with positional encodings tying each observation to when it was seen, making it difficult to reuse experience from earlier traversals of an environment. Systems that do reuse such experience usually construct an explicit representation, such as a map or a topological graph, and plan on it. We propose VNT-PA (Visual Navigation Transformer with Pose Attention), a transformer planner whose context is a set of depth keyframes indexed by camera pose. With camera poses as positional encoding, attention depends on the pose differences between keyframes rather than on their temporal order. VNT-PA is trained to imitate a shortest-path planner operating on the ground-truth scene mesh, predicting actions by querying the spatial context with only its current pose and the goal position. On point-goal navigation in HM3D validation scenes, VNT-PA reaches 93.3% success and 90.4% success weighted by path length (SPL), outperforming baselines that encode the same context as a temporal sequence or treat pose as an input feature, in both navigation performance and training efficiency. Because the spatial context is a pose-indexed set, frames from different trajectories can be fused at test time. The planner also degrades more gracefully under localization noise than a conventional baseline which plans on explicit maps. These results show that pose-stamped experience can serve directly as the environment representation for a learned planner, and that making attention depend on pose differences, rather than on temporal order, speeds up training and improves long-horizon navigation.

[AI-52] Information-Gain Rewards over Diversity-Pruned Tests: GT-Anchored Verifier Co-Training for Reliable Code Generation

链接: https://arxiv.org/abs/2609.21208
作者: Ana Nunez,Peyman Najafirad
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Self-play methods that co-train a single language model as both coder and test author promise to move code-generation RL beyond fixed test suites, but they suffer from two coupled pathologies: permissiveness collapse, where pass-rate rewards are maximised by trivial, non-discriminative tests, and concentration bias, where i.i.d. sampled tests cluster on modal inputs and inflate estimator variance. We introduce CoVer (Co-trained Coder and Verifier), a single-policy GRPO framework that addresses both failure modes. First, an information-gain (IG) reward scores each self-generated test by the mutual information between its pass/fail vector and a graded, ground-truth-anchored correctness signal y [0, 1] m, gated by the sign of their covariance so that only positively discriminative tests receive reward. Second, a three-stage diversity-aware selection step prunes a candidate pool to a behaviourally non-redundant suite (invalidity, input-string, execution-profile filtering), raising the effective sample size of the IG estimator at fixed execution budget. On five benchmarks (LiveBench, MBPP, LiveCodeBench, CodeContests, Code-Forces), CoVer raises one-shot pass@1 by +5.8 points at 7B and +7.1 points at 14B over the Qwen2.5-Instruct backbone, and achieves the highest macro-average among all compared methods at both scales. As a drop-in backbone inside the CodeT ranking pipeline, CoVer-7B adds +3.5 points, demonstrating the dual benefit of co-training for both generation and selection.

[AI-53] SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?

链接: https://arxiv.org/abs/2609.21190
作者: George Ma,Benjamin Mikek,Haoyu Li,Ferhat Erata,Yuhao Zhang,Zeren Shui,Behrooz Omidvar Tehrani,Jun Huan,Murali Krishna Ramanathan,Somayeh Sojoudi,Hao Zhou,Anoop Deoras
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注:

点击查看摘要

Abstract:Ensuring the correctness of LLM-generated code is a core challenge for modern software engineering. Benchmarks for agentic code generation check correctness with held-out test suites, which are inherently incomplete and increasingly susceptible to memorization. Formal verification avoids both problems, but existing work covers only standalone tasks whose specifications are given as input, not real issues, which touch large repositories and state intent in vague natural language. We present Benchproofer, a pipeline that turns a coding task with a known correct patch into a formally verified one: it writes a specification for the new code, summarizes the existing functions that code calls with axioms, and admits an instance only after mechanical and adversarial gates agree. Applying it to SWE-bench Verified yields SWE-Proof, 500 real issues whose correctness is formally verified rather than tested, and it extends to SWE-bench Pro. Across two frontier models, verification catches what tests miss: a quarter to a half of test-passing patches admit counterexamples, which a structured natural-language specification does not fix, while a correct formal one lifts resolution from 85% to 95% for Opus 4.8. Writing that specification is the hard part: models that must write their own gain nothing over an unaided baseline, and only 62% of their specifications pass our audit. The usual failure is faithfulness, a specification that constrains part of the required behavior and leaves the rest free. Specification quality still tracks the outcome, failing on 89% of unresolved instances against 47% of resolved ones, making faithful specification synthesis a concrete open problem.

[AI-54] Implicit Rule Induction with Test-Time Task Embeddings in ARC-like Tasks

链接: https://arxiv.org/abs/2609.21181
作者: Adrien Deliège,Claas Beger,Marc Van Droogenbroeck,Melanie Mitchell
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:The Abstraction and Reasoning Corpus and related benchmarks evaluate whether AI models can solve novel reasoning tasks, but often leave unclear whether success reflects inference of the intended underlying rule or reliance on shortcuts. We address this gap by studying test-time task embeddings in Vision ARC (VARC), a model in which a pre-trained backbone is complemented by a trainable embedding representing the transformation rule. In the original VARC, test-time training (TTT) is jointly applied to the backbone and task embedding. Here we introduce a novel two-step TTT protocol: first finetune only the task embedding (Embed-TTT), then freeze it and finetune the backbone. Across ARC-AGI-1, ConceptARC, and two controlled datasets with known rules, Embed-TTT consistently yields improved task embeddings, ones that align better with underlying task rules, improve embedding-based retrieval, and enable accurate linear probing of known rules. Qualitatively, Embed-TTT identifies more semantically meaningful relations between test and train tasks on ARC-AGI-1. We also show that optimizing only task embeddings (less than 0.01% of model parameters) already solves a non-trivial fraction of ARC-AGI-1, ConceptARC, and Mini-ARC tasks, while the full two-step pipeline improves final performance. Finally, we show that Embed-TTT recovers the underlying geometric structure of parametric rules and learns compositional capabilities that enable rule-wise interpolation, but not extrapolation. These findings support a clearer separation between rule induction and rule execution in ARC-like evaluations, motivating benchmarks that better distinguish in-distribution from out-of-distribution rules.

[AI-55] SpecOpt: Contact-Diff Reasoning for Agent ic Molecule Optimization Toward Binding Specificity

链接: https://arxiv.org/abs/2609.21165
作者: Thao Nguyen,Heng Ji
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Off-target protein binding is a major source of adverse effects for small-molecule drugs, yet most structure-based molecular design methods focus on generating selective compounds de novo rather than improving the selectivity of existing, well- characterized drugs. We introduce specificity optimization (SpecOpt), a molecular design task that seeks constrained structural modifications to an existing compound that increase its binding preference for an intended target over known off-targets while preserving its structural identity and drug-like properties. To enable systematic evaluation, we construct a ChEMBL-derived benchmark from compound-target interaction data, identifying intended targets through curated drug-mechanism annotations and off- targets through measured activities. We then develop an agentic framework that docks each compound against its intended target and off-targets, compares the resulting poses through residue-aware atom-protein contacts, and provides these differential interactions to a large language model to propose targeted structural modifications. Candidates are retained only if they satisfy molecular similarity, ADMET, and target-off-target docking selectivity criteria. On 915 compounds, the agent improves the target- off-target binding gap for 84.8% of compounds, shifting the mean gap from -0.72 to +0.47 kcal/mol while maintaining a mean Tanimoto similarity of 0.72 to the starting compounds. Ablation studies identify residue-specific contact information as the critical optimization signal: replacing residue identities with binary contact indicators eliminates improvement on all 29 ablation compounds. These results establish SpecOpt as a distinct molecular design problem and demonstrate residue-aware differential interactions as an effective signal for improving the specificity of existing compounds.

[AI-56] Can Agents Design Better Chips with a Higher Level Abstraction?

链接: https://arxiv.org/abs/2609.21157
作者: Zijian Ding,Yang Zou,Yizhou Sun,Jason Cong
类目: Artificial Intelligence (cs.AI); Hardware Architecture (cs.AR)
备注: 7 pages, ICCAD’26 special session

点击查看摘要

Abstract:Large Language Model (LLM) agents are increasingly being explored for chip design, but most existing approaches operate directly at RTL. We ask whether agents can design better chips by leveraging higher-level abstractions. We compare Direct RTL Design, Agent-based HLS Design, Post-Compiler HLS Refinement, and Post-HLS RTL Refinement, and combine Agent-based HLS Design with Post-HLS RTL Refinement as Agent-based HLS with RTL Refinement (AHRR). We use FPGAs as a practical, easy-to-deploy platform for end-to-end evaluation, but note that the design-flow tradeoffs we study are largely independent of the target technology. Across a diverse 11-tasks benchmark suite, AHRR achieves a 2.6 \times geometric-mean speedup over Direct RTL Design across our benchmark suite. Case studies show that HLS distills design knowledge into abstractions that agents can leverage, while RTL refinement recovers lower-level optimization opportunities. Together, these results make AHRR a promising workflow for agentic chip design. The code and evaluation artifacts are available at this https URL.

[AI-57] EnSol: an environment-aware graph neural network for molecular solubility prediction

链接: https://arxiv.org/abs/2609.21151
作者: Thao Nguyen,Saman Shafaei,Zhengyi Zhang,Huimin Zhao,Heng Ji
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Molecular solubility directly affects key aspects of molecular development such as reaction feasibility, formulation performance, separation efficiency, and solvent selection. However, experimental measurement across solutes, solvents, and temperatures remains costly and sparsely sampled. Existing computational models often rely on fixed-solvent assumptions, deterministic formulations, or simplified representations of solute-solvent interactions, limiting their ability to capture complex molecular interactions, continuous temperature effects, and experimental uncertainty. Here, we introduce EnSol, an environment-aware probabilistic framework for molecular solubility prediction. EnSol represents the solute and solvent as molecular graphs and learns separate representations for each before bringing them together through cross-attention to capture solute-solvent interactions. Temperature is incorporated directly into the solvent environment through feature-wise modulation, and a mixture density network predicts full solubility distributions to capture both temperature-dependent behavior and experimental uncertainty. On the independent SolProp and Leeds benchmark datasets, EnSol achieved Spearman correlations of 0.876 and 0.601, respectively, outperforming state-of-the-art solubility prediction models across both benchmarks. Beyond computational benchmarking, experimental validation across chemically diverse solute-solvent pairs showed that EnSol maintained strong predictive performance and supported reliable solvent ranking, achieving a Spearman correlation of 0.715. These results show that EnSol can support reliable solubility prediction and solvent selection across diverse chemical systems while accounting for predictive uncertainty.

[AI-58] nyCeNN-LM: Quality-Gated Conversion of Pretrained Attention with CeNN-Inspired Cellular-Recurrent Layers

链接: https://arxiv.org/abs/2609.21139
作者: Kabeh Mohsenzadegan,Vahid Tavakkoli,Kyandoghere Kyamakya
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Replacing attention in a pretrained language model is a compatibility problem: a plausible substitute may alter representations expected by later layers. TinyCeNN-LM introduces a \emphquality-gated post-training conversion framework using CeNN-inspired cellular-recurrent layers with bounded local processing, compact recurrent memory, routing, fusion, and accept-or-rollback validation. Three implementations are studied: Integrated Memory, MemoryFusion, and PDelta3-GDN2-CLVR+Local32. Strict PDelta3 conversion accepts a layer only when representation and NLL criteria pass fixed thresholds. On SmolLM2-135M, layers 0-2 are accepted with cumulative \Delta\mathrmNLL=+0.01209 , while layer 3 is rejected despite acceptable NLL because representation fidelity fails. On Qwen3.5-0.8B, full-attention layers 3, 7, and 11 are accepted with final \Delta\mathrmNLL=+0.02073 . Integrated Memory keeps perplexity within -0.07% to +0.93% while reducing total cache by up to 6.01% . A sampled 200-item downstream sanity check gives 28.5% – 32.0% overall accuracy for converted Qwen releases. The results support conservative, quality-gated structural conversion rather than universal attention replacement or speedup.

[AI-59] he Stochastic Shift: A New Evaluation Paradigm for Text-to-SQL with AI Operators

链接: https://arxiv.org/abs/2609.21133
作者: Tarfah Alrashed,Fatma Ozcan,Per Jacobsson,Tal Neiman,Xianshun Chen
类目: Databases (cs.DB); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:SQL has been augmented with AI operators, enabling modern data analytics platforms to derive insights from both structured and unstructured data. We observe that while current Text-to-SQL systems can successfully generate these AI-augmented queries, reliably evaluating their correctness remains a critical open challenge. Current metrics, which rely on exact query results and deterministic execution, systematically fail against the flexible, non-deterministic outputs of AI operators. In this paper, we formalize these unique evaluation failure modes and introduce a Multilayered Evaluation Framework that decouples deterministic database logic from flexible AI semantics. We test our approach across both industry (BigQuery) and academic (ThalamusDB) systems. We demonstrate that traditional Execution Accuracy severely penalizes valid queries, achieving as low as a 25% detection rate for correct translations. Furthermore, even a state-of-the-art LLM-based autorater falsely rejects 32% of accurate queries due to the complexity of judging both relational and AI components simultaneously. By validating the standard relational logic and the AI operations separately, our framework achieves state-of-the-art overall accuracy across both platforms (up to 97.2%), proposing a reliable standard for benchmarking AI-powered SQL generators.

[AI-60] Decoupling Internal Representational Changes and Causal Importance in Fine-Tuned Large Language Models

链接: https://arxiv.org/abs/2609.21113
作者: Lingfang Li,Procheta Sen,Shubham Das,Danushka Bollegala
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Fine-tuning has emerged as a widely adopted approach for adapting LLMs to a variety of downstream tasks. However, how it reshapes their internal mechanisms remains poorly understood. To address this, we investigate how fine-tuning alters internal representations in LLMs, including attention patterns and layer-wise activations, and examine whether these changes are linked to task-relevant components identified by EAP (e.g., attention heads and logit-level activations) that drive task performance. We find that EAP-identified components are concentrated within specific layers, indicating a degree of functional localisation in how models internalise task-specific behavior. Notably, the distribution of these components across layers is largely uncorrelated with the layers undergoing the most substantial representational changes during fine-tuning. Furthermore, we observe that overlap in EAP-identified components across tasks does not translate into cross-task performance transfer if the tasks are different in nature (e.g. classification vs. generative tasks). More specifically, fine-tuning on one task can lead to a degradation of performance on another when the two tasks exhibit a high degree of overlap in their EAP-identified components.

[AI-61] LoRA Enhanced Contrastive Learning with SAS Vision Transformers

链接: https://arxiv.org/abs/2609.21061
作者: Dan Zimmerman,Frank E. Bobe III,Amelia L. McCormack,Matthew Cook,Gregory D. Vetaw
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Automatic target recognition (ATR) with synthetic aperture sonar (SAS) supports advanced naval capabilities, but deep learning is constrained by scarce target imagery, background clutter, and human-in-the-loop assessment. We adapt DINOv3 Vision Transformer (ViT) models to underwater SAS ATR using a three-stage parameter-efficient framework. Stage 1 uses Low-Rank Adaptation (LoRA) while freezing the ViT backbone, bridging the gap between natural-image pretraining and underwater acoustic propagation. Stage 2 uses hard-negative mining to strengthen the decision boundary against acoustic mimics, including rocks and sediment formations resembling man-made targets. Stage 3 uses Supervised Contrastive Learning (SupCon) to separate target and clutter representations. We evaluate at-sea SAS data using a mission-level geographic split, compare all arms at 85 percent test recall, and repeat each comparison over three random seeds. LoRA accounts for the primary effect, increasing area under the precision-recall curve (AUPRC) from 0.300 to 0.679 +/- 0.027 using the same frozen backbone. Rank 4 achieves this result while training only 0.26 percent of weights. Neither refinement stage exceeds its matched control: hard-negative mining changes AUPRC by -0.0045 +/- 0.0119 versus an equal-size random curriculum, and SupCon changes AUPRC by +0.0002 +/- 0.0096 versus the preceding stage. These null results indicate that mining occurred on data the encoder had already fit and that supervised stages had already imposed most target-clutter geometry. One efficient adaptation stage is sufficient; stacked refinement is not.

[AI-62] PlantShade: Predicting Plant Shadows for Lighting-Aware Robotic Agricultural Operation IROS2026

链接: https://arxiv.org/abs/2609.21059
作者: Longchao Da,Xiaoou Liu,Xingjian Li,Lirong Xiang,Hua Wei
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: This paper has been accepted by the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026)

点击查看摘要

Abstract:Plant growth and agricultural production form the foundation of a country’s sustainable development and directly impact human livelihoods. Recent advances in frontier artificial intelligence have enabled scientific agriculture with strong potential to improve crop productivity. In this paper, we identify the importance and inherent complexity of plant shade simulation, as shading is a critical factor influencing plant growth. To advance this field and promote broader societal benefits, we focus on two main contributions. First, we introduce a comprehensive plant growth and shade dataset covering four plant species, including soybean, tomato, sugarbeet, and strawberry. The dataset includes top-down viewpoints with a supplementary light along a circular trajectory, casting dynamic shadows across multiple growth stages and diverse observation complexities. Second, we propose generative shade simulation based on diffusion models, enabling realistic shade generation for unseen plants and supporting downstream robotic tasks such as perception, lighting control, and view planning. The model incorporates temporal conditioning to facilitate flexible shade simulation across different time stages. We conduct both quantitative and qualitative evaluations to assess model performance. This work provides a foundational study for plant-aware shade modeling and has meaningful implications for broader agricultural and robotic applications.

[AI-63] How Much of a Real Workload Can LLM -Generated GPU Kernels Actually Reach?

链接: https://arxiv.org/abs/2609.21058
作者: Gaurav Agarwal,Ashish Garg,Isha Singhal
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI)
备注: Code, data and all 879 evaluations: this https URL

点击查看摘要

Abstract:Language models can now write GPU kernels that outperform PyTorch. We evaluate five model configurations on KernelBench level 1 and find that a frontier model produces correct kernels for 91.1% of problems and independently verified speedups on 22 of 56, including three convolutions, with a median of 1.235x. Open-weights models are far behind: the best reaches 30.4% correct with three verified speedups and solves zero convolutions. We then ask a question the literature does not: what fraction of a real model’s wall clock do such kernels govern? Profiling seven workloads across three domains, we find the addressable fraction ranges from 8.9% to 58.2%. On transformers, 80-86% of runtime is spent in cuBLAS GEMM and FlashAttention, bounding realistic end-to-end improvement at roughly 1%, and the fraction shrinks with model scale. On recommenders it is 58.2%, concentrated in a single embedding kernel. We introduce DLRM-Bench, 12 recommender kernel problems in KernelBench format, and measure a 41.7% win rate at a 1.552x median there, projecting 8.63% end-to-end. Separately, we show that KernelBench’s correctness check (this http URL with an absolute tolerance) is satisfied by a tensor of zeros on 4 of 60 level-1 problems. Two kernels in our own results exploited this before we detected them, including one scored at 283x that wrote 0.3% of its output buffer. We propose scale-invariant replacements and release all 879 evaluations. Comments: Code, data and all 879 evaluations: this https URL Subjects: Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI) Cite as: arXiv:2609.21058 [cs.DC] (or arXiv:2609.21058v1 [cs.DC] for this version) https://doi.org/10.48550/arXiv.2609.21058 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-64] Physically Based Rendering in the Latent Space

链接: https://arxiv.org/abs/2609.21054
作者: Vuk Radovanovic,Vishesh Gupta,Adrien Gruson,Binh-Son Hua
类目: Graphics (cs.GR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 15 pages, 12 figures. Pacific Graphics 2026, Journal Track (Computer Graphics Forum). Code: this https URL

点击查看摘要

Abstract:Image diffusion models have shown impressive image generation capabilities but are often hard to control, in contrast to classical computer graphics pipelines such as physically based rendering. However, we observe that there is a bridge between light transport phenomena and the distribution of latent space values produced by such models. Thus, we introduce physically based rendering in the feature space learned by the variational autoencoders in generative models, enabling light transport simulation in the latent space. This allows us to leverage physically based rendering techniques to output latent maps for physically guided content generation. We propose modifications to the rendering equation, which, when paired with a differentiable renderer, can yield an optimal set of scene parameters that require only minimal refinement to accurately render into the pretrained latent space. We train our method on a single rendered image, and then demonstrate the generalization of the method to scene geometry changes, lighting changes, and camera view changes.

[AI-65] CaLR: Causal Latent Revision for Robust Diffusion Reasoning

链接: https://arxiv.org/abs/2609.20981
作者: Wei Cai,Jian Zhao,Yuchen Yuan,Xuelong Li
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Autoregressive (AR) models suffer from local greediness, while diffusion language models (DLMs) often lack the strict causal structure required for reasoning. To combine the advantages and overcome the drawbacks of the dual, we propose Causal Latent Revision (CaLR), a framework that reformulates reasoning as constrained latent optimization. By adopting a causal topology matrix (CTM) from an expert model and implicit differentiation, CaLR performs gradient-guided ``thought revision" to enforce logical consistency, enabling dynamic self-correction of intermediate steps during parallel generation. Empirically, CaLR achieves SOTA DLM performance on complex benchmarks, surpassing strong AR baselines and demonstrating superior robustness in constrained tasks like Sudoku.

[AI-66] Attention-Aware Routing: Coupling Routing and Attention in MoEs

链接: https://arxiv.org/abs/2609.20974
作者: Despoina Kosmopoulou,Anastasios Tsetsilas,Efthymios Georgiou,Giannis Karamanolakis,Swastik Roy,Alexandros Potamianos
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:In Mixture-of-Experts language models, the router typically selects and weights experts based on the token’s hidden state, utilizing limited contextual information. We propose Attention-Aware Routing (AAR), which augments the router with temporal and spectral features extracted from a sliding window of attention weights that represent a summary of the model’s contextual state, disentangled from the hidden state. Keeping the base transformer entirely frozen, we train only the routing parameters, isolating routing as the sole variable. AAR improves GSM8K by +3.37 pp over a routing-only SFT baseline on OLMoE. Beyond performance, we show that routing and attention form a coupled circuit: routing changes at layer l propagate through the residual stream to amplify attention sinks at layer l+1, reshaping attention without any direct update to the attention mechanism itself. Further, AAR reduces long diverging generation, with incorrect answers getting shorter, while correct answers remain unchanged in length. Finally, AAR is strongly depth-sensitive: applying it indiscriminately across layers can degrade factual retrieval, whereas mathematical reasoning gains persist when it is introduced deeper in the network. This sensitivity exposes a retrieval–reasoning tension across depth and makes layer-selective AAR a controlled probe of the routing-relevant information carried by attention at different layers.

[AI-67] RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models

链接: https://arxiv.org/abs/2609.20971
作者: Chuxu Song,Jiuqi Wei,Zhencan Peng
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Long-context large language model inference is increasingly limited by prefill, where dense self-attention processes the entire prompt before generation begins. Sparse block selection can reduce this cost, but a block centroid may hide a highly relevant token among many irrelevant ones. We call this failure mode mean dilution and propose RBS-Attention, a training-free sparse-prefill method with two complementary selection branches. A centroid base branch captures average relevance, while a rescue branch uses the maximum key-block radius and its prompt-, layer-, and head-dependent distribution to identify blocks at risk of underestimation. Independently thresholding the two branches and combining their masks controls the contribution of rescue blocks while preserving regular block-sparse FlashAttention execution. On H100 GPUs, RBS-Attention achieves 20.65 \times standalone prefill-attention speedup, 11.92 \times vLLM prefill-attention speedup, and 5.97 \times end-to-end time-to-first-token speedup at 128K on Qwen3-30B-A3B-Instruct-2507-FP8. On the dense Qwen3-32B model, it obtains 88.65 overall RULER accuracy versus 89.52 for dense attention; LongBench-v2, InfiniteBench, and Video-MME provide additional quality evaluation. Supporting experiments measure actual retention, compare selectors at matched density, and characterize block-size, threshold, and memory behavior. Together, these results support radius-adaptive dual-branch selection as an effective approach to long-context prefill.

[AI-68] Bio-MF: Low-Latency and High-Fidelity EEG-to-fNIRS Cross-Modal Generation for Hybrid Motor-Imagery Brain–Computer Interfaces

链接: https://arxiv.org/abs/2609.20904
作者: Boyuan Zhao,Sifan Zhang,Luping Chen
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 10pages

点击查看摘要

Abstract:Hybrid motor-imagery brain-computer interfaces (MI-BCIs) combining EEG and fNIRS can outperform EEG-only systems by exploiting complementary electrophysiological and hemodynamic information. To obtain such hybrid information when paired EEG-fNIRS acquisition is unavailable or inconvenient, recent studies have focused on EEG-to-fNIRS cross-modal generation. However, existing methods still suffer from slow generation and often require pretraining, limiting their use in real-time MI-BCI scenarios. Although one-step generative models offer an attractive route to low-latency synthesis, removing the iterative refinement process can reduce generation fidelity and introduce non-physiological artifacts. To address these problems, this paper proposes Bio-MF, a latent-free one-step MeanFlow framework for EEG-conditioned fNIRS generation. Bio-MF performs direct signal-space x-prediction, converts this signal-space output into MeanFlow velocity supervision, and completes inference with one network evaluation. To preserve task-relevant hemodynamic structure under heterogeneous sensor layouts, Bio-MF integrates Spatial-Temporal Interactive 4D Encoding, cross-modal classifier-free guidance, and noise-level-gated FFT regularization. On Dataset 1, EEG + synthetic fNIRS improves ACC over EEG-only by 3.37 and 4.15 percentage points for HbR and HbO, respectively. On Dataset 2, the corresponding gains remain 2.98 and 2.50 percentage points under the unseen 64-channel EEG montage. On an RTX PRO 6000 GPU, Bio-MF generates one fNIRS trial in 7.0 ms, corresponding to an 857x speedup over the 1000-step SCDM latency. These results show that Bio-MF enables fast EEG-to-fNIRS synthesis while preserving task-relevant generation quality for downstream hybrid MI decoding. Our code is available at this https URL.

[AI-69] SpaceDiffusion: Over-the-Orbit Diffusion for Space Generate-and-Forward Communications

链接: https://arxiv.org/abs/2609.20899
作者: Jianhao Huang,Zhanwei Wang,Khaled B. Letaief,Kaibin Huang
类目: Information Theory (cs.IT); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Satellite communications are an essential component of sixth-generation (6G) mobile networks, which provide ubiquitous connectivity for global services. However, the satellite uplink remains a critical bottleneck for ground devices: their limited transmit power and antenna apertures result in low data rates and high packet errors. To overcome this bottleneck, this paper advocates a novel relaying paradigm termed generate-and-forward (GF) communications, where satellites exploit on-orbit generative artificial intelligence (AI) to robustly reconstruct corrupted data prior to forwarding. Specifically, we propose SpaceDiffusion, an over-the-orbit diffusion framework for satellite-assisted image transmission. The core of this framework is a channel-distortion-aware diffusion theory developed using the following approach. By formulating the recovery of compressed and lost image tokens as an inverse problem, this theory incorporates a channel-distortion correction term directly into the conventional denoising diffusion implicit model (DDIM) update. As a result, this design enables a single pretrained diffusion model to adapt dynamically to varying packet-loss patterns and compression distortions without retraining. Furthermore, we analytically characterize the progressive token-reconstruction error and derive a diffusion-step activation threshold that predicts when SpaceDiffusion is expected to outperform conventional decode-and-forward (DF) relaying. Building on these theoretical insights, we further develop an energy-aware early-exit policy to efficiently deploy SpaceDiffusion in orbit. Experimental results demonstrate that SpaceDiffusion achieves lower end-to-end latency compared to DF scheme with retransmission protocol and saves approximately 15 dB of uplink transmit power at a target perceptual quality.

[AI-70] Making Latent Evolution Explicit: Operator-Structured Transitions for World Action Models

链接: https://arxiv.org/abs/2608.27259
作者: Xiaoxiao Lu,Yunlong Dong,Jiahao Shi,Ye Yuan
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Systems and Control (eess.SY)
备注:

点击查看摘要

Abstract:World Action Models (WAMs) augment robot policies by predicting how task-relevant scene states may evolve under interaction. Recent WAMs increasingly perform such prediction in latent representation spaces, avoiding full appearance-level generation while preserving control-relevant information. Yet latent transitions are commonly realized with Transformer-based predictors whose inductive structure is centered on token interaction rather than temporal evolution. We study transition realization as an architectural choice distinct from predictive representation and prediction-policy coupling. We introduce the Latent Evolution Operator Network (LEON), which models latent evolution in a learned observable space through context-modulated operator-based propagation and additive forcing. Grounded in the controlled Koopman generator view of evolution, LEON organizes context-dependent transition variation around a shared evolution-operator structure while retaining a complementary path for additive change. Controlled dynamical systems verify the resulting evolution-specific inductive bias and the complementary roles of operator propagation and forcing. Across two WAM formulations that integrate latent prediction into the policy differently, LEON improves closed-loop performance and robustness while remaining effective under full transition replacement. These results establish transition realization as a consequential architectural choice in latent WAMs.

[AI-71] A Hybrid Computational Intelligence Framework for scRNA-seq Imputation: Integrating scRecover and Random Forests

链接: https://arxiv.org/abs/2511.16923
作者: Ali Anaissi,Deshao Liu,Yuanzhe Jia,Weidong Huang,Widad Alyassine,Junaid Akram
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Genomics (q-bio.GN)
备注:

点击查看摘要

Abstract:Single-cell RNA sequencing (scRNA-seq) enables transcriptomic profiling at cellular resolution but suffers from pervasive dropout events that obscure biological signals. We present SCR-MF, a modular two-stage workflow that combines principled dropout detection using scRecover with robust non-parametric imputation via missForest. Across public and simulated datasets, SCR-MF achieves robust and interpretable performance comparable to or exceeding existing imputation methods in most cases, while preserving biological fidelity and transparency. Runtime analysis demonstrates that SCR-MF provides a competitive balance between accuracy and computational efficiency, making it suitable for mid-scale single-cell datasets.

[AI-72] dSTAR: Strag gler Tolerant and Byzantine Resilient Distributed SGD

链接: https://arxiv.org/abs/2412.07151
作者: Jiahe Yan,Pratik Chaudhari,Leonard Kleinrock
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Machine Learning (cs.LG)
备注: 15 pages

点击查看摘要

Abstract:Distributed model training needs to be adapted to challenges such as the straggler effect and Byzantine attacks. When coordinating the training process with multiple computing nodes, ensuring timely and reliable gradient aggregation amidst network and system malfunctions is essential. To tackle these issues, we propose \textitdSTAR, a lightweight and efficient approach for distributed stochastic gradient descent (SGD) that enhances robustness and convergence. \textitdSTAR selectively aggregates gradients by collecting updates from the first (k) workers to respond, filtering them based on deviations calculated using an ensemble median. This method not only mitigates the impact of stragglers but also fortifies the model against Byzantine adversaries. We theoretically establish that \textitdSTAR is ((\alpha, f))-Byzantine resilient and achieves a linear convergence rate. Empirical evaluations across various scenarios demonstrate that \textitdSTAR consistently maintains high accuracy, outperforming other Byzantine-resilient methods that often suffer up to a 40-50% accuracy drop under attack. Our results highlight \textitdSTAR as a robust solution for training models in distributed environments prone to both straggler delays and Byzantine faults.

[AI-73] ResNLS: An Improved Model for Stock Price Forecasting

链接: https://arxiv.org/abs/2312.01020
作者: Yuanzhe Jia,Ali Anaissi,Basem Suleiman
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE)
备注: Accepted by Computational Intelligence 2023

点击查看摘要

Abstract:Stock prices forecasting has always been a challenging task. Although many research projects try to address the problem, few of them pay attention to the varying degrees of dependencies between stock prices. In this paper, we introduce a hybrid model that improves the prediction of stock prices by emphasizing the dependencies between adjacent stock prices. The proposed model, ResNLS, is mainly composed of two neural architectures, ResNet and LSTM. ResNet serves as a feature extractor to identify dependencies between stock prices, while LSTM analyzes the initial time series data with the combination of dependencies, which are considered as residuals. Our experiment reveals that when the closing price data for the previous 5 consecutive trading days is used as input, the performance of the model (ResNLS-5) is optimal compared to those with other inputs. Furthermore, ResNLS-5 demonstrates at least a 20% improvement over current state-of-the-art baselines. To verify whether ResNLS-5 can help clients effectively avoid risks and earn profits in the stock market, we construct a quantitative trading framework for back testing. The result shows that the trading strategy based on ResNLS-5 predictions can successfully mitigate losses during declining stock prices and generate profits in periods of rising stock prices.

[AI-74] Samsone: A Family of Open Small Audio Language Models for On-Device Inference INTERSPEECH2026

链接: https://arxiv.org/abs/2609.21666
作者: Piotr Masztalski,Michał K. Grzeszczyk,Olaf Sikorski
类目: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI)
备注: Accepted for Interspeech 2026

点击查看摘要

Abstract:The success of Large Audio Language Models has driven the development of massive multimodal networks exceeding billions of parameters. However, the demand for privacy-preserving, low-latency processing has shifted focus toward Small Audio Language Models (SALMs) capable of on-device execution. In this paper, we introduce Samsone, a family of SALMs designed for edge computing. Our core model, Samsone-134M, establishes a new state-of-the-art for its size class across multiple benchmarks. We further explore the scaling laws of SALMs by introducing Samsone-99M and Samsone-356M. Despite their compact footprint, the Samsone family delivers performance competitive with models orders of magnitude larger. To foster open research and reproducibility, we train Samsone on publicly available data. We release the training code, model weights, mobile-optimized checkpoints and provide an open-source Android application to demonstrate real-time on-device inference of Samsone.

[AI-75] OmniVChat: Synthesizing Benchmarking and Training for Native Audio-Visual Dialogue

链接: https://arxiv.org/abs/2609.21465
作者: Haolin He,Yunfei Chu,Qi Chen,Wen Huang,Yuan Feng,Muzhi Zhu,Zheqi Dai,Haoning Xu,Dongchao Yang,Chunyat Wu,Zining Liang,Zhengxi Liu,Xiquan Li,Xie Chen,Xize Cheng,Qize Yang,Jin Xu,Qiuqiang Kong
类目: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Image and Video Processing (eess.IV)
备注:

点击查看摘要

Abstract:We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In OmniVChat, omni models directly and simultaneously receive audio and video from a user and return text. The user’s query is embedded in the audio and video, without a separate text question, external captioning, or speech recognition. Direct audio-visual input reduces external latency and computation while preserving perceptual cues. However, research on OmniVChat faces two constraints: data availability and evaluation. Recordings of people using their own devices are scarce. Furthermore, a good reply often needs to account for the user’s surroundings, facial expressions, and nearby objects, and such responses can be expressed in many different ways, making keyword matching unreliable for evaluating reply quality. Recent progress in agent systems and video generation makes generation for comprehension viable, which means using synthesized dialogues for training and evaluation. Therefore, we present OmniVChat-Studio, a multi-agent data engine for synthesizing single- and multi-turn audio-visual dialogues. We use synthesized dialogues to build OmniVChat-Bench, an evaluation benchmark that evaluates omni models’ basic dialogue abilities across five ability categories. We also present OmniVChat-RL, a reinforcement learning reward design that jointly targets reply correctness, efficiency, and style in OmniVChat. Training Qwen3-Omni-Instruct with OmniVChat-RL on synthesized dialogues improves its performance on both OmniVChat-Bench and the human-recorded OmniVChat-Bench-Human. These gains validate the reward design and show transfer to real-world dialogues in training and evaluation.

[AI-76] Reinforcement learning for post-coronagraphic wavefront control

链接: https://arxiv.org/abs/2609.20880
作者: Manuela Castañeda-Medina(LIRA),Yann Gutierrez(LIRA),Johan Mazoyer(LIRA, CNRS),Baptiste Abeloos,Laurent Mugnier,Olivier Herscovici-Schiller
类目: Instrumentation and Methods for Astrophysics (astro-ph.IM); Earth and Planetary Astrophysics (astro-ph.EP); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Direct imaging of exoplanets is limited by the extreme contrast between the star and the planets, which is mitigated using a coronagraph. However, optical aberrations cause starlight leakage through the coronagraph, producing speckles that obscure the planetary signal. Achieving the required contrast levels demands wavefront control with subnanometric precision. Deep reinforcement learning offers a promising alternative to traditional focal-plane wavefront control techniques by enabling adaptive correction strategies learned directly from interaction with the system. In this work, we present a fully data-driven method for post-coronagraphic aberration correction in a simulated high-contrast imaging testbed. The agent controls a deformable mirror using observations consisting of focal-plane measurements (images) and physics-informed wavefront sensing information derived from these images. We evaluate different observation representations and control strategies, and the method is validated on simplified simulations of a high-contrast imaging testbed, where it successfully creates dark holes, i.e., regions of the focal plane in which residual starlight is strongly suppressed, while approaching the performance of conventional wavefront control methods.

机器学习

[LG-0] BrainWideBench: Benchmarking large-scale pretraining and across-animal transfer in multi-region neural recordings

链接: https://arxiv.org/abs/2609.22064
作者: Alexandre Andre,Shivashriganesh P. Mahato,Vinam Arora,Keshav Balaji,Divyansha Lachi,Nanda H. Krishna,Jingyun Xiao,Yizi Zhang,Ximeng Mao,Wenrui Ma,Han Yu,International Brain Laboratory,Daniel Birman,Niccolò Bonacchi,Gaelle A. Chapuis,Joana A. Catarino,Felicia Davatolhagh,Mayo Faulkner,Laura Freitas-Silva,Fei Hu,Julia M. Huntenburg,Anup Khanal,Inês Laranjeira,Petrina Lau,Guido T. Meijer,Nathaniel J. Miska,Jean-Paul Noel,Alejandro Pan-Vazquez,Georg Raiser,Cyrille Rossant,Karolina Z. Socha,Anne E. Urai,Miles J. Wells,Steven J. West,Olivier Winter,Blake Richards,Guillaume Lajoie,Cole Hurwitz,Mehdi Azabou,Matthew R. Whiteway,Liam Paninski,Eva L. Dyer
类目: Machine Learning (cs.LG); Neurons and Cognition (q-bio.NC)
*备注:

点击查看摘要

Abstract:Advances in large-scale neural recording have made it possible to collect data across many animals and distributed brain regions, raising the question of whether this scale can be exploited to learn general-purpose neural representations transferable across diverse downstream tasks. Yet, progress toward this goal has been limited by fragmented evaluation protocols and a narrow focus on individual task domains. Here, we present BrainWideBench, a benchmark for evaluating across-animal transfer on multi-region neural recordings, built on the International Brain Laboratory Brainwide Map dataset of neural and behavioral recordings spanning 276 brain regions from 139 mice performing a sensory-guided decision-making task. The benchmark is organized around three complementary task suites that evaluate whether learned representations support downstream decoding of behavior, can predict masked or future neural activity, and can recover biologically meaningful anatomical organization. With this benchmark, we systematically evaluate pretraining methods across transfer settings, including finetuning on downstream objectives and zero-shot generalization to unseen animals. Our results confirm pretraining improves performance over matched single-session baselines, but we show current methods exhibit heterogeneity in transfer capabilities: gains depend strongly on the alignment between pretraining objectives and downstream tasks. No single approach performs uniformly well across all three suites, and most methods are designed to only address a subset of them. Together, these findings suggest that learning representations that jointly generalize across behavior, dynamics, and anatomy remains an open challenge. By providing a unified and reproducible evaluation suite, BrainWideBench establishes a framework for measuring progress toward general-purpose models of the mouse brain.

[LG-1] Benchmarking World Models for Continual Learning on Compositional Tasks

链接: https://arxiv.org/abs/2609.22055
作者: Haoyu Zhou,Joe Watson,Anson Lei,Ingmar Posner
类目: Machine Learning (cs.LG); Robotics (cs.RO)
*备注:

点击查看摘要

Abstract:A desirable property of a world model is the ability to learn continually across tasks, adapting to new environments without forgetting what the agent has already learnt. In particular, the ability to retain and reuse knowledge obtained from prior experiences underpins an agent’s ability to efficiently adapt to novel environments, as the dynamics of the physical world can often be described in recurring mechanisms. However, the world model’s measure of adaptation entangles two abilities: the speed and capacity to learn unseen tasks, and the reuse of knowledge already acquired, since incoming tasks carry novel content alongside what recurs. In order to isolate knowledge reuse from prior experiences, we propose a compositional continual learning benchmark for world models in robot manipulation. Specifically, we design each task curriculum with compositional tasks that combine aspects of the tasks seen in the sequence. We further factorise this composition along the axes of action and perception to better understand how different input modalities bottleneck knowledge reuse. We evaluate state-of-the-art world models under canonical continual learning methods, alongside a modular world model whose dynamics backbone contains explicitly reusable components. Results show that modularity balances reuse against forgetting better than conventional methods, but none solve the problem fully, leaving clear room for continual world models built to reuse without forgetting. More details are available on our project website: this https URL.

[LG-2] Particle Competition and Cooperation for Robust Graph Convolutional Network Learning Under Label Noise

链接: https://arxiv.org/abs/2609.22053
作者: Fabricio Breve
类目: Machine Learning (cs.LG)
*备注: Submitted to Neurocomputing. Code and experimental results are publicly available at this https URL and this https URL

点击查看摘要

Abstract:Graph Convolutional Networks (GCNs) are highly sensitive to label noise, since corrupted supervision can propagate through the graph and degrade learned node representations. This work proposes PCC+GCN, a hybrid framework that uses Particle Competition and Cooperation (PCC) as a graph-based label-refinement stage before GCN training. PCC identifies suspicious labeled nodes through particle domination dynamics and determines whether their labels should be preserved, removed, or reassigned before GCN training. The framework also allows the graph used by PCC to be augmented with feature-based k -nearest-neighbor edges, while the GCN itself is trained on the original graph structure and node features. The proposed method was evaluated on ten graph datasets from the NoisyGL benchmark under conventional Uniform, Pair, and Random label noise, as well as under instance-dependent label noise. A detailed hyperparameter analysis was also conducted on Cora, CiteSeer, and PubMed. Under conventional noise, PCC+GCN achieved the highest overall average accuracy and the best average rank among the evaluated methods, with an average gain of 1.67 percentage points over the baseline GCN across the clean setting and all noisy scenarios. Under instance-dependent noise, PCC+GCN remained competitive with the best-performing robust methods while requiring substantially lower execution time, being the fastest robust method on eight of the ten datasets. The results indicate that PCC-based label refinement provides an effective and computationally efficient preprocessing strategy for improving GCN robustness under noisy supervision.

[LG-3] Available Guardrails: Certifying Selective Prediction across ML Systems

链接: https://arxiv.org/abs/2609.22048
作者: Parivesh Priye,Yufeng Wang,Haibin Ling,Michael Chaykowsky
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:A selective predictor acts as a safety gate: it returns an output only when the prediction appears sufficiently trustworthy. Deployments increasingly require this reliability to be certified at a target precision for every reporting unit of interest, such as a tool, policy label, or patient subgroup. The main difficulty is often not whether a granted certificate is valid, but whether finite calibration data can produce one at all. As the gate becomes safer or more fine-grained, some units may receive too little evidence to certify. We make this notion of availability computable through classical exact-binomial inversion and formulate reporting-partition selection, under a fixed group order, as a dynamic program that exposes the trade-off among safety, granularity, and served traffic. The resulting frontier reveals a large population opportunity that finite-sample estimation nearly erases: a truth-informed planner gains 0.157 mean coverage over support balancing, whereas a naive estimator recovers only 0.005 , making recovery from finite data the central challenge. Constructing candidate partitions on one planning split and selecting among them on another recovers part of this gap, improving mean coverage over support balancing by 0.060 , with the direction reproduced in 59 of 60 model effects across three intent-routing datasets and two architectures. A complementary validity-preserving lever, reallocating the familywise error budget across reporting units, recovers additional coverage both with population quantities and noisy estimates. The same frontier recurs, with predictor-specific ceilings, across LLM tool-calling, content moderation, lesion classification, and recommendation. Certified availability is therefore a plannable deployment resource that determines when a safety gate can be certified, at what granularity, and over how much traffic.

[LG-4] λ-Controlled GRPO: Turning Flow-Matching Ratio Instability into a Budgeted Resource

链接: https://arxiv.org/abs/2609.22041
作者: Yufeng Wang,Parivesh Priye,Meeshawn Marathe,Ramit Pahwa
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Reinforcement learning is increasingly used to align image generators with reward signals, and Flow-GRPO recently extended this paradigm to flow-matching models by treating the denoising sampler as a stochastic policy that can be optimized from reward feedback. Training in this setting is unstable in a way specific to multi-step denoising: the policy update changes systematically across denoising steps, with importance ratios drifting below one, becoming increasingly dispersed, clipping at different rates, and leaving fewer usable samples late in training. Prior work treats these effects as separate failure modes and addresses each with a hand-tuned stabilizer. We show instead that they arise from a single per-step quantity, which we call path variance. This quantity is determined exactly by the sampler’s Gaussian transition kernel and can be estimated cheaply during training. This reframes instability as a resource that can be measured and budgeted rather than a collection of symptoms to repair. Our method, \lambda -Controlled GRPO, calibrates importance-ratio behavior from this predicted law rather than from noisy empirical statistics, and allocates gradient effort across denoising steps according to their predicted cost. The two scales governing the update are fixed by standard policy choices rather than introduced as free tuning parameters. On a text-to-image model under two reward settings, rendering difficult target text scored by optical character recognition and matching human preferences scored by a preference model, \lambda -Controlled GRPO improves both text accuracy and preference reward over the strongest empirical stabilizer. It also keeps late-step path variance within its intended budget, precisely where the baseline systematically overshoots. The result is a Flow-GRPO update calibrated by its own transition law rather than stabilized after instability appears.

[LG-5] COMPLEX: A Closed-Form Certified Embedding of Multiparameter Persistence Modules

链接: https://arxiv.org/abs/2609.22012
作者: Sushovan Majhi,Atish Mitra,Žiga Virk,Pramita Bagchi
类目: Machine Learning (cs.LG); Algebraic Topology (math.AT)
*备注: 40 pages, 2 figures, 10 tables

点击查看摘要

Abstract:Every multiparameter persistence vectorization we know of carries a one-sided Lipschitz upper bound and nothing below it: without a lower gauge there is no sense in which the features are faithful, and no per-prediction guarantee can be built on them. This paper supplies the missing side. COMPLEX is a closed-form, training-free embedding of multiparameter modules – slice the module along a fixed near-diagonal net, embed each slice barcode by the certified PLACE/PALACE landmark map, concatenate. Under a checkable witnessing-slice coherence condition, holding on 100% of audited pairs on Orbit5k, a single slice carries a closed-form lower gauge: separated modules stay separated in the embedding. With the standard upper bound this gives, to our knowledge, the first two-sided distortion bound for a multiparameter feature map, making faithfulness measurable. Measuring it, we find the floor tight within a small factor of realized distances yet operationally local: an RBF-SVM reaches 91% where 1-NN reaches 78% on the same features. Local per-prediction certification therefore fails for a structural reason common to every landmark embedding whose lower gauge is witnessed by one coordinate. With no learned embedding and no held-out calibration – only a cross-validated SVM head – COMPLEX sets the state of the art on both Orbit benchmarks (91.95% on Orbit5k, 92.98% on Orbit100k), level with or above Euler-characteristic surfaces and above transformers and graphcode. On graphs it exceeds GRIL on all four shared molecular benchmarks with one fixed configuration, including the only multiparameter method to clear COX2’s majority baseline by more than three points. Closed-form selection – of the landmark radius, the kernel (certificate-preserving), and the bifiltration set – buys further accuracy; gradient-shaped adaptation buys none.

[LG-6] Assessment of Machine Learning-Based Critical Heat Flux Models in the CTF Subchannel Code for Square Rod Bundle Prediction

链接: https://arxiv.org/abs/2609.21995
作者: Aidan Furlong,Vinicius de Melo Monteiro,Robert Salko,Juliana Pacheco Duarte,Xu Wu
类目: Machine Learning (cs.LG)
*备注: 28 pages, 10 figures

点击查看摘要

Abstract:The prediction of critical heat flux (CHF), a key safety-related quantity in nuclear thermal hydraulics, remains an important challenge due to its direct relationship with fuel performance and reactor safety. Recent studies have demonstrated that relative to traditional empirical correlations and lookup tables (LUTs), machine learning (ML) methods can substantially improve CHF prediction accuracy. Most ML-based CHF models, however, have been developed and evaluated using tube databases, leaving their applicability to reactor-relevant rod bundle geometries largely unexplored. This study evaluates ML-based CHF models deployed within the CTF subchannel code using the Electric Power Research Institute (EPRI) rod bundle CHF database. Both pure and hybrid residual correction models are considered in local and semilocal formulations. The tube-trained ML CHF models generally transferred favorably to rod bundle applications and outperformed traditional CHF methods across most geometries and operating conditions. The local hybrid LUT model produced the strongest overall performance, and the semilocal pure ML model remained highly competitive. Comparison against the Bowring correlation, W-3 correlation, and 2006 Groeneveld LUT demonstrated that substantial improvements in rod bundle CHF prediction are possible even when models are trained exclusively on tube data. These findings provide one of the first large-scale assessments of ML-based CHF models in square rod bundles within a production-level subchannel analysis environment and support their broader application in reactor thermal hydraulic analysis. Comments: 28 pages, 10 figures Subjects: Machine Learning (cs.LG) Cite as: arXiv:2609.21995 [cs.LG] (or arXiv:2609.21995v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2609.21995 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-7] me series generation with spectrally aligned latent flow matching

链接: https://arxiv.org/abs/2609.21989
作者: Camilo Carvajal Reyes,Felipe Tobar
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Latent flow models have proven to be a reliable and cost-effective method for time series generation. However, the latent compression induces unwanted artefacts, such as a spectral mismatch with respect to the underlying dataset, thus hindering their use as training surrogates. In this article, we propose a spectrally-aligned latent-flow time series generator, where the latent space for flow matching is trained to preserve dynamical properties that are relevant for the suitability of synthetic samples. We find that incorporating fine-tuning losses based on canonical signal representations such as the Fourier, wavelet and signature transforms helps overcome these issues. The interpretability of these transformations allows us to ensure that the synthetic signals are aligned with the true ones in terms of relevant features, such as smoothness or targeted spectral content, as opposed to relying on pointwise reconstruction losses only. We compare the proposed aligned models against a base latent-flow model and the state of the art over real-world long-range univariate and multivariate benchmark datasets. Our quantitative results validate the superiority of the proposed method in terms of its performance on metrics reflecting signal realness and computational efficiency, while being aligned to the training set with respect to its local structure.

[LG-8] Multiplicative Optimism for Constant Regret in Games

链接: https://arxiv.org/abs/2609.21976
作者: Ashkan Soleymani,Georgios Piliouras
类目: Computer Science and Game Theory (cs.GT); Machine Learning (cs.LG); Optimization and Control (math.OC)
*备注:

点击查看摘要

Abstract:We introduce Multiplicatively Optimistic Regret Matching (MORM), an uncoupled learning rule for finite general-sum games. Under simultaneous full-information self-play, every player achieves external regret O(\sqrt n\log d) uniformly over all horizons, using only one-step optimism. The analysis combines a potential-based regret-matching argument with multiplicative stability and Hellinger control of strategy movement. A learning-rate safeguard additionally gives O(\sqrtT\log d) regret in the face of adversarial utilities.

[LG-9] RACER: Role-Aligned Competence Estimation for Human-AI Routing

链接: https://arxiv.org/abs/2609.21953
作者: Joshua Strong,Emma Sun,Alexander Capstick,Pramit Saha,Cheng Ouyang,J. Alison Noble
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Learning to defer asks a predictive system when to act autonomously and when to defer to a human expert. Population-adaptive deferral extends this problem to unseen experts using a small context set of expert behavior. Neural context encoders such as L2D-Pop can be query-dependent, but may learn routing shortcuts tied to absolute class coordinates. Identity-Free Deferral (IFD) removes such shortcuts through role-indexed classwise competence profiles, but its estimates are constant within each class and cannot capture instance-level expert specialization. We propose RACER—Role-Aligned Competence Estimation for Routing—a role-relative framework for estimating an unseen expert’s competence from context. RACER estimates the posterior-predictive probability that the expert is correct on a query under each candidate class role, then combines these estimates with the model posterior to obtain the Bayes-relevant expert-correctness probability. Nonparametric and neural kernel-pooling estimators use candidate-role relations, shared aggregation, and symmetric summaries, excluding absolute class-identity channels. We prove coherent class-relabelling invariance, derive a Bayes-aligned deferral surrogate, and give a plug-in regret bound relating routing regret to classifier and competence-estimation error. On controlled synthetic benchmarks, including a PathMNIST histopathology context-scaling study with simulated experts, RACER benefits from additional context under hidden subtype dependence and gives the strongest aggregate performance on a separately sampled unseen-expert split in the CIFAR-100 synthetic experiments. On the radiologist and human–AI chest-radiography benchmarks (VinDr-CXR and CheXpert), the RACER family is competitive or best in budget-swept deferral, with calibration results varying across metrics and datasets.

[LG-10] Learning to Move Cities: Deep Meta-Models and Reinforcement Policies for Calibration and Control in Urban Networks ITSC

链接: https://arxiv.org/abs/2609.21945
作者: Adewumi Augustine Adepitan,Christopher J. Haruna,Oluwasegun Adegoke,Ayooluwatomiwa Ajiboye,Oluwatobi Oluwasakin
类目: Machine Learning (cs.LG)
*备注: 7 pages, 2 figures. Accepted for publication in the Proceedings of the 2026 IEEE 29th International Conference on Intelligent Transportation Systems (ITSC), Naples, Italy. © 2026 IEEE. Personal use of this material is permitted; permission from IEEE must be obtained for all other uses

点击查看摘要

Abstract:Urban transportation networks present complex optimization challenges spanning calibration of high-fidelity simulators and real-time operational control. This paper presents a shared latent-space framework that connects simulator calibration and reinforcement learning control through a common learned representation of urban traffic dynamics. First, we develop a combinatorial MLP-autoencoder architecture that learns low-dimensional manifolds linking simulator inputs (origin-destination demand, network parameters) to outputs (travel times, congestion patterns), enabling efficient Bayesian optimization for calibration. This approach demonstrates superior sample efficiency compared to traditional dimension reduction methods, achieving better fit to observational data within fixed computational budgets. Second, we implement a deep Q-learning agent with experience replay and target networks to optimize dynamic traffic assignment through scheduling and routing adjustments. In empirical evaluations on benchmark networks, our approach reduces system-wide travel times by up to 51% compared to baseline operations. The learned latent representation is not only used to reduce the dimensionality of Bayesian calibration, but is also incorporated into the reinforcement learning state representation, allowing the control policy to operate on compressed and calibrated traffic dynamics. This shared latent-space formulation provides a unified pathway from simulator calibration to adaptive operational control within intelligent transportation systems. Our results highlight the transformative potential of deep learning methods in urban mobility planning and management, particularly for large-scale networks where traditional optimization approaches face computational bottlenecks.

[LG-11] End-to-End Hard-Label Cryptanalytic Model Extraction Using Efficient Sign Recovery

链接: https://arxiv.org/abs/2609.21941
作者: Akira Ito,Takayuki Miura,Yosuke Todo
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:The importance of deep neural networks (DNNs) is widely recognized, and the parameters obtained through training are regarded as valuable assets. Recently, attacks that extract these parameters using only oracle queries to a DNN have been actively studied at IACR conferences. The hard-label setting is the most challenging setting for model extraction, where an adversary can observe only the final output label, such as “dog” or “cat.” At Eurocrypt 2025, Carlini et al. proposed polynomial-time hard-label extraction of ReLU-based MLPs. However, one step of this attack process, i.e., sign recovery, requires a large number of queries and substantial computation. Implementing this step in a black-box setting remains difficult. Consequently, a fully black-box end-to-end demonstration on trained deep ReLU MLPs has remained a challenge. In this paper, we propose a new sign-recovery algorithm based on a completely different principle from the existing method. Our method requires no dedicated queries for sign recovery. In our experiments, it achieves higher sign-recovery accuracy than the existing method. Consequently, it enables efficient sign recovery even for trained models. With our sign-recovery algorithm, all steps of hard-label model extraction can be implemented in a black-box setting. By combining these implementations, we demonstrate end-to-end model extraction from models trained on MNIST and Fashion-MNIST, with width 16 and 4 or 6 hidden layers, achieving over 98% label agreement.

[LG-12] Joint Remaining Useful Life Prediction and Capacity Estimation of Lithium-Ion Batteries Using Partial-Charging Data

链接: https://arxiv.org/abs/2609.21932
作者: Khoa Tran,Ho-Si-Hung Nguyen,Phone Wai Yan Moe,Hung-Cuong Trinh,Thi-Hoang-Giang Tran
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Joint remaining useful life (RUL) prediction and capacity estimation require representations of both gradual degradation and recent battery behavior. This paper presents a cross-expert framework using partial-charging measurements without measured historical full-cycle capacity as an input. The RUL Expert encodes nominal 10-min segments from ten cycles sampled within a 30-cycle history using a pretrained gated recurrent unit (GRU) encoder, a two-dimensional convolutional neural network (2D-CNN), and a temporal GRU. The Capacity Expert processes statistical descriptors of nominal 40-min segments from ten consecutive cycles using a 2D-CNN and a Transformer. A feature-wise linear modulation module uses the short-term representation to condition the long-term representation for joint prediction. Training comprises supervised autoencoder pretraining, independent expert pretraining, and fusion training with frozen experts. On two public battery-aging datasets, the reference configuration achieves mean RUL root-mean-square errors of 143.69 and 161.10 cycles and capacity errors of 12.36 and 7.28mAh, respectively. On Dataset I, fusion reduces both mean errors relative to either standalone expert. The results demonstrate a trade-off between RUL and capacity accuracy: the proposed method attains the lowest reported RUL RMSE among the compared methods on both datasets, whereas several baselines yield lower capacity errors.

[LG-13] Kinks vs. Smoothness: Identifiability of Real Analytic nICA for Laplace-like Sources

链接: https://arxiv.org/abs/2609.21926
作者: Isaac Manring,Kejun Huang
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Many machine learning systems try to explain complex data - like images or financial time series - in terms of hidden, independent factors that generated them. Recovering the true underlying factors, rather than some scrambled version of them, is the central challenge of nonlinear Independent Component Analysis (nICA). We prove identifiability (exact recovery) up to trivial ambiguities for real analytic generating functions when source probability density functions have a finite number of discontinuities in the first derivative. The Laplace distribution is the most prominent example satisfying this assumption. Our proof relies on the contrast between kinks in the source distribution and the smoothness of real analytic functions. Real analytic functions comprise a broad class of generating mechanisms, and can be approximated with Normalizing Flows or Variational Autoencoders with standard activation functions (e.g., tanh, softplus, GELU), so our result applies with minimal changes to existing training pipelines. We perform experiments on real and synthetic data with both Normalizing Flows and Variational Auto-Encoders demonstrating their identifiability properties. In experiments on CelebA data we recover several interpretable latent factors controlling unique attributes across the dataset.

[LG-14] Beyond Kinematics: Benchmarking Simulation Fidelity for Muscle-Driven Imitation Learning ICRA2027

链接: https://arxiv.org/abs/2609.21909
作者: Ayah G. Ahmad,Claire E. Borden,Maegan Tucker
类目: Machine Learning (cs.LG); Robotics (cs.RO)
*备注: 8 pages, 5 figures, 2 tables, submitted to ICRA 2027

点击查看摘要

Abstract:In this work, we conduct a systematic comparison of two state-of-the-art motion-imitation reinforcement learning (MIRL) pipelines, one built on SCONE/HyFyDy and one built on MuJoCo/MyoSim. HyFyDy emphasizes physiological realism through detailed musculotendon modeling, while MuJoCo prioritizes computational efficiency and scalable policy learning. While recent work has demonstrated that both pipelines reproduce human kinematics with high fidelity, it remains unclear if they accurately capture the underlying neuromuscular behavior that produced the movement. This limitation is particularly important for robotic assistive-device design and control, where outcome measures such as muscle activation patterns and metabolic cost are often used as optimization targets. To conduct a systematic comparison, our work compares both pipelines using a common set of human motion-capture and electromyography (EMG) measurements. The results find that while both pipelines produce similar kinematics with relative accuracy, the muscle activations from HyFyDy are more aligned with the experimental EMG, as supported by the average pooled (RMSE, r) values for muscle activations from HyFyDy and MuJoCo: (0.164, 0.4) and (0.344, 0.11), respectively. While we conclude that the more advanced physiological realism of HyFyDy currently makes it more suitable for musculoskeletal modeling, both require further development to bring physiological realism to GPU-parallelizable simulation environments and advance robotic assistive device design.

[LG-15] Intervention Granularity Matters: Coherent Treatment Bundles in Counterfactual Simulation with Clinical World Models

链接: https://arxiv.org/abs/2609.21906
作者: Fangzhou Wang,Yixuan Yang,Camilla Balzarotti,Rishikesan Kamaleswaran
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Counterfactual simulation with a clinical world model means fixing a patient’s history, changing the treatment, and reading off the predicted response. Doing so requires deciding what counts as one intervention. In clinical settings, interventions are documented as bundles: a co-occurrence audit of 945,707 patient-hours from MIMIC-IV shows groups of components, such as every parameter of a dialysis circuit, that never appear apart, so an edit that changes one component on its own describes an hour that never occurs in the data. We hypothesize that the granularity at which an intervention is edited changes how a world model responds, and test this with Clin-JEPA, a latent world model of patient trajectories conditioned on hourly treatment text. At 1,019 documented onsets of invasive ventilation, we keep the patient’s history and other treatments fixed and compare editing one ventilator setting with editing the complete configuration recorded for a real patient with the most similar recent trajectory. The complete bundle moves the predicted next state further than any single setting, consistently across all five settings, and the difference remains after accounting for how much each edit changes the model’s input. Intervention granularity therefore materially affects the response of a clinical world model: single-component edits may understate treatment sensitivity, and bundle-aware editing may offer a better-supported basis for counterfactual treatment simulation.

[LG-16] ExpBoN: Exponential-Noise Best-of-n for Efficient Test-Time LLM Alignment

链接: https://arxiv.org/abs/2609.21899
作者: Yanxiao Liu,Sicheng Wan,Deniz Gündüz
类目: Machine Learning (cs.LG); Information Theory (cs.IT)
*备注:

点击查看摘要

Abstract:Best-of- n (BoN) sampling is a simple yet effective inference-time alignment method, but hard maximization provides only coarse control over the trade-off between reward and distribution shift. Soft Best-of- n (Verdun et al. 2025) provides smoother control and converges to the optimal distribution associated with KL-regularized reward maximization. In this paper, we introduce ExpBoN, an alternative soft BoN method based on the exponential-noise report-noisy-max mechanism. It admits an exact finite- n decomposition, which yields exponentially fast convergence in total variation, expected reward, and both directions of KL divergence. We provide comprehensive theoretical analyses of its convergence and regret behavior. We further integrate ExpBoN into the guided speculative inference (GSI) framework (Geuter, Mroueh, and AlvarezMelis 2025), resulting in ExpGSI, for efficient reward-guided LLM alignment. ExpGSI yields substantial reductions in computational cost while maintaining comparable accuracy. Experiments on MATH500, MMLU-STEM, and Minerva Math with the Qwen2.5-Math and Qwen3 model families show that ExpGSI reduces estimated computation by 14% - 39% across candidate budgets for Qwen2.5-Math and by up to 45% at n=16 for Qwen3. Overall, our results provide a theoretical and algorithmic foundation for exponential-noise BoN and efficient test-time LLM alignment.

[LG-17] LLM s as Feature Engineers for Text-and-Tabular Prediction

链接: https://arxiv.org/abs/2609.21894
作者: Merwan Barlier,Blaz Skrlj
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We introduce an iterative framework that automates the extraction of interpretable, schema-bound categorical features from unstructured text for tabular prediction models. To navigate the feature space, a generator LLM proposes semantic definitions, a separate extractor LLM materializes the features, and a downstream tabular model evaluates their predictive performance. We optimize this search by translating explicit model errors, such as AUC ranking inversions, into natural-language feedback, steering the LLM to resolve specific predictive failures. Evaluated across three public datasets, this error-driven loop accelerates feature discovery by up to 3\times compared to unguided search. Empirically, the generated features demonstrate strong multi-view complementarity, strictly outperforming any subset when combined with TF-IDF and dense embeddings. Finally, the framework guarantees instance-level interpretability: the discovered features dominate SHAP importance rankings and provide a fully transparent, semantic audit trail for every prediction.

[LG-18] Geometric Mean Pooling for Equal-Weight Multiplicative Coarse-Graining

链接: https://arxiv.org/abs/2609.21876
作者: Ang-Kun Wu,Fangdi Wen,Jingtao Zhang
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 17 pages, 6 figures

点击查看摘要

Abstract:As an alternative to the additive and extremal biases of average and max pooling, we introduce Geometric Mean Pooling (GMP), a signed pooling operator that combines the product of feature signs with the geometric mean of feature magnitudes. Motivated by local-to-global composition in quantum many-body physics, GMP retains both joint sign information and a characteristic multiplicative scale without introducing learnable pooling parameters. We show that non-overlapping hierarchical GMP preserves the corresponding global multiplicative statistic and evaluate it on synthetic sequence tasks, iterative coarse-graining, image classification, and molecular lipophilicity regression. On the synthetic tasks, GMP recovers product-based signals more accurately than average and max pooling and maintains predictive performance under the tested levels of multiplicative input noise. On image and molecular data, however, its effectiveness depends on the representation, target parameterization, and placement of local and global pooling. These results position GMP as a complementary, regime-dependent inductive bias for tasks in which equal-weight multiplicative composition is plausible, rather than as a universal replacement for standard pooling operators.

[LG-19] Watermarkable Multi-Draft Speculative Sampling via Poisson Processes

链接: https://arxiv.org/abs/2609.21858
作者: Yanxiao Liu,Sicheng Wan,Zhan Gao,Deniz Gündüz
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Large language models (LLMs) have achieved state-of-the-art performance across a wide range of tasks, motivating two important aspects of deployment: inference efficiency and output provenance, which can be tackled by speculative sampling and watermarking, respectively. However, recent works have shown that combining these two goals is highly nontrivial and can be potentially impossible. In this work, we develop a novel multi-draft speculative sampling algorithm based on Poisson processes that improves the frontier of this fundamental trade-off. The proposed algorithm has strong sampling efficiency on its own and, more interestingly, is naturally watermarkable: we can embed an unbiased watermark without degrading speculative acceptance. Moreover, our algorithm is based on an exact list-coupling-without-communication scheme, which yields a drafter invariance property that benefits both sampling and watermarking. It is the first multi-draft, drafter-invariant speculative sampling scheme that maintains both watermark strength and sampling efficiency, and we experimentally verify its strong performance in both aspects.

[LG-20] Adaptive Uncertainty-Aware Modeling and Stochastic Radial Basis Function Predictive Control for Personalized Fluid Resuscitation

链接: https://arxiv.org/abs/2609.21821
作者: Elham Estiri,Hossein Mirinejad
类目: ystems and Control (eess.SY); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:This paper presents a novel framework integrating Bayesian physiological modeling with optimal control strategies to achieve uncertainty-aware, personalized hemodynamic regulation during fluid resuscitation. An uncertainty-aware variational autoencoder state-space model (UVAE-SSM) was first developed to capture the dynamical relationship between mean arterial pressure (MAP) and fluid infusion using limited data, while explicitly modeling aleatoric uncertainty (i.e., randomness in the measurements, such as sensor noise). Then, a Bayesian nonlinear state-space model (BNSSM) was developed by utilizing Bayesian neural networks (BNNs) to capture epistemic uncertainty arising from physiological and patient-specific variability, enabling the creation of a virtual patient generator (VPG). Building on this uncertainty-aware modeling framework, a stochastic radial basis function model predictive control (sRBF-MPC) algorithm was designed to track the MAP target while satisfying physiological constraints. Finally, an online fine-tuning algorithm was developed to adapt the nominal UVAE-SSM using streaming VPG data, enabling progressive personalization during closed-loop therapy. Simulation results across unseen animal subjects and an independent human clinical dataset demonstrated the strong predictive accuracy and cross-population generalizability of the UVAE-SSM and BNSSM models. Closed-loop evaluations confirmed that the proposed sRBF-MPC framework achieved stable MAP regulation while providing better risk-aware control compared to quadratic MPC (Q-MPC) and stochastic quadratic MPC (sQ-MPC). Overall, the proposed framework accounts for inter- and intra-patient variability through online model adaptation, offering a promising step toward uncertainty-aware, personalized hemodynamic modeling and control in critical care.

[LG-21] RegKT: Interpretable and Robust Deep Knowledge Tracing With IRT-Regularizer

链接: https://arxiv.org/abs/2609.21791
作者: Samuel Girard,Juan D. Pinto,Jill-Jênn Vie,Amel Bouzeghoub
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:As deep learning models continue to advance, knowledge tracing models have achieved higher accuracy. However, these gains come at the cost of reduced interpretability, which is crucial for practitioners in educational settings to adopt new methodologies. Additionally, deep learning models are prone to overfitting, particularly when dealing with the small datasets that are common in educational applications. In this paper, we propose a novel regularization technique designed to enhance the robustness of deep-learning-based knowledge tracing models, while simultaneously improving their interpretability. Our method addresses both the interpretability and overfitting challenges, making it more feasible for real-world educational applications.

[LG-22] From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention

链接: https://arxiv.org/abs/2609.21788
作者: Sichang Su,Benjamin Yang,Zhiyun Deng,Boyuan Liang,Yip Fun Yeung,Zelin Wang,Lingfeng Sun
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注: Project page: this https URL

点击查看摘要

Abstract:A pretrained robot foundation policy may execute most of a long-horizon task yet repeatedly fail at a few critical subtasks. Collecting additional full-task demonstrations for supervised fine-tuning (SFT) requires operators to repeat behaviors the policy already performs well. Reinforcement learning (RL) fine-tuning offers a promising path to bridge this gap, but existing approaches struggle to solve long-horizon tasks using only sparse rewards. We present PARTS (Policy Adaptation with RL on Targeted Subtasks), a real-world subtask RL framework that concentrates practice at these bottlenecks while allowing training rollouts to proceed with minimal human intervention. The frozen pretrained policy supplies nominal actions throughout execution, while agent-generated selectors and success verifiers activate residual corrections and provide local outcome rewards. These rewards support learning from successful subtasks even when complete-task successes are scarce. Training combines online RL with success-reweighted retraining, and each retrained residual policy is redeployed to collect further experience. Humans identify bottlenecks during setup and perform physical resets when needed. On bimanual YAM and single-arm Franka tasks, PARTS improves complete-task success from 32% to 61% and from 50% to 95%, respectively, using tens of minutes of real-world RL rollouts per task on average. Compared with existing real-world RL fine-tuning methods, PARTS raises full-task success by more than 25% under the same robot-rollout budget while requiring less human involvement.

[LG-23] Bilevel Optimization of Topology and Hyperparameters (BOTH)

链接: https://arxiv.org/abs/2609.21758
作者: Suryanarayanan Manoj Sanu,Miguel Anibal Bessa,Alejandro Marcos Aragón
类目: Machine Learning (cs.LG)
*备注: Currently under submission to SMO journal

点击查看摘要

Abstract:Topology optimization (TO) represents a significant step towards automating the design process: given a working simulation, TO can produce a viable prototype at the press of a button by differentiating the simulation and iteratively improving the design. In practice, however, TO is riddled with magic numbers''---hyperparameters whose tuning significantly affects the outcome. Finding the right values typically requires not only deep problem-specific knowledge but also extensive trial-and-error. While practitioners can use surrogate-assisted hyperparameter optimization as an alternative, this approach requires strictly limiting the number of hyperparameters through careful problem formulation. Here, we propose differentiating TO itself using automatic differentiation. This yields hypergradients’’ that allow us to tune these hyperparameters in tandem with the primary optimization. We show that evaluating just one or two steps of TO is sufficiently informative and that the method scales favorably to thousands of hyperparameters at an expense comparable to only a few standard TO runs. We demonstrate this approach on stress-constrained and compliance problems, with the latter utilizing a neural parameterization of the density field.

[LG-24] GraphSkillEvo: Evolutionary Optimization of Graph-Structured Agent Skills

链接: https://arxiv.org/abs/2609.21749
作者: Rui Sun,Zhi Zheng,Zhenkun Wang,Zhichao Lu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Skills can improve the performance of Large Language Model (LLM) agents by providing task-specific procedural guidance, while skill optimization further improves their effectiveness through iterative refinement. However, existing skill optimization methods typically represent skills as unstructured natural-language instructions, creating two key challenges: 1) Unstructured skills often lack explicit workflow-level guidance and contain substantial redundancy, making them difficult for LLMs to execute; 2) the vast search space of unconstrained natural-language skills makes skill optimization ineffective. To address these challenges, we propose representing skills as graph-structured natural-language artifacts. In graph-structured skills, each node represents an execution step together with its operational guidance, while directed edges encode context-dependent transitions between steps. Compared to unstructured skills, graph-structured skills can provide clear workflow-level guidance. Moreover, the proposed graph-structured skill can also facilitate skill optimization. Building on this structured representation, we introduce GraphSkillEvo, a population-based evolutionary optimization framework with mutation and crossover operators for graph-structured skills. By maintaining multiple candidate skills and combining effective components, GraphSkillEvo enables broader and more comprehensive exploration of the structured skill space than purely LLM-based iterative self-refinement. Extensive experiments across five agent benchmarks demonstrate that GraphSkillEvo consistently outperforms the strong skill optimization baseline SkillOpt, improving average accuracy by 4.01% on GPT-5.4-nano and 1.76% on GPT-5.4. Our code is available at this https URL.

[LG-25] GEM-MPC: Balancing Exploration and Exploitation through Expert-Guided Planning

链接: https://arxiv.org/abs/2609.21735
作者: Alvaro Serra-Gomez,Thomas Moerland
类目: Machine Learning (cs.LG)
*备注: Preprint

点击查看摘要

Abstract:Effective exploration in high-dimensional continuous control remains a central challenge in reinforcement learning. Planning-based methods address this by combining online planning with learned policies and value functions, but their components can become misaligned during training: learned sampling policies may diverge from planner behavior, while planning distributions stored in replay become stale as the model and value function evolve. Reanalysis can refresh these targets, but at substantial computational cost. We propose GEM-MPC, an MPPI-based reinforcement learning method that improves the interaction between planning and learning. GEM-MPC uses MPPI to combine a policy trained to clone the planner with a KL-regularized policy that explores around it, providing complementary exploitation and guided exploration within planning. We further introduce Gated Prior Distillation, which selectively learns from stored planning distributions only when they provide a better target than the current prior, reducing the impact of stale planning data without requiring full reanalysis. Across continuous-control benchmarks, GEM-MPC consistently outperforms existing planning-based baselines under lower computational budgets.

[LG-26] SpecQuant: Speculative Decoding with Multi-Parent Quantization for Adaptive LLM Inference

链接: https://arxiv.org/abs/2609.21704
作者: Harish KB,Jagadeeswaran M,Pradheep P,Yuvanesh S,Sivakumar T
类目: Machine Learning (cs.LG)
*备注: 5 pages, 1 figure. Published in the 2026 Fifth International Conference on Power, Control and Computing Technologies (ICPC2T)

点击查看摘要

Abstract:Running large language models (LLMs) locally continues to be limited by restrictions of compute and memory on consumer hardware. The popular acceleration technologies, such as quantization, speculative decoding, and adaptive inferencing, offer substantial speed boosts but usually necessitate retraining, per architecture tuning, or draft models. SpecQuant is a trainingfree framework, that combines speculative decoding with multiparent quantization to perform adaptive, efficient inference of LLMs. SpecQuant derives multiple quantized variants (INT4, FP8, FP16) from a shared base model, and dynamically routes queries based on predicted complexity; lightweight variants are used for simple or factual tasks, and full-precision models are used for complex reasoning tasks or long-context inputs. The shared-weight design of SpecQuant ensures sufficient token acceptance for speculative decoding without compatibility issues using separate draft parent models. We evaluate SpecQuant on Qwen2.5 based models on the MMLU, AlpacaEval, and GSM8K datasets, or benchmarks, demonstrating 35-43% speedups without degrading accuracy greater than 2%, substantial within the LLM community. SpecQuant enables practical on-device LLM deployment across diverse hardware without special infrastructure or expertise.

[LG-27] Optimization Geometry of Equivalent Brownian RKHS Representations

链接: https://arxiv.org/abs/2609.21693
作者: Mahdi Mohammadigohari,Gustau Camps-Valls
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Equivalent finite parameterizations can represent the same functions and intrinsic norm yet induce different optimization algorithms. We study this effect in a controlled finite Brownian RKHS with nodal, increment, and spectral coordinates. Classical finite-element, RKHS-interpolation, Brownian-covariance, and mixed-boundary DCT identities make the shared hypothesis class, Brownian energy, approximation operator, and coordinate maps explicit. Our main results concern the optimization geometry of this fixed model. With mapped initialization, identical scalar steps, and identical minibatches, nodal and spectral GD/SGD have exactly the same mapped trajectories. Increment GD is an explicit Euler step for the constant Brownian/Sobolev metric, with factor 1/h . For Brownian-regularized least squares, \kappa_2(\mathbf H_\mathrminc)\le1+A/\rho , independently of grid resolution G for fixed A , \rho0 , and the stated normalization. Under the stated standard-Adam convention, the universal orthogonal equivariance group is exactly the signed permutations; the block DCT-VIII transform is not one. Float64 tests over five grids numerically verify the finite identities, mapped one-layer and recursive trajectories, conditioning predictions, and theorem-matched Adam separation. Thus coordinate effects are isolated without changing the represented functions, intrinsic regularizer, or approximation space.

[LG-28] Multi-Domain Clustering via Measure Quantization

链接: https://arxiv.org/abs/2609.21664
作者: Rafael Pereira Eufrazio,Eduardo Fernandes Montesuma,Charles Casimiro Cavalcante
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Clustering is a fundamental task in data analysis, typically addressed through centroid-based methods such as K-means. In this work, we present a general framework for multi-domain clustering via measure quantization: given samples from multiple domains, we learn a shared set of cluster prototypes by minimizing a probability metric, such as the Sinkhorn divergence or the Maximum Mean Discrepancy, between each domain’s probability measure and the measure of prototypes. Data points are then assigned to clusters either via nearest centroid, or via optimal transport, a collaborative strategy that couples all samples within a domain. A mini-batch optimization strategy makes both fitting and assignment scalable, reducing memory and computational cost while preserving clustering performance. Experimental results on 5 multi-domain benchmarks spanning image, audio and sensor data show that our Sinkhorn-based method consistently outperforms classical and multi-domain clustering baselines, and that this advantage persists when scaling to hundreds of thousands of samples.

[LG-29] Beyond Gaussian Worlds: Latent Geometry Matters for JEPAs

链接: https://arxiv.org/abs/2609.21656
作者: Léo Nicollier(CB, ATT),Enric Meinhardt-Llopis(CB),Marc Pic(ATT),Pablo Musé(CB, IFUMI),Gabriele Facciolo(CB)
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Recent Joint-Embedding Predictive Architectures (JEPAs) prevent representation collapse by constraining learned representations to follow a prescribed target distribution, such as an isotropic Gaussian or the uniform distribution on a hypersphere. Klindt et al. (2026) showed that, under their Euclidean assumptions, matching a Gaussian target can recover Gaussian latent variables up to a linear transformation, and that the Gaussian is the unique distribution with this guarantee. We extend their analysis to latent variables supported on embedded Riemannian manifolds and derive conditions on the latent geometry and positive-pair dynamics under which alignment and exact distribution matching guarantee linear recovery. In particular, when the latent variables are uniformly distributed on a sphere and the representations are matched to the same spherical distribution, every optimal representation recovers the latent state up to an orthogonal transformation. This shows that Gaussian uniqueness is not a universal property of distribution-matched JEPAs: non-Euclidean latent geometries can admit other linearly recoverable distributions. We further derive an approximate-recovery bound that is strictly tighter for the spherical world than for the Gaussian world. Experiments on Gaussian, spherical, and toroidal latent spaces show that geometrically compatible targets yield better linear recovery when optimization succeeds, whereas mismatched targets distort the latent structure. This advantage persists in high-dimensional Clifford-torus worlds.

[LG-30] Riemannian Neural Hamiltonian Flows: Geodesic Symplectic Transport and Interpretability

链接: https://arxiv.org/abs/2609.21647
作者: Vincent Souveton
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Hamiltonian normalizing flows are attractive generative models because their phase-space maps are invertible and volume preserving, but most neural constructions are formulated in Euclidean space. We introduce Riemannian Neural Hamiltonian Flows, which combine the fixed kinetic energy of a Riemannian manifold, a learned scalar potential, and an explicit geodesic leapfrog integrator. Our analysis explains how the learned Hamiltonian can be made interpretable. Every normalizable potential defines an implicit profile, and the position marginal initially accelerates along the relative score between that profile and the base. The matched potential is the interpretable specialization for which the implicit profile is the target. In the isotropic Gaussian case, the mechanism corresponds to a phase-space rotation. A local harmonic analysis extends this result around each mode of a general target on a manifold. The gap between the learned and the matched potential is the sum of a residual memory of the base and a bias of the model, and the two potentials agree when the position base has been transferred to the momentum. This can be achieved when the former is broader than the target. Numerical experiments on Euclidean, hyperbolic, and spherical spaces show competitive sample quality and numerical cost against a Riemannian continuous normalizing flow, and confirm the interpretability of the learned potential.

[LG-31] rading Depth for Time in Recurrent Transformers

链接: https://arxiv.org/abs/2609.21605
作者: Zeyi Huang,Xuehai He,Yong Jae Lee,Yelong Shen
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Recurrent Transformers increase computational depth through temporal recurrence, feeding each token’s high-level hidden state into the computation of the next. This raises a natural question: is additional computation better spent on more temporal steps or greater physical depth? We investigate this question using Latent Recurrent Transformers (LRTs), which retain one backbone forward pass per vocabulary token during decoding and provide a controlled setting for comparing these two ways of adding computation. Specifically, we insert a latent thought token between consecutive vocabulary tokens. Each thought token passes through the same L layers as a vocabulary token, sharing the backbone parameters and providing an additional stage of hidden-state refinement before predicting the next token. We compare this L -layer LRT against a 2L -layer LRT without thought tokens. Both execute 2L Transformer blocks per vocabulary token during decoding, but the thought-token model uses fewer parameters. On 16- and 20-layer mixture-of-experts NanoChat backbones, one thought token brings the shallower model within 0.006 and 0.004 bits per byte of its double-depth counterpart, recovering 67% and 81% of the improvement with approximately 48% fewer total parameters. These results suggest that temporal thinking offers a parameter-efficient alternative to increasing physical depth in recurrent Transformers.

[LG-32] Predictive Suppression Layers for Communication-Efficient Spiking Neural Networks

链接: https://arxiv.org/abs/2609.21583
作者: Aidin Attar,Michele Rossi
类目: Neural and Evolutionary Computing (cs.NE); Machine Learning (cs.LG); Signal Processing (eess.SP)
*备注: 6 pages, 5 figures. Accepted at the 1st Neuromorphic Physical Layer Signal Processing for Wireless Systems Workshop (NeuroPHY 2026), co-located with EWSN 2026

点击查看摘要

Abstract:Feedforward Spiking Neural Networks (SNNs) typically propagate every generated spike indiscriminately, disregarding whether the information is redundant from an information-theoretic perspective. This lack of selectivity induces high redundancy in inter-layer communication, creating an expensive overhead, e.g., in scenarios involving many-core neuromorphic hardware or communication-dominated Internet-of-Things (IoT) where features are transmitted wirelessly. To address this challenge, we trade localized processing for leaner network channels by introducing a minimal predictive coding framework for SNNs. We propose two layer variants sharing a predictor block: error units, which transmit signed spiking residuals, and predictive suppression, which uses residual magnitude to dynamically gate and forward only unpredictable, “surprising” activity. Evaluated on the N-MNIST and Spiking Heidelberg Digits (SHD) datasets using diagnostic metrics that decouple local processing from cross-layer communication, our new predictive coding layers achieve significant communication savings. Numerical results reveal a three-fold reduction in communicated activity, while increasing the task accuracy for both datasets. The latter finding is notable, and suggests that predictive coding layers not only minimize communication overhead, but also produce output feature vectors with a higher representation power.

[LG-33] MACE: Memory-Agent Co-Evolution with Adaptive Memory Graphs for Multi-Agent Systems

链接: https://arxiv.org/abs/2609.21533
作者: Kairui Yang,Minghao An,Xunkai Li,Ziheng Yi,Zekai Chen,Guangyuan He,Rong-Hua Li
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:LLM-based multi-agent systems generate collaboration traces that record how agents plan tasks, verify intermediate results, and repair failures. Reusing these procedures requires preserving an action’s prerequisites and the outputs needed by subsequent agents. Our empirical studies show that grouping these dependencies into functional memory units improves their retention, while connecting units increases retrieval of the units and links jointly required by a task. The preferred combination of units also changes between instructions and checklists, even when each combination’s content is fixed across formats. Updating choices from the outcomes of each combination and format pairing outperforms scoring combinations and formats separately. These findings motivate MACE, a memory-agent co-evolution framework that adapts memory organization and agent memory use through execution feedback. Its MemGoG structure represents functional units as subgraphs of related conditions, actions, and outputs, connecting them through support, conflict, and repair relations. MACE Loop selects task-relevant units and relations within a memory budget and provides each agent with instructions or checklists for its current operation. It records the selected units, presentation formats, agent outputs, and task outcomes to update unit scores and relations for retrieval and inform subsequent presentation choices. Across eight benchmarks, MACE outperforms ten baselines with an average score of 81.11%, compared with 78.97% for the strongest baseline, SAGE.

[LG-34] OpenMAS-GCom. A Diagnostic Benchmark for Graph-enhanced Multi-Agent Systems

链接: https://arxiv.org/abs/2609.21527
作者: Kairui Yang,Xunkai Li,Kaixiang Zhang,Minghao An,Zekai Chen,Yuxuan Ba,Rong-Hua Li
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Graph-enhanced multi-agent systems (G-MAS) coordinate large language model agents through communication graphs and role assignments, which determine how agents exchange information and divide responsibilities. However, final-score comparisons across systems combine differences in models, communication patterns, roles, and computation costs, making performance differences difficult to attribute to specific communication structures, role assignments, and information flows. To address this evaluation attribution problem, we introduce OpenMAS-GCom, a benchmark for diagnosing how these components affect G-MAS performance through controlled interventions. We represent systems through collaboration units, communication links, shared intermediate information, and execution rules. OpenMAS-GCom compares original systems with versions modified by changing one component while keeping tasks, models, prompts, and budget limits fixed. We rewire communication edges, remove specialist or critic agents, replace intermediate messages with incorrect content, and disable workers during execution. The benchmark evaluates 17 single-agent, ordinary multi-agent, and graph-enhanced configurations on 29 datasets across six domains. We add 400 G-MAS-Complex tasks requiring agents to combine information from multiple documents, resolve conflicting records, and return specified values with source identifiers. Experiments show larger mean losses after specialist removal than after critic removal, different performance degradation under incorrect messages and worker failures despite similar original scores, and different configurations achieving the highest accuracy and accuracy per token on G-MAS-Complex.

[LG-35] IncentRL: The Trade-Off Between Preference Guidance and Task Performance

链接: https://arxiv.org/abs/2609.21525
作者: Xuening Wu,Yanlan Kang,Shenqin Yin
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Preference-based reward shaping can guide reinforcement learning, but adding preference signals to the reward may unintentionally change the task being optimized. We address this problem with IncentRL, a framework that introduces preference guidance while explicitly characterizing its effect on external-task performance. IncentRL adds a Kullback–Leibler (KL) penalty between a specified outcome distribution and a preferred distribution. For finite discounted Markov decision processes with bounded shaping costs, we derive an external-value perturbation bound, establish a sufficient strict-action-gap condition for preserving the original optimal policy, and characterize the large-weight regime through discounted cumulative preference cost. Exact examples clarify the limits of these guarantees, including tied optima and support mismatch. We study a practical implementation using a hand-designed, distance-based outcome proxy, a fixed preference distribution, and score-weighted coefficient search. On MiniGrid DoorKey-8x8, the reported three-seed mean success rate after two million training steps reaches 98% with coefficient 0.01, compared with 90.5% for the reported zero-coefficient baseline, while the search progressively shifts toward smaller coefficients. Together, these results provide a principled view of the central trade-off in preference-based RL: using additional guidance to improve learning without excessively distorting the original task objective. The current experiments remain descriptive and do not yet isolate KL shaping from simpler alternatives.

[LG-36] What Must Survive? Exact Task-Information–State Frontiers for Resource-Sufficient Learning

链接: https://arxiv.org/abs/2609.21523
作者: Ronald Katende
类目: Machine Learning (cs.LG); Information Theory (cs.IT)
*备注: 9 pages, 0 figures

点击查看摘要

Abstract:A system may be compressed before its downstream task is fully known. We ask how much retained state is then necessary and how much can be saved by limited advance task information. For a finite family of linear tasks, a task message is revealed before state formation and the exact task only afterwards. For an advice alphabet of size K , the exact frontier is [ p^*(K)= \min_\substack\Pcal\text partition of \U\|\Pcal|\le K \max_C\in\Pcal\rank(T_C), ] with the b -bit frontier obtained by setting K=\min(2^b,|\U|) . Thus advance task information reduces state through partitions whose joint task operators have low rank. We also give an approximate singular-value frontier, a common-core lower bound and exact direct-sum law, and strong NP-hardness of finding an optimal advice partition. The hardness persists at every fixed positive approximation tolerance. Three examples illustrate the result. A well-conditioned softmax attention construction gives an exact 524,288\to1,024 coordinate frontier when nine bits resolve one of 512 continuations. A domain-decomposed digital twin yields an interface-plus-local-state law and a weighted partition problem for heterogeneous regions. A hierarchical multi-task model gives a two-stage frontier in which three bits reduce the required state from 3136 to 448 coordinates, with further task information approaching the irreducible 328 -coordinate single-task floor. Comments: 9 pages, 0 figures Subjects: Machine Learning (cs.LG); Information Theory (cs.IT) MSC classes: 68Q32, 68T07, 65Y20, 94A15 Cite as: arXiv:2609.21523 [cs.LG] (or arXiv:2609.21523v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2609.21523 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Ronald Katende [view email] [v1] Fri, 18 Sep 2026 09:12:54 UTC (9 KB) Full-text links: Access Paper: View a PDF of the paper titled What Must Survive? Exact Task-Information–State Frontiers for Resource-Sufficient Learning, by Ronald KatendeView PDFHTML (experimental)TeX Source view license Current browse context: cs.LG prev | next new | recent | 2026-09 Change to browse by: cs cs.IT math math.IT References Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading… BibTeX formatted citation loading… Data provided by: Bookmark checked="checked"class=“labs-tab-input”> Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?) IArxiv recommender toggle IArxiv Recommender (What is IArxiv?) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs. Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?) mathjaxToggle(); We gratefully acknowledge support from our major funders, member institutions, , and all contributors. About Help Contact Subscribe Copyright Privacy Accessibility Operational Status (opens in new tab) Major funding support from

[LG-37] ServeGuard: Verifiable Bounded-Residual Confinement of Operator-Invisible Channels Without Revealing the Certified Read Factor

链接: https://arxiv.org/abs/2609.21515
作者: Dominik Dahlem,Rui Vieira
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注: 30 pages, 2 figures, and 4 tables

点击查看摘要

Abstract:Third-party adapters for open-weight language models ship as opaque weight matrices; a recipient cannot check whether an adapter hides a backdoor without trusting the publisher or inspecting the weights, the publisher’s core asset. For one important class (payloads placed where a safety monitor is structurally blind), detection is unsound as a defense: every detector that factors through the declared monitor is invariant on its blind subspace, and honest and backdoored adapters overlap on every blind-subspace statistic we evaluate, because benign adaptation uses that subspace too. Rather than detect this channel, we make it structurally \emphabsent and prove that we did. The publisher builds the adapter to read the input only through directions the monitor covers and proves this in zero knowledge, revealing nothing about the read factor it certifies. The certificate is cheap because the expensive part, identifying the monitor’s blind spot, is a deterministic function of the \emphpublic base model, so only one linear identity is proved; the served residual is the base model’s own public floor, not a prover-chosen tolerance. The result is \emphServeGuard, a supply-chain primitive: the publisher ships a \emphproof-carrying adapter whose proof lets a consumer or regulator verify, without the certified read factor and without trusting the publisher, that the adapter carries no hidden channel of this class relative to the declared monitor; an admission-time typing guard binds the guarantee to the adapter bytes admitted at serving time. Across eight checkpoints up to 7B from four families, the monitoring budget is architectural: the measured frontier saturates at the value-path rank on grouped-query checkpoints but not on multi-head ones. On a 0.5B model confinement is nearly free for benign adaptation, making monitor quality the security lever.

[LG-38] Adaptive Rollout Truncation Based on Epistemic Uncertainty for Efficient Offline World Model Training IROS2026

链接: https://arxiv.org/abs/2609.21482
作者: Nikodem Sebastian Zymla,Laurin Thiele,Johannes Pitz
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注: 8 pages, 12 figures. Accepted at the IEEE/RSJ IROS 2026 Workshop “Rethinking Uncertainty for Modern Robotics Paradigms”

点击查看摘要

Abstract:Accurate neural world models are central to model-based robotics, where they enable robots to predict future states from previously observed trajectories. Multi-step autoregressive training improves long-horizon prediction, but fixed rollout horizons also increase computational cost and can amplify early training errors when the model is still inaccurate. Existing training schemes typically use the same rollout length throughout optimization, independent of the model’s current predictive reliability. We propose an epistemic uncertainty-driven adaptive rollout strategy for offline world model training following an auto-curriculum training scheme. Instead of always unrolling to a fixed horizon, the model terminates autoregressive rollouts once epistemic uncertainty exceeds a threshold calibrated from a warm-up phase. We study two uncertainty estimators: a five-head ensemble with a shared recurrent backbone and Monte Carlo Dropout. A two-stage warm-up procedure stabilizes uncertainty estimates before we enable adaptive truncation. Experiments on ANYmal-D and ANT show that ensemble-based adaptive truncation matches or improves the prediction accuracy of fixed-horizon training and the RWM-U baseline while requiring substantially fewer cumulative rollout steps. Training a world model on ANYmal-D following the presented approach reaches comparable final performance with the baselines with roughly 72% less rollout computation. These results indicate that epistemic uncertainty is useful not only for downstream policy regularization, but also for making world model training itself more compute-efficient. Comments: 8 pages, 12 figures. Accepted at the IEEE/RSJ IROS 2026 Workshop “Rethinking Uncertainty for Modern Robotics Paradigms” Subjects: Robotics (cs.RO); Machine Learning (cs.LG) Cite as: arXiv:2609.21482 [cs.RO] (or arXiv:2609.21482v1 [cs.RO] for this version) https://doi.org/10.48550/arXiv.2609.21482 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-39] Efficient Architecture Search under Leave-One-Subject-Out Evaluation

链接: https://arxiv.org/abs/2609.21457
作者: Heinke Hihn,Friedhelm Schwenker
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Deep neural architectures are widely used for signal processing in automated pain assessment systems. However, architecture design has remained largely a manual task despite the potential efficiency benefits of Neural Architecture Search (NAS). Embedding NAS in a Leave-One-Subject-Out (LOSO) evaluation is computationally demanding because a fully nested implementation requires N independent architecture searches and, assuming approximately linear training cost, scales as \mathcalO(N^2) . We propose a block-based, leakage-controlled approach that shares NAS runs between subjects, reducing the number of searches from N to B , where B \ll N , dubbed PainNAS. On the BioVid Heat Pain dataset, PainNAS yields comparable subject-level accuracy with substantially fewer parameters and FLOPs.

[LG-40] Understanding LLM Quantization through Activation-Guided Compensation and Orthogonal Residuals

链接: https://arxiv.org/abs/2609.21450
作者: Yamato Narita,Issei Sato
类目: Machine Learning (cs.LG)
*备注: 19 pages, 1 figure

点击查看摘要

Abstract:Post-training weight-activation quantization reduces the memory and inference costs of large language models, but aggressive W4A4 quantization remains difficult because activation outliers degrade effective quantization resolution. Although weight optimization, channel-wise scaling, and orthogonal rotation mitigate this problem, the error components they address and their relationship remain unclear. Using an exact decomposition of local weight-activation quantization error into an activation-guided weight compensation term and an orthogonal residual, we bound the residual using persistent channel-wise outlier and regular activation quantities. This decomposition clarifies which error components can be addressed by weight compensation and which require transformation design. We then use the residual bounds to derive practical guidelines for applying randomized Hadamard rotation, sign selection, and channel scaling. In particular, the analysis explains how random signs suppress constructive interference among persistent outlier channels, how sampling multiple sign patterns can improve transformation selection, and how second-moment balancing leads to an L_2 scaling rule while a further relaxation recovers SmoothQuant-style L_\infty scaling. We evaluate these guidelines through backpropagation-free configurations across eight Llama and Mistral models, obtaining performance competitive with gradient-trained SpinQuant.

[LG-41] FootQuery: Future-Touchdown-Guided Retrieval from Depth History for Perceptive Humanoid Locomotion

链接: https://arxiv.org/abs/2609.21447
作者: Tao Dong,Jia Yu,Yuxuan Fan,Linna Zhao,Jiaqi Gong,Andong Yang,Chao Gao,Guyue Zhou
类目: Robotics (cs.RO); Machine Learning (cs.LG); Systems and Control (eess.SY)
*备注: 9 pages, 11 figures

点击查看摘要

Abstract:Humanoid locomotion over complex terrain requires anticipating footholds that may no longer be visible at touchdown. Limited camera coverage and self-occlusion make it necessary to retrieve relevant terrain information from earlier observations. We present FootQuery, a perceptive locomotion framework that queries depth history using each foot’s predicted next touchdown. The policy predicts touchdown locations and uncertainty from proprioception and uses these distributions, together with per-foot features, to query sparsely sampled historical depth frames. During training, realized contacts are projected into historical images to supervise retrieval at the regions where those contacts were visible. The retrieved per-foot features are fused with global visual memory to generate control actions. A progressive force-assistance curriculum supports early exploration, while event-consistent tread-midline shaping encourages coordinated stair contacts. Deployment requires only proprioception and onboard depth images. In simulation, the complete framework outperforms its component ablations on the most challenging tested stairs, gaps, and platforms. Real-world experiments on a Unitree G1 demonstrate continuous traversal with a single policy across outdoor stairs and indoor routes combining stair ascent and descent, platforms, and gaps. These results support organizing visual history around anticipated contacts for perceptive humanoid locomotion.

[LG-42] Optimal Randomized Proper Online Learning

链接: https://arxiv.org/abs/2609.21445
作者: Zachary Chase,Idan Mehalel
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We prove that the optimal expected mistake bound of online learning a function class \mathcalH by a randomized proper learning algorithm is O(\mathttL(\mathcalH) \log T) , where \mathttL(\mathcalH) is the Littlestone dimension of \mathcalH and T is the time horizon. Our result improves upon the previously best known bound of O(\mathttL(\mathcalH) \log^6 T) given by Daskalakis and Golowich (STOC 2022), and is optimal up to a universal constant for worst-case classes.

[LG-43] Decision-Focused Learning for Mean-Variance Portfolio Optimization via KKT-Based Reformulation PRICAI2026

链接: https://arxiv.org/abs/2609.21427
作者: Kensei Nosaka,Shunnosuke Ikeda,Yuichi Takano
类目: Machine Learning (cs.LG)
*备注: 11 pages, 1 figure, 2 tables. Accepted at PRICAI 2026 (Pacific Rim International Conference on Artificial Intelligence)

点击查看摘要

Abstract:Mean-variance portfolio optimization (MVO) is a central framework in data-driven asset management. A widely adopted approach is a two-stage framework that first predicts expected returns and then solves the optimization problem based on these predictions, with the predictive models trained by minimizing prediction errors. However, this objective of prediction is not aligned with the quality of the downstream portfolio decision. Decision-focused learning (DFL), which directly minimizes the downstream decision loss within the learning process, has thus emerged as a promising direction. However, existing DFL approaches to MVO rely on surrogate losses or constraint relaxations for tractability, creating a structural mismatch between predictive model training and the constrained MVO solved at evaluation. We propose a single-level optimization formulation that incorporates the Karush-Kuhn-Tucker (KKT) optimality conditions of the lower-level MVO into the upper-level learning problem. This formulation explicitly preserves the budget and short-sale constraints while remaining tractable for standard nonlinear optimization solvers. Rolling-window experiments on real-world ETF (Exchange Traded Funds) data across two asset universes with different correlation structures show that our method achieved the best performance on multiple investment metrics and also demonstrated performance improvement due to the proposed regularization.

[LG-44] racing the Evidence Behind Zero-Shot Time-Series Forecasting: A Source-First Taxonomy and Audit Framework

链接: https://arxiv.org/abs/2609.21425
作者: Delun Kong,Wanyun Ling,Chenxi Liu,Ziyue Li
类目: Machine Learning (cs.LG)
*备注: 5 pages, 2 figures. Accepted to ACM AI Summit 2026 (Visionary Papers)

点击查看摘要

Abstract:Zero-shot time-series forecasting (TSF) is often described as forecasting without target-specific parameter updates, but that training-status condition does not specify what evidence the system may use. A frozen language model prompted with serialized values, a time-series model pretrained on broad forecasting corpora, and a retrieval-augmented forecaster may all satisfy the no-update condition while drawing on different transferable evidence. This paper argues that zero-shot TSF should therefore be governed as an evidence-access claim. We propose a source-first taxonomy that separates three primary evidence sources—frozen LLM prior reuse, parametric time-series pretraining, and retrieval-augmented external memory—from the architectures that implement them. After the source is identified, four additional audit questions remain: task interface, forecast object and scoring, prediction-time context, and resource budget. The resulting agenda is to make zero-shot leaderboards auditable by reporting evidence boundaries and interface assumptions alongside scores, so that benchmark progress reflects transferable forecasting capability rather than undisclosed changes in context, memory, or budget.

[LG-45] Probabilistic Forecasting of Business Process Executions with Neural Temporal Point Processes

链接: https://arxiv.org/abs/2609.21382
作者: Jiaxin Yuan,Daniela Grigori,Han van der Aa
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Operators of service-based systems act on forecasts of how a running execution will continue, and such a forecast is actionable only if its reliability is known. Mainstream deep-learning models for this task are discriminative and deterministic: they emit a single next activity and a single remaining-time estimate, without a distribution to reason over. We instead cast the problem as generative sequence modelling with marked temporal point processes, which define a joint density over the next mark and its inter-event time and therefore deliver predictive distributions by construction. Real event logs violate the simple-point-process assumption these models rest on, since consecutive events frequently carry identical timestamps; we handle such ties explicitly and combine a transformer encoder with a mixture decoder over inter-event times, trained by exact log-likelihood. On ten public logs, the resulting model matches discriminative baselines on point accuracy, dominates them on the calibration and sharpness of remaining-time distributions, and is the cheapest at inference, since a full predictive distribution is obtained in a single forward pass without sampling.

[LG-46] IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts

链接: https://arxiv.org/abs/2609.21346
作者: Ran Cheng,Longfei Xu,Zheng Liu,Kaikui Liu,Xiangxiang Chu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Mixture-of-Experts (MoE) scales capacity, but existing designs cannot set three quantities independently. For a single token, participation is how many experts contribute knowledge to its output, execution is how many are actually computed (compute cost), and materialization is how many expert-sized parameter sets must be built and stored (memory cost). Sparse routing keeps execution and materialization low, but shrinks participation: for each token, only a few experts contribute. Dense output-mixing restores full participation, but its execution grows with the number of experts. Parameter-merging keeps execution at one expert, but its materialization grows with the number of routing decisions. We propose IntBMoE, a block-conditioned MoE that decouples all three by pairing dense expert composition with sparse block execution. Its blocks come from a small learned codebook, one per entry. At each internal layer, a lightweight hypernetwork merges all expert bases in that layer’s pool into one composed expert. Participation is full, because every composed expert draws on the entire pool. Execution stays sparse, because a router sends each token to only a few blocks. Materialization is bounded, because the codebook, not the input, fixes how many blocks exist. Dual-Path Residual Gating (DPRG) further couples two independently composed paths through multiplicative gating. Experiments on image classification show consistent gains over representative sparse and dense MoE baselines. Additional experiments on language modeling and sequential recommendation validate its generalization beyond vision. IntBMoE is fully deployed in AMap’s generative recommendation system, serving hundreds of millions of users under a 60ms latency budget, with a 2.4% relative UVCTR gain in online A/B testing. Our code is available at this https URL.

[LG-47] Routine Blood Tests Outperform CRP for Distinguishing Bacterial From Viral Infection in Children

链接: https://arxiv.org/abs/2609.21332
作者: Mihaela Demireva,Zhecho Mitev,Djuna Chinareva-Klimentova,Svetoslav Ivanov,Georgi Nalbantov,Dimitar Mitev
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Acute infectious diseases are among the leading causes of medical consultations and hospitalizations in children worldwide. These infections are predominantly caused by viruses or bacteria, yet differentiating between the two remains a common clinical challenge. As a result, pediatricians often default to the safer option of prescribing antibiotics contributing to the growing problem of antimicrobial resistance. The objective is to assess the additional predictive value of CBC towards determining the current infection. This retrospective study used data from 906 pediatric patients aged between 2 and 14 years who were tested positive either for viral or bacterial infection between 2022 and 2026. Inclusion criteria further required availability of CBC results and CRP level measurements. These laboratory parameters as well as age were used as input features for several supervised classification models. Model performance was evaluated using AUC, sensitivity and specificity. The best performing model is XGBoost, which included all features, achieving out of-sample performance of AUC of 81.7% and sensitivity of 70.8%, specificity of 79.2%. All trained models outperform a CRP-based only decision-rule model in terms of AUC. We suggest that the decision to prescribe antibiotics should be based on a number of factors, including but not limited to CBC, some of which are not currently incorporated into routine practice.

[LG-48] An Introduction to Compression-Based Machine Learning

链接: https://arxiv.org/abs/2609.21309
作者: John Hurwitz,Edward Raff,Charles K. Nicholas
类目: Machine Learning (cs.LG)
*备注: To appear in The 13th IEEE International Conference on Data Science and Advanced Analytics (DSAA 2026)

点击查看摘要

Abstract:Any lossless compression algorithm (like gzip) may be converted into a machine learning method, via either Normalized Compression Distance or the Minimum Description Length principle. Any auto-regressive model may be converted into a lossless compression method via entropy coding. This seemingly circular dependence has unrealized potential in modern artificial intelligence and machine learning, and we survey and formalize the various strategies that have been used to leverage compression for machine learning. We introduce and empirically validate a design framework for compression-based ML, finding compression-based methods competitive with conventional baselines and decisively stronger on malware. We find that varying these design choices yields accuracy gains of up to 0.62.

[LG-49] Fast And Accurate Text Content File Type Identification

链接: https://arxiv.org/abs/2609.21306
作者: Manu Nandan,Michael Brautbar,Edward Raff
类目: Machine Learning (cs.LG)
*备注: To appear in The 13th IEEE International Conference on Data Science and Advanced Analytics (DSAA 2026)

点击查看摘要

Abstract:A common requirement across organizations is to have a tool that can identify file types based on their contents, particularly in the cybersecurity domain where magic numbers and file extensions can not be trusted. While existing tools work well in practice, there is plenty of room for improvement either in terms of computational load and time for detection in the case of model based tools like Magika or in terms of accuracy of detection in the case of file parsing tools that use programming language constructs. In this study, we propose a neural network model for identification of types of text content files, especially source code, that is more accurate and faster than other available tools. Our experiments on open-source files indicate that it is not only more accurate on average for text-content file-type identification, but also approximately four times faster than Magika, while being 28% smaller in size.

[LG-50] Identifying Security Platform Product Abuse with Machine Learning

链接: https://arxiv.org/abs/2609.21303
作者: Shaefer Drew,Michael Brautbar,Paul Knight,Edward Raff,Lana Peric-McDermott,Simran Sarin,Nickolas Machado,Hanna Albright,Vitaly Zaytsev
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注: To appear in The 13th IEEE International Conference on Data Science and Advanced Analytics (DSAA 2026)

点击查看摘要

Abstract:Product abuse is an individually rare, but growing, problem across the SaaS industry. Highly sophisticated threat actors can misuse security platforms within customer environments or conduct bypass experiments on the product itself. Threat actors can leverage living-off-the-land (LOTL) attacks to avoid using cumbersome, frequently detected malware. Remediating this threat requires collecting multiple data modalities across different types of databases, addressing a cold-start problem in the intrinsic rarity of such sophisticated but dangerous events, and designing within the constraints of real-world deployment (e.g., cost, user behavior, performance, etc). To wit, we provide the first study of such a whole-system defense, especially with respect to a deployed and operational capability. Our results show an increase in product abuse coverage by 35%, a 30% reduction in monthly alerts, and adaptability to changes in malicious actors’ behavior. We review both the constraints we considered in designing the system to meet operational requirements and a retrospective evaluation of the value of explainable features and counterfactual performance on previously identified attacks.

[LG-51] Multi-Subject Pretraining Enables Short-Calibration Personalization for Closed-Corpus Surface EMG Speech Decoding

链接: https://arxiv.org/abs/2609.21288
作者: Chenqian Le,Beatrice Fumagalli,Yasamin Esmaeili,Xupeng Chen,Tianyu He,Nikasadat Emami,Adeen Flinker,Yao Wang
类目: Machine Learning (cs.LG)
*备注: 20 pages

点击查看摘要

Abstract:Surface electromyography (sEMG)-based silent speech interfaces are limited by cross-user variability and calibration burden. We study a limited-data setting in which each of 27 speech-typical participants contributed less than 0.5 h of data (21.3 min on average) across Aloud and Mimed speech. Within a closed 50-sentence corpus, we used leave-one-subject-out evaluation, initializing from a released single-subject checkpoint, pretraining on non-held-out participants, and fine-tuning on the target participant. This pipeline achieved 21.7% character error rate (CER) and 31.9% word error rate (WER), compared with 49.3% CER without target-subject calibration and 68.0% CER for direct checkpoint fine-tuning. Multi-subject pretraining from random initialization followed by fine-tuning reached 44.9% CER and did not converge under the fixed schedule in 5 of 27 folds, indicating substantial optimization and accuracy benefits from checkpoint initialization. Macro-averaged CER declined from 74.4% with one pretraining participant to 21.7% with 26. Three minutes of target-subject calibration achieved 20.5% CER and 31.7% WER, with no statistically significant difference from the full approximately 13-min pool (21.7% CER and 31.9% WER). A subject-specific adapter provided no detectable benefit. Excluding the five evaluation sentences from all sEMG model-training data increased CER and WER to 78.6% and 99.9%. These results support short-calibration personalization in a standardized-montage, closed-corpus setting.

[LG-52] MIRCID: Inferred Hub-miRNAs Drive Cross-Task Improvements in Drug Mechanistic Modeling

链接: https://arxiv.org/abs/2609.21280
作者: Xin Cao,Yigang Chen,Jiatong Xu,Ziyue Zhang,Xiang Cheng,Shenyu Wang,Yangyi Zhang,Xiaoxuan Cai,Shidong Cui,Zihao Zhu,Xiang Ji,Hsi-Yuan Huang,Yang-Chi-Dung Lin,Hsien-Da Huang
类目: Machine Learning (cs.LG)
*备注: 25 pages, 6 figures, Advanced Science

点击查看摘要

Abstract:Drug mechanism-of-action (MoA) modeling commonly relies on perturbational transcriptomes, but matched microRNA (miRNA) measurements are often unavailable. Inferred regulatory features offer a scalable way to reuse these data. Here, we present MIRCID, a framework comparing gene expression with inferred transcription factor (TF) activity and miRNA expression across pathway classification and similarity-based MoA retrieval. HubmiRNet infers 414 pan-cancer hub miRNAs (HubmiRs) from 977 L1000 landmark genes, achieving a Pearson correlation coefficient of 87.72%; its 1,298-output variant also outperformed SiCmiR on the full-miRNA task (71.21% versus 67.30%). In the evaluated comparisons, miRNA augmentation provided more consistent gains than TF activity. Generic embedding controls showed model-dependent utility, while complementarity analyses identified a distinct, partially linearly recoverable representation that retained gene-derived structure. Illustrative rescue cases linked improved classification to biologically plausible miRNA patterns in samples with weak transcriptional signatures. These findings support inferred HubmiRs as a biologically informed recoding of transcriptomic data for perturbational drug modeling, while leaving recovery of measured perturbational miRNA responses to further validation.

[LG-53] Programming AMD XDNA NPUs with Open-source Compiler Tools: A FlashAttention Case Study

链接: https://arxiv.org/abs/2609.21264
作者: Erwei Wang,Ephrem Wu,Victor J. B. Jung,Jiajie Li,Andre Rosti,Joseph Melber,Samuel Bayliss
类目: Hardware Architecture (cs.AR); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Spatial NPUs such as AMD XDNA place compute tiles beside small local memories and leave data movement between them to software. Mapping a multi-stage workload onto such a device is largely a question of where the intermediate tensors live. We report what we learned making those choices for FlashAttention with the open-source IRON and MLIR-AIR flows. We compare four reference designs on XDNA 1 and XDNA 2: one runs each operator separately, two stream between operators on chip, and one fuses all three attention stages into a single kernel. The fused kernel holds the \boldsymbolQK^\mathsf T scores in compute-tile local memory and reduces partial results over the cascade interconnect, so the scores never return to shared MemTile memory. On XDNA 2, it reaches 3.62 TFLOP/s over complete end-to-end execution, twice the IRON design, with 5.3 to 7.2 times the energy efficiency of the integrated GPU on the same chip at 2K tokens and above. It covers twelve LLM configurations, from BERT to DeepSeek, up to 128K tokens. Roofline analysis at each memory level explains this result and shows when to stop. XDNA 1 has lower ridge points, so streaming on chip already reaches the compute-bound regime: the same fusion that doubles throughput on XDNA 2 is nearly wasted on XDNA 1. Comparing a mapping’s operational intensity against each level’s ridge point predicts which case applies before writing any code. Fuse until the mapping clears that ridge point, then stop. We release the reference designs as maintained open source. Subjects: Hardware Architecture (cs.AR); Machine Learning (cs.LG) Cite as: arXiv:2609.21264 [cs.AR] (or arXiv:2609.21264v1 [cs.AR] for this version) https://doi.org/10.48550/arXiv.2609.21264 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-54] Reliability-Centered Evaluation of Sparse Longitudinal CT Lesion-Size Forecasting with Conformal Interval Calibration and Gompertz-Inspired Regularization

链接: https://arxiv.org/abs/2609.21197
作者: Lingfei Kong
类目: Machine Learning (cs.LG); Quantitative Methods (q-bio.QM); Methodology (stat.ME)
*备注: 25 pages, 11 figures

点击查看摘要

Abstract:Sparse longitudinal CT follow-up limits lesion-size forecasting when only a few prior observations are available. We constructed a five-visit DLT-derived same-lesion trajectory benchmark from DeepLesion and Deep Lesion Tracker (DLT), yielding 205 trajectories from 129 patients. We compared an exploratory conventional sparse-to-final analysis with a primary fixed visit-index horizon design predicting the common log change from T3 to T4 while progressively adding earlier observations, evaluating predictive accuracy, uncertainty reliability, post-hoc conformal interval calibration, subgroup performance, and Gompertz-inspired trajectory regularization. The evaluated methods showed partially overlapping point-prediction accuracy but distinct uncertainty behavior. Mean held-out RMSE across ten training seeds was 0.4726, 0.4305, 0.4499, and 0.4513 for m = 1, 2, 3, 4, indicating the lowest mean RMSE at m = 2; additional history did not improve RMSE. At m = 4, raw Cohort-Level Feature GP coverage was near the 95% nominal level, whereas MC Dropout, Deep Ensemble, and residual-scale intervals were conservative. Patient-level conformal calibration generally produced near-nominal or conservative coverage at the cost of wider intervals. Patient-grouped development cross-validation selected lambda* = 0 for the Gompertz-inspired term. A global population reference frequently opposed lesion-level change directions, and prediction difficulty varied across anatomical subgroups. Overall, additional historical observations provided limited predictive benefit once the prediction horizon was controlled, while predictive accuracy, uncertainty reliability, and trajectory consistency did not necessarily improve together, and should be evaluated jointly in sparse longitudinal imaging.

[LG-55] rKV: Long-Context On-Device LLM s via Predictive Multi-Tier KV Caching

链接: https://arxiv.org/abs/2609.21172
作者: Zhihao Shu,Md Musfiqur Rahman Sanim,Jie Hu,Kun Yuan,Minghai Qin,Gagan Agrawal,Wei Niu
类目: Machine Learning (cs.LG); Systems and Control (eess.SY)
*备注:

点击查看摘要

Abstract:Large language models (LLMs) are moving onto mobile devices for increasingly diverse workloads over text, images, video, and audio. These applications often require long contexts, making the Key-Value (KV) cache a dominant memory bottleneck because it grows linearly with sequence length and is accessed at every decoding step. Prior work reduces KV-cache footprint through low-rank compression, token eviction, or flash offloading, but the resulting reconstruction overhead, irreversible token loss, or I/O stalls can offset the benefit of saving memory. We present TierKV, a mobile LLM inference framework built on Predictive Multi-Tier Cache Optimization (PMCO). Before decoding starts, PMCO predicts future cache demand from prefill hidden states and jointly assigns tokens to exact, low-rank, and flash-offloaded tiers under the device memory and accuracy budgets. This formulation retains access to the full context, removes the circular dependency of reactive eviction, and admits a closed-form solver that selects tier boundaries and per-layer ranks at runtime. Across eight text, vision, and audio models on three mobile SoCs, TierKV improves prefill throughput by up to 17.6x over existing mobile LLM frameworks, reduces RAM-resident KV cache by 12.5-34%, thereby enabling substantially longer contexts under the same memory budget, while incurring only minor accuracy degradation.

[LG-56] M2G-LLM : Enhancing Clinical Prediction via Multimodal Graph Reasoning and LLM Context Injection

链接: https://arxiv.org/abs/2609.21164
作者: Inyoung Choi,Sukwon Yun,Jiayi Xin,Jie Peng,Tianlong Chen,Qi Long
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Integrating diverse data modalities — such as clinical notes, laboratory results, and medical imaging — is essential for advancing clinical decision-making. While Large Language Models (LLMs) have shown remarkable performance in processing unstructured clinical text, their limited capacity to incorporate non-text modalities hinders their broader utility in healthcare applications. Here, we introduce M2G-LLM (Multimodal MedGraph-LLM), a novel framework that enhances LLMs with multimodal integration and alignment via Graph Neural Networks (GNNs). Our approach models temporal relationships between patient visits, propagates information across clinically similar patients, and aligns heterogeneous data sources to construct enriched multimodal context vectors. These vectors are injected into the intermediate layers of the LLM, enabling joint reasoning over textual and non-textual modalities. We evaluate M2G-LLM on the MIMIC-IV and MIMIC-CXR datasets, demonstrating improvements in clinical prediction tasks over strong baseline models. Our results highlight the promise of combining the language understanding of LLMs with the relational reasoning capabilities of GNNs for comprehensive, multimodal healthcare analysis.

[LG-57] HMB-GAN: Hybrid Multi-Bézier GAN for Vector Shape Synthesis

链接: https://arxiv.org/abs/2609.21158
作者: Elian Hugh Thiele-Evans,Binh Duong Pham,Hani Omar M Alharbi,Liibaan Aaden,Syed Umer Hasnain Zaidi,Prem Prakash Jayaraman,Muhammad Saeed,Boris Eisenbart
类目: Machine Learning (cs.LG)
*备注: 7 pages, 2 figures, 3 tables. Published in the 2026 IEEE International Conference on Quantum Software (QSW)

点击查看摘要

Abstract:We explore the use of hybrid quantum-classical generative adversarial networks for synthesising CAD-ready vector geometries. Unlike prior work that operates in rasterised or single-Bézier domains, we introduce HMB-GAN (Hybrid Multi-Bézier GAN), an end-to-end differentiable generative framework that constructs closed shapes through stitched multi-segment Bézier representations with geometric continuity enforced by construction. We compare a quantum-enhanced generator with a classical generator within this architecture and evaluate them across point cloud distribution metrics and geometric shape statistics. Results show that despite faster convergence, a reduction in model parameter count, and slightly improved performance on point cloud metrics, the quantum generator suffers from excessive simulator overhead and thus classically-simulated evaluation suffers from hardware constraints. These results demonstrate the feasibility of modelling structured geometries through hybrid quantum architectures whilst highlighting contemporary hardware limitations.

[LG-58] Diverse and Adaptable Arm Coordination for Octopus-Crawling via Diffusion-Based Uncertainty-Aware Optimization ICRA2027

链接: https://arxiv.org/abs/2609.21138
作者: Seung Hyun Kim,Heng-Sheng Chang,Kimia Kazemi,Prashant Mehta,Mattia Gazzola
类目: Robotics (cs.RO); Machine Learning (cs.LG); Optimization and Control (math.OC)
*备注: 8 pages, 5 figures, submitted to ICRA 2027

点击查看摘要

Abstract:Octopus crawling motivates soft robots that exploit redundancy, yet discovering and organizing diverse coordination modes for adaptation remains challenging. To address this, we introduce a Diffusion-based Uncertainty-aware Optimization (DUO) algorithm that learns demonstration-free crawling controllers for a simulated, muscle-actuated CyberOctopus. This work represents the first application of diffusion-based control to soft multi-arm robots in contact-rich simulations. By embedding a variety of locomotion behaviors within a shared control distribution, this approach enables the simulated octopus to navigate dynamic physical constraints, demonstrating that learned coordination diversity inherently facilitates robust adaptation. The main contributions include: (i) a symmetry-structured policy representation that folds radially equivalent controllers into a canonical directional sector, (ii) an online black-box optimization strategy, the DUO algorithm, that discovers and retains diverse coordination modes, and (iii) a control editing technique that adapts existing controllers to novel actuator constraints without retraining. These results show how learned coordination diversity makes motor abundance a practical resource for adaptation in soft multi-arm robots.

[LG-59] Layerwise Decoupling for Stable Structured Sparsification of Fully Connected Layers

链接: https://arxiv.org/abs/2609.21126
作者: Charles Kulick,Armenak Petrosyan,Sui Tang
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:We propose a decoupled, layerwise method for structurally sparsifying the fully connected layers of pretrained neural networks. Rather than penalizing all layers jointly, our approach extracts shallow two-layer subnetworks, normalizes the inner weights, and applies a structured group penalty to the outer weight matrix of each block, processing layers sequentially to prune neurons and reduce the width of each layer. We prove that the constrained decoupled objective is equivalent at optimality to a specific joint penalty on the inner and outer weights, for any positively homogeneous activation, and thus admits a clean projected and proximal formulation. Our central finding is that this decoupled reformulation is more robust than coupled methods. In numerical experiments it provides a wider usable range of the regularization strength and a lower rate of catastrophic over-pruning than the tested joint baseline while maintaining comparable accuracy. We establish these properties in controlled classification and sparse-recovery studies, and examine their scope in a high-dimensional PINN stress test and in the feed-forward layers of OPT-1.3B.

[LG-60] Signal-Centric Remote Sensing via Alternative Preprocessing and Acoustic Processing for ML-Driven Applications

链接: https://arxiv.org/abs/2609.21123
作者: Logan Luna,Sirio Jansen-Sánchez,Ilteris Demirkiran,Leo Ghelarducci
类目: Machine Learning (cs.LG); Robotics (cs.RO)
*备注: 7 pages, 9 figures, 3 tables. Published in Proc. IEEE SoutheastCon 2025, pp. 1078-1084, doi: https://doi.org/10.1109/SoutheastCon56624.2025.10971547

点击查看摘要

Abstract:The dominant method of processing sonar data is using image-based representations, requiring the preprocessing of image data on autonomous systems. We propose an alternative data processing method for remote sensing applications via the use of data in Comma-Seperated Value format. Experimentation on our alternative approach shows a reduction of processing time by 91.18%, an improvement in accurate object detection by Machine Learning, and an increase in SNR (Signal-to-noise ratio), PSNR (Peak signal-to-noise ratio), and other evaluation metrics.

[LG-61] alk to Me Jarvis: An Open-Source Edge-Deployable Voice Assistant Framework for Autonomous Racecars

链接: https://arxiv.org/abs/2609.21109
作者: Daniel Henel,Frederik Werner,Alexander Langmann,Johannes Betz
类目: Machine Learning (cs.LG); Robotics (cs.RO)
*备注:

点击查看摘要

Abstract:Recent advances in large language models have improved their effectiveness as back-end components for voice assistants, particularly in intent understanding and context-aware input classification. However, online-hosted models introduce network dependency and variable inference latency, limiting their suitability for time-critical autonomous driving applications. In this work, we address these issues by developing Jarvis, an offline voice assistant for high-level behavioral commands of autonomous vehicles. Its architecture integrates speech recognition and synthesis with natural language command classification into a lightweight, local framework. Jarvis core component is a text-to-command classifier, built using a domain-specific fine-tuning of the Mistral 7B model, demonstrating low-latency inference. Our experimental evaluation demonstrates that our solution outperforms larger online-hosted models, achieving 97.63 % intent recognition accuracy with an average processing latency of 1.39 s, making it well-suited for operations requiring quick response times. To support further research and fine-tuning, we provide an open-source implementation.

[LG-62] REFINEPPO: Learning Continuous Control Policies by Iterative Action Refinement

链接: https://arxiv.org/abs/2609.21108
作者: Sachini Weerasekara,Sagar Kamarthi,Jacqueline Isaacs
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Deep reinforcement learning (DRL) has achieved strong performance across a wide range of continuous-control problems. These continuous-control policies, however, are often defined as direct mappings from an observed state to an action or action distribution, requiring a single feed-forward network to construct an optimal control decision in one pass. While effective, this formulation leaves little opportunity for the policy to reconsider or progressively improve an action once an initial prediction has been formed. In this work, we explore an alternative approach: rather than learning only to directly predict an action, can a policy learn to iteratively improve one, and can this iterative process provide advantages during policy learning? We introduce Iterative Action Refinement (IAR), an iterative action-construction method that constructs control actions through a sequence of learned residual corrections. Starting from an initial proposal, a shared refinement network repeatedly conditions on the observed state and the current action proposal, allowing each refinement step to revise the action constructed by preceding steps. The final refined proposal is then used to determine the action executed by the agent. We integrate this iterative action-construction mechanism with Proximal Policy Optimization (PPO), yielding REFINEPPO. We evaluate REFINEPPO across 14 benchmark control tasks, complemented by controlled ablations of refinement depth and update schedules and analyses aimed at understanding why iterative refinement is effective. Across these environments, REFINEPPO matches or exceeds the performance of standard PPO while demonstrating faster convergence on several tasks.

[LG-63] oward individual-level calibration in affect recognition with perceptual adjustment queries

链接: https://arxiv.org/abs/2609.21073
作者: Xuanzhou Chen,Sankaraleengam Alagapan,Ashwin Pananjady
类目: Machine Learning (cs.LG)
*备注: 14 pages, 13 figures. Accepted as a poster at IEEE ACII 2026

点击查看摘要

Abstract:Behavioral tasks measuring facial affect perception assume that identical stimuli impose equivalent perceptual difficulty across participants. However, this assumption is systematically violated by individual differences in perceptual sensitivity. Using an affective perception task as our testbed, we propose a framework to normalize for perceptual difficulty that directly estimates each participant’s Just Noticeable Difference (JND) along the facial affect spectrum via cognitively lightweight perceptual adjustment queries (PAQs). We use these PAQ-inferred JNDs to re-express stimulus distances, constructing difficulty-equated tasks in perceptual space. We validate the framework in a Two-Alternative Forced-Choice (2AFC) task using two complementary behavioral measures: binary metacognitive difficulty judgments and response time variance decomposition. We find that PAQ calibration significantly equalizes perceived task difficulty at an individual level when compared to both the non-calibrated baseline and population-level Weibull calibration, while also reducing mean response time and between-subject variance in response time. These results establish PAQ as a principled and practical instrument for individualized perceptual calibration in facial affect recognition.

[LG-64] FedeRag e: Provably Convergent Agnostic Federated Learning under General Client Drift

链接: https://arxiv.org/abs/2609.21057
作者: Herlock Rahimi,Dionysis Kalogerias
类目: Machine Learning (cs.LG); Distributed, Parallel, and Cluster Computing (cs.DC); Signal Processing (eess.SP); Systems and Control (eess.SY)
*备注:

点击查看摘要

Abstract:Federated learning (FL) enables collaborative model training without sharing raw data, but its performance degrades under non-IID data and stochastic client participation. Remedies built on classical Federated Averaging (FedAvg) typically presuppose that client participation probabilities are known to the server, which is rarely the case in deployed systems. We first discuss and then characterize the optimization problem that \emphdistributionally agnostic FedAvg actually solves when participation is entirely unknown, possibly highly skewed, and of variable size across rounds: uniform aggregation is shown to minimize a well-defined stochastic objective, weighted by the participation-induced marginal, at a standard \mathcalO(1/\sqrtT) rate for convex and possibly nonsmooth losses. Building on this characterization, we propose \emphFederated Risk-Averse Averaging (\textscFedeRage), a risk-averse extension of FedAvg that embeds the \emphConditional Value-at-Risk (CVaR) into the local objective within a natural distributionally robust optimization (DRO) framework. \textscFedeRage implicitly upweights high-loss and infrequently participating clients while adding only a \emphsingle scalar per-client, and admits an \mathcalO(\kappa/\sqrtT) rate in which the factor \kappa is the upper bound on the ``price" of risk aversion. In contrast with aggregation-alignment schemes based on optimal transport, which require the availability distribution as an input, \textscFedeRage remains agnostic to it. Several experiments on three heterogeneous benchmarks indicate consistent improvements over state-of-the-art methods in accuracy, fairness, and convergence speed.

[LG-65] A Lightweight Plug-in Gate for Transformer-Based Time-Series Forecasters

链接: https://arxiv.org/abs/2609.21044
作者: Hongkai Zhuang,Tao Huang,Chen Hou
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Covariate-rich time-series forecasting requires deciding how external variables enter the target forecasting path. Existing Transformer-based forecasters usually build a covariate representation and pass it to the encoder without an explicit admission stage. This paper studies pre-encoder covariate admission as an input-side interface that regulates that representation immediately before encoder processing. We implement the interface with a lightweight representation-level pre-encoder gate that assigns sigmoid scores to representation units, and we also study a usage-regularized variant that penalizes average admission. The interface is evaluated as a plug-in module for TimeXer, Inverted Transformer (iTransformer), and Patch Time Series Transformer (PatchTST) under a zero-extra-tuning protocol, where each gated model inherits the corresponding baseline configuration. Experiments on the Electricity Transformer Temperature minute-level (ETTm1 and ETTm2) datasets, Traffic, Energy, and influenza-like illness (ILI) include paired forecasting comparisons, gate-placement ablation, initialization ablation, controlled covariate-admission analysis, and a variance inflation factor (VIF)-informed permutation feature importance (PFI) diagnostic case study. In the tested settings, the gate is competitive with the corresponding baselines, and the usage penalty reduces average admission scores while keeping forecasting errors close to the unpenalized TimeXer setting.

[LG-66] Stiefel-AdamW: Geometry-Aware AdamW for Linear Factorization Blocks

链接: https://arxiv.org/abs/2609.21039
作者: Emanuele Zangrando,Marco Sutti,Francesco Tudisco
类目: Machine Learning (cs.LG); Numerical Analysis (math.NA)
*备注:

点击查看摘要

Abstract:A pervasive structural pattern in modern deep learning is the linear factorization block: a submodule of the form W = BA in which two parameter matrices are multiplied directly, with no intervening nonlinearity. Such blocks appear in LoRA adapters, low-rank compressed layers, query-key products of self-attention, and share a common pathology: the factorization is non-unique, which can destabilize training and limit usable learning rates. Despite this, factorization blocks are typically optimized with standard Euclidean methods that ignore the underlying geometry. We introduce Stiefel-AdamW, a near drop-in replacement for AdamW for use wherever such blocks appear. By constraining one factor on the Stiefel manifold while leaving the other Euclidean, Stiefel-AdamW relaxes the full \mathrmGL(\mathbbR^r) gauge symmetry to a compact orthogonal symmetry, ruling out factor blow-up while retaining the coordinate-wise diagonal preconditioning that gives AdamW its practical strength. Moment estimation is performed in the ambient Euclidean space, with geometry entering only through a tangent-space projection and a manifold retraction. The implementation overhead over AdamW is minimal, and we show that the resulting optimizer inherits both the stability benefits of Riemannian methods and standard convergence guarantees. We validate Stiefel-AdamW on LoRA-style fine-tuning of GPT2, ViT, and Mistral 7B and on full pretraining of GPT2 on OpenWebText, showing consistent improvements over strong baselines at essentially no additional cost over AdamW.

[LG-67] On the Limits of Maximal Coding Rate Reduction for Out-of-Distribution Generalisation

链接: https://arxiv.org/abs/2609.21001
作者: Menghui Zhou,Gaoshan Bi,Vitaveska Lanfranchi,Po Yang
类目: Machine Learning (cs.LG)
*备注: 31 pages, 8 figures, including appendices

点击查看摘要

Abstract:Substantial efforts have been devoted to making deep learning objectives, representations, and architectures interpretable, with the goal of improving the safety, robustness, and generalisation of learning systems in diverse real-world applications. The recently proposed maximal coding rate reduction ( \mathrmMCR^2 ) offers a promising information-theoretic framework for learning structured, discriminative representations of class-wise submanifolds and has inspired interpretable white-box architectures. However, we observe that \mathrmMCR^2 can completely fail under distribution shift, motivating our study of its out-of-distribution (OOD) generalisation limits. We establish two limitations of \mathrmMCR^2 for OOD generalisation. First, the \mathrmMCR^2 objective alone can admit complete prediction failure: a representation based entirely on unstable environmental features can achieve the global coding optimum yet fail completely after correlation reversal, despite an available perfectly stable feature. This exact-optimum example includes test inputs that cannot occur during training. Even when every possible test input can also occur during training, coding quality can be arbitrarily close to optimal while prediction error is arbitrarily close to 100%. Second, directly incorporating the invariance principle underlying widely successful invariant risk minimisation (IRM) and risk extrapolation (REx) does not eliminate this failure. The failing representation admits the same optimal coding operator across training environments, showing that shared coding optimality does not ensure stable prediction. Reliable OOD guarantees for \mathrmMCR^2 therefore require additional new assumptions or learning principles that establish stable predictive relationships across environments.

[LG-68] MOSAIC-SR: Transformer-Guided Symbolic Regression for Scientific Equation Recovery

链接: https://arxiv.org/abs/2609.20997
作者: Peiyi Zheng,Yanming Kang,Hans De Sterck,Giang Tran
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Symbolic regression aims to recover closed-form equations from observations, providing interpretable models for scientific discovery. Existing approaches struggle to combine flexible structural search with efficient inference. Search-based methods can refine expression structure but often rely on costly combinatorial optimization with random initialization. Pretrained neural models generate formulas almost instantly, but their predictions often contain symbolic errors. We introduce MOSAIC-SR, which uses a pretrained Transformer to propose multiple initial sketches. These sketches initialize searches in several promising regions, avoiding random starts in the vast expression space. Each search jointly recovers structure and constants through scale-aware constant optimization and local symbolic repair. We evaluate MOSAIC-SR on the SRSD-Feynman dataset with and without dummy variables and on six additional benchmarks. MOSAIC-SR obtains the highest symbolic solution rate on every dataset while ranking among the top two methods in predictive accuracy. This advantage persists in the presence of irrelevant dummy inputs. The results show that learned priors can focus search on promising equation structures, and that numerical optimization and symbolic repair are important for recovery.

[LG-69] From Stress to Affect: Multimodal Deep Learning for Physiological Emotion Recognition Across Wearable Sensor Modalities

链接: https://arxiv.org/abs/2609.20991
作者: Desta Haileselassie Hagos,Saurav Keshari Aryal,Legand L. Burge
类目: Machine Learning (cs.LG)
*备注: Under review

点击查看摘要

Abstract:Physiological emotion recognition using wearable sensors has important applications in mental health monitoring, affective computing, and human-computer interaction. However, existing studies typically evaluate a single model, sensing configuration, or dataset, limiting our understanding of how these factors influence recognition performance. We present a comparative study of temporal deep learning architectures for physiological emotion recognition using two multimodal wearable datasets: WESAD and EmoWear. Bidirectional long short-term memory (LSTM), temporal convolutional network (TCN), and Transformer models are evaluated under wrist-only, chest-only, and multimodal sensing configurations using participant-independent leave-one-subject-out cross-validation (LOSO-CV). We also investigate soft-voting ensembles, sensor ablation, sampling frequency, and gradient-based saliency. The Transformer achieved the highest multimodal accuracy on WESAD (99.02% +/- 0.51%), whereas the LSTM achieved the best multimodal accuracy on EmoWear for both arousal (91.80% +/- 1.06%) and valence (89.96% +/- 0.36%). These results show that relative architecture performance depends on dataset characteristics rather than one architecture being uniformly superior. Multimodal sensing consistently outperformed wrist-only and chest-only configurations across both datasets. Sampling-frequency analysis showed that 4 Hz provides a practical operating point, with performance comparable to higher frequencies at substantially lower training cost. These findings provide guidance for selecting architectures, sensing modalities, and sampling frequencies for wearable physiological emotion recognition.

[LG-70] ASGARD: Action-Space Guard for UAV Resilience via Reinforcement Learning

链接: https://arxiv.org/abs/2609.20982
作者: Mohsen Salehi,Karthik Pattabiraman
类目: Machine Learning (cs.LG); Cryptography and Security (cs.CR); Robotics (cs.RO)
*备注:

点击查看摘要

Abstract:Reinforcement learning (RL) controllers have been recently adopted for Unmanned Aerial Vehicles (UAV) navigation and control. However, they are susceptible to action-space attacks that overwrite the action commands after the policy generates them and before the actuators execute them. While most existing defenses target attacks on the policy’s inputs, those addressing action-space attacks retrain the policy at training time and are not resilient to corrupted actions at runtime. We propose ASGARD, a two-phase teacher-student pipeline for making RL-based UAV control resilient to action-space attacks. In the teacher phase, an encoder combines the UAV’s physical state with action-attack-related privileged information to produce an action-attack-aware latent that trains the RL control policy and a monitor that outputs corrected action commands to the actuators. In the student phase, both the encoder and the monitor are trained via supervised learning from their teacher counterparts to run on-board using only the UAV’s physical state history. We evaluate ASGARD across attack scenarios targeting different action commands on UAV. We find that ASGARD is resilient to action-space attacks and completes the missions despite the attack. We further find that ASGARD generalizes to unseen attacks and remains resilient against stealthy attacks.

[LG-71] Generative inversion for early ranking of competing geologic interpretations

链接: https://arxiv.org/abs/2609.20978
作者: Harun Ur Rashid,Daniel O’Malley
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:High-consequence subsurface decisions are often made under severe data scarcity. Experts may arrive at competing interpretations of the same subsurface system, yet early in a project there is rarely a practical way to determine which one is most realistic. This uncertainty can persist until several wells are drilled, often costing millions of dollars. Existing approaches for evaluating geologic interpretations rely either on subjective judgment or on dense data that are rarely available in early-stage investigations. We present a workflow that addresses this challenge by translating competing geologic interpretations into alternative spatial priors and ranking them according to their consistency with hydraulic-head observations. For each interpretation, a text-to-image foundation model generates an ensemble of 1600 geologic images, and a separately trained variational autoencoder provides an interpretation-specific latent representation. A supervised inverse network maps the head observations into this latent space, and the frozen decoder produces an image that is mapped to a log-conductivity field. Steady-state flow simulation then provides predicted heads, and the resulting mismatch is converted into a Gaussian-form compatibility score. We evaluate the framework using a synthetic benchmark based on the Johansen Formation and three interpretations of decreasing consistency with the reference representation. Across 925 test cases, the mean head RMSE increases from 0.197 for the Precise \ Accurate interpretation to 0.227 for the Accurate interpretation and 0.280 for the Mismatched interpretation. We subsequently apply the workflow to two published conceptual models of the Culebra Dolomite Member at the Waste Isolation Pilot Plant. The revised model receives a compatibility weight of 0.991, compared with 0.009 for the original model, consistent with the independent evidence.

[LG-72] From Switching to Dynamic Regret: A Simple Reduction via Unbiased Random Sequences

链接: https://arxiv.org/abs/2609.20968
作者: Yibo Wang,Wenhao Yang,Sifan Yang,Yuanyu Wan,Lijun Zhang
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:In non-stationary online learning, dynamic regret has attracted increasing attention as a measure of how well an online learner performs against a time-varying comparator sequence. Despite considerable advances, attaining optimal bounds for strongly convex and exp-concave losses often involves intricate analysis. In this paper, we present a \textitsimple framework that reduces dynamic regret minimization to switching regret minimization. As a result, we can derive dynamic regret bounds by using off-the-shelf algorithms with switching regret guarantees. The key idea of our reduction is to construct, for \textitany comparator sequence, an auxiliary random sequence that is unbiased at each round, with the controlled variance and a manageable number of switches. Combining this construction with suitable surrogate losses, we can decompose dynamic regret into the expected switching regret against the random sequence and its controlled variance. Theoretically, for strongly convex and exp-concave losses, we establish the \widetildeO(T^1/3P_T^2/3) dynamic regret bounds, where T denotes the time horizon and P_T denotes the path-length of the comparator sequence. Moreover, for general convex losses, the same reduction also recovers the O(\sqrtT(1+P_T)) dynamic regret bound. Notably, all our findings match the minimax optimal results for these three types of losses, highlighting the versatility of our proposed framework.

[LG-73] Efficient Bayes-Adaptive Reinforcement Learning with Temporal Logic Specifications

链接: https://arxiv.org/abs/2609.20954
作者: Jonathan Hau,Alessandro Abate
类目: Machine Learning (cs.LG)
*备注: ©~2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works

点击查看摘要

Abstract:We present a novel end-to-end model-based Reinforcement Learning (RL) algorithm for efficient policy synthesis under given Linear Temporal Logic (LTL) specifications (e.g., safety or reachability) in unknown environments. To do so, a Limit-Deterministic Büchi Automaton (LDBA) representation of the LTL task is synchronised with a Bayes-Adaptive Markov Decision Process (BAMDP) representation of the environment, which allows us to leverage an enhanced exploration-exploitation trade-off that is achieved via Bayesian RL, as opposed to traditional non-Bayesian approaches. We further propose a novel Bayes-Adaptive Monte-Carlo Planning (BAMCP) algorithm to allow for approximate Bayes-optimal strategy synthesis in the synchronised BAMDP construct. A range of finite- and infinite-horizon task experiments demonstrate the effectiveness of our approach in terms of both property satisfaction and sample efficiency, when compared to traditional model-free approaches. Additional ablation studies also successfully highlight the value of the novel BAMCP algorithm in comparison to classical BAMCP for LTL task satisfaction. Finally, we also showcase a successful application of our approach for \textitcautious RL, namely to reduce the number of task violations incurred during policy training.

[LG-74] When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation

链接: https://arxiv.org/abs/2609.20942
作者: Sy-Tuyen Ho,Minghui Liu,Furong Huang
类目: Machine Learning (cs.LG)
*备注: Under Review

点击查看摘要

Abstract:Large language models (LLMs) increasingly participate in scientific evaluation, both as automated reviewers and as assistants to human reviewers. As model-generated reviews enter public data and future training corpora, AI peer review can become recursive: later reviewers learn from judgments produced by earlier models. We study one step of this feedback loop in a controlled setting. Starting from Llama 3.1 8B, we first fine-tune a reviewer on official ICLR reviews from 2018–2023 and then train four successor models on ICLR 2024 data with systematically varied mixtures of official and model-generated reviews. Our study shows that introducing synthetic reviews compresses rating distributions and reduces both same-paper and corpus-level semantic diversity. We call this pattern \textbfscientific-judgment collapse . To mitigate this failure mode, we introduce \textbfTrustReviewer , an open-source LLM-based system for generating peer reviews of AI and machine learning papers. TrustReviewer intervenes at two complementary stages. For training-time prevention, we train the core reviewer in a single stage on a curated corpus designed to reduce low-quality and semantically degenerate supervision. For test-time correction, paired activation steering aims to further mitigate residual tendencies toward collapsed judgments without further training or additional expert annotation. Together, these results characterize a concrete risk of recursive reviewer training and provide practical interventions for preserving judgment diversity and improving recommendation alignment in AI-assisted scientific evaluation. Comments: Under Review Subjects: Machine Learning (cs.LG) Cite as: arXiv:2609.20942 [cs.LG] (or arXiv:2609.20942v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2609.20942 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-75] Do Quantum Models Scale Like LLM s?

链接: https://arxiv.org/abs/2609.20912
作者: David S. Berman,Ying-Jer Kao,Roger G. Melko,Alexander G. Stapleton
类目: Machine Learning (cs.LG); Quantum Physics (quant-ph)
*备注: 10 pages, 6 figures

点击查看摘要

Abstract:In this work, we study the neural scaling laws of RydbergGPT, an autoregressive transformer model trained on qubit projective measurement data gathered from interacting Rydberg atom arrays. The quantum system is known to exhibit a finite-size remnant of a critical point as the laser detuning parameter is varied. We find that near the critical point the transformer loss as a function of training dataset size is well described by a power-law with a loss floor correction. However, away from criticality the quality of the power-law description is substantially reduced. We then compare the statistical structure of both Rydberg measurements and natural-language corpora using an entropy-normalised, finite sample corrected mutual information “two-point” function. We find that near-critical statistics of the two point functions are closest to those observed in natural-language, whilst other qubit configurations far from the critical point have two-point functions that decay more rapidly. This supports the hypothesis that multi-scale dependence contributes to stable neural scaling, and that scaling behaviour should be viewed as a property of the model-data pair.

[LG-76] Continuous Delayed-Memory Stochastic Gradient Descent and Continuous-Time Reinforcement Learning from History of Astrophysical Time Series Studies STOC

链接: https://arxiv.org/abs/2609.20906
作者: Debartha Paul,Juncheng Yi
类目: Machine Learning (cs.LG); Applications (stat.AP)
*备注: Keywords: Stochastic process, Stochastic gradient descent, Continuous-Delayed-Memory Stochastic Gradient Descent, Stochastic Delay Differential Equation, Reinforcement Learning, Adjoint method

点击查看摘要

Abstract:Quasars are luminous objects in the universe that exhibit stochastic brightness variations encoding information about the supermassive black holes powering them, and modeling these variations from ground-based survey data time series, known as light curves, is a statistical challenge. This paper reviews how stochastic differential equations (SDEs) have been adapted with neural network parameterizations to overcome this challenge in history. We create the Continuous-Delayed-Memory Stochastic Gradient Descent which depend on the past state of the discrete iteration process. We performed the simulation on some 2-dimensional landscape and observed some wider-exploration and more precise convergent behavior compared to Vanilla SGD by adjusting hyperparameters. Besides, we proposed a reinforcement learning structure with continuous time policy gradients for exploratory policies without solving HJB PDE, and we show that its optimality conditions recover the Gibbs policy of previous works.

[LG-77] Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding

链接: https://arxiv.org/abs/2609.20888
作者: Themistoklis Haris,Henry Li,Maryam Karimzadehgan
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Massive KV caches can cause severe memory-bandwidth bottlenecks during long-context decoding. Sparse attention methods mitigate this via selective loading, but that comes at a cost: rigid heuristics drop necessary context, leading to quality degradation. We introduce \textbfElastic Threshold Attention (ETA), an end-to-end trainable architecture that achieves hardware-accelerated decoding speed without sacrificing dense model quality. ETA predicts dynamic, contextual thresholds directly from query representations, allowing the model to allocate dense-like context to difficult retrieval or reasoning steps while pruning routine tokens. To learn this policy from scratch without representation collapse, ETA \emphmultiplicatively suppresses sub-threshold logits toward zero during training rather than deleting them. Training against this smooth uniform attention floor provides a distributed probability reservoir that \textbfcauses localized attention sinks on initial tokens to disappear. It also enables the model to hard-prune uninformative KV blocks at inference time and absorb incidental tokens co-admitted by coarse GPU block selection. As a result, a 1.45B pretrained ETA model rivals dense attention across language modeling, commonsense reasoning, and long-context needle retrieval at \approx 85% training sparsity and \approx 38% active decode density. At inference time, we implement a custom decode kernel in Triton that screens KV blocks in O(1) time using cached geometric-probabilistic bounds, delivering up to 2.5\times wall-clock decode speedups over FlashAttention-2 on sequences up to 512K tokens. Finally, we introduce an offline calibration algorithm for domain-specific deployments that freezes per-head constant thresholds to eliminate predictor overhead, cutting attention compute by an additional 27% .

[LG-78] Sparse Priors for Efficient Distribution Learning

链接: https://arxiv.org/abs/2609.20883
作者: Saumya Goyal,Barnabás Póczos
类目: Machine Learning (cs.LG); Statistics Theory (math.ST)
*备注:

点击查看摘要

Abstract:Despite the widespread use and success of generative AI techniques today, theoretical guarantees on learning a distribution supported in d dimensions from n samples degrade as O(n^-1/\Theta(d)) , though shown to be minimax optimal. We hypothesize that present bounds are too pessimistic because smoothness assumptions are not enough to capture the structure of distributions that often appear in real applications. Consequently, we introduce the class of sparse priors and define the “Sparse Dimension” as a measure of sparsity of a prior over the space of all distributions. We show that distribution learning under a k -sparse prior achieves a Bayesian risk lower bound of \Omega(\sqrtk/n) under common distance metrics, and show a matching (up to logarithmic terms asymptotically in n,k ) upper bound for the TV distance under mild additional assumptions. We show the statistical equivalence of distribution learning and learning to sample in the Bayesian setting so that our results apply to learning to sample as well. While k can still depend on the dimension d , or a notion of intrinsic dimension, our results show that learning under an appropriate prior overcomes the curse of dimensionality with respect to the dependence on n .

[LG-79] he Refutation Gap: Certifying Both Halves of an Optimality Claim

链接: https://arxiv.org/abs/2609.20873
作者: Rohan Pandey
类目: Logic in Computer Science (cs.LO); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Synthesis pipelines increasingly claim not just that a program is correct, but that it is optimal. Such a claim has two halves with radically different verification stories. The upper bound, “a program of size m exists”, is witnessed by an artifact that can be re-executed, proved equivalent to its specification, and shipped with a machine-checked certificate. The lower bound, “no program of size m-1 exists”, has no witness and is discharged by running a solver until it reports UNSAT. Combinatorial optimization has known this asymmetry for decades and has largely addressed it: certifying algorithms make it explicit (McConnell et al., 2011), and pseudo-Boolean proof logging can certify optimality end to end with a formally verified checker (Bogaerts et al., 2023; Koops et al., 2025). That discipline has not reached circuit minimization. We call this the refutation gap: published gate counts for minimal XOR circuits provide no certificate for either half of the claim, and neither did 121 optimality results we ourselves produced. We close the gap with a pipeline that synthesizes minimal linear straight-line programs over GF(2), where every decisive UNSAT answer emits a DRAT proof checked by an independent third-party checker. We certify all 121 optimality results established by the project, across n = 6 to 9: 111 carry independently verified refutations, 10 are closed by a free counting bound, and none disagrees with the uncertified value. The median proof is 1.1 MB and checking costs 1.9x solving. We give five case studies where verification caught defects that code review did not, report two interface obstacles that push practitioners toward the uncertified path, and describe an adversarial audit that revealed a failure tail we were about to attribute to the problem was actually caused by our own budget. Subjects: Logic in Computer Science (cs.LO); Machine Learning (cs.LG) Cite as: arXiv:2609.20873 [cs.LO] (or arXiv:2609.20873v1 [cs.LO] for this version) https://doi.org/10.48550/arXiv.2609.20873 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-80] Schedule optimization for tau-leaping in masked discrete diffusion

链接: https://arxiv.org/abs/2609.21960
作者: Cecilia Secchi,Giacomo Zanella
类目: atistics Theory (math.ST); Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Masked discrete diffusion models are commonly accelerated using the so-called tau-leaping discretization method, which reveals several coordinates in parallel at each sampling step. The sampler replaces the joint conditional law of each revealed block by a product distribution, incurring a factorization error \varepsilon_\textfact present even with perfectly learned predictors. We analyze the standard sampler on N coordinates with K sampling steps, whose random block sizes depend on a denoising schedule. Our analysis uses an exact integral representation of \varepsilon_\textfact in terms of a distribution-dependent dependence density \rho , which records how conditional dependence evolves as the revealed fraction of coordinates grows. We develop estimators for this profile and quantify how estimation errors affect schedule selection. We derive recursive stationarity equations for the finite- K optimization problem and, under a monotonicity condition, characterize its unique optimizer. In the joint limit N,K\to\infty , we obtain an explicit characterization of the optimal limiting smooth schedule and quantify the cost of random block sizes relative to a deterministic planner. When \rho_N converges uniformly to a strictly positive continuous profile, optimizing over fixed smooth schedules can improve the leading constant but not the N/K scaling of \varepsilon_\textfact . By contrast, if \rho_N degenerates, suitable schedules can improve the asymptotic order relative to the uniform schedule. Examples based on stationary processes and exchangeable mixtures illustrate these two regimes.

[LG-81] Riemannian Simultaneous Inference for Tangent Vector Field Regression

链接: https://arxiv.org/abs/2609.21910
作者: Xiaotian Chang,Yangdi Jiang,Qirui Hu
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Methodology (stat.ME)
*备注:

点击查看摘要

Abstract:We consider nonparametric tangent vector field regression on a Riemannian manifold without boundary. Because responses at different points lie in different tangent spaces, the proposed kernel estimator first parallel transports nearby responses to the target tangent space and then forms a volume-corrected local average. We first derive its uniform second-order bias, finite-bandwidth covariance, and stochastic rate. For simultaneous inference, the tangent norm is written as a supremum over the unit tangent bundle. Exact covariance whitening gives a unit-variance Gaussian field whose correlation length is of order h along the base manifold and of order one along the fibre. Its local covariance geometry leads to a Gumbel limit with an explicit intrinsic constant. Combining this limit with Gaussian approximation and cross-fitted covariance estimation yields a feasible simultaneous confidence tube for the regression field. We further discuss improved finite-sample inference with bandwidth selection and high-order bias corrections. Simulations on various manifolds support the proposed inference procedure. A randomized reconstruction of global wind data illustrates how the tube’s cross-sections describe spatially varying uncertainty.

[LG-82] Near-Optimal Acceleration for Smooth ell_p / ell_q Nondual Convex First-Order Oracle Optimization

链接: https://arxiv.org/abs/2609.21880
作者: David Martínez-Rubio,Brian Bullins,Cristóbal Guzmán,Mathieu Molina
类目: Optimization and Control (math.OC); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We study the optimization of convex objectives with (L,\kappa-1) -Hölder-continuous gradients in \ell_q over R B_p^d , 1\kappa\le 2 . (MG26) provides selectors with a movement bound for the problem of chasing high-dimensional convex nested sets for every pq and generally reduces Lipschitz convex optimization to bounds on the movement of selectors. We couple that movement with Hölder descent yielding a polynomial-runtime first-order method whose feasible output, in the high-dimensional regime T\le d and for p\min\q,2\ , has error \widetilde O_\kappa,p,q!\left( \fracLR^\kappaT^\kappa(1+1/p-(1/q-1/2)_+)-1 \right), after T queries to a first-order oracle, solving the COLT 2015 open problem of (Guz15), up to logarithmic factors. At (p,q)=(1,2) , the rate is \widetildeO(LR^\kappa/T^2\kappa-1) , including \widetildeO(LR^2/T^3) cubic decay in the smooth case. Subjects: Optimization and Control (math.OC); Machine Learning (cs.LG) Cite as: arXiv:2609.21880 [math.OC] (or arXiv:2609.21880v1 [math.OC] for this version) https://doi.org/10.48550/arXiv.2609.21880 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: David Martínez-Rubio [view email] [v1] Fri, 18 Sep 2026 15:07:30 UTC (43 KB)

[LG-83] Complete Neural Electronic Initialization Accelerates Materials DFT

链接: https://arxiv.org/abs/2609.21759
作者: Felix Ærtebjerg,Jonas Elsborg,Arghya Bhowmik
类目: Materials Science (cond-mat.mtrl-sci); Machine Learning (cs.LG); Computational Physics (physics.comp-ph)
*备注: 34 pages, 4 figures, 15 tables

点击查看摘要

Abstract:We present the first complete machine learning method for accelerating plane-wave density functional theory (DFT) in materials under the projector augmented wave (PAW) formalism. We formalize seven criteria that a \textitComplete Neural Electronic Initializer must satisfy for practical end-to-end PAW DFT acceleration. Applying these criteria to prior work reveals two missing structure-dependent components, augmentation occupancies and spin initialization, that prevent existing methods from providing complete reference-free initialization. Controlled ablations show that omitting these components can eliminate or reverse the acceleration obtained via models that only predict the smooth valence density. We satisfy these missing requirements by introducing AugNet, the first general equivariant model for PAW augmentation occupancies, and the first general spin density model for materials, which predicts the smooth spin-difference density and spin-difference PAW augmentation occupancies using predicted magnetic moments to constrain the global magnetic state. Combined with existing valence density models, these components satisfy all seven criteria and form a fully reference-free electronic initializer for materials DFT, requiring no electronic quantities from a converged target calculation. Our method reduces end-to-end DFT wall time by up to ~25% on unseen structures while preserving converged energies.

[LG-84] Single-Loop Stochastic Projected Damped Extrag radient Methods for Stochastic Nonconvex–(Strongly) Concave Minimax Optimization

链接: https://arxiv.org/abs/2609.21747
作者: Huiling Zhang,Minhao Zhang,Zi Xu
类目: Optimization and Control (math.OC); Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:We develop single-loop stochastic projected damped extragradient methods for stochastic nonconvex–(strongly) concave minimax optimization, with complexity guarantees for both game stationarity (GS) and optimization stationarity (OS). Our approach combines a stochastic projected damped extragradient (SPDE) method with a recursive variance-reduced variant, VR-SPDE, both of which retain a single-loop structure. Under an unbiased stochastic gradient oracle with uniformly bounded variance, SPDE finds an \varepsilon -game-stationary point with stochastic first-order oracle (SFO) complexities of O(\kappa\varepsilon^-4) and O(\varepsilon^-5) in the nonconvex–strongly concave and nonconvex–concave settings, respectively, where \kappa=L/\mu . Under an additional mean-square Lipschitz condition on the stochastic gradients, VR-SPDE improves these GS complexities to O(\kappa^3/2\varepsilon^-3) and O(\varepsilon^-9/2) , respectively. For an \varepsilon -optimization-stationary point, SPDE achieves SFO complexities of O(\kappa\varepsilon^-4) and O(\varepsilon^-6) , while VR-SPDE achieves O(\kappa^3/2\varepsilon^-3) and O(\varepsilon^-6) , in the two settings, respectively. These OS guarantees match the best-known bounds achieved by multi-loop methods while preserving a single-loop implementation. To the best of our knowledge, our results provide the best-known SFO complexity guarantees among single-loop stochastic first-order methods for the respective stationarity criteria and problem classes.

[LG-85] Bayesian classification of astronomical spectra with class uncertainties

链接: https://arxiv.org/abs/2609.21694
作者: Simon Barton,Martin Sahlén,Andreas Korn,Christian Glaser
类目: Instrumentation and Methods for Astrophysics (astro-ph.IM); Machine Learning (cs.LG); Data Analysis, Statistics and Probability (physics.data-an)
*备注:

点击查看摘要

Abstract:Context: We developed a probabilistic machine learning method with the aim of performing the O(10)-way classification of low- and high-resolution spectra of stellar and extragalactic targets for the upcoming 4MOST survey. In fulfilment of the survey requirements, this method should be able to express uncertainty in the input data as well as uncertainty introduced in its prediction. Aims: Four different methods are explored: (1) convolutional neural networks (CNNs), (2) the Dirichlet distribution, (3) Monte Carlo dropout (MCD), (4) Bayesian neural Networks (BNNs) + variational inference (VI). Training and validation was performed using labelled spectra from the SDSS database and a custom 4MOST mock dataset. All the methods were compared in terms of the same metrics: accuracy, area under the curve (AUC), expected calibration error (ECE), Shannon entropy, negative log-likelihood (NLL), Brier score, training time, and inference time. Methods: A CNN with simple architecture and about 20,000 parameters was trained to achieve classification accuracies of 91.5% on SDSS data and 92.8% on 4MOST mock data. The direct Dirichlet prediction and VI models tested provide uncertainties on class membership probabilities, but they confuse classes more often. The MCD on a CNN is found to be the most suitable; it boosts the point-estimate accuracies to 92.6% and 93.9%, while still providing fast training and sufficiently fast inference. Compared to a standard CNN, the method additionally provides well-calibrated uncertainties at marginal extra cost.

[LG-86] Periodic Neural Mapping for Unsteady Rotor-Blade Pressure and Aeroelastic Load Prediction

链接: https://arxiv.org/abs/2609.21590
作者: Lionel Salesses,Joachim Dominique,Tariq Benamara,Théo Flament,Franck Mastrippolito
类目: Fluid Dynamics (physics.flu-dyn); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Accurate prediction of unsteady aerodynamic loads remains a major challenge in turbomachinery design. High-fidelity Computational Fluid Dynamics (CFD) simulations are expensive, while aeroelastic Quantities of Interest (QoI) depend sensitively on the temporal evolution of the pressure field. This work introduces periodic Fourier Neural Mapping (p-FNM), a neural-operator framework for predicting unsteady pressure distributions on turbine rotor blades simulated using the chorochronic numerical hypothesis. The architecture embeds temporal periodicity into the model and learns a continuous mapping from operating conditions and time to pressure fields. Unlike sequential latent-space approaches, p-FNM predicts pressure fields independently at any time, avoiding error accumulation while preserving temporal continuity. The model is evaluated on a database of unsteady rotor-blade simulations and compared with a reduced-order baseline based on a variational autoencoder and recurrent neural network, refered as the Temporal Prediction Model (TPM). Performance is assessed for pressure fields and Generalized Aerodynamic Forces (GAFs), the primary aeroelastic QoI. Across all training datasets, p-FNM consistently outperforms TPM. On the largest dataset, p-FNM achieves a pressure-field mean absolute percentage error of 0.46% and a GAF-magnitude prediction error of 4.42%, corresponding to improvements of 60.7% and 77.6%, respectively. The minimum weighted phase error reaches 0.060 rad, demonstrating accurate preservation of the temporal characteristics of the aerodynamic response. The results show that GAF prediction is more challenging than pressure-field prediction and that temporal coherence is critical for accurately predicting spectral aerodynamic quantities. These findings demonstrate the potential of periodic neural operators for reduced-order modeling and aeroelastic analysis in turbomachinery.

[LG-87] Weighted Quantum Signal Processing: Low-Depth Polynomial Approximation with Applications to Kolmogorov-Arnold Networks

链接: https://arxiv.org/abs/2609.21567
作者: Rohit Sarma Sarkar,Rupayan Bhattacharjee,Elias F. Combarro,Michele Grossi,Lirandë Pira,Carmen G. Almudéver,Sergi Abadal,Eduard Alarcon
类目: Quantum Physics (quant-ph); Computational Complexity (cs.CC); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Quantum Signal Processing is a powerful quantum framework for generating and approximating univariate polynomials. However, QSP is often limited by circuit-depth bottlenecks and parity constraints on the class of realizable polynomials. In this work, we introduce Weighted Quantum Signal Processing, an extension of QSP in which a weight function is assigned to the central rotation operator. This formulation provides a deeper understanding of QSP, which emerges as the special case of WQSP with unit weights. The choice of weights determines the structure and expressive capabilities of WQSP circuits. When the weights are natural numbers greater than one, WQSP reduces to a pruned version of QSP, revealing parameter redundancies in the standard framework. Through appropriate selection of integer weights, WQSP achieves linear-to-exponential reductions in the number of parameters required to realize arbitrary bounded univariate polynomials while preserving approximation quality. For generic weights, we establish corresponding approximation error bounds and show that, in many cases, the approximation is exact. We analyze WQSP from both a deterministic perspective, where polynomial generation is formulated as the solution of a linear system, and a quantum machine learning perspective, where WQSP serves as a structured and expressive quantum learning model. We further employ this learning framework to parameterize learnable activation functions in Kolmogorov–Arnold Networks for multivariate function approximation. Our results show that WQSP provides a compact, flexible, and theoretically grounded framework for realizing arbitrary univariate polynomials while requiring significantly fewer trainable parameters than conventional QSP. This yields expressive and parameter-efficient neural architectures, highlighting the potential of WQSP as a scalable primitive for quantum-enhanced machine learning.

[LG-88] Improving the Predictive Performance of Bootstrap Aggregating by Dirichlet Resampling

链接: https://arxiv.org/abs/2609.21454
作者: Quoc Viet Le,Joonha Park
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Methodology (stat.ME)
*备注: 29 pages (10 main text, 19 pages appendix), 21 tables, 3 algorithms. No figures

点击查看摘要

Abstract:We revisit Breiman’s observation that reducing inter-tree correlation without weakening individual trees can improve random forests. Building on this principle, we introduce two variants: Dirichlet-Multinomial Bagging Random Forest (DM) and Dirichlet-Weighted Random Forest (DW). Both modulate sample reweighting via a concentration parameter \alpha0 . We provide a simple theoretical criterion that clarifies when these variants behave indistinguishably from standard random forests, and we use it to guide a lightweight tuning strategy. In a controlled evaluation on public classification benchmarks, DM and DW are consistently competitive and often stronger than other random-forest (RF) baselines, with negligible additional runtime.

[LG-89] Brownian Heads for Deep ReLU Representations: Activation Mass and the Cost of Same-Sample Selection

链接: https://arxiv.org/abs/2609.21422
作者: Mahdi Mohammadigohari,Nicole Mücke
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Deep representation learning often selects hidden features and fits the final predictor on the same sample, so fixed-feature analysis performed after selection can omit selection cost. We study the conditional empirical Rademacher complexity of deep ReLU representations followed by bounded-norm predictors in additive or Lévy-Brownian RKHSs, termed Brownian heads. For a fixed representation, we derive an exact dual identity and sharp bounds in terms of activation mass, the average norm of the observed hidden vectors. Under same-sample selection, the representation supremum induces a quadratic Rademacher process. Brownian layer-cake and Gaussian-projection identities reduce it to coordinatewise or signed projected threshold traces, separating realized scale from selection complexity. For samples with pairwise-distinct inputs, explicit scalar ReLU families match the finite-trace and VC rates up to universal constants at the realized trace-and-envelope level. Induced-norm contraction also yields architecture-level bounds for rectangular, rank-deficient ReLU networks. Experiments verify the sharp bounds and rates, exhibit a selection gap at fixed activation mass, and assess the predictive feasibility of Brownian heads.

[LG-90] Sparse Identification for Automatic Large-Scale Screening: A Constraint-Aware Framework with Ultra Fast Decoding Algorithm

链接: https://arxiv.org/abs/2609.21321
作者: Jianing Li,Li Chai,Yingcheng Lai
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Signal Processing (eess.SP)
*备注: This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible

点击查看摘要

Abstract:In the early stages of a pandemic, identification of a small number of infected individuals through large-scale screening is critical for pandemic control, yet remains challenging under limited reagents and testing capacity. Existing group testing methods suffer from either high computational complexity or low identification accuracy. Even worse, no available methods provide theoretically rigorous analysis for sparse identification with hard constraints caused by the sample usage constraint and the dilution effect existing ubiquitously in practical applications. In this article, we propose the Logic Screening method (LoSc), an ultra fast, accurate, and theoretically grounded framework for large-scale screening. LoSc introduces a novel decoding algorithm with a very simple selection strategy, achieving identification of all positives with only O(klogn) pooled tests. The decoding relies only on logical operations, enabling direct hardware implementation and yielding ultra fast computational implementation. Moreover, LoSc explicitly incorporates dilution and sample usage constraints into pooling designs, and establishes theoretical guarantees to guide optimal pooling configurations. Extensive simulations confirm the superior effectiveness, efficiency, and scalability. We believe LoSc offers a fast and reliable solution for automatic large-scale screening.

[LG-91] Diagonalized Attention for Individualized Regression: Latent-Row Localization and Prediction

链接: https://arxiv.org/abs/2609.21320
作者: Borui Peng,Liwei Lin,Feifei Wang,Long Feng
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Modern text and image representations are often matrix-valued, with rows corresponding to tokens, patches, or other local feature vectors. Predictive information is often sparse but sample-specific, making classical sparse regression methods with a common support poorly suited to this heterogeneity. This paper formalizes an individualized sparse regression framework for matrix-valued covariates in which each observation has its own rows of interest, while the associated regression effects are shared across the population. To estimate this model, we introduce a diagonalized attention mechanism that uses query–key scores to localize sample-specific signal rows and a value matrix for downstream regression. The proposed method has a parameter dimension independent of sample size and can identify rows of interest for new observations without their responses. We establish existence theorems showing that, under suitable score-separation and concentration conditions, single-head and multi-head diagonalized attention models recover the latent rows with high probability, yielding prediction risk bounds. Our theory therefore provides a statistical explanation of how attention-based scoring localizes sample-specific signals in heterogeneous matrix-valued data. Simulations demonstrate strong prediction and localization in regression and misspecified classification across varying sample sizes, dimensions, and signal cardinalities. Real sentiment analyses show improved classification accuracy and interpretable token selection.

[LG-92] From Trainability Diagnostics to Optimization Claims: Boundaries and Controls in Variational Quantum Optimization

链接: https://arxiv.org/abs/2609.21243
作者: Pilsung Kang
类目: Quantum Physics (quant-ph); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Barren plateau diagnostics characterize whether gradient signal remains available for training, but surviving signal need not translate into successful optimization. We study this trainability–optimization gap at the level of optimizer steps. Treating coefficient-weighted Hamiltonian-term gradients as task-like components, we introduce step-level diagnostics and derive an exact bridge between signed termwise organization, directional activity, and first-order descent. Resolving this bridge into standard first-order geometry shows that the apparent organization–activity factors are not independent optimization axes and that, at fixed state and update norm, the raw gradient maximizes first-order descent of the summed objective. We compare vanilla gradient descent, a deterministic Hamiltonian-term PCGrad variant, and probe-gated LSO-PCGrad on transverse-field Ising model instances with hardware-efficient and Hamiltonian variational ansatzes, together with matched controls for update norm and probe budget. Blind projection can improve an organization diagnostic while worsening final energy and first-order predictability. After conditioning on standard first-order geometry, residual term-space composition shows no reproducible material incremental association with realized descent, while optimizer-relative update norm shows positive material associations in some settings without cross-regime reproducibility. Matched controls provide no resolved final-energy benefit attributable to the projected direction, and the improvement of LSO-PCGrad is more consistent with probe-based search and step-norm adaptation than with Hamiltonian-term projection itself. These results show that gradient-structure diagnostics can characterize trainability and update geometry without serving as standalone evidence of optimization benefit, which requires controls matched on update norm and search budget.

[LG-93] riply-Scalable Equivariant Gaussian Process Modeling

链接: https://arxiv.org/abs/2609.21085
作者: Tim Steinert,David Ginsbourger
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Statistics Theory (math.ST); Applications (stat.AP); Computation (stat.CO)
*备注:

点击查看摘要

Abstract:Gaussian processes (GPs) provide principled probabilistic predictions while encoding prior knowledge, including equivariances. Yet, their use in large-scale scientific problems is limited by computational cost. Equivariant neural networks are common but typically lack the uncertainty quantification offered by GPs, which is valuable in applications such as molecular research. High-dimensional inputs and large symmetry groups further demand scalability. We establish results pertaining to the interplay of GP equivariance and conditioning and leverage them to obtain equivariant sparse GPs through suitable mean functions and covariance kernels. We instantiate this framework with a flexible class of integration-free equivariant kernels, yielding scalable and data-efficient GP inference. In particular, we introduce triply scalable equivariant Gaussian processes. We employ equivariant sparse variational Gaussian processes for \mathrmSO(2) -equivariant vector fields and molecular property prediction. Alongside the SVGP, we develop a matrix-free equivariant full-GP implementation that combines an exact Kronecker reduction with preconditioned conjugate-gradient solves, enabling fast and scalable evaluation of the full joint predictive density. We further compare different approaches for selecting inducing points in the equivariant sparse GP models. Our test cases include synthetic \mathrmSO(2) -equivariant fields as well as the prediction of electric dipole moments of N-methylformamide based on quantum chemistry simulations, achieving accurate, uncertainty-aware predictions at a fraction of the computational cost of classical GP inference. Subjects: Machine Learning (stat.ML); Machine Learning (cs.LG); Statistics Theory (math.ST); Applications (stat.AP); Computation (stat.CO) Cite as: arXiv:2609.21085 [stat.ML] (or arXiv:2609.21085v1 [stat.ML] for this version) https://doi.org/10.48550/arXiv.2609.21085 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-94] A Smoothed Discrepancy Principle for Random Feature Methods and Neural Networks

链接: https://arxiv.org/abs/2609.21017
作者: Mike Nguyen,Nicole Mücke
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Statistics Theory (math.ST)
*备注:

点击查看摘要

Abstract:We study data-driven early stopping for spectral regularisation methods in the classical non-parametric regression setting. Building on the discrepancy principle, we propose a multi-scale stopping rule that applies to general kernel estimators and show that, unlike previous approaches, it achieves full adaptivity over all smoothness levels in the well-specified case. A key contribution of our work is an extension based on random feature approximations, which reduces computational cost on large datasets while preserving minimax-optimal statistical guarantees. Our procedure not only selects an optimal stopping time but also provides a fully data-driven choice of the number of random features needed to achieve optimal rates. Through the established connection between random features and neural networks in the neural tangent kernel regime, our method further yields a principled, data-driven recommendation for the network width. We prove that the resulting simultaneously chosen width and stopping time allow neural networks to attain minimax-optimal learning rates without prior knowledge of smoothness or capacity parameters.

[LG-95] Aggregated Posterior Predictive Checks for Generative Modeling

链接: https://arxiv.org/abs/2609.20999
作者: Shweta Dutta,Gemma E. Moran
类目: Methodology (stat.ME); Machine Learning (cs.LG); Statistics Theory (math.ST); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Latent variable generative models are commonly fit using simple priors over latent variables, but draws from these priors often fail to produce realistic data. This failure is due to a mismatch between the prior and the aggregated posterior, the distribution of latent variables induced by the fitted model and the data. This mismatch is often viewed as evidence that the prior is misspecified and should be replaced. Alternatively, in modern generative models, a two-stage strategy is increasingly used where first, the model is fit, and second, the aggregated posterior is estimated (van den Oord et al.,2017; Rombach et al., 2022.). Synthetic data are then obtained by sampling from this aggregated posterior instead of the prior. To check such procedures, we introduce the aggregated posterior predictive check (APPC). Theoretically, we establish sufficient conditions under which the APPC is asymptotically calibrated. For probabilistic principal component analysis, we show that the APPC can remain calibrated under a misspecified latent prior when pervasive factors permit recovery of the signal space. Experiments with variational autoencoders show that aggregated posterior sampling improves generation for heavy-tailed and clustered data relative to Gaussian prior sampling while performing comparably to models with more flexible latent priors.

[LG-96] Complex Problem Solving in Large Language Models : A Statistical Control Survey and Diagnostic Framework

链接: https://arxiv.org/abs/2609.20973
作者: Jiazhang Cai,Tao Wang,Ruidong Zhang,Siyuan Li,Terry Ma,Luyang Fang,Haoran Lu,Huimin Cheng,Yingchuan Zhang,Shushan Wu,Rui Xie,Lin Tang,Chao Huang,Rongjie Liu,Ziyu Liu,Meizhi Yu,Yongkai Chen,Yifan Zhou,Zeliang Sun,Chang Liu,Zhen Xiang,Wei Xiao,Zixin Rao,Xinyi Liu,Yutong Hu,Mengrui Zhang,Jing Zhang,Weidi Luo,Jincheng Yu,Zhengliang Liu,Weihang You,Hanqi Jiang,Yi Pan,Junhao Chen,Xinliang Li,Tianming Liu,Wenxuan Zhong,Ping Ma
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: 82 pages, 7 figures. Submitted to Artificial Intelligence Review

点击查看摘要

Abstract:Complex problem solving (CPS) with large language models (LLMs) is often framed as a matter of stronger reasoning or longer generation. Yet early-step error amplification, prompt brittleness, and failures to revise incorrect commitments are difficult to explain by missing knowledge or expressive capacity alone. This survey interprets CPS as a sequential estimation-and-decision problem over a latent solution state. A controller maintains a belief about an unobserved solution trajectory, updates it as noisy intermediate evidence arrives, and decides whether to commit, verify, branch, roll back, or abstain to minimize expected loss. Reasoning supplies candidate transitions and interpretations, whereas process control shapes and evaluates those proposals and regulates subsequent transitions and observations. Within this framework, we organize existing methods around five components: explicit state representation, transition structuring, validation and constraint enforcement, search and rollback, and uncertainty management. We also interpret evaluation metrics according to the statistical quantities they estimate. The framework further yields a diagnostic hypothesis: interventions should be most effective when they target the error or uncertainty component implicated by an observed failure. We distinguish systematic, stochastic, and irreducible error together with epistemic and aleatoric uncertainty, and call this alignment problem-control fit and its failure control mismatch. For example, additional sampling may reduce sampling variability while leaving a shared systematic error unchanged. This perspective clarifies what current methods estimate and control, what remains uncontrolled, and why reliable validation, targeted recovery, calibrated uncertainty, and matched-budget evaluation are central open problems.

[LG-97] Extreme classification: beating chance with one training example from each class

链接: https://arxiv.org/abs/2609.20897
作者: Kevin Bleakley(LMO, CELESTE),Aaditya Ramdas
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We study a minimal classification problem: Given independent labeled observations X\sim P and Z\sim Q from two unknown distributions P,Q , and given an independent target Y drawn with equal probability from P or Q , can one classify Y strictly better than chance whenever P\neq Q ? The one-nearest-neighbor rule succeeds for every pair of multivariate Gaussian distributions with distinct means and a common positive-definite covariance matrix but can perform strictly worse than chance even for smooth densities on the real line. We construct a fixed randomized kernel rule whose expected accuracy is exactly 1/2+\operatornameMMD_k^2(P,Q)/4 , and obtain characteristic kernels on countably generated measurable spaces from countable families of measurable binary questions. We also prove that a deterministic order rule on \mathbb R beats chance for every pair of distinct Borel probability measures. A measurable encoding then gives a deterministic distribution-free rule which beats chance on every countably generated measurable space, in particular every separable metric space. Finally, we show that no rule works for every distinct pair of distributions and every unknown unbalanced class prior; under adaptive target-class selection, every rule other than a fair coin is strictly worse than chance for some finitely supported pair.

[LG-98] Automated Physics-Informed Neural-Networks-Based Calibration of Highly Segmented Silicon Telescopes

链接: https://arxiv.org/abs/2609.20868
作者: M. Rejmund,A. Lemasson,P. Morfouace,D. Ramos,J. Taieb,J. D. Frankland
类目: Instrumentation and Detectors (physics.ins-det); Machine Learning (cs.LG); Nuclear Experiment (nucl-ex)
*备注:

点击查看摘要

Abstract:Transfer and multi-nucleon transfer reactions are essential tools for probing nuclear structure and reaction dynamics, requiring precise determination of the identity, energy, and emission angles of reaction products. The increasing granularity of modern silicon telescope arrays enhances experimental capabilities but challenges detector calibration, as conventional channel-by-channel approaches become inefficient and difficult to scale. In this work, we present a fully automated, physics-informed calibration framework based on neural networks, specifically designed for highly segmented silicon detector arrays. The method formulates calibration as a global optimization problem, in which detector gains and geometrical corrections are determined simultaneously by minimizing the width of the reconstructed excitation energy under two-body kinematics constraints. The approach relies exclusively on experimental data and well-established physical principles, without requiring explicit modeling of detector response. A distinctive feature is the use of multiple neural network sub-models sharing a common loss function with embedded physics constraints, enabling coherent and self-consistent calibration across all detector channels. This strategy ensures scalability, robustness, and reproducibility, making it particularly suitable for next-generation detector systems with increasing complexity. The performance of the method is demonstrated using experimental data from the Particle-Identification Silicon-Telescope Array (PISTA) in high-resolution fission studies in inverse kinematics. The results show excellent agreement with theoretical kinematics, high-quality particle identification, and a significant improvement in calibration efficiency. The proposed framework provides a general and adaptable solution for the calibration of complex detector systems in modern nuclear physics experiments. Subjects: Instrumentation and Detectors (physics.ins-det); Machine Learning (cs.LG); Nuclear Experiment (nucl-ex) Cite as: arXiv:2609.20868 [physics.ins-det] (or arXiv:2609.20868v1 [physics.ins-det] for this version) https://doi.org/10.48550/arXiv.2609.20868 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Antoine Lemasson [view email] [v1] Tue, 15 Sep 2026 12:37:50 UTC (2,423 KB)

[LG-99] Reconstruction of 4D Mitral Regurgitation Hemodynamics from Sparse Planar Data using Deep Operator Networks with Test-Time Adaptation

链接: https://arxiv.org/abs/2609.20857
作者: Jakob Marcel Hoffmann,Yosuke Hasegawa,Alexander Stroh
类目: Medical Physics (physics.med-ph); Machine Learning (cs.LG); Fluid Dynamics (physics.flu-dyn)
*备注:

点击查看摘要

Abstract:Quantifying mitral regurgitation severity remains limited by the assumptions of clinical flow convergence methods, while high-fidelity simulation and volumetric velocimetry are too slow for routine use. We investigate whether a learned solution operator can reconstruct transient three-dimensional transvalvular hemodynamics from the sparse observation an in-vitro experiment actually provides: a single planar velocity slice and two boundary pressure traces. A Deep Operator Network is pretrained on an experimentally benchmarked URANS database spanning eleven mitral regurgitation orifice phantoms, learning a mapping from a masked two-component planar velocity snapshot to the surrounding volumetric field, and is subsequently adapted to unseen target cases by fine-tuning on their sparse measurements. Adaptation reliably corrects the flow topology within the supervised plane, reorienting a strongly eccentric jet that the pretrained operator predicts as straight, and yields full-field predictions in minutes rather than the days required by the underlying simulations. Its influence decays sharply with distance from that plane, however: measured against phase-resolved particle image velocimetry, the reconstruction error rises from 24.6% at 2mm to 52.6% at 6mm, and the resulting mismatch between corrected and uncorrected layers degrades physical consistency. Single-plane supervision thus constrains the observed plane far more effectively than the surrounding volume, which we identify as the principal obstacle to coherent 4D reconstruction from sparse planar data.

附件下载

点击下载今日全部论文列表