Powered by
1st International Workshop on Trustworthy and Responsible aUtonomous SysTems (TRUST 2026), October 12–16, 2026,
Munich, Germany
1st International Workshop on Trustworthy and Responsible aUtonomous SysTems (TRUST 2026)
Frontmatter
Title Page
Article: asews26trustforeword-fm000-p (type: Frontmatter) doi:
Welcome from the Chairs
Welcome to TRUST 2026, the 1st International Workshop on Trustworthy and Responsible aUtonomous SysTems, co-located with the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026) in Munich, Germany, on October 12, 2026. As autonomous systems and AI agents increasingly move beyond the role of assistants and take on responsibilities akin to team members, and in some cases managerial functions such as coordination, delegation, and decision-making, TRUST focuses on how these systems can themselves be systematically engineered to be trustworthy and responsible, emphasizing a shift from Agents for Software Engineering to Software Engineering for Autonomous Systems.
Article: asews26trustforeword-fm001-p (type: Frontmatter) doi:
TRUST 2026 Organization
Organizing Committee
• Gianmario Voria, University of Salerno, Italy — Organizing Chair
• Yutaro Kashiwa, Nara Institute of Science and Technology, Japan — Organizing
Chair
• Fabio Palomba, University of Salerno, Italy — Organizing Chair
• Patrizio Pelliccione, Gran Sasso Science Institute, Italy — Organizing Chair
Article: asews26trustforeword-fm002-p (type: Frontmatter) doi:
Papers
Dynamic Deception: When Pedestrians Team Up to Fool Autonomous Cars
Masoud Jamshidiyan Tehrani,
Marco Gabriel,
Jinhan Kim, and
Paolo Tonella
(USI Lugano, Switzerland)
Many adversarial attacks on autonomous-driving perception models fail to cause system-level failures once deployed in a full driving stack. The main reason for such ineffectiveness is that once deployed in a system (e.g., within a simulator), attacks tend to be spatially or temporally short-lived, due to the vehicle’s dynamics, hence rarely influencing the vehicle behaviour. In this paper, we address both limitations by introducing a system-level attack in which multiple dynamic elements (e.g., two pedestrians) carry adversarial patches (e.g., on clothes) and jointly amplify their effect through coordination and motion. We evaluate our attacks in the CARLA simulator using a state-of-the-art autonomous driving agent. At the system level, single-pedestrian attacks fail in all runs (out of 10), while dynamic collusion by two pedestrians induces full vehicle stops in up to 50% of runs, with static collusion yielding no successful attack at all. These results show that system-level failures arise only when adversarial signals persist over time and are amplified through coordinated actors, exposing a gap between model-level robustness and end-to-end safety.
Article Search
Article: asews26trustmain-p1-p (type: Full Paper) doi:10.1145/3843782.3844650
Software Project Management with LLM-Based Automation: Coordination, Validation, and Governance in Practice
Ronnie de Souza Santos,
Italo Santos, and
Cleyton Magalhães
(University of Calgary, Canada; University of Hawaii at Manoa, USA; Rural Federal University of Pernambuco, Brazil)
As software engineering has evolved, development environments have increasingly integrated automated support for tasks such as testing, analysis, and code generation, requiring project management to coordinate both human work and automated workflows. In this paper, we investigate how software project managers perceive changes in their work and learning demands associated with LLM-based automation. Motivated by a gap in software engineering research, which has largely focused on task-level and developer-centered uses of LLMs, we adopted an exploratory case study approach to capture managerial perspectives on automation in practice. Based on the experience of software project managers working in a large, multi-project software organization, our analysis indicates that LLM-based automation influences planning, estimation, coordination, monitoring, and governance activities rather than introducing new formal management practices. Participants described LLMs as becoming embedded in everyday project work, producing uneven effects on productivity, increasing the need for review and validation, and reducing visibility into task execution. These effects contribute to greater reliance on managerial judgment and coordination. Learning demands were perceived as experiential and incremental, centered on understanding LLM capabilities and limitations, critically assessing generated artifacts, and guiding responsible use within teams. The findings provide empirical evidence on how software project management work is adapted in contexts where LLM-based automation is integrated into ongoing software development practice.
Article Search
Article: asews26trustmain-p2-p (type: Full Paper) doi:10.1145/3843782.3844651
From Regulation to Innovation: Mining High-Impact AI Patents Grounded in the EU AI Act and Korea’s AI Basic Act
Jimin Kim,
Ho Jun Kang, and
Jaehyun Lee
(Software Policy & Research Institute, Republic of Korea)
Artificial intelligence (AI) increasingly operates in domains that impact human life, safety, and fundamental rights, and emerging regulation singles out such high-impact domains for binding obligations. Beyond legislative and use-case definitions, this study empirically maps and analyzes the high-impact AI landscape. A unified taxonomy of 15 high-impact AI domains is derived from the heightened-obligation categories of the EU AI Act and Korea's AI Basic Act. Patents are then matched to each domain, rendering the 15 technological landscapes mutually comparable. The matching procedure combines domain-specific Cooperative Patent Classification (CPC) codes with keyword queries and AI-specific filter words, and yields 81,702 high-impact AI patents from USPTO publications filed between January 2018 and July 2026. Embedding-based density clustering then identifies the principal technological topics within each domain, and a topic-level change-point analysis traces how their distribution shifts over time. The resulting map is convergent in capability but sharply uneven in maturity: the 15 domains draw on a shared set of AI capabilities—prediction, recognition, monitoring, decision support, and autonomous control—yet range from healthcare (27,515 patents) to essential public services (372 patents). By linking regulatory categories to patent evidence, this study offers a reproducible framework for monitoring high-impact AI innovation and evidence that domains unevenly supplied with technology may require the obligations to be implemented at different depths and on different timelines.
Article Search
Article: asews26trustmain-p3-p (type: Full Paper) doi:10.1145/3843782.3844652
Closed-World Trustworthy Copilots for Legal and Regulatory Text
Monalisha Ojha and
Pragya Shrey Mishra
(University of Mannheim, Germany; Pillai College of Engineering, India)
Applying large language models (LLMs) to regulated domains such as financial and legal compliance is fundamentally a software engineering problem: an answer produced from a model’s parameters is, at the surface, indistinguishable from one grounded in the authoritative source, yet only the latter can be audited. Prior studies report legal-fact hallucination rates of 58–88% for general-purpose models and 17–33that trustworthiness here should be engineered into the surrounding system rather than sought from a stronger model, through confinement, which restricts the model to a single sealed corpus, and measurement, which quantifies per-claim groundedness. We present the closed-world copilot, a pipeline that seals one regulatory document into a structure-aware index, enforces a citation-and-verbatim-quote contract on every claim, and runs an independent leakage verifier combining a deterministic quote check, an isolated entailment judge, and a parametric-leakage probe. On the EU Markets in Crypto-Assets (MiCA) regulation, confinement improves groundedness and abstention over a baseline LLM and vanilla retrieval-augmented generation (RAG), and a synthetic-amendment test shows that an unconfined model reports memorized and potentially stale values. A second study on the Contract Understanding Atticus Dataset, reproduces this result and quantifies its cost in coverage. We position the approach as a trust layer above retrieval rather than a replacement for it.
Article Search
Article: asews26trustmain-p8-p (type: Full Paper) doi:10.1145/3843782.3844654
Automated Counterfactual Scenario Generation for Fairness Assessment of Large Language Models: The ScenGen Framework
Alessandra Parziale,
Antonio D'Auria,
Daniele Pio Scaparra, and
Valeria Pontillo
(Gran Sasso Science Institute, Italy; University of Salerno, Italy)
Large Language Models (LLMs) are increasingly adopted in decision-support applications, making fairness assessment essential. Existing approaches mainly rely on static benchmarks, which are difficult to extend across tasks, domains, and sensitive attributes, limiting the systematic and reproducible evaluation of model fairness. This paper presents ScenGen, an automated framework for generating counterfactual scenarios for LLM fairness assessment. Counterfactual scenarios consist of prompt pairs that differ only in a sensitive attribute, enabling the evaluation of whether this variation influences model decisions. Starting from a taxonomy that specifies tasks, sensitive attributes, and counterfactual substitutions, the framework automatically generates neutral scenarios and counterfactual prompt pairs, enabling reusable and reproducible fairness evaluations across models and application domains.
The framework was evaluated by generating 564 counterfactual prompt pairs across three task families and executing them on three LLMs over two independent runs, yielding 3,384 evaluated prompt pairs. We assessed the scenarios using the Decision Flip Rate, which measures changes in model decisions between original and counterfactual prompts. The results show that the generated scenarios successfully reveal decision instability across models, tasks, and sensitive attributes, demonstrating the effectiveness of the proposed approach.
Article Search
Article: asews26trustmain-p32-p (type: Full Paper) doi:10.1145/3843782.3844655
On Resilience in Multi-agent Systems for Software Engineering
Sandra Mitrović,
Cezar Sas, and
Matteo Salani
(IDSIA USI-SUPSI, Switzerland)
LLM-based Multi-Agent Systems (MAS) are increasingly adopted across domains, including Software Engineering (SE). As these systems become more prevalent, ensuring their resilience, i.e., the ability to withstand, adapt to, and recover from disruptions, becomes increasingly important. However, resilience remains largely unexplored in the context of MAS for SE. This paper examines resilience in LLM-based MAS, with a particular focus on SE. We propose a conceptual resilience evaluation framework and use it to assess resilience capabilities of several MAS for SE, identifying current gaps.
Our evaluation shows that resilience is mostly overlooked in LLM-based MAS for SE, with the existing mechanisms protecting the produced artifacts, not the MAS itself.
Article Search
Article: asews26trustmain-p41-p (type: Short Paper (5 pages)) doi:10.1145/3843782.3844656
Leveraging Large Language Models for Automated Discriminatory Test Generation
Sadia Afrin Mim,
Fatema Tuz Zohra, and
Brittany Johnson
(George Mason University, USA)
Machine learning (ML) models can exhibit systematic discrimi-
natory behavior based on demographic attributes (e.g., race and
gender) due to imbalanced training data and the complexity of
underlying algorithms. To identify such discriminatory behavior,
researchers have proposed automated discriminatory test genera-
tion techniques that mostly rely on specialized search algorithms
(e.g., Genetic Algorithm). Recent advances in Large Language Mod-
els (LLMs) have demonstrated remarkable capabilities across a wide
range of software engineering tasks through natural language based
prompts, motivating us to investigate their potential for automated
discriminatory test generation. We developed an LLM-assisted au-
tomated discriminatory test generation approach for identifying
discriminatory behavior in ML model outcomes using GPT-4o Mini
and Claude Sonnet 4 model. Our evaluation shows that LLM-
assisted approach, particularly with Claude, achieve higher success
rates in identifying discriminatory behavior compared to existing
state-of-the-art discriminatory test generation methods.
Article Search
Article: asews26trustmain-p69-p (type: Full Paper) doi:10.1145/3843782.3844658
Explaining LLM Adaptation Decisions in Self-Adaptive Systems via Faithful Input Attribution
Balint Mate,
Vincenzo Scotti,
Diego Perez-Palacin, and
Raffaela Mirandola
(FZI Research Center for Information Technology, Germany; KIT, Germany; Linnaeus University, Sweden)
Large language models (LLMs) are increasingly investigated as decision-making components in MAPE-K self-adaptive loops. For such autonomous behavior to be trustworthy, stakeholders need evidence with which to assess and justify confidence in its decisions. We envision a modular and extensible platform for explaining adaptation decisions generated by LLMs. Each platform module checks one candidate hypothesis about what a decision depends on, operationalizes it with one or more explanation methods, and evaluates the explanation for faithfulness and plausibility. We identify four initial hypotheses with their respective modules and implement one module, which uses replay-based input attribution to trace each recorded action to the information that supports it, separating which component to act on from what to do to it. We evaluate the module with three LLM planners on a well-established self-adaptive-system exemplar. The results distinguish system-wide dependence in component selection from stronger component grounding in action specification.
Article Search
Article: asews26trustmain-p73-p (type: Full Paper) doi:10.1145/3843782.3844659
Metamorphic Stress Testing of Vision-Language Components for Time-Critical Autonomous Systems
Arman Chhetri and
Yinxi Liu
(Rochester Institute of Technology, USA)
Vision-language models (VLMs) are increasingly integrated as perception and decision components in time-critical autonomous systems such as driving advisories and robot control, where a component must return an acceptable decision within an allocated timing budget. Existing adversarial-vision testing asks only whether a VLM returns the correct decision, but a free-form generative VLM that returns a correct answer only after exhausting its timing slack is still unsuitable for a latency-sensitive integration. Semantic-only testing is insufficient: VLM acceptance testing for such systems must assert a joint behavior-and-timing contract, and metamorphic stress testing is a practical way to check that contract before integration. The contract requires a usable decision and completion within the timing budget left for the VLM after the rest of the pipeline has run. We encode this as a metamorphic relation that pairs a benign visual input with a bounded (epsilon-ball) perturbation under a fixed prompt, checking both sides of the contract over repeated benign and perturbed decodings.
We apply the method on LLaVA-1.5-7B for a constrained Go/Stop driving-decision scene, reporting distribution-based metrics on this single scene. On this scene, a single bounded, image-only perturbation jointly violates both contracts: it collapses the behavioral decision (keyword-valid rate 1.00 to 0.19) and suppresses end-of-sequence termination, amplifying mean component response latency 18.6×; 81% of perturbed decodings exceed the benign p95 timing slack. We use unrestricted free-form generation as the vulnerable baseline and evaluate two operational controls: a tight output-length cap restores the component timing bound to within the benign p99 latency while truncating almost no benign responses, and a structured short-answer prompt restores keyword-valid short answers and timing (0.19 to 1.00) but yields incorrect Go decisions in 82% of benign trials on this red-light scene (versus 1.2% under free-form). Repeating the same protocol on Qwen2.5-VL-7B with RECALLED shows a different failure shape: keyword-valid stays 1.00 but length stretches 28.1×, latency amplifies 18.7×, and 73% of trials violate the benign p95 slack (near-cap 0.01). We report a component-level, single-scene result and conclude that resource-robustness and timing budgets belong in VLM acceptance testing alongside semantic correctness. We release the trial data, reuse matrix, configuration templates, and analysis scripts.
Article Search
Artifacts Available
Article: asews26trustmain-p76-p (type: Full Paper) doi:10.1145/3843782.3844660
Fairness Guardrails for Autonomous Agents: A Trustworthiness-by-Design Vision from Social Cybersecurity
Maria Teresa Baldassarre,
Vita Santa Barletta,
Vito Bavaro, and
Ronnie de Souza Santos
(University of Bari, Italy; University of Bari Aldo Moro, Italy; University of Calgary, Canada)
Autonomous GenAI agents are increasingly operating as active participants in adversarial socio-technical environments, assuming evaluative and governance responsibilities whose consequences affect different user populations. Large Language Models (LLMs) deployed in these roles routinely exhibit systematic group unfairness. We frame this as a software engineering (SE) failure: the SE community treats fairness as a post-hoc evaluation metric rather than a structural requirement, so external mitigations such as prompt-based debiasing fail under adversarial pressure. This vision paper argues that the ongoing shift from Agents for Software Engineering to Software Engineering for Autonomous Systems requires trustworthiness to be specified, enforced, and continuously assured as a first-class structural property. We present the Fair Social Honeypot Framework, an architecture for Online Social Networks (OSNs) that implements a Trustworthiness-by-Design pattern, treating fairness, runtime monitoring, and human oversight as first-class structural requirements. Three integrated mechanisms realize this pattern: decoupled Fairness Guardrail modules, inference-time activation steering, and structured Human-in-the-Loop oversight. The Guardrail modules and activation steering can be applied as general-purpose SE components to any autonomous system that generates content or makes classification decisions over diverse populations under adversarial conditions. We conclude with a research roadmap and a mixed-methods empirical agenda, ranging from in-vitro experiments to in-vivo deployment.
Article Search
Article: asews26trustmain-p89-p (type: Short Paper (5 pages)) doi:10.1145/3843782.3844661
AgentChaosBench: Benchmarking External Fault Detection and Localization in Agentic Systems
Chenkai Zhang,
Yiran Li,
Yifang Tian,
Michail Bachras, and
Hans-Arno Jacobsen
(University of Toronto, Canada)
Reliability in LLM-based agentic systems is a property of the whole execution (its tool calls, model calls, guardrails, and inter-agent messages), not of the final answer alone, yet evaluating only task outcomes reveals little about how or why a run fails. We present AgentChaosBench, a benchmark for detecting and localizing runtime faults in agentic systems from their execution telemetry. We run five heterogeneous applications that coordinate agents over the Agent-to-Agent protocol and call tools through the Model Context Protocol, and inject ten types of operational fault (unavailable or slow tools, corrupted or oversized responses, and delayed, looped, or misrouted delegations and bypassed guardrails) at their tool, model, guardrail, and inter-agent boundaries, alongside a no-fault control. The resulting dataset contains 275 sanitized traces: 250 faulty executions spanning ten fault types and 25 no-fault controls. Each faulty trace is aligned with the no-fault execution of the same input; fault-type labels and, where applicable, location labels are held out from diagnosis. On structured single-trace inputs, a first set of zero-shot LLM baselines shows the task is far from solved: local detectors up to 14B parameters reach only 13.6–19.2% top-1 fault-type accuracy and the frontier DeepSeek-v4-pro only 24.8%, while jointly identifying the fault type and its location tops out at 22%; reference-dependent faults (above all a bypassed guardrail) stay near-unsolved from a single trace. An aligned reference improves selected relative faults but does not resolve guardrail bypass. The held-out labels and compact prediction format support reproducible comparison of LLM-based and non-LLM diagnosis methods.
Article Search
Article: asews26trustmain-p97-p (type: Full Paper) doi:10.1145/3843782.3844662
Can Open-Weight LLMs Predict Test Outcomes from Traces? A Cross-Subject Boundary Study on Tests4Py
Giuliano Rasper,
Tural Mammadov,
Marius Smytzek, and
Andreas Zeller
(Saarland University, Germany; CISPA Helmholtz Center for Information Security, Germany)
Determining whether a test execution should pass or fail remains
a core obstacle to end-to-end test automation. Recent work has
focused mainly on generating assertions from source context, leav-
ing open how much pass/fail signal can be recovered directly from
runtime behavior, especially under subject transfer. We present a
cross-subject boundary study of this question on Python unit tests
from the Tests4Py benchmark. Our pipeline collects execution traces
via runtime instrumentation, compresses them through redundancy
pruning and two minimization strategies (holistic filtering and key-
point amplification), splits long traces into chunks whose partial
predictions are aggregated per test, and fine-tunes open-weight
large language models (LLMs) to classify tests from previously un-
seen subjects as passing or failing. We evaluate CodeT5-small, Phi-
3.5-mini-instruct, and Llama-3.1-8B under varying context budgets
and overlap settings. The best completed configuration, CodeT5-
small with keypoint amplification and no overlap, reaches a macro-
averaged failing-class F1 score of 𝐹 1fail = 0.365 across the three
held-out subjects of our fixed protocol, compared with 𝐹 1fail = 0.025
for an empirical-distribution reference; since the released traces
contain no exception events, this signal cannot stem from spotting
recorded exceptions. Larger context windows improve the holistic
filtering configurations, whereas overlap has mixed effects.
Within this single-protocol setting, runtime traces contain a non-
trivial cross-subject signal. However, performance remains too low
and too unstable across subjects for stand-alone deployment. In
this setting, progress appears to depend more on representation,
aggregation, and robustness to subject shift than on model scale
alone.
Article Search
Article: asews26trustmain-p99-p (type: Full Paper) doi:10.1145/3843782.3844663
proc time: 0.25