Scope 

 

Artificial intelligence components are increasingly embedded in information systems that span centralized cloud platforms, resource-constrained edge devices, and cyber-physical environments. However, their evaluation often stops at offline accuracy or an isolated benchmark. Once deployed, models interact with changing data, prompts, software dependencies, accelerators, communication networks, operational policies, and human decision-makers. Consequently, results that appear convincing in the laboratory can be difficult to reproduce, compare, diagnose, or defend under real operating conditions.

This workshop defines lifecycle assurance as the systematic creation, linkage, and maintenance of empirical evidence from experiment design through deployment, operation, update, rollback, and retirement. Its focus is on methods and artifacts that make claims independently examinable, including versioned datasets and transformations, configuration identity, provenance records, frozen baselines, quantified uncertainty, machine-readable evidence packages, validation gates, operational telemetry, and traceable change histories. Reproducibility is understood as deterministic repetition where feasible and as statistically characterized repeatability where stochastic models, adaptive behavior, heterogeneous hardware, or variable runtime conditions make exact replication impossible.

The scientific focus is the joint evaluation of AI behavior and systems behavior in heterogeneous cloud edge environments. Relevant work may examine the interactions among predictive quality, latency, energy use, memory footprint, data movement, reliability, security, robustness, and operational cost. The scope also includes portability across hardware and software stacks, drift and anomaly detection, degraded operating modes, controlled recovery, human oversight, and evidence-based update or rollback decisions. Generative and agentic AI systems are included when they are studied as operational systems with traceable prompts, tools, policies, dependencies, and runtime behavior rather than as isolated model demonstrations.

The workshop welcomes methodological contributions, system architectures, benchmark and reporting protocols, reproducibility studies, tools, datasets, negative results, and quantified industrial or public-sector case studies. Submissions should identify the evidence artifact or evaluation protocol they contribute to and should either address at least two lifecycle stages or establish a measurable connection between a laboratory claim and field behavior. Papers whose contribution is limited to a new model, algorithm, architecture, or application without a substantive reproducibility or lifecycle-assurance component will be considered outside the workshop scope.

The workshop is organized around one research object: a reproducible evidence chain that connects laboratory claims to deployment decisions, operational observations, controlled changes, and system retirement.

 

 

Topics of interest 

 

Topics of interest include (but are not limited to):

 

•         Reproducible evaluation protocols for deterministic, stochastic, generative, and agentic AI systems

•         End-to-end provenance linking data, transformations, models, prompts, tools, policies, software, and hardware

•         Machine-readable experiment records, evidence packages, validation gates, and assurance cases

•         Model and adaptation passports, AI bills of materials, configuration identity, and dependency tracking

•         Frozen baselines, statistical testing, uncertainty reporting, and evidence aggregation

•         Multi-objective benchmarking of quality, latency, energy, memory, data movement, reliability, and cost

•         Portability and repeatability across heterogeneous cloud, edge, accelerator, and software environments

•         MLOps, LLMOps, and EdgeOps methods supporting traceable AI lifecycles

•         Runtime observability, telemetry design, and linkage between development and operational evidence

•         Data, concept, model, and system drift; anomaly detection; and uncertainty-aware monitoring

•         Controlled model and policy updates, rollback, supersession, decommissioning, and evidence retention

•         Reliability, resilience, degraded-mode operation, recovery, and fail-safe behavior in AI-enabled systems

•         Security, integrity, auditability, and dependency risks in tool-using and agentic AI systems

•         Human-in-the-loop validation, decision rights, escalation paths, and accountable operational oversight

•         Digital twins, simulation, synthetic data, replication studies, negative results, and quantified field case studies

 

Organizing Committee 

 

·       Krzysztof Wołk, Academy of Social Sciences, Poland, This email address is being protected from spambots. You need JavaScript enabled to view it.

 

Program Committee 

 

·       In construction