Skip to content

MSc dissertation · Aston University · 2026

Measuring the Unmeasured

An empirical instrument for auditing the code-quality and governance behaviour of agentic AI coding workflows.

MSc Artificial Intelligence and Business Strategy · Supervisors: Julien Barney and Kate Sugden · Aligned to the Aston–Capgemini Centre of Excellence for Enterprise AI

Abstract

The problem it addresses

Agentic AI coding tools went from research demonstration to mainstream developer infrastructure in under three years. Enterprise adoption is outpacing the evidence needed to govern it: procurement decisions rest on vendor benchmarks that measure whether generated code passes tests, while staying silent on the properties that determine what it costs to own — security exposure, structural complexity, redundancy, and above all fidelity to the specification that was actually requested.

This dissertation designs, builds and applies an instrument that quantifies five metrics across five workflow conditions — a hand-coded human baseline and four commercial agentic tools — against three fixed specifications spanning distinct task domains. The design is pre-registered, and the analysis is blinded by construction: one capture contract forces every condition, human keystrokes and agent tool-calls alike, into a single comparable shape, and the metric code is never told which condition produced what it is scoring.

The contribution is threefold: a reusable, vendor-agnostic measurement instrument and its capture contract; a pre-registered cross-vendor comparison that foregrounds governance rather than functional success; and a framing of specification fidelity as a first-class, measurable quality metric relevant to enterprise AI adoption.

Findings

Three results worth the read

An agent overrode an unambiguous specification

Given a brief for a command-line tool, under a clean workspace with an explicit instruction prohibiting pipeline output, one commercial agent shipped a data pipeline instead. Its pretrained architectural bias beat the brief.

Tool quality depends on the task

Every testable metric showed a significant condition-by-specification interaction. The defensible claim is never "tool X beats tool Y", only "tool X beats tool Y for this kind of work".

The validation found a defect in the instrument

Two independent raters disagreed with the tool on one item, and they were right: it could not see a whole class of web framework. The erratum, the repair and the withdrawn claim are all published in the text.

The instrument is open source. So is the argument.

Everything the dissertation measures can be run on your own code, and the statistics reproduce from the public repository. The errata are in the text rather than quietly corrected, including the one that overturned a claim the study had advanced as its cleanest result.

Read the full dissertation

54 pages. Free, and yours to keep.

Include your country code.