Back to InsightsData Annotation

    Data Annotation Projects: How to Manage Them Without Losing Control

    Data annotation is often treated as a mechanical precursor to machine learning. In practice, it is one of the most consequential phases of any applied AI programme.

    5 February 202610 min readData Understood
    Data Annotation Projects: How to Manage Them Without Losing Control

    Data annotation is often treated as a mechanical precursor to machine learning. In practice, it is one of the most consequential phases of any applied AI programme.

    Industry research shows that a significant proportion of machine learning effort is spent before any model is trained. A widely cited MIT study found that data scientists can spend up to 80 percent of their time on data preparation and labelling rather than modelling itself.

    When annotation efforts struggle, it is rarely due to lack of effort or tooling. More often, issues arise because the nature of the task, the roles involved, and the way quality is managed were not fully aligned from the outset.

    From working with teams delivering annotation at scale, certain principles consistently separate effective programmes from those that introduce downstream risk.

    Recognising when annotation is judgement-heavy

    Not all annotation tasks are the same.

    Where labels are objective and unambiguous, annotation can be treated as a throughput exercise. Where interpretation, uncertainty, or contextual judgement are involved, the work becomes fundamentally different.

    This distinction is especially important in regulated and data-intensive environments such as financial services, healthcare, and the public sector, where annotation decisions can directly affect downstream model behaviour, explainability, and trust.

    In judgement-heavy tasks:

    • Annotators require clear guidance and concrete examples
    • Disagreement is expected, particularly in early phases
    • Speed emerges gradually as shared understanding develops

    Planning annotation as if it were purely mechanical often leads to hidden quality issues that surface later as unexplained model behaviour or loss of confidence in AI outputs.

    This is why annotation quality must be treated as part of a broader data foundation, alongside capabilities such as data cleaning and data integration.

    Being explicit about roles and responsibilities

    Well structured annotation programmes are clear about who does what.

    Typical roles include:

    • Annotators, who execute the work to defined guidance
    • Domain experts, who define what correctness means
    • Quality functions, which measure consistency and agreement
    • Delivery management, which controls pace and escalation

    When these responsibilities are implicit or overlapping, inconsistency and rework tend to follow.

    Clear ownership boundaries are a form of governance, not bureaucracy. They sit naturally alongside broader data governance practices and help ensure that annotation decisions remain consistent as volume grows.

    Defining and measuring quality early

    Quality is harder to measure than volume, but far more important.

    Effective annotation programmes agree early on:

    • How quality will be measured
    • What acceptable agreement looks like
    • How reference examples will be used

    Research into human labelling tasks shows that inter-annotator agreement can vary significantly even among experts, particularly when tasks involve interpretation rather than objective classification.

    Agreement metrics and sampling approaches provide a shared language for discussing quality. Without them, disagreements quickly become subjective and difficult to resolve.

    Importantly, quality should stabilise before teams consider increasing scale.

    Allowing for ambiguity in the data

    In real-world datasets, some items are genuinely unclear.

    Studies of human annotation consistently show that a meaningful proportion of data points contain legitimate ambiguity, even when clear guidelines are provided.

    Forcing confident labels in ambiguous cases can damage both data quality and downstream models.

    Mature annotation workflows explicitly allow for:

    • Ambiguous or unclear outcomes
    • Escalation paths for difficult cases
    • Throughput measured on items processed, not only confidently labelled outputs

    Treating uncertainty as a valid outcome improves consistency and leads to more trustworthy training data.

    Planning for iteration rather than perfection

    Annotation guidance rarely starts perfect.

    Strong programmes expect an initial period where:

    • Guidance is refined
    • Examples are clarified
    • Edge cases are surfaced and discussed

    This stabilisation phase is essential. It reduces pressure on annotators and leads to more consistent outcomes over time.

    From an infrastructure perspective, this phase benefits from robust data infrastructure that supports versioning, auditability, and traceability of annotation decisions.

    Focusing on capacity and learning over fixed outcomes

    For judgement-heavy tasks, it is often more effective to plan around capacity and learning rather than fixed output targets.

    This approach allows teams to:

    • Adjust guidance based on evidence
    • Improve consistency before scaling
    • Maintain alignment between scientific intent and delivery reality

    It also reduces the risk of prioritising speed over quality when the two are in tension.

    Organisations that take this approach tend to see better downstream results in analytics, visualisation, and AI adoption, where trust in the underlying data is critical.

    A practical example in financial services

    Consider a financial services organisation building a model to classify customer interactions for conduct risk monitoring.

    At first glance, the annotation task appears straightforward. Historical call transcripts and digital communications are labelled to indicate whether they contain potential conduct issues such as mis-selling, vulnerable customer indicators, or unsuitable advice.

    In practice, the task is judgement-heavy.

    Many interactions sit in grey areas. A customer may express dissatisfaction without making a complaint. An adviser may follow policy but phrase something ambiguously. Context and intent matter.

    If this work is treated as commodity labelling, teams often push volume early. Disagreement between annotators is high, quality metrics fluctuate, and model performance becomes unpredictable.

    A more effective approach recognises annotation as expert-guided work. Clear guidance is developed with compliance and risk teams. Early disagreement is expected and used to refine definitions. Quality is measured consistently before scale is introduced.

    The result is slower progress in the early stages, but higher quality data and more reliable models that are easier to explain to regulators and internal stakeholders.

    Closing thoughts

    Data annotation sits at the intersection of human judgement, data quality, and delivery pressure.

    Treating it as a governed process rather than a mechanical step is key to producing reliable training data and trustworthy AI systems.

    The principles above are not complex, but they require discipline to apply. Teams that invest early in clear ownership, explicit quality definitions, and realistic delivery assumptions consistently see better outcomes downstream.

    If you would like to discuss how Data Understood supports organisations with data foundations, governance, and AI-ready datasets, you can explore our services or get in touch directly.

    For organisations interested in understanding data culture and decision-making readiness alongside technical capability, our DIAlog diagnostic provides a structured starting point.

    Data AnnotationMachine LearningData QualityAI GovernanceEnterprise Data

    About Data Understood

    Data Understood is a Data and AI consultancy based in Dundee, working with ambitious organisations across Scotland and the UK. Our articles are written from work delivered with clients.

    More about Data Understood

    Common questions

    Common questions on this topic

    Practical answers to the questions leadership teams ask us about AI readiness and delivery.

    Where to go next

    Related services, sectors and client work

    If this article reflects where your organisation is, these pages cover the same ground in delivery terms.

    Let's talk about your Data & AI priorities.

    Whether you're developing a data strategy, modernising your data platform, preparing for AI or tackling a specific business challenge, start with a 30-minute conversation with Inez.

    Free. 30 minutes. Directly with Inez. No pitch.