Decision Brief

Item This lesson
Decision Determine whether the available data can support the intended workflow and control design.
Output A data-readiness inventory and remediation plan.

The Readiness Problem

“Do we have data?” is too weak a question. A team can have thousands of files and still lack usable evidence. Readiness depends on whether the data is accessible, permitted, current, representative, interpretable, and owned.

For retrieval, the critical asset may be a small set of authoritative documents. For classification, it may be historical examples with consistent operational labels. For extraction, it may be the range of document formats and the source evidence needed to validate each field.

Six Readiness Dimensions

Dimension Questions
Availability Can the project access the records in the pilot environment?
Authorization Is the proposed use allowed by policy, contract, consent, and role?
Quality Are records complete enough and labels consistent enough for the task?
Representativeness Do examples cover normal, ambiguous, rare, and high-consequence cases?
Traceability Can an output be connected back to the source and version used?
Ownership Who corrects, refreshes, approves, and retires the data?

A “no” does not always mean stop. It may mean reduce scope, create a source-of-truth process, or use a human-led pilot while the data is repaired.

Data Shape Must Match the Task

  • Prompted drafting: approved facts, audience, examples, and editorial constraints.
  • Retrieval: authoritative sources, permissions, versioning, and citation locations.
  • Classification: stable labels, boundary examples, and corrected historical decisions.
  • Extraction: document variety, field definitions, source locations, and validation rules.
  • Forecasting: governed historical data, assumptions, time cutoffs, and scenario ownership.
  • Agent-like actions: tool permissions, system state, action logs, and recoverability.

Detect False Readiness

Common warning signs include:

  • labels reflect whichever person handled the work rather than a stable policy;
  • only successful or easy cases were retained;
  • documents have no owner or review date;
  • sensitive fields are mixed into broad exports;
  • historical decisions contain bias or outdated policy;
  • the team cannot explain what a missing value means;
  • there is no safe way to reproduce the data used for a result.

Harborline Example

Harborline has two leave documents: a current policy and a stale FAQ. The volume of documents is not the issue. The readiness problem is source authority and conflict handling. A knowledge assistant is not ready until the policy is marked authoritative, the stale FAQ is excluded or clearly flagged, and HR owns refresh decisions.

Its support data has a different problem: historical tickets exist, but rerouted cases reveal inconsistent labels. Before training or evaluating a classifier, the team needs a corrected gold sample tied to the new queue taxonomy.

Practice: Complete the Inventory

For each data or knowledge source, record:

Field Required answer
Source Name, system, and version
Intended use Retrieval, classification, extraction, drafting, evaluation, or monitoring
Owner Role accountable for accuracy and access
Sensitivity Data classification and prohibited environments
Quality risk Missing, stale, inconsistent, biased, or unrepresentative content
Evidence How outputs will link back to the source
Remediation Fix, exclude, narrow scope, or add human review
Review date When readiness must be reassessed

Finish with one of three decisions: ready for a narrow pilot, ready after remediation, or not suitable under the current design.

Continue Learning