Decision Brief
| Item | This lesson |
|---|---|
| Decision | Determine whether the available data can support the intended workflow and control design. |
| Output | A data-readiness inventory and remediation plan. |
The Readiness Problem
“Do we have data?” is too weak a question. A team can have thousands of files and still lack usable evidence. Readiness depends on whether the data is accessible, permitted, current, representative, interpretable, and owned.
For retrieval, the critical asset may be a small set of authoritative documents. For classification, it may be historical examples with consistent operational labels. For extraction, it may be the range of document formats and the source evidence needed to validate each field.
Six Readiness Dimensions
| Dimension | Questions |
|---|---|
| Availability | Can the project access the records in the pilot environment? |
| Authorization | Is the proposed use allowed by policy, contract, consent, and role? |
| Quality | Are records complete enough and labels consistent enough for the task? |
| Representativeness | Do examples cover normal, ambiguous, rare, and high-consequence cases? |
| Traceability | Can an output be connected back to the source and version used? |
| Ownership | Who corrects, refreshes, approves, and retires the data? |
A “no” does not always mean stop. It may mean reduce scope, create a source-of-truth process, or use a human-led pilot while the data is repaired.
Data Shape Must Match the Task
- Prompted drafting: approved facts, audience, examples, and editorial constraints.
- Retrieval: authoritative sources, permissions, versioning, and citation locations.
- Classification: stable labels, boundary examples, and corrected historical decisions.
- Extraction: document variety, field definitions, source locations, and validation rules.
- Forecasting: governed historical data, assumptions, time cutoffs, and scenario ownership.
- Agent-like actions: tool permissions, system state, action logs, and recoverability.
Detect False Readiness
Common warning signs include:
- labels reflect whichever person handled the work rather than a stable policy;
- only successful or easy cases were retained;
- documents have no owner or review date;
- sensitive fields are mixed into broad exports;
- historical decisions contain bias or outdated policy;
- the team cannot explain what a missing value means;
- there is no safe way to reproduce the data used for a result.
Harborline Example
Harborline has two leave documents: a current policy and a stale FAQ. The volume of documents is not the issue. The readiness problem is source authority and conflict handling. A knowledge assistant is not ready until the policy is marked authoritative, the stale FAQ is excluded or clearly flagged, and HR owns refresh decisions.
Its support data has a different problem: historical tickets exist, but rerouted cases reveal inconsistent labels. Before training or evaluating a classifier, the team needs a corrected gold sample tied to the new queue taxonomy.
Practice: Complete the Inventory
For each data or knowledge source, record:
| Field | Required answer |
|---|---|
| Source | Name, system, and version |
| Intended use | Retrieval, classification, extraction, drafting, evaluation, or monitoring |
| Owner | Role accountable for accuracy and access |
| Sensitivity | Data classification and prohibited environments |
| Quality risk | Missing, stale, inconsistent, biased, or unrepresentative content |
| Evidence | How outputs will link back to the source |
| Remediation | Fix, exclude, narrow scope, or add human review |
| Review date | When readiness must be reassessed |
Finish with one of three decisions: ready for a narrow pilot, ready after remediation, or not suitable under the current design.