# Data workflow — SkillDB starter bundle

Development preview · Evaluation pending

Plan a repeatable batch import, define what each row means, and make bad data visible.

For a small analytical data pipeline. Select the simplest suitable storage and runtime; no infrastructure is created.

Selection review: 2026-10-04. Selection reviewed for topic coverage and overlap. Code examples and task outcomes have not been evaluated.
Maintainer provenance: private repository latentsmurf/SkillDB, revision 0a89389d0ec3e4f12e694baab71bdb56134eeb7e. Maintainer access required; this is not a public source link. Catalog pages may change after the selection review.

## Export and source rights

Reference manifest only: selection notes, skill IDs, and public catalog links. No skill bodies, executable code, credentials, or host configuration are included. Downloading does not install or enable skills.

Source licenses have not been verified for redistribution. Catalog pages provide a public reference; access does not grant redistribution rights. Verify the original license and attribution requirements before copying or packaging skill content.

This is a reading and review plan, not an agent skill or host configuration. Start with the public catalog pages and review available instructions before use.

## Why these skills fit together

Data Pipeline Architecture covers movement and recovery, Data Modeling covers grain and relationships, and Data Quality covers validation. All discuss contracts; write one shared contract rather than maintaining three versions.

## 1. Data Pipeline Architecture

Role: Pipeline and recovery plan

Covers batch versus streaming choices, idempotency, schema changes, and recoverable failures.

Review note: The scale and latency examples are illustrative. Do not adopt streaming infrastructure or exactly-once claims without testing the actual source and sink.

Skill ID: data-engineering-skills/data-pipeline-architecture.md
Public catalog: https://skilldb.dev/skills/data-engineering-skills/data-pipeline-architecture
Maintainer provenance path (private repository; maintainer access required): packs/data-engineering-skills/data-pipeline-architecture.md

## 2. Data Modeling

Role: Row grain and relationships

Explains grain, keys, and model choices so the pipeline produces data that supports the intended questions.

Review note: Enterprise modeling approaches are alternatives, not a checklist to implement. A small workflow may need only a clearly documented table.

Skill ID: data-engineering-skills/data-modeling.md
Public catalog: https://skilldb.dev/skills/data-engineering-skills/data-modeling
Maintainer provenance path (private repository; maintainer access required): packs/data-engineering-skills/data-modeling.md

## 3. Data Quality

Role: Validation and failure visibility

Adds completeness, uniqueness, freshness, and reconciliation checks around the pipeline contract.

Review note: Choose thresholds from the dataset and business rules. The example percentages and composite scores are not validated targets for your project.

Skill ID: data-engineering-skills/data-quality.md
Public catalog: https://skilldb.dev/skills/data-engineering-skills/data-quality
Maintainer provenance path (private repository; maintainer access required): packs/data-engineering-skills/data-quality.md

## Suggested evaluation — not yet run

Import a small orders dataset into a reporting table, then repeat the import with duplicate rows, missing fields, and a late correction.

- [ ] Define row grain and a repeatable key before implementing transformations.
- [ ] Run the same import twice and verify that intended totals remain stable.
- [ ] Make rejected records, freshness, and reconciliation results visible; test a recovery or backfill.

Compare the same task with and without these references using the same environment and checks. Record outcomes before claiming an improvement.
