Skip to main content

Development preview · Evaluation pending

Data workflow starter bundle

Plan a repeatable batch import, define what each row means, and make bad data visible.

For a small analytical data pipeline. Select the simplest suitable storage and runtime; no infrastructure is created.

Download the reference plan

Reference manifest only: selection notes, skill IDs, and public catalog links. No skill bodies, executable code, credentials, or host configuration are included. Downloading does not install or enable skills.

Markdown is a reading plan; JSON is a SkillDB reference manifest. Neither is a host-specific install package.

Why these skills fit together

Data Pipeline Architecture covers movement and recovery, Data Modeling covers grain and relationships, and Data Quality covers validation. All discuss contracts; write one shared contract rather than maintaining three versions.

  1. 1. Pipeline and recovery plan

    Data Pipeline Architecture

    Covers batch versus streaming choices, idempotency, schema changes, and recoverable failures.

    Review note: The scale and latency examples are illustrative. Do not adopt streaming infrastructure or exactly-once claims without testing the actual source and sink.

    Data Engineering · Open catalog reference

  2. 2. Row grain and relationships

    Data Modeling

    Explains grain, keys, and model choices so the pipeline produces data that supports the intended questions.

    Review note: Enterprise modeling approaches are alternatives, not a checklist to implement. A small workflow may need only a clearly documented table.

    Data Engineering · Open catalog reference

  3. 3. Validation and failure visibility

    Data Quality

    Adds completeness, uniqueness, freshness, and reconciliation checks around the pipeline contract.

    Review note: Choose thresholds from the dataset and business rules. The example percentages and composite scores are not validated targets for your project.

    Data Engineering · Open catalog reference

Try a representative task

Import a small orders dataset into a reporting table, then repeat the import with duplicate rows, missing fields, and a late correction.

  • Define row grain and a repeatable key before implementing transformations.
  • Run the same import twice and verify that intended totals remain stable.
  • Make rejected records, freshness, and reconciliation results visible; test a recovery or backfill.

This evaluation has not been run. Compare the same task with and without these references using the same environment and checks, and record outcomes before claiming an improvement.

Selection review: . Selection reviewed for topic coverage and overlap. Code examples and task outcomes have not been evaluated. Links open public catalog pages, which may change. Downloads retain the reviewed commit as maintainer provenance; the source repository is private and requires maintainer access.

Source licenses have not been verified for redistribution. Catalog pages provide a public reference; access does not grant redistribution rights. Verify the original license and attribution requirements before copying or packaging skill content.