Write your awesome label here.
Data

Design a data estate against carbon, energy, water and waste — from sourcing through to disposal.

Thirteen modules following the data lifecycle: source, model, store, process, place, serve, index, retain, dispose and telemetry — each measured against SCI, and each ending in the design gate that stage cannot be signed off without.
10h 40m
of learning
13
modules
30 questions
final assessment
80%
to pass
Credential
on completion

Green Data Practitioner

What
How
Tools
Who

How the course works

One lifecycle, one expression

SCI runs through all thirteen modules, with E split into stored, moved and scanned, I as the carbon-aware term, M as media and servers, and R as the functional unit you choose.

A design gate at every stage

Each of Modules 3 to 11 ends with the question to require before sign-off, and Module 12 assembles the nine into a review artefact.

Patterns with the trade stated

Around fifty catalog entries, each naming the term it moves and whether it saves once or on every hour, query or run.

What this course helps with

See where a data estate actually costs

Stored, moved and scanned bytes run on three different clocks, and the mechanism that dominates a service is frequently not the one it is named after.

Design it in rather than remediate it

Almost every expensive data decision is cheap when it is taken and paid at a later stage — the course puts the question before the decision.

Cover what the estate-level courses leave short

Warehouse layout, index and vector economics, retention authority and telemetry get a module each rather than a section.

Who this is for

  • Data engineers and analytics engineers

  • Database, warehouse and platform engineers

  • Data architects and technical leads responsible for a data estate

  • ML and AI engineers who own retrieval, feature or vector stores

  • Assumes working familiarity with data systems; no sustainability background required

Related courses:

For estate-wide architecture across six layers, see Green and Efficient Architect; for training-side model decisions, see Green AI for ML Practitioners; for producing the underlying figures, see SCI for AI Practitioner.

What you will work with

The three mechanisms

Why a small hot table can cost more than a large cold archive, and which lever works on each.

The copy count

One 1.2 TB warehouse, nine physical copies, and the policies that created them separately.

The Green Data Blueprint

Four blocks — assessment, measurement plan, pattern register, governance wiring — bound by a traceability test.

What you will be able to do

  • Separate stored, moved and scanned bytes as three cost mechanisms, and identify which dominates a service
  • Write the SCI expression for a data service and choose a functional unit that changes when the decision changes
  • Rate a data estate by mechanism with evidence levels, and record UNKNOWN as a measurement action
  • Design schema, partitioning and clustering against the queries a table will actually receive
  • Count every physical copy a storage policy creates, and tier retention and DR posture by criticality
  • Replace full reloads with incremental processing, and bound retries, idle clusters and zombie pipelines
  • Place flexible data work in low-intensity windows, and report an intensity reduction as distinct from an efficiency gain
  • Read query history for the scanning that dashboards and scheduled jobs generate without a person attached
  • Size an embedding estate, decide index rebuild cadence, and test quantisation against a stated recall bar
  • Measure dark data, establish the authority to delete, and verify that deletion reaches every copy
  • Set telemetry verbosity, sampling, cardinality budgets and retention tiers from measured read rates
  • Assemble the nine design gates into a review, contracts and provider asks, and a Green Data Blueprint

Course Lessons

The full course

Every module ends with an application exercise and a worked solution.

01

Where ML’s Footprint Is Decided

A model’s footprint is created twice: once when it is trained, and again on every inference it serves. This module sets the decision metric and the amortisation rule that every later module is judged against.
module-01

02

Reuse, Adapt, or Train: The First Decision

The largest reduction available on the training side is not to train, yet the question is usually settled by habit before anyone computes it. This module compares six routes to a working model and what to ask of a base model before inheriting its cost.
module-02

03

Sizing the Model and the Data Budget

How large a model, trained on how much data, sets the energy term more directly than any later choice. This module covers compute-optimal sizing and how architecture changes training and serving energy independently of parameter count.
module-03

04

Data Efficiency: What the Model Is Trained On

What the training tokens actually are changes both how many are needed and how good the resulting model is. This is the layer where the footprint lever and the quality lever most often point in the same direction, and this module is careful about where that alignment breaks.
module-04

05

The Training Run: Execution, Stopping and Waste

This module examines the compute that produces nothing: epochs past convergence, crashes without a checkpoint, and experiments that repeat work already done. Waste at this layer is invisible in a cost report, because a failed run and a successful one look identical on the bill.
module-05

06

Search and Evaluation: The Two Multipliers

Earlier modules reduce the cost of one run; this one addresses how many runs happen, which is usually the larger number. It covers search budgets, comparing strategies by cost per answer, and evaluation as a cost that recurs for the life of the model.
module-06

07

Hardware-Aware Training

Compute demanded is one quantity; how efficiently it becomes energy is another. This module covers precision, batch size, accelerator fit and utilisation as the conversion factor sitting between the two.
module-07

08

When and Where to Train

This is the one lever that changes carbon without changing energy, by moving a run to a different hour or region. Training is often a strong candidate, and this module covers the constraints that decide it, along with water, preemptible capacity and embodied allocation.
module-08

09

From Weights to Deployable Artifact

This module deliberately spends training cost once in order to reduce serving cost on every inference that follows. It covers parameter-efficient adaptation, post-training and compression, and the repayment calculation that decides whether the trade is worth making.
module-09

10

Lifecycle, Retirement and Disclosure

Retraining cadence multiplies the training footprint, and retirement is the point at which every amortisation estimate becomes a measured fact. This closing module covers cadence, drift, registry and model card records, disclosure duties, and the capstone footprint report.
module-10

11

Final Assessment

A final assessment of 30 scenario-judgement questions - 80% to pass, 60 minutes, retake available.

Assessment and Credential

How you are assessed

  • A short knowledge check after every module, drawn at random from that module’s pool
  • A final assessment of 30 scenario-judgement questions from a pool of 130+
  • 80% to pass, 60 minutes, retake available
  • Questions test judgement in situations you will meet — not recall of figures that date

What you earn

GSF Certified
Green Data Practitioner
  • Issued on passing the final assessment
  • Valid for 24 months, refreshed by a short update module
  • Shareable credential for your profile, CV and tender responses

Time to complete

Reading and exercises
10h 40m across 13 modules
Knowledge checks
About 5 minutes per module
Final assessment
60 minutes
Typical completion
2–3 weeks at a module per session, entirely self-paced

Join the Pilot

The GSF Academy is currently in pilot. If your organization would like to enrol learners and issue GSF certifications to your teams, tell us a little about what you need and we'll follow up to discuss access, pricing, and enrolment.

Request Access for Your Organization

First name
Last name
Work email
Tell us about your organisation
Thank you.

We will be in touch shortly