Write your awesome label here.
Design a data estate against carbon, energy, water and waste — from sourcing through to disposal.
Thirteen modules following the data lifecycle: source, model, store, process, place, serve, index, retain, dispose and telemetry — each measured against SCI, and each ending in the design gate that stage cannot be signed off without.
10h 40m
13
30 questions
80%
Credential
Green Data Practitioner
How the course works
One lifecycle, one expression
SCI runs through all thirteen modules, with E split into stored, moved and scanned, I as the carbon-aware term, M as media and servers, and R as the functional unit you choose.
A design gate at every stage
Each of Modules 3 to 11 ends with the question to require before sign-off, and Module 12 assembles the nine into a review artefact.
Patterns with the trade stated
Around fifty catalog entries, each naming the term it moves and whether it saves once or on every hour, query or run.
What this course helps with
See where a data estate actually costs
Stored, moved and scanned bytes run on three different clocks, and the mechanism that dominates a service is frequently not the one it is named after.
Design it in rather than remediate it
Almost every expensive data decision is cheap when it is taken and paid at a later stage — the course puts the question before the decision.
Cover what the estate-level courses leave short
Warehouse layout, index and vector economics, retention authority and telemetry get a module each rather than a section.
Who this is for
- Data engineers and analytics engineers
- Database, warehouse and platform engineers
- Data architects and technical leads responsible for a data estate
- ML and AI engineers who own retrieval, feature or vector stores
- Assumes working familiarity with data systems; no sustainability background required
Related courses:
For estate-wide architecture across six layers, see Green and Efficient Architect; for training-side model decisions, see Green AI for ML Practitioners; for producing the underlying figures, see SCI for AI Practitioner.
What you will work with
The three mechanisms
Why a small hot table can cost more than a large cold archive, and which lever works on each.
The copy count
One 1.2 TB warehouse, nine physical copies, and the policies that created them separately.
The Green Data Blueprint
Four blocks — assessment, measurement plan, pattern register, governance wiring — bound by a traceability test.
What you will be able to do
Course Lessons
The full course
01
Where ML’s Footprint Is Decided
A model’s footprint is created twice: once when it is trained, and again on every inference it serves. This module sets the decision metric and the amortisation rule that every later module is judged against.
module-01
02
Reuse, Adapt, or Train: The First Decision
The largest reduction available on the training side is not to train, yet the question is usually settled by habit before anyone computes it. This module compares six routes to a working model and what to ask of a base model before inheriting its cost.
module-02
03
Sizing the Model and the Data Budget
How large a model, trained on how much data, sets the energy term more directly than any later choice. This module covers compute-optimal sizing and how architecture changes training and serving energy independently of parameter count.
module-03
04
Data Efficiency: What the Model Is Trained On
What the training tokens actually are changes both how many are needed and how good the resulting model is. This is the layer where the footprint lever and the quality lever most often point in the same direction, and this module is careful about where that alignment breaks.
module-04
05
The Training Run: Execution, Stopping and Waste
This module examines the compute that produces nothing: epochs past convergence, crashes without a checkpoint, and experiments that repeat work already done. Waste at this layer is invisible in a cost report, because a failed run and a successful one look identical on the bill.
module-05
06
Search and Evaluation: The Two Multipliers
Earlier modules reduce the cost of one run; this one addresses how many runs happen, which is usually the larger number. It covers search budgets, comparing strategies by cost per answer, and evaluation as a cost that recurs for the life of the model.
module-06
07
Hardware-Aware Training
Compute demanded is one quantity; how efficiently it becomes energy is another. This module covers precision, batch size, accelerator fit and utilisation as the conversion factor sitting between the two.
module-07
08
When and Where to Train
This is the one lever that changes carbon without changing energy, by moving a run to a different hour or region. Training is often a strong candidate, and this module covers the constraints that decide it, along with water, preemptible capacity and embodied allocation.
module-08
09
From Weights to Deployable Artifact
This module deliberately spends training cost once in order to reduce serving cost on every inference that follows. It covers parameter-efficient adaptation, post-training and compression, and the repayment calculation that decides whether the trade is worth making.
module-09
10
Lifecycle, Retirement and Disclosure
Retraining cadence multiplies the training footprint, and retirement is the point at which every amortisation estimate becomes a measured fact. This closing module covers cadence, drift, registry and model card records, disclosure duties, and the capstone footprint report.
module-10
11
Final Assessment
A final assessment of 30 scenario-judgement questions - 80% to pass, 60 minutes, retake available.
Assessment and Credential
Time to complete
Join the Pilot
The GSF Academy is currently in pilot. If your organization would like to enrol learners and issue GSF certifications to your teams, tell us a little about what you need and we'll follow up to discuss access, pricing, and enrolment.
Thank you.
We will be in touch shortly
We will be in touch shortly
