The Casper project investigates how the execution of large scalable batch processing applications can be aligned with the availability of low-carbon energy. It is guided by the vision of Stepwise Performance-, Interruption-, Resource-, and Carbon-Aware Schedulers (SPIRCS) for scalable batch data processing on elastic compute clusters.
We develop methods for profiling and predicting the performance, scalability, interruption behaviour, resource requirements, energy consumption, and carbon emissions of individual data processing steps, as well as for their profile-informed, carbon-aware execution. We implement these methods in prototypes for the Spark dataflow and Nextflow workflow systems as well as Kubernetes clusters. We assess the prototypes in experiments on public- and private-cloud infrastructure.
Alongside this technical work, Casper aims to raise awareness of cloud application carbon footprints and contributes to wider research on demand-side management, including exploring the relationship between spare cloud capacity and the availability of low-carbon energy.
The project’s partners provide perspectives on complementary routes to real-world impact: AWS as a major public-cloud provider, the BBC as a significant cloud user, and HU Berlin as a provider of private-cloud infrastructure and expertise in scientific workflows.