Context & Motivation
Modern distributed storage infrastructures are made up of thousands of hard disk drives (HDDs) operating continuously. While compute and network energy costs are increasingly well understood, HDD/SSD power consumption under real-world workloads remains difficult to measure and predict accurately. Understanding and modeling this consumption is critical for:
- Reducing the environmental footprint of large-scale storage systems
- Improving capacity planning and thermal management
- Enabling proactive power optimization in production environments
To date, no robust, interpretable multi-parameter model exists that can accurately estimate disk energy consumption from observable parameters in a production environment. This internship aims to fill that gap.
Your Mission
As part of our R&D team, you will design, build, and validate a multi-parameter model for estimating HDD/SSD power consumption. Your work will include:
1. State of the Art
Review existing literature and open-source projects related to HDD/SSD power modeling, thermal behavior, and energy measurement in storage systems (including tools such as PowerAPI).
2. Physical Modeling
Develop a physics-based multi-parameter model of temperature and power draw, capturing relationships between:
- Drive temperature
- Fan speed and airflow
- Read/write throughput and IOPS
- Idle vs. active state transitions
3. Interpretable Machine Learning Model
Design a multi-parameter ML model (e.g., gradient boosting, linear regression with feature engineering) that is both accurate and interpretable — enabling engineers to understand which parameters drive consumption under different workload profiles.
Instrument real drives in a controlled laboratory environment and under
production-representative workloads to collect ground-truth measurements.
- Calibrate and validate the model
- Quantify model accuracy across workload types
- Identify parameters with the greatest predictive value
Expected Outcomes & Valorization
- A scientific publication or technical white paper
- Integration into the open-source PowerAPI project
- Direct integration into Scality's internal monitoring and capacity planning tooling
Technical Stack
- Python (primary language for data collection, modeling, and analysis)
- scikit-learn (ML modeling and evaluation)
- PowerAPI ecosystem (https://powerapi.org/)
- Linux system tooling for hardware instrumentation (smartctl, lm-sensors, etc.)
- Jupyter Notebooks for exploratory analysis and result visualization
Candidate Profile
- Final-year student in a Master's program or Engineering school (Bac+5)
- Strong interest in physical modeling and/or applied machine learning
- Comfortable working with real hardware and experimental data
- Autonomous, curious, and rigorous in your approach to problem-solving
- Able to communicate results clearly in written and spoken English
Work Environment
- Mentorship from senior R&D engineers with expertise in distributed systems and performance engineering
- Access to real production-grade storage hardware and laboratory infrastructure
- A high-trust environment where your findings will directly influence engineering decision
Why Join Scality?
- Work on a concrete research problem with real-world industrial impact
- Contribute to open-source energy-efficiency tooling used beyond Scality
- Be part of a team building infrastructure trusted by Fortune 500 companies
- Potential to publish research or continue as a full-time engineer after the internship