| Titre : |
Experimental benchmarking of reactive autoscaling strategies in cloud-native kubernetes deployments |
| Type de document : |
document multimédia |
| Auteurs : |
Radouane Djoudi, Auteur ; Mohamed Amine Merad, Auteur ; Mustapha Bouakkaz, Directeur de thèse |
| Editeur : |
Laghouat : Université Amar Telidji - Département d'informatique |
| Année de publication : |
2026 |
| Importance : |
52 p. |
| Accompagnement : |
1 disque optique numérique (CD-ROM) |
| Note générale : |
Option : Distributed networks, systems and applications |
| Langues : |
Anglais (eng) |
| Mots-clés : |
Kubernetes Reactive autoscaling HPA KEDA Prometheus adapter Cloudnative Latency Resource allocation QoS |
| Résumé : |
Modern cloud-native applications require precise and efficient autoscaling mechanisms to balance quality of service (QoS) with infrastructure cost. Yet distributed systems engineers have little rigorous empirical guidance for selecting and configuring the most suitable strategy under complex, variable load conditions.
This thesis presents a controlled comparative evaluation of four reactive autoscaling strategies within a Kubernetes container environment: two strategies based on physical resource metrics via the native HPA controller (CPU utilization and memory utilization), one strategy based on a custom application-level metric (active HTTP requests) exposed via the Prometheus Adapter, and one strategy based on queue depth via KEDA. These strategies were tested on a web application simulating a multi-tenant architecture, deployed on a local Kubernetes cluster, using a single composite five-phase load profile (ramp-up, sustained load, sudden spike, recovery, and ramp-down, generated via k6) applied identically to each strategy to ensure comparative rigor.
Experimental results reveal a decisive performance gap between the four strategies during the sudden-spike phase. The HPA-CPU strategy demonstrated clear superiority, achieving a reaction time of only 35 seconds, a stable P95 latency of 332 ms, and a 0.00% error rate. In comparison, HPA-Memory reacted considerably slower (3 min 53 s) with an error rate below 5%. The Prometheus Adapter–based strategy failed entirely to trigger scaling throughout the test, due to a slow metric-discovery cycle. Notably, KEDA, driven by queue depth, recorded the slowest reaction time of all (4 min 15 s) and experienced a critical queue overshoot of 86 times the configured threshold (theoretically requiring 87 replicas, while only 5 were ever deployed).
These findings highlight a fundamental trade-off between reaction speed and resource stability: strategies based on native system metrics, particularly CPU, deliver markedly faster reactivity than those relying on application-level metrics or external queues. This study offers a reproducible experimental framework and an empirical baseline to guide future predictive autoscaling approaches. |
| note de thèses : |
Mémoire de master en informatique |
Experimental benchmarking of reactive autoscaling strategies in cloud-native kubernetes deployments [document multimédia] / Radouane Djoudi, Auteur ; Mohamed Amine Merad, Auteur ; Mustapha Bouakkaz, Directeur de thèse . - Laghouat : Université Amar Telidji - Département d'informatique, 2026 . - 52 p. + 1 disque optique numérique (CD-ROM). Option : Distributed networks, systems and applications Langues : Anglais ( eng)
| Mots-clés : |
Kubernetes Reactive autoscaling HPA KEDA Prometheus adapter Cloudnative Latency Resource allocation QoS |
| Résumé : |
Modern cloud-native applications require precise and efficient autoscaling mechanisms to balance quality of service (QoS) with infrastructure cost. Yet distributed systems engineers have little rigorous empirical guidance for selecting and configuring the most suitable strategy under complex, variable load conditions.
This thesis presents a controlled comparative evaluation of four reactive autoscaling strategies within a Kubernetes container environment: two strategies based on physical resource metrics via the native HPA controller (CPU utilization and memory utilization), one strategy based on a custom application-level metric (active HTTP requests) exposed via the Prometheus Adapter, and one strategy based on queue depth via KEDA. These strategies were tested on a web application simulating a multi-tenant architecture, deployed on a local Kubernetes cluster, using a single composite five-phase load profile (ramp-up, sustained load, sudden spike, recovery, and ramp-down, generated via k6) applied identically to each strategy to ensure comparative rigor.
Experimental results reveal a decisive performance gap between the four strategies during the sudden-spike phase. The HPA-CPU strategy demonstrated clear superiority, achieving a reaction time of only 35 seconds, a stable P95 latency of 332 ms, and a 0.00% error rate. In comparison, HPA-Memory reacted considerably slower (3 min 53 s) with an error rate below 5%. The Prometheus Adapter–based strategy failed entirely to trigger scaling throughout the test, due to a slow metric-discovery cycle. Notably, KEDA, driven by queue depth, recorded the slowest reaction time of all (4 min 15 s) and experienced a critical queue overshoot of 86 times the configured threshold (theoretically requiring 87 replicas, while only 5 were ever deployed).
These findings highlight a fundamental trade-off between reaction speed and resource stability: strategies based on native system metrics, particularly CPU, deliver markedly faster reactivity than those relying on application-level metrics or external queues. This study offers a reproducible experimental framework and an empirical baseline to guide future predictive autoscaling approaches. |
| note de thèses : |
Mémoire de master en informatique |
|  |