Adaptive CPU Scheduling using a Q-Learning Meta-Scheduler
DOI:
https://doi.org/10.65091/icicset.v3i1.96Abstract
Conventional CPU scheduling policies such as FirstCome First-Served (FCFS), Shortest Job First (SJF), Round Robin
(RR), and Priority scheduling apply fixed, pre-defined rules that
are optimal only under specific workload conditions and cannot
adapt to dynamic and heterogeneous execution environments. This
paper presents an adaptive meta-scheduler that employs tabular
Q-learning to dynamically select among classical scheduling
policies based on the runtime state of the system. A discreteevent CPU simulation environment models process lifecycle
behaviour, and a Gymnasium-compatible reinforcement learning
(RL) environment exposes a discretised five-dimensional state
representation, a four-action policy-selection space, and a reward
function that penalises accumulated waiting and turnaround
time while rewarding process completions and CPU utilisation.
The agent is evaluated on four workload profiles—CPU-bound,
I/O-bound, bursty, and mixed—generated from Poisson arrivals
and exponential or uniform burst distributions. Experimental
results show that the agent learns a meaningful state-dependent
policy, exhibiting a clear preference for Round Robin under highvariance workloads. Coarsening the state representation from
3,125 to 243 discrete states raises state-space coverage from 2.34%
to 12.35% for a fixed training budget, and the resulting policy
remains competitive with the strongest static baseline on bursty
and CPU-bound workloads, with a gap of 5.6% and 13.8% in
average waiting time respectively. The framework is modular
and reproducible, with a full-stack implementation comprising a
FastAPI backend, PostgreSQL persistence, and a React dashboard.