Hungry Bird
Play Flappy Bird yourself, then watch a Q-Learning AI start from zero โ dying on the first pipe โ and slowly master the game through pure trial and error!
How RL Masters Games
State
What the agent observes: bird height, vertical speed, distance to next pipe gap, gap position.
Actions
Just two: flap or don't flap. Simple actions, but the timing makes all the difference!
Rewards
+1 per frame alive, +10 per pipe passed, -100 for collision. The agent learns to maximise total reward.
Q-Table
Maps every (state, action) pair to an expected reward. Updates after every frame using the Bellman equation.
Step 1 โ You Play First
Click canvas or press Space to flap๐ฎ Your Score
Score as high as you can! The AI starts knowing nothing โ it will fail over and over. But after hundreds of tries, it will beat your score easily. That's the power of RL!
Step 2 โ AI Training
๐ค AI Stats
๐ Score per Episode
Step 3 โ Full Learning Curve
Step 4 โ Design the Reward Function
๐๏ธ Reward Values
๐๏ธ Reward Training
๐ข Alive only: alive=5, pipe=0, crash=-10
๐ต Pipe focus: alive=0, pipe=50, crash=-100
๐ด Survival: alive=2, pipe=10, crash=-200
๐ Reward Training Progress
Game AI Badge!
You trained a Q-Learning agent to play Hungry Bird from scratch!
Optional. Stays on this device only โ not sent to WhizzStep.
Key Concepts Mastered
๐ What the Agent Sees
The discretised observation: bird height bucket, velocity bucket, horizontal distance to gap, gap position bucket.
๐ฏ Designing the Goal
The reward function defines what the agent optimises. A poorly designed reward leads to unexpected, often hilarious, behaviour.
๐ฒ Try vs Use
Epsilon-greedy: high ฮต early = try random actions. Low ฮต later = use learned knowledge. The schedule matters!
๐ Getting Better
When the Q-table stabilises and scores plateau at a high level, the agent has converged to a near-optimal policy.
๐ Real Examples
DeepMind's AlphaGo used RL + MCTS to beat the world Go champion. OpenAI Five beat world champions at Dota 2.
โณ Delayed Feedback
In games like Go, you only find out if you won at the very end โ no reward for 200+ moves. Hard for RL to handle!
About this lab
Learning objective: Guide a simulated bird toward food and see how a simple reward-driven strategy changes its path.
What this simplifies: A small game-like simulation is used to make reward-driven behaviour visible and intuitive.
Privacy: No learner input leaves the device.
Teacher prompt: Ask the class why this simulation might mislead someone who takes it too literally.
Reflect: What is one thing this activity showed you that you did not expect?
โ Back to all Labs