Start with a safe test environment

Use an isolated learning environment containing no real customer data. Record normal application behavior and keep a cleanup plan.

Choose one reversible failure

Select a small failure you understand, such as stopping a test application process. Write down what you expect the user to see and which signal should reveal the interruption. Avoid changing several things at once.

Observe before you repair

Capture the time, the symptom and the relevant application logs. Distinguish an observation from your explanation of its cause. A failing page is evidence; its cause remains a hypothesis until you check it.

Restore service

Use your documented recovery step. Confirm that the application responds again and that expected data remains available. Record the action and recovery time.

Explain what you learned

Write a short incident note: expected behavior, observed behavior, cause, recovery and prevention idea. Ask another learner to follow your explanation, then revise the deployment plan.

A successful lab produces repeatable evidence. It does not establish that a production service can recover from every failure.

This original lesson is part of the Primus Learning live product-verification pilot.