Lessons in reliability, load testing, observability and engineering leadership from supporting three Black Friday and Cyber Monday events.
Some of my clearest memories from Groupon are not of writing code or reviewing an architecture document.
They are of late nights in the office during Black Friday and Cyber Monday—sitting with other engineers, watching traffic climb on our dashboards and waiting for the peak.
Weeks of preparation had led to those hours. Runbooks had been reviewed. Load tests had been executed and questioned. Capacity assumptions had been challenged. Dashboards and alerts had been improved. Plan B and Plan C had been discussed.
Then came the waiting.
There was always a degree of tension because production systems do not necessarily behave according to a plan. Something could slow down or fail abruptly. At the same time, there was confidence—not because we expected everything to behave perfectly, but because considerable thought had gone into preparing for the unexpected.
Those nights were demanding, but they were also some of the most memorable and enjoyable parts of the work. There was anticipation, focus and a shared sense of responsibility. Tracking the traffic together and waiting for the peak created a kind of camaraderie that is difficult to reproduce during an ordinary working day.
Over the years, I supported three Black Friday and Cyber Monday events at Groupon. The systems I worked on changed, but each event taught me something new about reliability and the responsibilities of a senior engineer.
Three events, three perspectives
I started with Orders and Payments.
These systems sit close to a customer’s ability to complete a transaction. Their technical health is directly connected to the customer experience and the business outcome. A platform can appear broadly available while a problem in one critical part of the purchase journey still creates significant impact.
I later spent a year working on Card Linked Offers. It presented a different operational context, with a different set of flows and dependencies. However, the underlying responsibility remained the same: understand the system well enough to anticipate where pressure might appear and prepare the team for situations in which reality differed from the plan.
For the following couple of years, I worked on the Groupon API—an orchestration and proxy layer carrying critical platform traffic.
Working at that layer provided a broader perspective. An orchestration layer sits between consumers and multiple downstream capabilities. A problem originating elsewhere can surface through it, while its own behaviour can influence a significant part of the overall customer journey.
Moving between these areas taught me that reliability cannot be treated as an isolated property of one service.
A customer request crosses application boundaries. It encounters networks, dependencies, infrastructure and several layers of software. Customers do not experience those components separately. They experience one journey.
That makes architecture important, but it makes context equally important.
Black Friday was a business event before it was an engineering event
One of the most important lessons I learned was that Black Friday and Cyber Monday were not primarily infrastructure exercises.
They were major business events with significant engineering consequences.
As engineers, it was natural to begin by thinking about capacity, latency, Kubernetes, services and dashboards. But those discussions only made sense after understanding what mattered to the business and the product.
Before deciding where to increase capacity or concentrate monitoring, we needed to understand questions such as:
- Which customer journeys were most important during the event?
- Where could additional traffic produce the greatest impact?
- Which dependencies required the closest attention?
- If a trade-off became necessary, what needed to be protected first?
- Which failures could be tolerated temporarily, and which could not?
As a senior engineer, I could not remain confined to a narrow technical view. I needed enough product and business context to connect an engineering decision with its consequences.
The objective was not simply to keep every graph green. It was to protect the important customer journeys and ensure that the team could detect and respond to meaningful degradation quickly.
Business and product priorities gave direction to engineering priorities.
The event started early in the fourth quarter
Black Friday readiness did not begin during Black Friday week. Preparation usually started early in the fourth quarter.
Traffic during the event could rise to several times its normal level. Preparing for that increase required more than multiplying an existing capacity number.
Systems change during a year. Infrastructure evolves. Dependencies change. Traffic patterns shift. Teams modify services and deploy new capabilities. A system that handled the previous year’s event successfully could not automatically be assumed to behave the same way again.
Preparation included reviewing the stability and capacity of the Kubernetes stack, examining latency across availability zones, planning load tests, improving Grafana dashboards and alerts, and updating runbooks and fallback options.
The process forced us to make assumptions explicit:
- What traffic pattern were we preparing for?
- Which components were most likely to become constrained?
- How would the Kubernetes stack behave under sustained pressure?
- Could latency across availability zones become significant?
- How quickly would we detect an abnormal change?
- What would happen if the expected scaling path was insufficient?
- What was Plan B?
- If Plan B did not work, what was Plan C?
A calm peak event is rarely the result of a calm preparation period. Behind it are weeks of reviews, experiments, disagreements and decisions about risks that may never materialize.
Much of this work becomes invisible when it succeeds. That invisibility is often the result we are trying to achieve.
Load testing was engineering work—not a checkbox
Load testing was essential, but it was never as simple as repeating the previous year’s test.
The platform evolved, and so did the tools and processes used to generate and analyse load. In some years, parts of the load-testing setup had to be rebuilt from scratch.
That increased the preparation time considerably. It could be frustrating to recreate something that had worked before, but it also prevented us from relying on outdated assumptions.
The previous year’s successful test represented the previous year’s system.
A useful load test is not merely one that produces a large number of requests. It should help answer meaningful questions:
- Does the system degrade gradually or abruptly?
- Which resource or dependency becomes constrained first?
- Do our dashboards reflect what the platform is actually experiencing?
- Are we testing an isolated component or the important end-to-end path?
- Does the test still represent the current architecture and traffic pattern?
- What does the result imply for capacity and fallback planning?
A graph going up while everything remains green may be reassuring, but it is not the only measure of a useful test.
The more valuable outcome is a better understanding of where assumptions stop being true.
This experience also taught me that load-testing infrastructure should be treated as a maintained engineering capability. Test scenarios, assumptions and tooling become less useful when they are allowed to decay between major events.
The real output of a load test is not only a result. It is an improved mental model of the system.
Observability improved one event at a time
Grafana dashboards, graphs and alerts improved year after year.
More graphs did not automatically provide more clarity. During peak traffic, engineers need to distinguish normal variation from meaningful degradation and decide where to investigate next.
A useful operational view should help answer:
- Is there an actual customer or system impact?
- Where is the change occurring?
- Is the effect isolated or expanding?
- Does the signal require action?
- What should the responder examine next?
Each preparation cycle gave us an opportunity to reconsider what we measured and how the information was presented. We could apply what we had learned during the previous event while accounting for changes in the current system.
Latency required particular attention. Aggregate capacity could appear sufficient while the behaviour of requests across availability zones or dependencies introduced a different constraint.
This reinforced a broader lesson for me: observability is not decoration added after a platform has been built. It is part of the platform’s ability to operate safely.
A system that is running but cannot explain its condition leaves engineers guessing. During peak traffic, guessing is expensive.
Plan B and Plan C had to be executable
Peak preparation also meant accepting that the preferred operating path might not work.
We prepared runbooks and discussed fallback options before the event. One contingency was retaining the ability to scale out servers manually if required.
Manual scaling was not necessarily the most elegant option, but fallback plans do not become valuable by being sophisticated. They become valuable by being understandable, available and executable when time and attention are limited.
Under pressure, there is an enormous difference between inventing a response and following one that has already been considered.
A useful runbook cannot predict every incident. It can, however, reduce unnecessary cognitive load. It should help the responder understand:
- What condition should trigger the action?
- Who owns the decision?
- What needs to be changed?
- How will the team verify that the action helped?
- What happens if it does not help?
- How can the change be reversed safely?
Preparing Plan B and Plan C also improves architecture discussions.
Instead of discussing only how the system should behave, the team begins discussing how it might fail. Instead of assuming automation will always respond correctly, the team considers which actions should remain possible manually. Instead of treating dependencies as permanently healthy, the team considers degraded behaviour.
Plan B is only real when the team knows when and how to use it. Plan C matters when the limits of Plan B have already been considered.
What I learned from SRE
I learned enormously from the Site Reliability Engineering teams during this period.
Many of the SRE engineers around me were among the sharpest people from whom I had the opportunity to learn. What impressed me was not only their knowledge of infrastructure, but their method of approaching uncertainty.
They challenged optimistic assumptions. They looked for gaps between what a design promised and what production might actually do. They asked how a failure would become visible and what action a particular signal should trigger.
They treated monitoring, capacity, latency and recovery as integral parts of the system—not operational work to be added later.
Most importantly, they demonstrated that preparing for failure is not negativity. It is professional responsibility.
That mindset influenced the way I approached engineering beyond Black Friday and Cyber Monday. Designing the normal path is only part of the job. Senior engineers must also think about degraded behaviour, detection, recovery and the people who will operate the system under pressure.
Reliability is also a human capability
Some of my strongest memories are not of a particular graph or technical decision. They are of people sitting together late at night, watching traffic and waiting for the expected peak.
There was tension because the stakes were real. There was also excitement because everyone had contributed to the preparation, and this was the moment when our assumptions met production reality.
Engineers from different areas were aligned around the same outcome. Communication became direct. The business significance of the event was visible, and everyone understood their part in the response.
Groupon went through different phases as a company, but the passion around this time of year remained high. The architecture evolved, teams changed and responsibilities moved, yet Black Friday and Cyber Monday continued to bring a distinctive energy.
The experience reminded me that reliability is both technical and human.
Capacity, Kubernetes stability, availability-zone latency, load tests, Grafana dashboards and runbooks all matter. But the ability of people to communicate clearly, remain calm and make sound decisions matters just as much.
What stayed with me
Supporting three peak seasons changed my understanding of senior engineering.
The responsibility was not merely to write code or solve the most difficult technical problem. It was to reduce uncertainty for both the system and the team.
Five lessons have stayed with me:
Start with the business journey. Understand what needs to be protected before deciding how to protect it.
Revalidate assumptions every year. The previous system’s success does not guarantee the current system’s readiness.
Treat load testing and observability as maintained capabilities. They should evolve with the platform.
Make fallback plans simple and executable. A basic recovery option that the team can use safely is better than an elaborate plan that exists only in a document.
Build reliability into the team, not only the software. Clear ownership, preparation and calm communication are part of system resilience.
We could never predict every possible failure. That was not a realistic objective.
The goal was to detect reality quickly, protect what mattered most and respond without allowing uncertainty to become panic.
A stable system during a peak event can look uneventful from the outside. Behind that quietness is usually a considerable amount of thoughtful engineering.
The visible sale lasted only a few days. The lessons from preparing for it have stayed with me for years.

Comments
Post a Comment