Indie Dev Workflow
When to Pause Feature Work for a Support Spike
A practical guide for indie developers and small SaaS teams on when a surge in support tickets means you should stop building features, and when to keep shipping.
Pause feature work when a support spike points to a problem that is actively hurting customers. That means lost or exposed data, a core workflow broken for many users, or a single root cause, often your own recent release, driving most of the new tickets. Keep building when the spike comes from expected activity like a launch, from cosmetic issues, or from questions a quick doc update can answer.
On a small team, the hard part is usually deciding in the moment. Too often that decision comes from whoever is most stressed. The rest of this guide shows how to make the call deliberately and how to get back to building afterward.
What established teams do
Several well-documented practices show how experienced teams decide when to stop normal work.
Basecamp's Shape Up drops everything only for real crises. In the Betting Table chapter of Shape Up, Basecamp writes that "there is nothing special about bugs that makes them automatically more important than everything else." It describes a true crisis as one where "data is being lost, the app is grinding to a halt, or a huge swath of customers are seeing the wrong thing," and says that in that case they will "drop everything to fix it." The book also says "crises are rare" and that most bugs can wait six weeks or more. Non-urgent fixes go into cool-down periods, get pitched as planned work, or get saved for a yearly "bug smash."
Google's SRE error budget policy freezes changes when reliability falls below target. In the example error budget policy in the Google SRE Workbook, a team that has missed its reliability target (SLO) over the past four weeks will "halt all changes and releases other than P0 issues or security fixes until the service is back within its SLO." The policy also lists cases where feature work can continue, such as when another team's service caused the outage. It says the policy is not meant as punishment. It gives the team permission to focus on reliability when the data says that matters more.
Atlassian ties response urgency to severity. Atlassian's severity levels define:
- SEV 1: a critical incident with very high impact, such as data loss, a security breach, or a customer-facing service down for everyone.
- SEV 2: a major incident, such as a service down for a subset of customers.
- SEV 3: a minor incident that causes slight inconvenience.
SEV 1 and SEV 2 incidents trigger an immediate on-call response at any hour. SEV 3 incidents can wait for working hours. Atlassian also encourages teams to adapt the levels to their own business.
Toyota lets any worker stop the line. In Toyota's production system, a worker who spots an abnormality can stop the line by pulling a cord. This keeps defects from moving downstream and makes problems visible so they can be kept from happening again. For a software team, the parallel is simple: stop shipping new work on top of something that is broken.
All four share one idea. Decide your stop criteria before the emergency, base them on how much customers are affected rather than how many tickets arrive, and keep a full stop for situations that need one.
Signs you should pause feature work
The criteria below are our recommendations, adapted from the practices above for teams of one to ten people.
- Data loss, data exposure, or a security issue. This matches Atlassian's SEV 1 examples and the security exception in Google's policy. Stop and fix it.
- A core workflow is broken for many customers. If people can't log in, pay, sync, or do the main thing your product exists for, more features won't help anyone.
- One root cause explains most of the spike. If 30 tickets trace back to one bug, fixing that bug is the fastest way to reduce your support load. Answering each ticket one by one only treats the symptoms.
- Your own recent release caused it. Shipping more changes on top of a bad release makes the problem harder to diagnose. Google's change freeze exists for this reason.
- Response times are slipping past what you've promised. If you publish response times or have contracts with B2B customers, missing them becomes a trust problem and can become a revenue problem.
- High-value or at-risk accounts are affected. For a small B2B SaaS, a few key customers can make up a large share of revenue.
Signs you should keep building
- The spike is expected. A launch, a price change, or a press mention brings questions. That's healthy demand, not a fire.
- The issue is cosmetic or minor. Typos, layout quirks, and edge cases with workarounds are SEV 3 work. Schedule them.
- Customers are asking how to do something, not reporting that it's broken. A clearer help article, a saved reply, or a tweak to onboarding usually fixes this faster than an engineering pause.
- The cause is outside your control. If an upstream provider is down, Google's policy treats that as a case where feature work can continue. Tell customers what's happening and watch the provider's status page.
- Only one customer is affected and a workaround exists. Help them well, log the bug, and keep going.
Decide how much to pause
Pausing doesn't have to mean the whole team stops. Pick the smallest response that fixes the problem.
| Situation | Suggested response |
|---|---|
| Crisis (data, security, core workflow down for many) | Full stop. Everyone works on the fix and on customer communication. |
| Clear root cause, significant impact | Partial pause. One or two people fix the cause while the others handle tickets. |
| High volume, low severity | No engineering pause. Add saved replies, update docs, and set aside a fixed block of time for the backlog. |
| Minor or isolated issues | Log them and schedule them for a cool-down or planned bug-fix time. |
This table is our recommendation, not an industry standard.
If you're a solo founder, "partial pause" might mean spending mornings on the fix and afternoons on replies, rather than switching between the two all day.
Write your policy before the next spike
Google's error budget policy works because it is written down and agreed on in advance. A small team can do the same in a short document:
- Stop triggers: the conditions from the "pause" list above that apply to your product.
- Who decides: on a team of three, name one person so you don't argue about it mid-incident.
- What continues: for example, security fixes and the fix for the incident itself.
- Exit criteria: what has to be true before feature work resumes.
- Follow-up: Google's example requires a postmortem when one incident uses more than 20% of the four-week error budget. A lighter version for small teams is a short written note on what happened and one concrete action to prevent it.
What to do during the pause
- Group the tickets. Sort incoming messages by cause. The biggest group is usually your priority.
- Send one clear, honest update. Tell affected customers what you know, what you're doing, and when you'll update them next. Don't promise a fix time you can't keep.
- Write one good answer and reuse it. Draft one careful reply for the main issue and adapt it for each customer, so later replies don't get rushed.
- Fix the cause, not just the tickets. Answering every message without fixing the bug only delays the next wave.
- Update your docs. If customers kept asking the same question, put the answer where they'll find it next time.
AI drafting tools can help with step 3. SupportMe, for example, drafts replies in your writing style and learns from your edits, and nothing sends without your approval. A tool like this can speed up repetitive replies during a spike. It doesn't fix the underlying bug, and it shouldn't stand in for deciding whether to pause.
Resume feature work on purpose
End the pause when your exit criteria are met, not when the inbox simply feels quieter. Reasonable criteria include:
- The fix is shipped and confirmed working.
- New tickets about the issue have dropped back to normal levels.
- Affected customers have received a final update.
- The follow-up action is written down and scheduled.
A hypothetical example
This scenario is illustrative, not a real case.
A two-person SaaS team ships a billing update on Tuesday. By Wednesday morning, support email is about three times its usual volume. Most messages say customers were charged twice.
That meets three stop triggers: it involves money, it affects many customers, and their own release caused it. One founder rolls back the change and investigates. The other sends affected customers one clear explanation, then processes refunds.
A smaller cluster of tickets in the same inbox asks where the new invoice download button went. That isn't a crisis. A short help article covers it, and the button's placement goes on the list for their next planned cleanup. Feature work resumes on Friday, after refunds are confirmed and a test for duplicate charges is added.
Conclusion
Pause feature work when a support spike shows real customer harm: data or security problems, a broken core workflow, or a single cause, often your own release, behind most of the tickets. Keep building when the spike is expected, minor, or answerable with better docs. Basecamp, Google SRE, Atlassian, and Toyota all follow the same pattern: decide your stop criteria in advance, base them on severity rather than volume, and make the pause as small as the problem allows. Writing your policy down now means that during the next spike you only have to follow it.
References
- Basecamp, Shape Up, "The Betting Table": https://basecamp.com/shapeup/2.2-chapter-08
- Google SRE Workbook, "Example Error Budget Policy": https://sre.google/workbook/error-budget-policy/
- Atlassian, "Understanding incident severity levels": https://www.atlassian.com/incident-management/kpis/severity-levels
- Toyota, "Toyota Production System": https://global.toyota/en/company/vision-and-philosophy/production-system/
Tags
Related posts
Indie Dev Workflow
How to Measure How Much Build Time Support Costs You
A practical method for solo developers and small SaaS teams to measure how much time customer support really takes, including interruptions and lost focus, and turn it into usable numbers.
8 min read
Indie Dev Workflow
A Support Workflow for Notifying Customers When Fixes Ship
A practical workflow for small teams to track who reported a bug, confirm the fix is live, and notify every affected customer by email or app store review reply.
9 min read
Indie Dev Workflow
A Support Workflow for Side Projects With a Day Job
A practical support workflow for side-project builders with full-time jobs: one inbox, fixed support windows, clear triage rules, reusable answers, and a weekly review that keeps support from spreading into every hour.
10 min read