olOpsLyft Docs

Investigate a cost spike

Go from an anomaly or an unexpected number to the resource, the owner, and the fix.

Start from an anomaly

Open Monitor → Anomalies

Start with Fix these first. It ranks running anomalies by their share of the burn rate.

Read the row

Note the type (new spend, spike, step up, recurring), the account, the usage type, how long it's been running, and the before → after level.

Ask Iris why

Select ✦ on the row, or Investigate in Iris on the home banner. Ask at High depth: What caused this, which resources are behind it, and who owns them?

Find the resources

Select View in Cost Explorer on the chart Iris adds, switch to the Engineering lens, and drill down to Resource ID. Open the resource in Live inventory to see its tags and owner.

Check what changed

Open Change history for the window around the start date to see resources that were added or changed.

Assign and ticket it

Assign the anomaly to the owner and select + Ticket. If it's expected, mark it as a false positive instead.

Start from a number

If you see an unexpected number on a canvas or in Cost Explorer:

  1. Ask Iris Why did [service] spike on [date]?
  2. Group by Charge Type to rule out one-off purchases, credits, or tax.
  3. Switch the cost type between Billed and Amortized. Upfront commitment purchases spike billed cost but not amortized.
  4. Check data freshness. The last day or two may be incomplete.

Common causes

PatternLikely cause
Single-day spike on the 1stMonthly subscriptions, support, or marketplace charges posting at month start
New spend on an AI modelA team started using a new model
Step up after a deployA config change, such as more replicas or memory
Data transfer spikeCross-region pulls or a NAT gateway change
Recurring spikeA scheduled job or batch

On this page