正在加载内容...

963963 Chat News Review Independent coverage of news

Common Mistakes When Evaluating Observability

By Laura Bennett · · 1213 words
Common Mistakes When Evaluating Observability

Observability: If a metric has no owner, it will drift until it causes an incident. Observability: The cheapest optimisation is usually removing work nobody asked for. Observability: Aggregating at write time trades flexibility for predictable read cost.

If a metric has no owner, it will drift until it causes an incident. This is most visible in observability. Consider observability specifically. The cheapest optimisation is usually removing work nobody asked for. Observability: Aggregating at write time trades flexibility for predictable read cost.

Schema Migration: A queue smooths spikes but also hides how far behind you are. Retries without jitter turn a small outage into a large one. That applies to schema migration as well. In practice, schema migration behaves differently: Separating the reads from the writes buys room to change either side.

If the rollback plan needs a meeting, it is not a rollback plan. The same reasoning holds for data pipelines. For data pipelines, the constraint matters more than the feature list. Small pages that stay small are easier to keep fast than large ones made fast. Teams working on data pipelines usually discover this the hard way. Write the invariant down; otherwise it lives only in someone's memory.

Data Pipelines: The interesting number is not the average, it is the 99th percentile. Data Pipelines: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Data Pipelines: Every abstraction you add is a place where behaviour can differ from intent.

Teams working on search indexing usually discover this the hard way. Serving static bytes is the cheapest thing you can do at the edge. A schema is an interface; changing it is a migration, not an edit. This is most visible in search indexing. Consider search indexing specifically. Track the denominator as carefully as the numerator.

If a metric has no owner, it will drift until it causes an incident. This is most visible in cloud infrastructure. Consider cloud infrastructure specifically. The cheapest optimisation is usually removing work nobody asked for. Cloud Infrastructure: Aggregating at write time trades flexibility for predictable read cost.

Packaging is only one part of the record. A card or bank statement may show a payment descriptor that differs from the storefront name, and an order confirmation may name the seller in its subject line or sender address. Shipping alerts can appear on a shared email account, text-message preview or carrier account. Review which contact details are attached to checkout and whether the carrier offers delivery alerts. Do not count on a “discreet” parcel to hide an order from someone who can see those accounts or payment records.

Data Pipelines: If a metric has no owner, it will drift until it causes an incident. Data Pipelines: The cheapest optimisation is usually removing work nobody asked for. Data Pipelines: Aggregating at write time trades flexibility for predictable read cost.

Storage Tiers: You can often replace a coordination problem with an idempotency key. Storage Tiers: Anything that grows without a bound will eventually hit one. Storage Tiers: Documentation that is not tested tends to describe the previous version.

Listening is part of the conversation. Ask what the other person understands, and invite them to describe their own boundaries without treating the exchange as a negotiation in which every limit must be traded away. Open questions such as “What would help you feel comfortable?” can clarify expectations. If a question feels intrusive, either person can decline to answer it.

Consider schema migration specifically. Serving static bytes is the cheapest thing you can do at the edge. Schema Migration: A schema is an interface; changing it is a migration, not an edit. Track the denominator as carefully as the numerator. That applies to schema migration as well.

Serving static bytes is the cheapest thing you can do at the edge. That applies to schema markup as well. In practice, schema markup behaves differently: A schema is an interface; changing it is a migration, not an edit. Track the denominator as carefully as the numerator. The same reasoning holds for schema markup.

Cost Controls: Periodic jobs should be safe to run twice, because they will be. Cost Controls: You rarely need a new component to fix a boundary problem. Cost Controls: The signal you want is often already logged, just not aggregated.

Crawl Budget: The first thing to settle is the failure mode, not the happy path. Crawl Budget: Measurements taken once are anecdotes; you need a baseline that repeats. Crawl Budget: Costs usually concentrate in a small number of operations, so find those first.

In practice, queue design behaves differently: Configurations should be reviewable in a diff, not only in a console. The best time to add an index is before the table gets large. The same reasoning holds for queue design. For queue design, the constraint matters more than the feature list. Failures are usually correlated, so plan for the shared dependency.

Search Indexing: Serving static bytes is the cheapest thing you can do at the edge. Search Indexing: A schema is an interface; changing it is a migration, not an edit. Search Indexing: Track the denominator as carefully as the numerator.

Release Process: You can often replace a coordination problem with an idempotency key. Release Process: Anything that grows without a bound will eventually hit one. Release Process: Documentation that is not tested tends to describe the previous version.

Edge Caching: Serving static bytes is the cheapest thing you can do at the edge. Edge Caching: A schema is an interface; changing it is a migration, not an edit. Edge Caching: Track the denominator as carefully as the numerator.

A design that cannot be rolled back is a design that cannot be changed safely. That applies to backup strategy as well. In practice, backup strategy behaves differently: Latency budgets are easier to defend when every hop has a stated ceiling. Caching helps only until the invalidation rules become the bottleneck. The same reasoning holds for backup strategy.

In practice, cost controls behaves differently: A queue smooths spikes but also hides how far behind you are. Retries without jitter turn a small outage into a large one. The same reasoning holds for cost controls. For cost controls, the constraint matters more than the feature list. Separating the reads from the writes buys room to change either side.

Access Control: A design that cannot be rolled back is a design that cannot be changed safely. Access Control: Latency budgets are easier to defend when every hop has a stated ceiling. Access Control: Caching helps only until the invalidation rules become the bottleneck.

Search Indexing: Configurations should be reviewable in a diff, not only in a console. Search Indexing: The best time to add an index is before the table gets large. Search Indexing: Failures are usually correlated, so plan for the shared dependency.

The interesting number is not the average, it is the 99th percentile. The same reasoning holds for schema migration. For schema migration, the constraint matters more than the feature list. Adding a cache in front of a slow query is a fix; fixing the query is a cure. Teams working on schema migration usually discover this the hard way. Every abstraction you add is a place where behaviour can differ from intent.

Related reading