可观测与监测

原名:observability-and-instrumentation

为生产行为提供可见性和可诊断性。在添加日志、指标、追踪或告警时使用;在交付需验证运行效果的线上功能时使用;在生产问题无法从现有数据判断原因时使用。

分类
开发提效
版本
v1.0.0
作者
郭哲(@user7gkg2q)
下载
2
收藏
0
发布
2026-08-08
更新
2026-08-14
TRACE 评分
4 / 5

内容概览

Code you can't observe is code you can't operate. Observability is the ability to answer "what is the system doing and why?" from the outside, using the telemetry the code emits. Instrumentation is not a post-launch add-on — it's written alongside the feature, the same way tests are. If a feature ships without telemetry, the first user-reported bug becomes archaeology instead of a query. - Building any feature that will run in production - Adding a new service, endpoint, background job, or external integration - A production incident took too long to diagnose ("we couldn't tell what happened") - Setting up or reviewing alerting rules - Reviewing a PR that adds I/O, retries, queues, or cross-service calls NOT for: - Diagnosing a failure happening right now — use the debugging-and-error-recovery skill (observability is what makes that skill fast next time) - Profiling and optimizing measur…

查看 SKILL 详情