homelab K3s 从默认 Flannel 换到 Cilium 之后冒出三个问题:Cloudflared QUIC 握手超时、Pod 访问不到节点物理 IP、ZITADEL 报 Master Key 长度错误,三个成因不同,解法也不同。
Posts for: #troubleshooting
Spring AI 2.0.0-M2 的 Ollama think 字段问题:排查过程与 Interceptor 临时方案
disableThinking() 本该只关推理,结果 think 字段泄漏进了 options map,Ollama 返回 HTTP 400。上游修复还没进正式版,我用一个 ClientHttpRequestInterceptor 在请求发出前把它摘掉。
Reactive Redis 与 Lettuce SharedLock 的连锁问题:make coverage 卡死排查
集成测试卡在 make coverage 阶段,第一眼看到的是连接池超时,真正让进程挂住的是 Reactive Redis 下 Lettuce SharedLock 的自旋,两层问题得分开修。
VS Code 跑 mirrord 遇到 WebSocket 403:从 IDE 报错追到 K8s impersonation 的鉴权链路
把 mirrord OSS 接入 VS Code 时遇到一个 WebSocket 403,最后追到 K8s impersonation 的两条 authorization 链路,记录了排查、原理和最小修复。
How a Performance Optimization Caused Cascading Redis Timeouts in Spring WebFlux
A seemingly harmless removal of publishOn(Schedulers.boundedElastic()) led to cascading Redis timeouts in production. This post explains how Spring’s @Cacheable blocks the Netty event loop when used with RedisCacheManager, and why BlockHound failed to catch it.
postgresql 在 prometheus stack 中没有采集到 metrics 的排查
我在 homelab 的 k8s 集群中使用 helm 部署了 postgresql,但是 prometheus stack 没有采集到 postgresql 的指标数据。怎么排查这个问题呢?
Netflix 的 Java 21 虚拟线程死锁排查
Netflix 在 Java 21 上用 virtual thread 碰到的一个死锁案例,记录他们从 closewait 一路查到 carrier thread 被 pin 的过程。