背景#

homelab 里跑着一个 LiteLLM 网关,公网域名 llm.meirong.dev,后面是 DGX Spark 和 Mac 上自托管的模型。准备给几位朋友各发一把虚拟 key(virtual key),发之前想确认:每把 key 什么时候被调用、用了多少 token、请求从哪个 IP 来,事后都能查到。

这次有三条限制:LiteLLM 用官方镜像,不自己构建;不存对话内容;Cloudflare 和 Tunnel 的配置不动。

spend log 里已经有的东西#

LiteLLM 接上 Postgres 之后,每次调用(包括失败的)都往 LiteLLM_SpendLogs 表写一行。跟这次有关的列:

列 内容
api_key key 的 sha256,不是原文
metadata->>'user_api_key_alias' 建 key 时给的 key_alias
"startTime"、model_group、status 调用时间、模型别名、success 或 failure
prompt_tokens、completion_tokens、total_tokens token 数
metadata->>'user_agent' 客户端 User-Agent
requester_ip_address 来源 IP

用量看 token 那几列。spend 列按 LiteLLM 内置的价目表折算成钱,自托管模型不在表里,这张表里自托管模型的每一行都是 0。对话内容默认不落库:store_prompts_in_spend_logs 没开时 messages 和 response 两列是空的,全表查下来没有一行例外。

用量这一半是现成的,要补的是 IP。

requester_ip_address 记成了网关地址#

按 IP 分组看最近 7 天的记录:

kubectl -n databases exec deploy/apps-pg -- psql -U postgres -d litellm -At -F ' | ' -c "
select requester_ip_address, count(*) from \"LiteLLM_SpendLogs\"
where \"startTime\" > now() - interval '7 days' group by 1 order by 2 desc;"
10.42.0.152 | 43
10.42.1.76 | 13
 | 2
10.42.1.47 | 1

10.42.1.0/24 是 worker 节点的 pod 网段,这两个地址是在集群内直接走 Service 调网关的 pod,记的就是 pod IP。空的两行是失败的请求。剩下的 10.42.0.152 是控制面节点 k8s-node 的 CiliumInternalIP(kubectl get ciliumnode k8s-node -o yaml 的 spec.addresses 里能看到),这 7 天里从公网进来的调用都记成了它。

写这个字段的代码在 litellm/proxy/litellm_pre_call_utils.py,镜像里的 litellm 包版本是 1.103.1:

requester_ip_address = ""
if True:  # Always set the IP Address if available
    # logic for tracking IP Address

    # logic for tracking IP Address
    if (
        general_settings is not None
        and general_settings.get("use_x_forwarded_for") is True
        and request is not None
        and hasattr(request, "headers")
        and "x-forwarded-for" in request.headers
    ):
        requester_ip_address = request.headers["x-forwarded-for"]
    elif (
        request is not None
        and hasattr(request, "client")
        and hasattr(request.client, "host")
        and request.client is not None
    ):
        requester_ip_address = request.client.host
data[_metadata_variable_name]["requester_ip_address"] = requester_ip_address

默认取 request.client.host,也就是 TCP 连接的对端。公网请求到 LiteLLM 之前要经过这几跳:

flowchart TB C["调用方"] -->|"HTTPS"| E["Cloudflare 边缘
写入 CF-Connecting-IP
XFF 末尾追加客户端 IP"] E -->|"Tunnel"| T["cloudflared pod"] T --> G["Cilium Gateway 的 Envoy
XFF 末尾追加 cloudflared 的 pod IP"] G --> L["LiteLLM
TCP 对端是 10.42.0.152"]

到最后一跳,连接是网关的 Envoy 发起的,LiteLLM 看到的对端只能是集群内地址。Tunnel 用通配规则把 *.meirong.dev 交给 Gateway 的做法,在 external-dns 那篇 里有写。

读 CF-Connecting-IP,不开 use_x_forwarded_for#

LiteLLM 自带的开关是 general_settings.use_x_forwarded_for: true。照上面那段代码,打开后它把整个 X-Forwarded-For 头原样存成一个字符串,不从里面挑某一项。这条头在链路上被追加两次:Cloudflare 的文档说,请求里已经有 X-Forwarded-For 时,它把客户端 IP 追加在后面;Envoy 的文档说,use_remote_address 为 true 时它把下游对端追加上去,Cilium 为这个 Gateway 生成的 CiliumEnvoyConfig 里正是 useRemoteAddress: true。按两份文档推,存下来的是「调用方自己填的值, 真实 IP, cloudflared 的 pod IP」,最左边那项调用方想写什么都行,没法拿来按 IP 聚合。

CF-Connecting-IP 只有一个地址,由 Cloudflare 边缘写入。客户端自己带上这个头会怎样,经公网试了一次:

curl -s -w '\n%{http_code}\n' https://llm.meirong.dev/v1/chat/completions \
  -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' \
  -H 'CF-Connecting-IP: 192.0.2.1' \
  -d '{"model":"studio/qwen3.8-27b","max_tokens":8,"messages":[{"role":"user","content":"say ok"}],"chat_template_kwargs":{"enable_thinking":false}}'
error code: 1000

403

请求在 Cloudflare 就被拒了,到不了源站。error 1000 的文档把「请求带了 CF-Connecting-IP 头」列为成因之一。所以在公网这条路上,源站收到的这个头可以直接信。

LiteLLM 没有「从指定请求头取客户端 IP」的配置项,在镜像里搜 cf-connecting-ip 一处都没有。不改镜像的前提下,只能写一个 callback。

用 async_pre_call_hook 改写这个字段#

LiteLLM proxy 处理一个请求的顺序是:add_litellm_data_to_request 先填 metadata(requester_ip_address 就在这时写入),同时把去掉鉴权头的请求头存进 data["proxy_server_request"]["headers"];接着建 logging 对象;然后依次调各个 callback 的 async_pre_call_hook。hook 拿到的 data 里,请求头和 metadata 都在,把那个字段改掉就行:

CLIENT_IP_HEADER = "cf-connecting-ip"
# 不同路由把 metadata 放在不同的键下(chat 是 metadata,部分新路由是 litellm_metadata),两个都看。
METADATA_KEYS = ("metadata", "litellm_metadata")


def client_ip_from_request(data: dict) -> Optional[str]:
    """从 add_litellm_data_to_request 留下的请求快照里取 CF-Connecting-IP,规范化成 IP 字面量。"""
    snapshot = data.get("proxy_server_request")
    headers = snapshot.get("headers") if isinstance(snapshot, dict) else None
    if not isinstance(headers, dict):
        return None
    raw = next(
        (v for k, v in headers.items() if isinstance(k, str) and k.lower() == CLIENT_IP_HEADER),
        None,
    )
    if not isinstance(raw, str):
        return None
    try:
        return str(ipaddress.ip_address(raw.strip()))
    except ValueError:
        return None


class ClientIp(CustomLogger):
    async def async_pre_call_hook(self, user_api_key_dict, cache, data, call_type):
        try:
            client_ip = client_ip_from_request(data)
            if client_ip is None:
                return None
            for key in METADATA_KEYS:
                metadata = data.get(key)
                if isinstance(metadata, dict) and "requester_ip_address" in metadata:
                    metadata["requester_ip_address"] = client_ip
        except Exception as exc:  # 记账字段不许把请求搞挂
            verbose_proxy_logger.warning("client_ip: 改写失败,保留原值: %s", exc)
            return None
        # 原文件这里还有一段:首次改写时打一条 WARNING,略
        return data


client_ip_handler = ClientIp()

这段是原文件删掉 import、文件头说明、类型注解和只为打日志服务的几行之后的样子。设计上有三处取舍:

  • 头缺失或者值不是合法 IP 时什么都不改。集群内 pod 走 Service、tailnet 设备走 NodePort 进来的请求默认不带这个头,LiteLLM 原来记的对端在这两条路上本来就是真实地址。反过来,这两条路不经过 Cloudflare,调用方要是自己带上这个头,hook 一样会采信,记下的就是它填的值。
  • LiteLLM 按路由把 metadata 放在 metadata 或 litellm_metadata 下(_get_metadata_variable_name 决定,assistants、threads 这类走后者),两个都看。
  • 异常一律吞掉、保留原值。这个字段只用于记账,hook 出错不该把请求带挂。

部署沿用这个网关已有的做法:.py 放进 ConfigMap,subPath 挂到容器的 /app/ 下,和 config.yaml 在同一个目录,LiteLLM 解析 callbacks 时按 config 文件所在目录找模块:

litellm_settings:
  callbacks: ["codex_compat.codex_compat_handler", "client_ip.client_ip_handler"]

subPath 挂载收不到 ConfigMap 的更新,LiteLLM 也只在启动时加载 callback,所以 pod 模板上另挂一个 checksum/client-ip-py 注解,值是脚本的哈希。脚本一改注解就变,ArgoCD 同步时 pod 跟着滚动重启。

本地和公网各验证了什么#

上线前,用和生产同一个 digest 的镜像、一个本地 Postgres、一个配了 mock_response 的假模型起了一套,直接带各种请求头打本地 proxy,再查 LiteLLM_SpendLogs。地址用的是文档保留段:

请求 记下的 requester_ip_address
chat,CF-Connecting-IP: 203.0.113.7,同时伪造 X-Forwarded-For 203.0.113.7
chat,不带这个头 docker 网桥地址,即原来的对端
流式,CF-Connecting-IP: 2001:db8::5 2001:db8::5
CF-Connecting-IP: not-an-ip docker 网桥地址
/v1/responses、/v1/embeddings 头里的地址
超过 rpm_limit 被拒的 chat(429) 头里的地址
超过 rpm_limit 被拒的 /v1/responses(429) 空串

最后一行不带这个头也是空串,是 LiteLLM 自己的行为,hook 管不到。

Cloudflare 那一层本地模拟不了,上线后经公网又测了一轮:

请求 结果
正常调用 200,记下的地址和 https://llm.meirong.dev/cdn-cgi/trace 返回的 ip= 一致
伪造 X-Forwarded-For 200,记下的仍是真实地址
伪造 True-Client-IP 200,同上
伪造 CF-Connecting-IP 403 error code: 1000,表里没有这一行
从 tailnet 走 NodePort 200,记下的是 Mac 的 tailnet 地址

每个 pod 第一次改写时打一条 WARNING(真实 IP 换成了占位符):

LiteLLM Proxy:WARNING: client_ip.py:97 - client_ip: 本层已生效(首次改写,后续不再打印)—— acompletion 的对端 10.42.0.152 记为 <真实 IP>

这两轮覆盖不到 LiteLLM 升级。hook 依赖两个内部细节:proxy_server_request.headers 里有完整的请求头,以及在 pre-call hook 里改过的 metadata 会被 spend log 读到。两者都不是公开接口,升级后变了也不会报错,表现只是 IP 又回到 10.42.0.152。所以每次升级后要看一眼那条 WARNING,再查最新几行的 requester_ip_address。

按 key 聚合,以及发 key 时要带的参数#

日常要看的是 key × IP 的聚合:

kubectl -n databases exec deploy/apps-pg -- psql -U postgres -d litellm -c "
select coalesce(nullif(metadata->>'user_api_key_alias', ''), left(api_key, 12)) as key,
       requester_ip_address as ip,
       count(*) as calls, count(*) filter (where status = 'failure') as failed,
       sum(total_tokens) as tokens, max(\"startTime\")::timestamp(0) as last_seen
from \"LiteLLM_SpendLogs\" where \"startTime\" > now() - interval '7 days'
group by 1, 2 order by calls desc;"

没填 key_alias 的 key,user_api_key_alias 是空的,查询里拿 api_key 的前 12 位顶上。现有的 18 把 key 里有 17 把没有别名,只能对着哈希认。

拉长时间窗口按 key 一聚合,一把原本只给 oracle 集群上 calibre 元数据作业建的 key 出现了 5 个来源:作业所在节点的 tailnet 地址、Mac、worker 上 xiaogpt 的 pod、在网关 pod 里用 127.0.0.1 调的两次,以及改造前经网关进来、记成 10.42.0.152 的 12 次调用。前四个都认得出来,最后那批是从哪来的,现在已经查不到了。

发给外部的 key,建的时候带上这几个参数:

curl -s -H "Authorization: Bearer $MASTER_KEY" -H 'Content-Type: application/json' \
  -d '{"key_alias":"ext-alice","models":["studio/qwen3.8-27b"],"duration":"30d",
       "rpm_limit":60,"max_parallel_requests":2}' \
  http://<网关的内网地址>/key/generate
  • key_alias:聚合查询里显示的就是它。LiteLLM 要求别名唯一,重名会直接拒绝。
  • duration:到期自动失效,30d 生成的 expires 就是 30 天之后,过期后调用方拿到 401 expired_key。
  • rpm_limit、max_parallel_requests:超了返回 429 throttling_error。本地用 max_parallel_requests: 1 同时发 4 个请求,1 个 200、3 个 429。

限额别指望 max_budget。它按 spend 累计,自托管模型不在内置价目表里,除非照 custom pricing 文档 在 model_info 下给它配上 input_cost_per_token、output_cost_per_token,否则 spend 一直是 0:本地建了一把 max_budget: 0.001 的 key 连发 6 次,全是 200,spend 还是 0。

小结#

LiteLLM 的 spend log 已经按 key 记了调用时间和 token,放在 Cloudflare Tunnel 和 Cilium Gateway 后面时,缺的只是真实来源 IP。默认记的 TCP 对端是网关地址,自带的 use_x_forwarded_for 会把整条可伪造的 XFF 存成字符串;一个读 CF-Connecting-IP 的 pre-call hook 就能补上,公网这条路上调用方伪造不了这个头。对外发的 key 带上 key_alias、duration、rpm_limit 和 max_parallel_requests,没给自托管模型配单价时,按金额的 max_budget 不起作用。hook 靠的是 LiteLLM 的内部细节,每次升级后要重新确认一遍。

相关文章#