免费跑 Codex,竟然不用 API 钱 | Codex Runs Local, Free

 

Codex 是能在终端读代码、改文件、跑命令的 AI 编程小工;llama.cpp 是本地“发动机”,在你 M2 上免费跑开源模型;CLIProxyAPI 是“翻译台”,把接口调成 Codex 认得的样子;cc-switch 是“档位开关”,把这条免费链路存成 provider 随时切——全程不登录付费账号、不用 key。再让一段精简 Python 脚本自动起服务、生成配置。全开源、零 API 费。唯一的小情绪是:16GB 喂本地模型有点紧,它偶尔慢悠悠憋半天才吐一行💤,重活还得让给云端。
Codex is a terminal AI coding buddy; llama.cpp is the local engine; CLIProxyAPI is the translator; cc-switch is the gear shifter. No paid account, no API key, zero API cost. The only mood swing: 16GB is a tight squeeze, so heavy jobs may still be better left to the cloud.

① 模型怎么选 / Choosing the model

M2 16GB 优先:
gpt-oss:20b:首选,20B MoE、约 3.3B 激活、128K context。
Qwen3-Coder:编程强,Q4_K_M。
Gemma 4 12B:更省内存。

首选:
gpt-oss-20b-mxfp4.gguf,约 12–13GB。
内存紧时换 gpt-oss-20b-Q4_K_M.gguf,约 11.6GB。

https://huggingface.co/ggml-org/gpt-oss-20b-GGUF
https://huggingface.co/unsloth/gpt-oss-20b-GGUF

② 下载模型 / Download the model

最省事:

llama-server -hf ggml-org/gpt-oss-20b-GGUF --jinja -ngl 999 --ctx-size 32768

手动:

pip install -U "huggingface_hub[cli]"

hf download ggml-org/gpt-oss-20b-GGUF gpt-oss-20b-mxfp4.gguf --local-dir ~/models

③ 四个角色 / Four roles

llama.cpp = 本地引擎;CLIProxyAPI = Responses 网关;cc-switch = provider 管理;Codex CLI = 本地终端入口。四者都不靠付费 API 计费,唯一成本就是本机算力/电费。Codex CLI 本身就是运行在本机终端的 coding agent。(GitHub)

④ 安装 / Install

第一步 / Step 1

brew install llama.cpp && npm install -g @openai/codex

第二步 / Step 2

brew install --cask cc-switch

第三步 / Step 3

# 到 https://github.com/router-for-me/CLIProxyAPI 的 releases 下载 macOS arm64 版,解压后:

sudo mv ./cli-proxy-api /usr/local/bin/ && chmod +x /usr/local/bin/cli-proxy-api

第四步 / Step 4

下载 GGUF 到 ~/models/。

⑤ Python 一键起服务 / Start everything with Python

脚本只做三件事:启动 llama-server、启动 CLIProxyAPI、生成 cc-switch provider 配置。

#!/usr/bin/env python3

"""M2 16GB 免费跑 Codex:起 llama.cpp + CLIProxyAPI,并生成 cc-switch provider 配置。

   不登录付费账号、不用 API key,只用本机算力。"""

import json

import subprocess

import time

import urllib.request

from pathlib import Path


# ---- 配置区(按实际改)----

MODEL_PATH   = Path.home() / "models" / "gpt-oss-20b-mxfp4.gguf"  # 内存紧可换 Q4_K_M 版

MODEL_ALIAS  = "local-coder"

LLAMA_PORT   = 8080

GATEWAY_PORT = 8317

GATEWAY      = f"http://127.0.0.1:{GATEWAY_PORT}"

GATEWAY_CONFIG   = Path.home() / ".cliproxyapi" / "config.yaml"

CCSWITCH_SNIPPET = Path.cwd() / "ccswitch_provider_snippet.json"

procs = []



def start_llama_server():

    """①启动 llama.cpp 服务;-ngl 999 用 M2 的 Metal 全量加速,--jinja 支持工具调用"""

    print("→ 启动 llama-server(Metal 加速)…")

    procs.append(subprocess.Popen(

        ["llama-server", "-m", str(MODEL_PATH), "--alias", MODEL_ALIAS,

         "--host", "127.0.0.1", "--port", str(LLAMA_PORT),

         "--ctx-size", "32768", "-ngl", "999", "--jinja"],

        stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL))



def start_gateway():

    """②写 CLIProxyAPI 配置(只挂本地 llama.cpp)并启动 —— 字段名以其 releases 文档为准"""

    GATEWAY_CONFIG.parent.mkdir(parents=True, exist_ok=True)

    GATEWAY_CONFIG.write_text(f"""\

port: {GATEWAY_PORT}

providers:

  - name: local-llamacpp

    type: openai

    base_url: http://127.0.0.1:{LLAMA_PORT}/v1

    models: ["{MODEL_ALIAS}"]

""")

    print("→ 启动 CLIProxyAPI(纯本地路由,无付费账号)…")

    procs.append(subprocess.Popen(

        ["cli-proxy-api", "--config", str(GATEWAY_CONFIG)],

        stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL))



def wait_ready(url: str, timeout: int = 240):

    for _ in range(timeout):

        try:

            urllib.request.urlopen(f"{url}/v1/models", timeout=2)

            return

        except Exception:

            time.sleep(1)

    raise RuntimeError(f"{url} 在 {timeout}s 内没起来")



def write_ccswitch_snippet():

    """③生成 cc-switch 可录入的 provider 配置(关键:wire_api=responses、无尾斜杠、占位key)"""

    snippet = {

        "name": "Local-llamacpp (free)",

        "app": "codex",

        "base_url": f"{GATEWAY}/v1",          # 结尾不要多带 /

        "api_key": "sk-local-no-key-needed",  # 本地网关不校验,占位即可

        "model": MODEL_ALIAS,

        "wire_api": "responses",              # Codex 必需,最容易漏

        "requires_openai_auth": False,

    }

    CCSWITCH_SNIPPET.write_text(json.dumps(snippet, indent=2, ensure_ascii=False))

    print(f"→ 已生成 cc-switch 配置:{CCSWITCH_SNIPPET}")

    print("  在 cc-switch → Codex → Add Provider(自定义)里照此填写,再设为当前 provider。")



def cleanup():

    for p in procs:

        p.terminate()



if __name__ == "__main__":

    if not MODEL_PATH.exists():

        raise SystemExit(f"❌ 找不到模型:{MODEL_PATH},请先下载 GGUF 到该路径")

    try:

        start_llama_server()

        wait_ready(f"http://127.0.0.1:{LLAMA_PORT}")

        start_gateway()

        wait_ready(GATEWAY)

        write_ccswitch_snippet()

        print("\n✅ 服务已就绪:按生成的片段在 cc-switch 里录入并激活 provider,")

        print("   然后开新终端跑 `codex`,即用本地免费模型,零 API 费。")

        print("(服务在后台,Ctrl+C 结束并清理)")

        while True:

            time.sleep(3600)

    except KeyboardInterrupt:

        pass

    finally:

        cleanup()

        print("已清理后台服务。")

⑥ 运行 / Run

python free_codex_pipeline.py

然后:cc-switch → Codex → Add Provider(自定义),照生成的片段填写并激活。

⑦ 开新终端 / Open a new terminal

切换后必须重启进程:

cd 你的项目目录 && codex

⑧ 切 Claude Code / Switch to Claude Code

换之前先 Ctrl+C,脚本会清理 llama.cpp 和 CLIProxyAPI;再用 Claude Code 的 Anthropic 兼容方式,把 ANTHROPIC_BASE_URL 指向本地网关。

补充:llama.cpp 已原生支持 Anthropic /v1/messages,也可直连:

ANTHROPIC_BASE_URL=http://127.0.0.1:8080

三个坑 / Three traps:
① wire_api 必须是 responses;② cc-switch 保存/切换会重写 Codex 配置,全程 GUI 或全程手动,别混用;③ 切换后开新终端,base_url 结尾别带 /。CLIProxyAPI 字段名以其 releases 文档为准。

古人云:工欲善其事,必先利其器。
但这次器就在你桌上,何必每次都向云端伸手?

相关链接 (Related Links):
https://openai.com/codex/
https://github.com/openai/codex
https://github.com/ggml-org/llama.cpp
https://huggingface.co/ggml-org/gpt-oss-20b-GGUF
https://huggingface.co/unsloth/gpt-oss-20b-GGUF
https://github.com/router-for-me/CLIProxyAPI
https://help.router-for.me/
https://help.router-for.me/agent-client/codex
https://ccswitch.io/
https://github.com/farion1231/cc-switch
https://huggingface.co/models?library=gguf


评论

爽文共赏/Popular Posts

5分钟拿下原生IP + 免费节点 手残变大佬 | From Noob to Pro in 5 Minutes: 10 Real IPs + Daily Fresh Free Nodes

免费节点无限续?OneBox+subs-check,付费党沉默了 | Free Nodes Never Die? OneBox + subs-check, Paid Users Left Speechless

一口气搭好!Cloudflare融合Railway与Galaxy打造纯净原生免费VPS | Build a Pure Free VPS with Cloudflare + Railway + Galaxy in One Go

3 Steps to AI Video Fame✨WebSocket+Copilot Auto-Update, Newbies Win Too | 3 步 AI 视频封神✨WebSocket+Copilot 自动更,小白也能卷赢

Stop Note-Taking Pain—AI One-Click, Efficiency Skyrockets | 别再苦抄笔记!用AI一键效率爆表💥

Game-Changing! Get AI Parenting+Math Enlightenment for Free | 0成本get AI育儿+数学启蒙双buff,躺赢养娃