
“Python 多线程是假的。” “GIL 让 Python 永远只能跑一个核。” “想并行?上多进程。”
这些说法在工程师群体中流传甚广,但它们大多只对了一半。更准确的说法是:CPython 的 GIL 不是 Python 语言的缺陷,而是一种实现层面的工程取舍。它改变了并发的成本结构,但没有封死并行的所有路径。 真正的问题从来不是“GIL 好不好”,而是“你的任务到底受什么限制,以及你愿意为并行支付多少复杂度”。
这篇文章不打算重复“I/O 密集用线程、CPU 密集用进程”的口号。它想拆开 CPython 的运行时,看看 GIL 到底保护了什么,asyncio 为什么会被一个 time.sleep 卡死,多进程的序列化开销从哪里来,以及 Python 3.13/3.14 之后出现的自由线程和子解释器,究竟改变了哪些工程约束。
CPython 使用引用计数作为主要的内存管理机制。每个 Python 对象都有一个 ob_refcnt 字段,记录有多少个引用指向它。当引用计数归零时,对象立即被释放。
引用计数的更新必须原子。如果两个线程同时修改同一个对象的 ob_refcnt,计数就会错乱,导致对象被提前释放或永久泄漏。在 CPython 的设计中,保证这种原子性的最便宜方式,就是一把全局锁——同一时刻只允许一个线程执行 Python 字节码。
这就是 GIL(Global Interpreter Lock)的起源。它不是为了让 Python 线程安全,而是为了让 CPython 的引用计数实现足够简单。GIL 保护的范围包括:
关键点在于:GIL 是 CPython 的实现细节,不是 Python 语言规范。 Jython 没有 GIL,PyPy 有 GIL 但正在实验 STM,GraalPy 没有 GIL。当你写 import threading 时,你依赖的是 CPython 的运行时行为,而不是语言承诺。
GIL 的释放时机也很重要。线程在执行以下操作时会释放 GIL:
time.sleep()这意味着多线程在 I/O 密集场景下确实能并发。但在 CPU 密集场景下,线程会频繁争抢 GIL,最终退化为单核执行,甚至因为上下文切换而比单线程更慢。
先看一个最容易被误解的实验:CPU 密集任务用多线程到底会发生什么。
import threading
import time
def cpu_bound(n: int) -> int:
total = 0
for i in range(n):
total += i * i
return total
def run_single():
start = time.perf_counter()
cpu_bound(20_000_000)
print(f"single: {time.perf_counter() - start:.3f}s")
def run_threads():
threads = [
threading.Thread(target=cpu_bound, args=(20_000_000,))
for _ in range(4)
]
start = time.perf_counter()
for t in threads:
t.start()
for t in threads:
t.join()
print(f"threads: {time.perf_counter() - start:.3f}s")
if __name__ == "__main__":
run_single()
run_threads()在一台多核机器上,run_threads 通常不会比 run_single 快 4 倍,甚至可能更慢。原因是四个线程在争抢同一把 GIL,Python 字节码始终只有一个线程在执行。线程切换本身还有开销。
但把任务换成 I/O 等待,结论就完全不同:
import threading
import time
import urllib.request
URLS = [
"https://example.com",
"https://httpbin.org/get",
"https://www.python.org",
] * 5
def fetch(url: str) -> bytes:
with urllib.request.urlopen(url, timeout=10) as resp:
return resp.read()
def run_sequential():
start = time.perf_counter()
for url in URLS:
fetch(url)
print(f"sequential: {time.perf_counter() - start:.3f}s")
def run_threaded():
threads = [threading.Thread(target=fetch, args=(url,)) for url in URLS]
start = time.perf_counter()
for t in threads:
t.start()
for t in threads:
t.join()
print(f"threaded: {time.perf_counter() - start:.3f}s")在 I/O 阻塞期间,线程会释放 GIL,其他线程可以继续执行。因此多线程能把等待时间重叠起来。这是线程模型在 Python 中最有价值的场景。
但线程带来的真正风险不是性能,而是共享状态。GIL 不保证你的业务逻辑原子性。看下面这段代码:
import threading
counter = 0
def worker():
global counter
for _ in range(100_000):
counter += 1
threads = [threading.Thread(target=worker) for _ in range(4)]
for t in threads:
t.start()
for t in threads:
t.join()
print(counter) # 通常小于 400000counter += 1 不是原子操作。它大致对应:读取 counter、加一、写回。GIL 可能在字节码之间切换线程,导致两个线程读到相同的旧值。修复方式是显式加锁:
import threading
counter = 0
lock = threading.Lock()
def worker():
global counter
for _ in range(100_000):
with lock:
counter += 1锁会带来新的成本:竞争、死锁风险、调试难度。更安全的做法是避免共享可变状态,改用 queue.Queue 传递消息,或者让每个线程处理独立的数据分片。
多进程是 CPython 中实现真正 CPU 并行的经典方案。每个进程有独立的 Python 解释器和独立的内存空间,因此有独立的 GIL。四个进程可以真正同时跑在四个核上。
from multiprocessing import Pool
import time
def cpu_bound(n: int) -> int:
total = 0
for i in range(n):
total += i * i
return total
if __name__ == "__main__":
start = time.perf_counter()
with Pool(4) as pool:
results = pool.map(cpu_bound, [20_000_000] * 4)
print(f"multiprocessing: {time.perf_counter() - start:.3f}s")这段代码通常能接近线性加速。但多进程的代价隐藏在细节里:
第一,进程启动成本。 在 Linux 上默认使用 fork,子进程继承父进程内存,启动较快。在 Windows 和 macOS 上默认使用 spawn,子进程重新导入模块,启动慢得多。这就是为什么多进程代码必须放在 if __name__ == "__main__": 保护之下。
第二,序列化开销。 pool.map 会把参数和返回值通过 pickle 序列化,在进程间传输。如果任务本身很短,序列化可能比计算还贵。如果数据很大,内存带宽会成为瓶颈。
第三,共享状态困难。 多进程之间不能直接共享 Python 对象。multiprocessing.Queue 和 Pipe 通过序列化传递消息;Manager 提供代理对象,但每次访问都要经过进程间通信。真正高效的共享内存需要 multiprocessing.shared_memory:
from multiprocessing import shared_memory
import numpy as np
# 创建共享内存块
shm = shared_memory.SharedMemory(create=True, size=8 * 1024 * 1024)
arr = np.ndarray((1024, 1024), dtype=np.float64, buffer=shm.buf)
# 子进程可以通过名称附加到同一块内存
# shm_child = shared_memory.SharedMemory(name=shm.name)
# arr_child = np.ndarray((1024, 1024), dtype=np.float64, buffer=shm_child.buf)共享内存避免了序列化,但要求数据是固定布局的二进制格式,且需要手动管理生命周期和同步。它适合 NumPy 数组、图像缓冲区、科学计算中间结果,不适合任意 Python 对象。
asyncio 的核心是事件循环。它在单线程中运行,通过协作式调度在多个协程之间切换。协程只在 await 处让出控制权。这种模型没有线程切换开销,也没有 GIL 争抢,在高并发 I/O 场景下非常高效。
import asyncio
import aiohttp
async def fetch(session: aiohttp.ClientSession, url: str) -> str:
async with session.get(url) as resp:
return await resp.text()
async def main():
urls = ["https://example.com"] * 20
async with aiohttp.ClientSession() as session:
tasks = [fetch(session, url) for url in urls]
results = await asyncio.gather(*tasks)
return results
# asyncio.run(main())asyncio.gather 会并发调度所有任务。当某个任务等待网络响应时,事件循环切换到其他任务。整个过程中只有一个线程,没有锁竞争,没有 GIL 切换。
但 asyncio 有一个致命约束:任何阻塞调用都会卡住整个事件循环。 下面这段代码会让所有并发任务停摆:
import asyncio
import time
async def bad_task():
time.sleep(5) # 阻塞整个事件循环
return "done"
async def main():
await asyncio.gather(bad_task(), bad_task(), bad_task())time.sleep 是同步阻塞调用,它不会让出控制权。事件循环被卡住,其他协程无法执行。正确做法是使用 asyncio.sleep,或者把阻塞调用放到线程池中:
import asyncio
import time
from concurrent.futures import ThreadPoolExecutor
def blocking_io():
time.sleep(5)
return "done"
async def main():
loop = asyncio.get_running_loop()
with ThreadPoolExecutor() as pool:
result = await loop.run_in_executor(pool, blocking_io)
return resultPython 3.9 之后可以使用更简洁的 asyncio.to_thread:
result = await asyncio.to_thread(blocking_io)asyncio 不是并行,它是单线程内的并发。它适合大量网络连接、文件描述符、定时任务。它不适合 CPU 密集计算。如果你在协程里做矩阵运算,事件循环同样会被卡死。
Python 3.13 引入了实验性的 free-threaded 构建,通过 --disable-gil 配置选项编译。Python 3.14 将自由线程推进到官方支持的可选构建。这意味着 CPython 正在认真探索“没有 GIL 的 Python”。
自由线程的核心挑战不是“把锁删掉”,而是重新设计整个运行时的并发安全机制:
在自由线程构建中,某些在 GIL 下“碰巧原子”的操作不再安全。例如 list.append 在 GIL 下是原子的,但在自由线程下需要显式同步。这意味着一批依赖 GIL 隐式保护的旧代码可能出现竞争条件。
检查当前解释器是否启用 GIL:
import sys
print(sys._is_gil_enabled()) # Python 3.13+ 自由线程构建返回 False自由线程的工程意义在于:它让 CPU 密集任务可以用线程实现并行,而不必承担多进程的序列化和内存隔离成本。但代价是单线程性能回退、内存占用增加、C 扩展生态需要时间适配。短期内,它不会替代多进程,而是成为第三种选择。
除了多进程和自由线程,CPython 还有第三条并行路径:子解释器。PEP 684 让每个子解释器拥有独立的 GIL。PEP 734 在 Python 3.14 中引入了 concurrent.interpreters 标准库模块。
子解释器的优势是启动成本低于进程,隔离性强于线程。每个解释器有独立的模块命名空间、独立的内置对象、独立的 GIL。它们共享同一个进程地址空间,但默认不共享 Python 对象。
from concurrent import interpreters
interp = interpreters.create()
interp.run("print('hello from subinterpreter')")子解释器的限制在于 C 扩展兼容性。许多 C 扩展依赖全局状态,无法在多个解释器中安全加载。NumPy 等大型扩展正在逐步支持多阶段初始化,但生态完全适配仍需时间。子解释器目前更适合纯 Python 工作负载,或者作为隔离执行环境,而不是通用的并行计算方案。
并发模型的选择不应该基于信仰,而应该基于对瓶颈的判断。下面是一棵简化的决策树:
任务特征 | 首选方案 | 备选方案 | 主要代价 |
|---|---|---|---|
I/O 密集,连接数极高 | asyncio | 线程池 | 阻塞调用会卡死事件循环 |
I/O 密集,代码同步 | 多线程 | asyncio | 共享状态需要加锁 |
CPU 密集,数据可独立 | 多进程 | 自由线程 | 序列化、启动成本 |
CPU 密集,需共享大数组 | 多进程 + 共享内存 | C 扩展 | 手动内存管理 |
需要强隔离 | 多进程 | 子解释器 | 通信成本 |
短任务、低延迟 | 线程或 asyncio | 进程池 | 进程启动太慢 |
测量永远优先于猜测。cProfile 可以定位 CPU 热点,py-spy 可以在不修改代码的情况下采样运行中的进程,time.perf_counter 可以测量端到端延迟。先确认任务是 CPU 密集还是 I/O 密集,再选择模型。
一个实用的诊断代码:
import time
import threading
def diagnose(func, *args, n=4):
# 单线程
start = time.perf_counter()
func(*args)
single = time.perf_counter() - start
# 多线程
threads = [threading.Thread(target=func, args=args) for _ in range(n)]
start = time.perf_counter()
for t in threads:
t.start()
for t in threads:
t.join()
threaded = time.perf_counter() - start
print(f"single: {single:.3f}s")
print(f"threaded x{n}: {threaded:.3f}s")
print(f"speedup: {single * n / threaded:.2f}x")如果 speedup 远小于 1,说明任务受 GIL 限制,线程没有帮助。如果 speedup 接近 n,说明任务在等待 I/O,线程有效。如果 speedup 在 1 附近徘徊,说明任务混合了 CPU 和 I/O,需要进一步拆分。
Python 的并发故事,本质上是一个关于取舍的故事。GIL 用全局锁换来了引用计数的简单和 C 扩展的稳定;多进程用序列化和内存隔离换来了真正的 CPU 并行;asyncio 用单线程事件循环换来了高并发 I/O 的低开销;自由线程和子解释器则在尝试用新的运行时设计,重新划分这些取舍的边界。
理解 GIL 是理解 CPython 的钥匙,但不应该成为思维枷锁。真正专业的 Python 工程师不会说“多线程没用”,也不会说“多进程一定快”。他们会先问:这个任务的瓶颈是 CPU、I/O、内存带宽,还是锁竞争?然后选择匹配的模型,并测量结果。
原创声明:本文系作者授权腾讯云开发者社区发表,未经许可,不得转载。
如有侵权,请联系 cloudcommunity@tencent.com 删除。