pathlib 完全指南

pathlib 是 Python 3.4 引入的标准库模块,提供面向对象的文件系统路径操作。相比 os.path 的字符串拼接方式,pathlib 的 Path 对象更直观、更安全,跨平台兼容性更好。 Path() 构造函数的参数: pathlib 重载了 / 运算符来进行路径拼接,这是最推荐的方式: joinpath 的参数: with_name 的参数: with_suffix 的参数: write_text 的参数: read_text 的参数: write_bytes 的参数: read_bytes 无参数,返回 bytes。 Path.open

分享

官方文档:https://docs.python.org/3/library/pathlib.html
适用版本:Python 3.12(2026-05-07 核实)

pathlib 是 Python 3.4 引入的标准库模块,提供面向对象的文件系统路径操作。相比 os.path 的字符串拼接方式,pathlibPath 对象更直观、更安全,跨平台兼容性更好。

Path 对象创建

基本构造

from pathlib import Path

# 从字符串创建
p = Path("/home/user/documents")
p = Path("relative/path/to/file.txt")

# Windows 风格(反斜杠)也被支持
p = Path(r"C:\Users\user\documents")

# 传入多个部分,自动拼接
p = Path("/home", "user", "documents", "file.txt")
# 等同于 Path("/home/user/documents/file.txt")

Path() 构造函数的参数:

参数 类型 默认值 说明
*pathsegments strPath 必填(至少一个) 路径片段,自动用系统分隔符连接

特殊构造方法

from pathlib import Path

# 当前工作目录
cwd = Path.cwd()          # 等同于 os.getcwd()

# 用户主目录
home = Path.home()        # 等同于 Path(os.path.expanduser("~"))

# 从环境变量构建
import os
config_dir = Path(os.environ.get("XDG_CONFIG_HOME", Path.home() / ".config"))

平台特定的 Path 子类

from pathlib import PurePosixPath, PureWindowsPath, PosixPath, WindowsPath

# 纯路径(只做字符串操作,不访问文件系统)
posix = PurePosixPath("/home/user/file.txt")
win = PureWindowsPath(r"C:\Users\user\file.txt")

# 具体路径(实际访问文件系统)
# 在 Linux/macOS 上:Path() 返回 PosixPath
# 在 Windows 上:Path() 返回 WindowsPath

路径拼接

/ 运算符

pathlib 重载了 / 运算符来进行路径拼接,这是最推荐的方式:

from pathlib import Path

base = Path("/home/user")

# Path / str
p = base / "documents" / "file.txt"
# PosixPath('/home/user/documents/file.txt')

# Path / Path
sub = Path("documents")
p = base / sub / "file.txt"

# str / Path 也可以(左侧是字符串)
p = "/home/user" / Path("file.txt")

# 警告:如果右侧是绝对路径,会丢弃左侧
p = base / "/etc/passwd"
# 结果是 PosixPath('/etc/passwd'),而非预期的拼接

joinpath() 方法

from pathlib import Path

p = Path("/home/user").joinpath("documents", "file.txt")
# 等同于 Path("/home/user") / "documents" / "file.txt"

joinpath 的参数:

参数 类型 默认值 说明
*pathsegments strPath 必填 要连接的路径片段,可传入多个

路径分解

from pathlib import Path

p = Path("/home/user/documents/report.final.pdf")

p.parent        # PosixPath('/home/user/documents')
p.parents[0]    # PosixPath('/home/user/documents')
p.parents[1]    # PosixPath('/home/user')
p.parents[2]    # PosixPath('/home')
p.name          # 'report.final.pdf'(含扩展名的文件名)
p.stem          # 'report.final'(不含最后一个扩展名)
p.suffix        # '.pdf'(最后一个扩展名)
p.suffixes      # ['.final', '.pdf'](所有扩展名)
p.parts         # ('/', 'home', 'user', 'documents', 'report.final.pdf')
p.root          # '/'(根路径)
p.anchor        # '/'(根 + 盘符,Windows 上为 'C:\\')
p.drive         # ''(Windows 上为 'C:')

修改路径组成部分

from pathlib import Path

p = Path("/home/user/documents/report.pdf")

# 更换文件名
p.with_name("summary.pdf")
# PosixPath('/home/user/documents/summary.pdf')

# 更换扩展名
p.with_suffix(".txt")
# PosixPath('/home/user/documents/report.txt')

# 去掉扩展名
p.with_suffix("")
# PosixPath('/home/user/documents/report')

# 更换父目录(Python 3.12+)
p.with_segments("/tmp", p.name)
# PosixPath('/tmp/report.pdf')

with_name 的参数:

参数 类型 默认值 说明
name str 必填 新的文件名(含扩展名),不能包含路径分隔符

with_suffix 的参数:

参数 类型 默认值 说明
suffix str 必填 新的扩展名(需以 . 开头),传入 "" 则去掉扩展名

文件读写操作

文本读写

from pathlib import Path

p = Path("data.txt")

# 写入文本
p.write_text("Hello, World!\n第二行", encoding="utf-8")

# 读取文本
content = p.read_text(encoding="utf-8")

write_text 的参数:

参数 类型 默认值 说明
data str 必填 要写入的字符串内容
encoding strNone None(系统默认) 编码格式,建议显式指定 "utf-8"
errors strNone None 编码错误处理策略,如 "replace", "ignore"
newline strNone None 换行符模式,与 open()newline 参数相同

read_text 的参数:

参数 类型 默认值 说明
encoding strNone None(系统默认) 编码格式,建议显式指定
errors strNone None 编码错误处理策略
newline strNone None(Python 3.13+) 换行符模式

二进制读写

from pathlib import Path

p = Path("data.bin")

# 写入二进制
p.write_bytes(b"\x00\x01\x02\x03")

# 读取二进制
data = p.read_bytes()

write_bytes 的参数:

参数 类型 默认值 说明
data bytes 或类字节类型 必填 要写入的二进制数据

read_bytes 无参数,返回 bytes

使用 open() 打开文件

from pathlib import Path

p = Path("data.txt")

# 与内置 open() 用法相同,返回文件对象
with p.open("r", encoding="utf-8") as f:
    content = f.read()

with p.open("w", encoding="utf-8") as f:
    f.write("新内容")

# 追加模式
with p.open("a", encoding="utf-8") as f:
    f.write("\n追加内容")

Path.open() 的参数:

参数 类型 默认值 说明
mode str "r" 打开模式,与内置 open() 相同
buffering int -1 缓冲策略
encoding strNone None 文本模式的编码
errors strNone None 编码错误处理
newline strNone None 换行符处理

目录操作

创建目录

from pathlib import Path

p = Path("/home/user/new_dir")

# 创建单级目录
p.mkdir()

# 创建多级目录(parents=True 自动创建父目录)
deep = Path("/home/user/a/b/c")
deep.mkdir(parents=True, exist_ok=True)

mkdir 的参数:

参数 类型 默认值 说明
mode int 0o777 权限位(Unix),受 umask 影响
parents bool False True 时自动创建所有父目录
exist_ok bool False True 时若目录已存在不抛异常

删除目录

from pathlib import Path

# 只能删除空目录
p = Path("/home/user/empty_dir")
p.rmdir()

# 删除非空目录需要 shutil
import shutil
shutil.rmtree("/home/user/full_dir")

遍历目录

from pathlib import Path

p = Path("/home/user/documents")

# iterdir():遍历直接子项(不递归)
for item in p.iterdir():
    if item.is_file():
        print(f"文件: {item.name}")
    elif item.is_dir():
        print(f"目录: {item.name}")

iterdir 无参数,返回一个生成器,每次产出一个 Path 对象。遍历顺序不保证。

glob 模式匹配

from pathlib import Path

p = Path("/home/user/documents")

# glob:在当前目录下匹配(支持 * 和 ?)
for txt_file in p.glob("*.txt"):
    print(txt_file)

# 匹配直接子目录下的所有 Python 文件
for py_file in p.glob("*/*.py"):
    print(py_file)

# rglob:递归匹配所有层级(相当于 glob("**/*.txt"))
for txt_file in p.rglob("*.txt"):
    print(txt_file)

# 只匹配目录
for d in p.glob("*/"):
    print(d)

glob 的参数:

参数 类型 默认值 说明
pattern str 必填 glob 模式,* 匹配任意文件名字符,** 匹配任意层级目录
case_sensitive boolNone None(跟随系统) 是否区分大小写(Python 3.12+)

rglob 的参数:

参数 类型 默认值 说明
pattern str 必填 glob 模式,自动在前面加 **/ 前缀
case_sensitive boolNone None 是否区分大小写(Python 3.12+)

路径检查

from pathlib import Path

p = Path("/home/user/file.txt")

p.exists()      # 路径是否存在(文件或目录)
p.is_file()     # 是否是普通文件
p.is_dir()      # 是否是目录
p.is_symlink()  # 是否是符号链接
p.is_absolute() # 是否是绝对路径
p.is_relative_to("/home")  # 是否相对于给定路径(Python 3.9+)

stat() 获取文件元信息

from pathlib import Path
import datetime

p = Path("file.txt")
info = p.stat()

info.st_size    # 文件大小(字节)
info.st_mtime   # 最后修改时间(Unix 时间戳)
info.st_ctime   # 创建时间(Windows)/ inode 变更时间(Unix)
info.st_mode    # 文件权限和类型

# 转换时间戳为 datetime
mtime = datetime.datetime.fromtimestamp(info.st_mtime)
print(f"最后修改: {mtime}")

stat 无参数,返回 os.stat_result 对象。若路径不存在则抛出 FileNotFoundError

文件重命名、移动与删除

rename 重命名 / 移动

from pathlib import Path

p = Path("old_name.txt")

# 重命名(返回新路径的 Path 对象)
new_p = p.rename("new_name.txt")

# 移动到其他目录
new_p = p.rename("/home/user/documents/old_name.txt")

rename 的参数:

参数 类型 默认值 说明
target strPath 必填 目标路径,如目标已存在行为取决于平台

replace 强制替换

from pathlib import Path

# replace 会强制覆盖目标文件(原子操作)
p = Path("source.txt")
p.replace("destination.txt")  # 即使 destination.txt 存在也会覆盖

replace 的参数:

参数 类型 默认值 说明
target strPath 必填 目标路径,若存在则原子性地替换
from pathlib import Path

p = Path("file.txt")
p.unlink()                    # 文件不存在时抛出 FileNotFoundError
p.unlink(missing_ok=True)     # 文件不存在时静默忽略(Python 3.8+)

unlink 的参数:

参数 类型 默认值 说明
missing_ok bool False True 时若文件不存在不抛异常

symlink_to 创建符号链接

from pathlib import Path

link = Path("/home/user/link.txt")
link.symlink_to("/home/user/original.txt")

# 读取符号链接目标
target = link.resolve()       # 解析为绝对路径
target = link.readlink()      # 读取链接目标(Python 3.9+)

symlink_to 的参数:

参数 类型 默认值 说明
target strPath 必填 符号链接指向的目标路径
target_is_directory bool False Windows 上创建目录符号链接时需设为 True

路径转换

from pathlib import Path

p = Path("/home/user/file.txt")

# 转为字符串
str(p)                # '/home/user/file.txt'
p.as_posix()          # '/home/user/file.txt'(强制 POSIX 格式)

# 转为绝对路径(不解析符号链接)
p.absolute()

# 解析为规范绝对路径(解析符号链接和 .., .)
p.resolve()

# 计算相对路径
p.relative_to("/home/user")  # PosixPath('file.txt')

# 转为 URI
p.as_uri()            # 'file:///home/user/file.txt'

resolve 的参数:

参数 类型 默认值 说明
strict bool False True 时若路径不存在抛 FileNotFoundError

relative_to 的参数:

参数 类型 默认值 说明
other strPath 必填 基准路径,若当前路径不在其下抛 ValueError
walk_up bool False True 时允许使用 .. 向上跳(Python 3.12+)

os.path 的对比迁移表

os.path 写法 pathlib 写法
os.path.join(a, b) Path(a) / b
os.path.abspath(p) Path(p).resolve()
os.path.dirname(p) Path(p).parent
os.path.basename(p) Path(p).name
os.path.splitext(p)[0] Path(p).stem
os.path.splitext(p)[1] Path(p).suffix
os.path.exists(p) Path(p).exists()
os.path.isfile(p) Path(p).is_file()
os.path.isdir(p) Path(p).is_dir()
os.path.expanduser("~") Path.home()
os.getcwd() Path.cwd()
os.rename(src, dst) Path(src).rename(dst)
os.remove(p) Path(p).unlink()
os.rmdir(p) Path(p).rmdir()
os.makedirs(p, exist_ok=True) Path(p).mkdir(parents=True, exist_ok=True)
glob.glob("*.txt") Path(".").glob("*.txt")
open(p, "r") Path(p).open("r")Path(p).read_text()

os 函数的互操作

很多接受 str 路径的 os 函数也接受 Path 对象(Python 3.6+ 通过 os.fspath 协议):

import os
from pathlib import Path

p = Path("/home/user/file.txt")

os.stat(p)          # 直接传 Path 对象
os.chmod(p, 0o644)
os.environ["PATH"]  # 环境变量仍是字符串

# 需要字符串时用 str() 或 os.fspath()
os.system(f"cat {p}")      # f-string 中 Path 自动转字符串
subprocess.run(["cat", str(p)])

最佳实践

1. 始终显式指定编码

from pathlib import Path

# 不好:依赖系统默认编码,在不同系统上行为不一致
content = Path("file.txt").read_text()

# 好:显式指定 UTF-8
content = Path("file.txt").read_text(encoding="utf-8")

2. 用 with_suffix 安全地修改扩展名

from pathlib import Path

def convert_path(src: Path, new_suffix: str) -> Path:
    return src.with_suffix(new_suffix)

# 批量转换
src_dir = Path("input")
dst_dir = Path("output")
dst_dir.mkdir(exist_ok=True)

for csv_file in src_dir.glob("*.csv"):
    dst_file = dst_dir / csv_file.with_suffix(".parquet").name
    process(csv_file, dst_file)

3. 用 resolve() 消除路径歧义

from pathlib import Path

def safe_open(base_dir: Path, user_input: str) -> str:
    # 防止路径遍历攻击(../../etc/passwd)
    target = (base_dir / user_input).resolve()
    base_dir = base_dir.resolve()
    if not target.is_relative_to(base_dir):
        raise PermissionError(f"禁止访问 base_dir 之外的路径: {target}")
    return target.read_text(encoding="utf-8")

4. 处理大量文件时用生成器

from pathlib import Path

# glob 和 rglob 返回生成器,不会一次性加载所有路径到内存
def count_lines(directory: Path) -> int:
    total = 0
    for py_file in directory.rglob("*.py"):
        try:
            total += len(py_file.read_text(encoding="utf-8").splitlines())
        except (PermissionError, UnicodeDecodeError):
            continue
    return total

踩坑与注意事项

踩坑 1:Windows 路径分隔符

from pathlib import Path

# 在 Windows 上,Path 使用反斜杠,but as_posix() 转换为正斜杠
p = Path("C:/Users/user/file.txt")   # 正斜杠在 Windows 上也可以
p = Path(r"C:\Users\user\file.txt")  # 反斜杠(raw string)

str(p)        # 'C:\\Users\\user\\file.txt'(Windows 上)
p.as_posix()  # 'C:/Users/user/file.txt'(跨平台字符串)

# 不要用普通字符串拼接路径!
bad = "C:\users\new_file.txt"   # \n 和 \u 会被解释为转义字符
good = Path(r"C:\users\new_file.txt")

踩坑 2:mkdirparentsexist_ok 参数缺失

from pathlib import Path

# 错误:父目录不存在时抛 FileNotFoundError
Path("/tmp/a/b/c").mkdir()

# 错误:目录已存在时抛 FileExistsError
Path("/tmp/existing").mkdir()

# 正确:几乎总是应该使用这两个参数
Path("/tmp/a/b/c").mkdir(parents=True, exist_ok=True)

踩坑 3:rename 的跨设备移动限制

from pathlib import Path
import shutil

src = Path("/tmp/file.txt")

# 错误:跨设备(跨磁盘)rename 会抛 OSError
src.rename("/mnt/other_disk/file.txt")

# 正确:跨设备移动用 shutil.move
shutil.move(str(src), "/mnt/other_disk/file.txt")

踩坑 4:glob 的大小写敏感性

from pathlib import Path

# 在 Windows(不区分大小写的文件系统)上:
p = Path("C:/Users")
list(p.glob("*.TXT"))  # 可能匹配 .txt 文件,也可能不匹配,取决于 Python 版本

# Python 3.12+ 用 case_sensitive 参数明确指定
list(p.glob("*.txt", case_sensitive=False))  # 总是不区分大小写

踩坑 5:iterdir 不排序

from pathlib import Path

# iterdir 的顺序是文件系统顺序,不是字母序
p = Path(".")
files = list(p.iterdir())   # 顺序不确定

# 需要排序时:
files = sorted(p.iterdir())                           # 按文件名排序
files = sorted(p.iterdir(), key=lambda f: f.stat().st_mtime)  # 按修改时间

踩坑 6:resolve() 在路径不存在时的行为

from pathlib import Path

# Python 3.6+ 默认 strict=False:即使路径不存在也不报错
p = Path("/nonexistent/path/file.txt").resolve()

# 需要确保路径存在时使用 strict=True
try:
    p = Path("/nonexistent/path").resolve(strict=True)
except FileNotFoundError:
    print("路径不存在")

常见陷阱

陷阱:Path / 运算符右侧使用绝对路径会丢弃左侧

现象: Path('/home/user') / '/etc/passwd' 结果是 PosixPath('/etc/passwd'),左侧路径被忽略。
原因: / 运算符等价于 Path.joinpath(),若右侧是绝对路径,行为与 os.path.join 一致——直接返回右侧路径。
解决:.lstrip('/') 处理右侧路径,或确保右侧始终是相对路径:

# 意外:结果是 /etc/passwd
Path('/home/user') / '/etc/passwd'

# 正确:用 relative 路径
Path('/home/user') / 'etc/passwd'  # /home/user/etc/passwd

陷阱:Path.glob('**/*.py') 在某些 Python 版本中不递归根目录

现象: Path('.').glob('**/*.py') 没有返回当前目录下的直接 .py 文件,只返回子目录中的。
原因: Python 3.11 之前 ** 不匹配当前目录自身,只匹配子目录。
解决: 同时使用 glob('*.py')glob('**/*.py'),或升级到 Python 3.12+(** 语义统一)。

陷阱:Windows 路径与 POSIX 路径混用导致跨平台失败

现象: 代码在 macOS/Linux 正常,Windows 上路径拼接出现反斜杠/正斜杠混用报错。
原因: 硬编码 / 分隔符的字符串路径在 Windows 可能不被某些 API 接受,或 str(path) 输出反斜杠让字符串处理出错。
解决: 始终用 Path 对象操作路径,不手动拼接字符串;需要字符串时用 path.as_posix() 获取 POSIX 格式。


参见

内置函数完全参考
contextlib完全指南

阅读更多

Web 安全基础

1. HTML 转义(服务端渲染必须): 2. CSP(Content Security Policy): 3. HttpOnly Cookie:防止 JS 读取会话 Cookie: 4. 前端框架防护: 攻击者在第三方网站构造一个表单,诱导已登录用户提交,浏览器会自动携带目标站的 Cookie。 触发条件: 1. 用户已登录目标网站(Cookie 有效) 2. 目标 API 仅凭 Cookie 识别用户身份 3. 请求来源未验证 1. CSRF Token(推荐): 2. SameSite Cookie: 3. 验证 Origin/Referer 头:

By yellowdog

HTTP 协议深度指南

HTTP(HyperText Transfer Protocol)是 Web 的基础传输协议,基于 TCP/IP,采用请求/响应模型。 相关文档:Web安全基础(/web-an-quan-ji-chu/) FastAPI完全指南(/fastapi-wan-quan-zhi-nan/) Nginx完全指南(/nginx-wan-quan-zhi-nan/) 幂等性:多次执行相同请求,服务器状态结果相同。PUT /users/1 多次执行结果一致;POST /users 每次创建新资源,非幂等。 浏览器直接从本地缓存读取,不向服务器发送请求。 缓存命中时,状

By yellowdog

系统设计基础

SLA 对照表: 选择建议:无状态服务(Web 层、API 层)优先水平扩展;数据库初期垂直扩展,达到瓶颈后考虑分库分表或读写分离。 缓存穿透(查询不存在的 key,每次都打到 DB): 缓存击穿(热点 key 过期,瞬间大量请求打到 DB): 缓存雪崩(大量 key 同时过期,或缓存服务宕机): 令牌桶 Python 实现: Redis 实现分布式限流(滑动窗口): URL 命名规则: Cursor 分页响应格式: 雪花算法结构(64 bit): 定义:分布式系统不能同时满足以下三个特性: 在分布式环境中 P 是必须保证的,所以实际是 CP vs AP

By yellowdog

算法思路与模板

二分查找要求序列有序,每次将搜索范围缩减一半,时间复杂度 O(log n)。 两个指针从两端向中间收缩,常用于有序数组。 滑动窗口维护一个满足条件的区间 left, right,right 不断向右扩张,条件不满足时收缩 left。 滑动窗口通用框架: 1. 确定"子问题":原问题可以分解为哪些规模更小的同类问题 2. 定义 dpi 或 dpij 的含义,要足够清晰 3. 推导状态转移方程 4. 确定初始状态(边界条件) 5. 确定计算顺序(确保依赖的子问题先计算) 每件物品最多选一次。dpj = 容量为 j 时的最大价值,逆序遍历容量防止重复选取。 每

By yellowdog