Python · 4 分钟阅读
`difflib`:比较序列并生成差异
目录
- 1. 三个核心 API
- 2.
SequenceMatcher:算相似度 - 3.
unified_diff:生成 unified diff - 4.
HtmlDiff:可点亮的网页报告 - 5. 实战:Nginx 配置差异报告
- 6. 找最相似的字符串
- 7. 易错点
- 8. 小结
difflib 是 Python 标准库的 文本/序列对比 工具集。常用于:
- 生成 unified / context diff(与
git diff风格类似) - 计算两段文本的 相似度
- 生成可读的 HTML 差异报告
- 在一组候选里找最相似的字符串
详细实战:业务服务监控(Nginx 配置对比)。
1. 三个核心 API
| API | 作用 |
|---|---|
SequenceMatcher |
任意可哈希序列的相似度比对 |
Differ / unified_diff / ndiff |
输出可读的差异文本 |
HtmlDiff |
输出 HTML 格式的差异(适合发邮件/网页) |
2. SequenceMatcher:算相似度
from difflib import SequenceMatcher
a = "Python is great"
b = "Python is awesome"
m = SequenceMatcher(None, a, b)
print(f"ratio = {m.ratio():.3f}") # 0.741 左右
print("opcodes:", m.get_opcodes())
# [('equal', 0, 9, 0, 9), ('replace', 9, 14, 9, 16), ('equal', 14, 15, 16, 17)]
| 操作码 | 含义 |
|---|---|
equal |
两边完全相同 |
replace |
两边内容不同 |
delete |
仅在第一段中有 |
insert |
仅在第二段中有 |
ratio()=2 * matches / len(a) + len(b),范围[0, 1]。常用于 搜索
提示 / 去重 / 抄袭检测。
2.1 quick_ratio / real_quick_ratio
m.quick_ratio() # 快但粗略
m.real_quick_ratio() # 最快
m.ratio() # 最准也最慢
大量匹配时,先用
quick_ratio过滤,再对命中项跑ratio。
3. unified_diff:生成 unified diff
from difflib import unified_diff
text1 = """\
Python is great.
It is easy to learn.
"""
text2 = """\
Python is awesome.
It is easy to learn.
"""
for line in unified_diff(
text1.splitlines(),
text2.splitlines(),
fromfile="v1",
tofile="v2",
lineterm="",
):
print(line)
输出类似:
--- v1
+++ v2
@@ -1,2 +1,2 @@
-Python is great.
+Python is awesome.
It is easy to learn.
3.1 关键参数
| 参数 | 含义 |
|---|---|
fromfile / tofile |
文件名占位符(仅显示用) |
n |
上下文行数(默认 3) |
lineterm |
行尾字符,文本流里设 "" 避免多换行 |
4. HtmlDiff:可点亮的网页报告
from difflib import HtmlDiff
html = HtmlDiff().make_file(
text1.splitlines(keepends=True),
text2.splitlines(keepends=True),
fromdesc="备份",
todesc="当前",
)
Path("diff.html").write_text(html, encoding="utf-8")
生成的 diff.html 在浏览器里可以红绿高亮对比。
5. 实战:Nginx 配置差异报告
import sys
from pathlib import Path
import difflib
def read_lines(path: str) -> list[str]:
p = Path(path)
if not p.exists():
sys.exit(f"file not found: {p}")
return p.read_text(encoding="utf-8").splitlines(keepends=True)
def main() -> int:
if len(sys.argv) != 3:
print("usage: diff.py <file1> <file2>")
return 2
a = read_lines(sys.argv[1])
b = read_lines(sys.argv[2])
# 1) HTML 报告
Path("diff.html").write_text(
difflib.HtmlDiff().make_file(a, b, fromdesc="a", todesc="b"),
encoding="utf-8",
)
# 2) 粗略统计差异行数
changes = [
l for l in difflib.unified_diff(a, b, lineterm="")
if l.startswith(("+ ", "- "))
]
print(f"diff lines: {len(changes)}")
return 0
if __name__ == "__main__":
sys.exit(main())
想把差异消息告警出去时,先 比较文件哈希(
hashlib.sha256)确定“变了没”,
真有变化再走difflib出报告——见 业务服务监控。
6. 找最相似的字符串
from difflib import get_close_matches
candidates = ["apple", "banana", "grape", "orange"]
print(get_close_matches("pineapple", candidates, n=2, cutoff=0.4))
# ['apple', 'grape']
| 参数 | 含义 |
|---|---|
n |
返回前 N 个 |
cutoff |
相似度阈值(0~1),低于则丢弃 |
7. 易错点
- 行尾换行:
splitlines()会去掉\n,而HtmlDiff期待keepends=True,
否则行尾判断会出问题。 ratio不是归一化“距离”:SequenceMatcher内部算法对超长字符串可能
慢;用quick_ratio预过滤。unified_diff不输出“没有差异”:两个文件完全一致时它返回空,别把它
当成“有变化就一定非空”的判断。- 大文件 diff 慢:先哈希比一下“变了没”,没变就别 diff。
8. 小结
SequenceMatcher:序列相似度,0~1。unified_diff/ndiff:可读差异文本,模仿diff工具。HtmlDiff:网页报告,发邮件/附档方便。get_close_matches:搜索/纠错的常用工具。- 业务上 先哈希后 diff,避免每次都做全量比对。