R / Richie全部文章 ↑

Python · 4 分钟阅读

`difflib`:比较序列并生成差异

目录


difflib 是 Python 标准库的 文本/序列对比 工具集。常用于:

  • 生成 unified / context diff(与 git diff 风格类似)
  • 计算两段文本的 相似度
  • 生成可读的 HTML 差异报告
  • 在一组候选里找最相似的字符串

详细实战:业务服务监控(Nginx 配置对比)。


1. 三个核心 API

API 作用
SequenceMatcher 任意可哈希序列的相似度比对
Differ / unified_diff / ndiff 输出可读的差异文本
HtmlDiff 输出 HTML 格式的差异(适合发邮件/网页)

2. SequenceMatcher:算相似度

from difflib import SequenceMatcher

a = "Python is great"
b = "Python is awesome"

m = SequenceMatcher(None, a, b)
print(f"ratio = {m.ratio():.3f}")    # 0.741 左右
print("opcodes:", m.get_opcodes())
# [('equal', 0, 9, 0, 9), ('replace', 9, 14, 9, 16), ('equal', 14, 15, 16, 17)]
操作码 含义
equal 两边完全相同
replace 两边内容不同
delete 仅在第一段中有
insert 仅在第二段中有

ratio() = 2 * matches / len(a) + len(b),范围 [0, 1]。常用于 搜索
提示
/ 去重 / 抄袭检测。

2.1 quick_ratio / real_quick_ratio

m.quick_ratio()         # 快但粗略
m.real_quick_ratio()    # 最快
m.ratio()               # 最准也最慢

大量匹配时,先用 quick_ratio 过滤,再对命中项跑 ratio。


3. unified_diff:生成 unified diff

from difflib import unified_diff

text1 = """\
Python is great.
It is easy to learn.
"""
text2 = """\
Python is awesome.
It is easy to learn.
"""

for line in unified_diff(
    text1.splitlines(),
    text2.splitlines(),
    fromfile="v1",
    tofile="v2",
    lineterm="",
):
    print(line)

输出类似:

--- v1
+++ v2
@@ -1,2 +1,2 @@
-Python is great.
+Python is awesome.
 It is easy to learn.

3.1 关键参数

参数 含义
fromfile / tofile 文件名占位符(仅显示用)
n 上下文行数(默认 3)
lineterm 行尾字符,文本流里设 "" 避免多换行

4. HtmlDiff:可点亮的网页报告

from difflib import HtmlDiff

html = HtmlDiff().make_file(
    text1.splitlines(keepends=True),
    text2.splitlines(keepends=True),
    fromdesc="备份",
    todesc="当前",
)
Path("diff.html").write_text(html, encoding="utf-8")

生成的 diff.html 在浏览器里可以红绿高亮对比。


5. 实战:Nginx 配置差异报告

import sys
from pathlib import Path
import difflib


def read_lines(path: str) -> list[str]:
    p = Path(path)
    if not p.exists():
        sys.exit(f"file not found: {p}")
    return p.read_text(encoding="utf-8").splitlines(keepends=True)


def main() -> int:
    if len(sys.argv) != 3:
        print("usage: diff.py <file1> <file2>")
        return 2

    a = read_lines(sys.argv[1])
    b = read_lines(sys.argv[2])

    # 1) HTML 报告
    Path("diff.html").write_text(
        difflib.HtmlDiff().make_file(a, b, fromdesc="a", todesc="b"),
        encoding="utf-8",
    )

    # 2) 粗略统计差异行数
    changes = [
        l for l in difflib.unified_diff(a, b, lineterm="")
        if l.startswith(("+ ", "- "))
    ]
    print(f"diff lines: {len(changes)}")
    return 0


if __name__ == "__main__":
    sys.exit(main())

想把差异消息告警出去时,先 比较文件哈希(hashlib.sha256)确定“变了没”,
真有变化再走 difflib 出报告——见 业务服务监控。


6. 找最相似的字符串

from difflib import get_close_matches

candidates = ["apple", "banana", "grape", "orange"]
print(get_close_matches("pineapple", candidates, n=2, cutoff=0.4))
# ['apple', 'grape']
参数 含义
n 返回前 N 个
cutoff 相似度阈值(0~1),低于则丢弃

7. 易错点

  • 行尾换行:splitlines() 会去掉 \n,而 HtmlDiff 期待 keepends=True,
    否则行尾判断会出问题。
  • ratio 不是归一化“距离”:SequenceMatcher 内部算法对超长字符串可能
    慢;用 quick_ratio 预过滤。
  • unified_diff 不输出“没有差异”:两个文件完全一致时它返回空,别把它
    当成“有变化就一定非空”的判断
    。
  • 大文件 diff 慢:先哈希比一下“变了没”,没变就别 diff。

8. 小结

  • SequenceMatcher:序列相似度,0~1。
  • unified_diff / ndiff:可读差异文本,模仿 diff 工具。
  • HtmlDiff:网页报告,发邮件/附档方便。
  • get_close_matches:搜索/纠错的常用工具。
  • 业务上 先哈希后 diff,避免每次都做全量比对。