Python · 3 分钟阅读
案例研究:文本统计工具
目录
本章以一个 文本统计 CLI 为例,串联前面学到的:
- 文件 IO(逐行读取)
- 字符串处理
- 集合 / 计数器
- CLI 解析
- 大文件分块
最终产出一个能跑的命令行工具 textstat,统计:
- 总字符数 / 行数 / 单词数
- 词频 Top-N
- 出现过的字符集合
1. 项目结构
textstat/
├── pyproject.toml
├── README.md
└── src/
└── textstat/
├── __init__.py
├── cli.py
├── stats.py
└── normalize.py
1.1 pyproject.toml
[project]
name = "textstat"
version = "0.1.0"
requires-python = ">=3.10"
dependencies = []
[project.scripts]
textstat = "textstat.cli:main"
[build-system]
requires = ["hatchling"]
build-backend = "hatchling.build"
2. 文本标准化
2.1 设计
为了统计“英语文本里的词频”,我们需要把句子规整成可比较的 token:
- 全部转小写;
- 保留 字母、空格、连字符
-、撇号'; - 其余字符当分隔符。
不同业务有不同的规则。比如:
- 中文文本应当用
jieba之类的分词库,而不是按空格;- 代码 / 日志统计词频意义不大,可能直接按 token。
2.2 normalize.py
"""文本标准化:转小写 + 保留有效字符。"""
from __future__ import annotations
VALID_CHARS = set(
"abcdefghijklmnopqrstuvwxyz"
"0123456789"
" -'"
)
def normalize(text: str) -> str:
"""把 text 转小写,只保留有效字符。"""
return "".join(c for c in text.lower() if c in VALID_CHARS)
进一步用
str.translate()会更快一些,但在 1 MB 级别文件上用生成器
表达式已经够快。
2.3 分词
def tokenize(text: str) -> list[str]:
"""用空白切分。"""
return text.split()
3. 统计逻辑:stats.py
from __future__ import annotations
from collections import Counter
from dataclasses import dataclass
from typing import Iterable
from .normalize import normalize, tokenize
@dataclass
class TextStats:
chars: int
lines: int
words: int
top_words: list[tuple[str, int]]
unique_chars: set[str]
def compute_stats(text: str, top_n: int = 10) -> TextStats:
"""对一段文本计算统计信息。"""
if not text:
return TextStats(0, 0, 0, [], set())
# 基础计数
chars = len(text)
# splitlines 不带换行符;用 '\n' split 更稳,但 splitlines() 更通用
lines = text.count("\n") + (0 if text.endswith("\n") else 1)
norm = normalize(text)
words = tokenize(norm)
counter = Counter(words)
return TextStats(
chars=chars,
lines=lines,
words=len(words),
top_words=counter.most_common(top_n),
unique_chars=set(norm),
)
def stream_chunks(fp, chunk_size: int = 64 * 1024) -> Iterable[str]:
"""生成器:按块从文件中读文本,避免一次性加载大文件。"""
while True:
chunk = fp.read(chunk_size)
if not chunk:
return
yield chunk
大文件思路:分块读取 → 逐块统计 → 累加
Counter;行数也可以每块用
chunk.count("\n")累加。
4. 大文件分块版
from collections import Counter
from typing import IO
def compute_stats_file(fp: IO[str], top_n: int = 10) -> TextStats:
"""流式统计:常驻内存只有一块。"""
chars = 0
lines = 0
counter: Counter[str] = Counter()
unique: set[str] = set()
for chunk in stream_chunks(fp):
chars += len(chunk)
lines += chunk.count("\n")
norm = normalize(chunk)
unique.update(norm)
counter.update(norm.split())
return TextStats(
chars=chars,
lines=lines,
words=sum(counter.values()),
top_words=counter.most_common(top_n),
unique_chars=unique,
)
注意:分块统计时,行数 = 全部换行符数 +(最后一块不以
\n结尾时 +1)。
上面代码里没做这个修正,简化了;如要严谨可记录“是否以换行结尾”。
5. CLI 入口:cli.py
from __future__ import annotations
import argparse
import sys
from pathlib import Path
from .stats import compute_stats_file, compute_stats
def parse_args(argv: list[str] | None = None) -> argparse.Namespace:
p = argparse.ArgumentParser(
prog="textstat",
description="统计文本文件的字符 / 行 / 词 / 高频词",
)
p.add_argument("path", type=Path, help="要分析的文本文件")
p.add_argument("--top", type=int, default=10, help="输出前 N 个高频词")
p.add_argument("--show-chars", action="store_true", help="输出字符集合")
return p.parse_args(argv)
def main(argv: list[str] | None = None) -> int:
args = parse_args(argv)
if not args.path.exists():
print(f"file not found: {args.path}", file=sys.stderr)
return 2
with args.path.open("r", encoding="utf-8", errors="replace") as f:
# 演示用整文件版;如果文件 > 几百 MB 换 compute_stats_file(f, args.top)
text = f.read()
stats = compute_stats(text, top_n=args.top)
print(f"chars : {stats.chars}")
print(f"lines : {stats.lines}")
print(f"words : {stats.words}")
print(f"top {args.top}:")
for word, n in stats.top_words:
print(f" {word:<15} {n}")
if args.show_chars:
print("unique chars:")
print(" " + "".join(sorted(stats.unique_chars)))
return 0
if __name__ == "__main__":
raise SystemExit(main())
5.1 运行
pip install -e .
textstat README.md
textstat README.md --top 5 --show-chars
5.2 演示输入 / 输出
$ echo "A long time ago, in a galaxy far, far away..." | textstat -
chars : 46
lines : 1
words : 10
top 10:
a 2
far 2
ago 1
away 1
galaxy 1
in 1
long 1
time 1
6. 单元测试
from textstat.normalize import normalize, tokenize
from textstat.stats import compute_stats
def test_normalize():
assert normalize("I'd like a Copy!") == "i'd like a copy"
def test_tokenize():
assert tokenize("a b c") == ["a", "b", "c"]
def test_compute_stats():
text = "A long time ago, in a galaxy far, far away..."
s = compute_stats(text)
assert s.chars == len(text)
assert s.lines == 1
assert s.words == 10
assert s.top_words[0] in {("a", 2), ("far", 2)}
7. 设计回顾
| 步骤 | 决策 | 理由 |
|---|---|---|
| 输入 | 显式 encoding="utf-8" |
避免平台默认编码差异 |
| 标准化 | 集合查找 + 生成器 | 简单可读,足够快 |
| 统计 | Counter 统计词频 |
标准库够用,无须第三方 |
| 大文件 | 分块 + 累加 Counter |
内存常驻 O(块大小) |
| 输出 | argparse + 表格打印 |
易扩展 |
8. 进阶方向
- 中文支持:接入
jieba,按词统计。 - 多文件 / 目录:扩展为
textstat dir/*.txt,对每个文件分别输出。 - JSON / CSV 输出:加
--format json选项,方便二次处理。 - 性能优化:用
str.translate()加速标准化;用re.finditer(r"\w+", text)
替代split(),并支持 Unicode 词边界。 - TUI / Web 界面:在
compute_stats之上加一个交互层。
9. 常见问题
- 统计行数不对? 末尾是否有换行会让结果差 1。明确说明你的工具“是否把
文件末尾空行算作一行”。 - 词频不直观? 停用词(
the/a/of…)拉低了有意义词的排名。可以
维护一个停用词表过滤。 - 中文输出乱码? 终端设 UTF-8:
export LANG=en_US.UTF-8或chcp 65001。 - 统计太慢? 优先排查
normalize:每次循环都做if c in set;改用
str.translate预编译表,能提速 5~10x。