R / Richie全部文章 ↑

Python · 3 分钟阅读

案例研究:文本统计工具

目录


本章以一个 文本统计 CLI 为例,串联前面学到的:

  • 文件 IO(逐行读取)
  • 字符串处理
  • 集合 / 计数器
  • CLI 解析
  • 大文件分块

最终产出一个能跑的命令行工具 textstat,统计:

  • 总字符数 / 行数 / 单词数
  • 词频 Top-N
  • 出现过的字符集合

1. 项目结构

textstat/
├── pyproject.toml
├── README.md
└── src/
    └── textstat/
        ├── __init__.py
        ├── cli.py
        ├── stats.py
        └── normalize.py

1.1 pyproject.toml

[project]
name = "textstat"
version = "0.1.0"
requires-python = ">=3.10"
dependencies = []

[project.scripts]
textstat = "textstat.cli:main"

[build-system]
requires = ["hatchling"]
build-backend = "hatchling.build"

2. 文本标准化

2.1 设计

为了统计“英语文本里的词频”,我们需要把句子规整成可比较的 token:

  • 全部转小写;
  • 保留 字母、空格、连字符 -、撇号 ';
  • 其余字符当分隔符。

不同业务有不同的规则。比如:

  • 中文文本应当用 jieba 之类的分词库,而不是按空格;
  • 代码 / 日志统计词频意义不大,可能直接按 token。

2.2 normalize.py

"""文本标准化:转小写 + 保留有效字符。"""
from __future__ import annotations

VALID_CHARS = set(
    "abcdefghijklmnopqrstuvwxyz"
    "0123456789"
    " -'"
)


def normalize(text: str) -> str:
    """把 text 转小写,只保留有效字符。"""
    return "".join(c for c in text.lower() if c in VALID_CHARS)

进一步用 str.translate() 会更快一些,但在 1 MB 级别文件上用生成器
表达式已经够快。

2.3 分词

def tokenize(text: str) -> list[str]:
    """用空白切分。"""
    return text.split()

3. 统计逻辑:stats.py

from __future__ import annotations

from collections import Counter
from dataclasses import dataclass
from typing import Iterable

from .normalize import normalize, tokenize


@dataclass
class TextStats:
    chars: int
    lines: int
    words: int
    top_words: list[tuple[str, int]]
    unique_chars: set[str]


def compute_stats(text: str, top_n: int = 10) -> TextStats:
    """对一段文本计算统计信息。"""
    if not text:
        return TextStats(0, 0, 0, [], set())

    # 基础计数
    chars = len(text)
    # splitlines 不带换行符;用 '\n' split 更稳,但 splitlines() 更通用
    lines = text.count("\n") + (0 if text.endswith("\n") else 1)

    norm = normalize(text)
    words = tokenize(norm)

    counter = Counter(words)
    return TextStats(
        chars=chars,
        lines=lines,
        words=len(words),
        top_words=counter.most_common(top_n),
        unique_chars=set(norm),
    )


def stream_chunks(fp, chunk_size: int = 64 * 1024) -> Iterable[str]:
    """生成器:按块从文件中读文本,避免一次性加载大文件。"""
    while True:
        chunk = fp.read(chunk_size)
        if not chunk:
            return
        yield chunk

大文件思路:分块读取 → 逐块统计 → 累加 Counter;行数也可以每块用
chunk.count("\n") 累加。


4. 大文件分块版

from collections import Counter
from typing import IO


def compute_stats_file(fp: IO[str], top_n: int = 10) -> TextStats:
    """流式统计:常驻内存只有一块。"""
    chars = 0
    lines = 0
    counter: Counter[str] = Counter()
    unique: set[str] = set()

    for chunk in stream_chunks(fp):
        chars += len(chunk)
        lines += chunk.count("\n")
        norm = normalize(chunk)
        unique.update(norm)
        counter.update(norm.split())

    return TextStats(
        chars=chars,
        lines=lines,
        words=sum(counter.values()),
        top_words=counter.most_common(top_n),
        unique_chars=unique,
    )

注意:分块统计时,行数 = 全部换行符数 +(最后一块不以 \n 结尾时 +1)。
上面代码里没做这个修正,简化了;如要严谨可记录“是否以换行结尾”。


5. CLI 入口:cli.py

from __future__ import annotations

import argparse
import sys
from pathlib import Path

from .stats import compute_stats_file, compute_stats


def parse_args(argv: list[str] | None = None) -> argparse.Namespace:
    p = argparse.ArgumentParser(
        prog="textstat",
        description="统计文本文件的字符 / 行 / 词 / 高频词",
    )
    p.add_argument("path", type=Path, help="要分析的文本文件")
    p.add_argument("--top", type=int, default=10, help="输出前 N 个高频词")
    p.add_argument("--show-chars", action="store_true", help="输出字符集合")
    return p.parse_args(argv)


def main(argv: list[str] | None = None) -> int:
    args = parse_args(argv)
    if not args.path.exists():
        print(f"file not found: {args.path}", file=sys.stderr)
        return 2

    with args.path.open("r", encoding="utf-8", errors="replace") as f:
        # 演示用整文件版;如果文件 > 几百 MB 换 compute_stats_file(f, args.top)
        text = f.read()

    stats = compute_stats(text, top_n=args.top)

    print(f"chars : {stats.chars}")
    print(f"lines : {stats.lines}")
    print(f"words : {stats.words}")
    print(f"top {args.top}:")
    for word, n in stats.top_words:
        print(f"  {word:<15} {n}")
    if args.show_chars:
        print("unique chars:")
        print("  " + "".join(sorted(stats.unique_chars)))

    return 0


if __name__ == "__main__":
    raise SystemExit(main())

5.1 运行

pip install -e .
textstat README.md
textstat README.md --top 5 --show-chars

5.2 演示输入 / 输出

$ echo "A long time ago, in a galaxy far, far away..." | textstat -
chars : 46
lines : 1
words : 10
top 10:
  a            2
  far          2
  ago          1
  away         1
  galaxy       1
  in           1
  long         1
  time         1

6. 单元测试

from textstat.normalize import normalize, tokenize
from textstat.stats import compute_stats


def test_normalize():
    assert normalize("I'd like a Copy!") == "i'd like a copy"


def test_tokenize():
    assert tokenize("a  b c") == ["a", "b", "c"]


def test_compute_stats():
    text = "A long time ago, in a galaxy far, far away..."
    s = compute_stats(text)
    assert s.chars == len(text)
    assert s.lines == 1
    assert s.words == 10
    assert s.top_words[0] in {("a", 2), ("far", 2)}

7. 设计回顾

步骤 决策 理由
输入 显式 encoding="utf-8" 避免平台默认编码差异
标准化 集合查找 + 生成器 简单可读,足够快
统计 Counter 统计词频 标准库够用,无须第三方
大文件 分块 + 累加 Counter 内存常驻 O(块大小)
输出 argparse + 表格打印 易扩展

8. 进阶方向

  • 中文支持:接入 jieba,按词统计。
  • 多文件 / 目录:扩展为 textstat dir/*.txt,对每个文件分别输出。
  • JSON / CSV 输出:加 --format json 选项,方便二次处理。
  • 性能优化:用 str.translate() 加速标准化;用 re.finditer(r"\w+", text)
    替代 split(),并支持 Unicode 词边界。
  • TUI / Web 界面:在 compute_stats 之上加一个交互层。

9. 常见问题

  • 统计行数不对? 末尾是否有换行会让结果差 1。明确说明你的工具“是否把
    文件末尾空行算作一行”。
  • 词频不直观? 停用词(the / a / of …)拉低了有意义词的排名。可以
    维护一个停用词表过滤。
  • 中文输出乱码? 终端设 UTF-8:export LANG=en_US.UTF-8 或 chcp 65001。
  • 统计太慢? 优先排查 normalize:每次循环都做 if c in set;改用
    str.translate 预编译表,能提速 5~10x。