R / Richie全部文章 ↑

Python · 9 分钟阅读

I/O:输入与输出

程序要能运行,必须能 与外界沟通:从控制台 / 文件 / 网络 / 数据库读入数据,再把结果输出到这些地方。本章用一份"可运行示例清单"覆盖日常 I/O 编程的主要场景。

推荐使用现代 API:

  • 文件:pathlib + open()(参见下文 §2「文件读写」)
  • 进程:subprocess
  • 网络:requests / httpx
  • 序列化:json、pickle(仅限可信数据)

1. 控制台 I/O

1.1 打印

print("Hello, world!")
print("a", "b", "c", sep="-", end="\n\n")

1.2 读取一行

name = input("请输入你的名字: ")      # 始终返回 str
age  = int(input("请输入年龄: "))     # 必要时手动转 int

input() 抛 EOFError 时表示到达文件末尾(管道输入很常见)。捕获示例:

try:
    line = input("> ")
except EOFError:
    return

1.3 字符串格式化

方式 推荐度 示例
f-string ★★★ f"name={name}, age={age}"
str.format() ★★ "name={n}".format(n=name)
% 格式化 ★ "name=%s" % name(老代码)
x = 1 / 81
print(f"value: {x:.5f}")        # 0.01235
print(f"hex  : {255:#x}")        # 0xff
print(f"pad  : {42:>5}")         # '   42'

转换说明符速查:

说明符 含义
d 整数
o / x / X 八 / 十六进制
e / E 科学记数法
f 浮点
s 字符串
% % 字符

简单场景直接用 f-string;str.format() 用于动态模板;% 格式化只在维护老代码时用。


2. 文件读写

文件读写是几乎所有程序都会做的事:配置、日志、数据持久化、网络 IO 都离不开它。Python 的 open() + with 提供了简洁、安全的 API;pathlib 进一步把"路径 + 文件"做成面向对象。

2.1 打开文件

file_object = open(file, mode="r", encoding=None, newline=None)
参数 作用
file 文件路径(str 或 Path)
mode 模式(见下表),默认 "r"
encoding 文本模式下的编码,建议显式传 "utf-8"
newline None 时自动转换;"" 保留原始换行

模式速查:

模式 含义
r / rt 只读(默认)
w / wt 写入(先清空)
a 末尾追加
x 独占创建(已存在则抛 FileExistsError)
b 二进制(与上述组合)
+ 可读可写(与上述组合)

完整示例:

from pathlib import Path

p = Path("data/notes.txt")
p.parent.mkdir(parents=True, exist_ok=True)   # 自动建目录

with p.open("w", encoding="utf-8") as f:
    f.write("Hello, world!\n")

with 语句保证文件一定被关闭,即便中途抛异常。

2.2 读取文件

一次性读全部:

text = Path("data.txt").read_text(encoding="utf-8")

适合:文件较小(< 几十 MB);要交给正则 / 解析器处理。

逐行迭代(推荐):

with Path("data.txt").open("r", encoding="utf-8") as f:
    for i, line in enumerate(f, start=1):
        print(i, line.rstrip())
  • 内存常驻只有一行,GB 级文件也安全。
  • line 末尾会带 \n,通常 rstrip() 一下。

read() / readline() / readlines():

方法 行为
f.read(n) 读 n 个字符(文本模式)或字节(二进制)
f.readline() 读下一行(含 \n)
f.readlines() 把所有行读到列表(小心大文件)
with open("data.txt", encoding="utf-8") as f:
    head = f.readline()       # 第一行
    body = f.read()            # 剩余全部

pathlib 风格:

操作 写法
读全部文本 p.read_text(encoding="utf-8")
读全部字节 p.read_bytes()
一次性写出 p.write_text(text, encoding="utf-8")
一次性写字节 p.write_bytes(data)

大文件慎用 read_*(),会一次性加载到内存。

2.3 写入文件

with open("out.txt", "w", encoding="utf-8") as f:
    f.write("first line\n")
    f.writelines(["second\n", "third\n"])

注意:

  • writelines 不会自动加 \n,要自己写。
  • 文本模式默认换行转换;想完全控制用 newline=""。

追加 / 独占创建:

with open("log.txt", "a", encoding="utf-8") as f:
    f.write("appended\n")

# 不覆盖已有文件
with open("config.yaml", "x", encoding="utf-8") as f:
    f.write("k: v\n")

把字符串插到文件开头:

文件开头插入没有 O(1) 方案,最简单的办法是 整体读 → 修改 → 整体写:

def insert_title(title: str, fname: str = "story.txt") -> None:
    p = Path(fname)
    p.write_text(title + "\n\n" + p.read_text(encoding="utf-8"),
                 encoding="utf-8")

大文件场景可以把 read_text 换成按块流式处理。

2.4 二进制文件

def is_gif(path: str | Path) -> bool:
    with open(path, "rb") as f:
        return f.read(4) == b"GIF8"

二进制模式下:

  • 拿到的不是 str 而是 bytes。
  • 不会做任何换行符转换。
  • f.read(n) 是 n 个 字节。

2.5 实用技巧

保留原始换行:

with open("data.txt", encoding="utf-8", newline="") as f:
    for line in f:        # 行尾就是文件原始换行符
        ...

newline="" 关闭自动转换,避免在 Windows 上行尾 \r\n 被吃一半。

跳过表头:

with open("data.csv", encoding="utf-8") as f:
    next(f)                                # 跳过表头
    for line in f:
        ...

容错读取非 UTF-8 文件:

with open("legacy.txt", encoding="utf-8", errors="replace") as f:
    text = f.read()
  • errors="strict"(默认):坏字节抛 UnicodeDecodeError
  • errors="ignore":丢弃坏字节
  • errors="replace":替换为 ?

临时文件:

import tempfile

with tempfile.NamedTemporaryFile("w", delete=False, suffix=".txt") as f:
    f.write("tmp")
    print("path:", f.name)

想要更稳的临时文件生命周期可以用 tempfile.TemporaryDirectory()。

2.6 常见错误

错误信息 原因 解决
FileNotFoundError 路径不存在 先 Path.parent.mkdir(parents=True)
PermissionError 没权限 ls -l / 调整权限
UnicodeDecodeError 编码不一致 显式 encoding="utf-8",或 errors 容错
写入文件没有内容 没调用 flush() 或没 close() 用 with,或 f.flush()
readlines() 卡住大文件 一次性读入 改成 for line in f

2.7 选型清单

场景 推荐
小文件、一次性读写 Path.read_text / write_text
大文件 / 边读边处理 with open(...) as f: for line in f
写日志 open(..., "a"),按行 write
配置文件不存在则创建 open(..., "x")
二进制数据 / 图片 / 网络下载 open(..., "rb") / "wb"
临时文件 tempfile

3. 路径处理

from pathlib import Path

p = Path("/usr/local/bin/python3")
print(p.name)        # python3
print(p.parent)      # /usr/local/bin
print(p.suffix)      # ''
print(p.exists())    # True / False
# 旧 API
os.getcwd()                        # 当前工作目录
os.listdir(".")                    # 列出
os.path.exists("a.txt")            # 存在判断

# 新 API(推荐)
p = Path("a.txt")
p.exists()
p.parent.glob("*.txt")             # 通配

os.path 还能用,但 pathlib 是面向对象、跨平台、与 open 协作更顺。


4. 序列化

4.1 JSON(推荐)

import json

data = {"name": "Alice", "scores": [90, 88, 75]}

# 写文件
with open("out.json", "w", encoding="utf-8") as f:
    json.dump(data, f, ensure_ascii=False, indent=2)

# 读文件
with open("out.json", "r", encoding="utf-8") as f:
    obj = json.load(f)

# 字符串 <-> 对象
text = json.dumps(data, ensure_ascii=False)
obj  = json.loads(text)

ensure_ascii=False 保留中文字符;indent=2 让输出可读。

4.2 pickle(仅可信数据)

import pickle

data = {"users": [{"id": 1, "name": "alice"}]}

with open("data.pkl", "wb") as f:
    pickle.dump(data, f)

with open("data.pkl", "rb") as f:
    obj = pickle.load(f)

⚠️ 不要加载不信任来源的 pickle 文件——反序列化可以执行任意代码。

4.3 shelve:带 key 的简单持久化

import shelve

with shelve.open("store") as db:
    db["user:1"] = {"name": "alice"}
    print(db["user:1"])

底层是 dbm,轻量、键值访问,但 不适合并发写。


5. 数据库:sqlite3

import sqlite3

with sqlite3.connect("app.db") as conn:
    conn.execute("""
        CREATE TABLE IF NOT EXISTS users (
            id   INTEGER PRIMARY KEY,
            name TEXT NOT NULL,
            email TEXT UNIQUE
        )
    """)
    conn.execute(
        "INSERT OR REPLACE INTO users(name, email) VALUES (?, ?)",
        ("Alice", "alice@example.com"),
    )
    conn.commit()

    for row in conn.execute("SELECT id, name, email FROM users"):
        print(row)
  • 内存数据库 :memory: 可做单元测试。
  • 真正高并发写需要换 MySQL / PostgreSQL / sqlmodel + psycopg。

6. 进程 / Shell

import subprocess

r = subprocess.run(
    ["ls", "-l"],
    capture_output=True,
    text=True,
    timeout=5,
    check=True,
)
print(r.stdout)
选项 作用
capture_output 捕获 stdout / stderr
text=True 字符串返回(默认字节)
timeout=5 秒级超时
check=True 非零退出码抛 CalledProcessError

永远用列表传参(["ls", "-l"]),不要拼字符串。shell=True 留给必须用 && / | 的场景,且要小心注入。


7. 网络 I/O

7.1 读取网页(requests)

import requests

resp = requests.get("https://example.com", timeout=5)
resp.raise_for_status()
print(resp.text[:200])

requests 的好处:自动重定向、连接池、Cookie、Session、SSL 校验。

7.2 用 with 控制 HTTP 连接(httpx)

import httpx

with httpx.Client() as client:
    r = client.get("https://example.com", timeout=5)
    print(r.status_code, r.text[:80])

httpx 同步 / 异步一套 API,更适合现代服务。异步版用 httpx.AsyncClient + await。

7.3 浏览器自动化

webbrowser 只是把 URL 丢给系统默认浏览器:

import webbrowser
webbrowser.open("https://www.python.org")

真正在代码里控制浏览器(点击、填表、抓数据)请用 Playwright 或 Selenium。


8. 常见问题

  • 打印中文乱码? 终端 LANG=zh_CN.UTF-8;Python 3 默认 utf-8,多数情况没问题。
  • 读文件时夹带 \r\n? 用 open(..., newline="") 或 splitlines()。
  • json.dump 报错 TypeError: Object of type X is not JSON serializable? 写自定义 default=:
    json.dump(obj, f, default=lambda o: o.__dict__)
    
  • pickle 跨语言/跨版本? 不能。跨语言用 json;跨版本用 json 或 msgpack。
  • 大文件占内存? 用流式 for line in f,或 Path.read_bytes() + 自定义分块解析。

9. 小结

场景 推荐工具
控制台 I/O input / print / f-string
文本文件 pathlib + open(..., encoding)
二进制文件 open(..., "rb")
序列化 json(首选) / pickle(可信)
简单 KV 存储 shelve / sqlite3
进程 / Shell subprocess.run
HTTP requests / httpx
浏览器自动化 Playwright / Selenium
跨平台路径 pathlib