Python · 5 分钟阅读
Python:用 Selenium 爬取网易云音乐
目录
- 1. 环境准备
- 2. 最小可运行示例
- 3. 更稳妥的写法:
selenium.webdriver.remote.webdriver通用基类 - 4. 常见反爬与缓解
- 5. 比 Selenium 更轻的方案
- 6. 调试技巧
- 7. 合规清单(务必阅读)
⚠️ 重要声明
- 网易云音乐的页面结构、接口参数经常变化,不保证示例长期可用。
- 抓取网易云音乐的歌曲/歌单/音频等受其 服务条款 约束,
请勿用于商业用途 或绕过付费 / 版权限制。- 大量并发抓取可能触发反爬(封 IP、要求登录、返回假数据)。
- 抓取前请认真阅读目标站点的
robots.txt和相关法律法规。
本章定位为 Selenium 入门练习,示例仅演示访问首页 + 解析歌单标题的最小流程。
1. 环境准备
1.1 安装 Selenium
pip install -U selenium
# 或 uv
uv add selenium
1.2 下载 WebDriver
Selenium 需要 与浏览器版本一致 的 WebDriver。
| 浏览器 | WebDriver |
|---|---|
| Chrome | https://googlechromelabs.github.io/chrome-for-testing/ |
| Edge | https://developer.microsoft.com/microsoft-edge/tools/webdriver/ |
| Firefox | https://github.com/mozilla/geckodriver/releases |
| Safari | 内置于 macOS(需要 safari --enable 在 Safari 中打开 Develop 菜单) |
推荐做法:用 webdriver-manager 自动管理
pip install webdriver-manager
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()))
1.3 WebDriver 放在哪?
| 方式 | 路径 | 优缺点 |
|---|---|---|
| 全局 PATH | /usr/local/bin 或 %PATH% |
多个项目共享,最常用 |
| 项目目录 | ./drivers/chromedriver |
多版本管理方便,但部署复杂 |
| 自动管理 | webdriver-manager |
强烈推荐,免去手动同步 |
2. 最小可运行示例
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
def make_driver() -> webdriver.Chrome:
options = webdriver.ChromeOptions()
options.add_argument("--no-sandbox")
options.add_argument("--disable-dev-shm-usage")
# 打开无头模式:服务器 / CI 环境用
# options.add_argument("--headless=new")
return webdriver.Chrome(
service=Service(), # 让 webdriver-manager 自动选 driver
options=options,
)
def main() -> None:
driver = make_driver()
try:
driver.get("https://music.163.com/")
print("title:", driver.title)
# 等到 id="g_iframe" 出现(歌单内容通常在 iframe 中)
WebDriverWait(driver, 10).until(
EC.presence_of_element_located((By.ID, "g_iframe"))
)
driver.switch_to.frame("g_iframe")
# 找歌单里的链接 / 标题(选择器随时可能失效)
titles = driver.find_elements(By.CSS_SELECTOR, "a.msk")
for t in titles[:5]:
print(t.get_attribute("title") or t.text)
finally:
driver.quit()
if __name__ == "__main__":
main()
要点:
- 永远用
try / finally或with包住driver,保证driver.quit()关闭浏览器。 WebDriverWait显式等待比time.sleep稳定很多。- 网易云的很多内容在 iframe 中,记得
switch_to.frame。 - 关闭浏览器时浏览器进程可能不退出,再加一句
driver.quit()即可。
3. 更稳妥的写法:selenium.webdriver.remote.webdriver 通用基类
如果你的脚本会在 Chrome / Edge / Firefox 之间切换,建议用 webdriver-manager
管理多浏览器,并抽象一个工厂函数:
from selenium import webdriver
from webdriver_manager.chrome import ChromeDriverManager
from webdriver_manager.firefox import GeckoDriverManager
from selenium.webdriver.chrome.service import Service as ChromeService
from selenium.webdriver.firefox.service import Service as FirefoxService
def make_driver(browser: str = "chrome", headless: bool = True) -> webdriver.Remote:
if browser == "chrome":
options = webdriver.ChromeOptions()
if headless:
options.add_argument("--headless=new")
return webdriver.Chrome(
service=ChromeService(ChromeDriverManager().install()),
options=options,
)
if browser == "firefox":
options = webdriver.FirefoxOptions()
if headless:
options.add_argument("-headless")
return webdriver.Firefox(
service=FirefoxService(GeckoDriverManager().install()),
options=options,
)
raise ValueError(f"unsupported browser: {browser}")
4. 常见反爬与缓解
| 现象 | 排查 / 缓解 |
|---|---|
selenium.common.exceptions.WebDriverException |
WebDriver 与浏览器版本不匹配。pip install -U webdriver-manager |
| 打开空白 / 一直转圈 | 检测 UA / 用 WebDriverWait 等元素 |
| “检测到自动化工具” | 用 undetected-chromedriver,加 --disable-blink-features=AutomationControlled |
| 频繁封 IP | 限速 + 代理池(前提是合法) |
| 登录后才返回真实数据 | 用 Selenium 自动登录,但请勿绕过付费内容 |
# 让 Selenium 看起来更像真实用户
options.add_experimental_option("excludeSwitches", ["enable-automation"])
options.add_experimental_option("useAutomationExtension", False)
options.add_argument(
"user-agent=Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) "
"AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0.0.0 Safari/537.36"
)
5. 比 Selenium 更轻的方案
当目标页面有公开 API 时,优先用 HTTP 而非浏览器:
| 方案 | 适合 |
|---|---|
requests |
接口是公开的、或只是 GET/POST JSON |
httpx |
同步/异步都要用 |
playwright |
浏览器自动化,但 API 更现代、速度更快 |
pyppeteer |
Chrome DevTools Protocol(已被 playwright 取代) |
用浏览器爬一切是“最慢、也最容易被识别”的方案。除非页面是 重度 SPA / 强
JavaScript 渲染,否则应优先尝试直接请求 API。
6. 调试技巧
- 截图:
driver.save_screenshot("debug.png") - 查看页面源码:
print(driver.page_source) - 浏览器控制台:
driver.get_log("browser") - 断点:在脚本里
import pdb; pdb.set_trace()后手工driver.get(...)。 - 用
--user-data-dir复用同一个浏览器配置,避开重复登录。
7. 合规清单(务必阅读)
- ✅ 只抓取 公开 且 允许爬取 的内容。
- ✅ 在请求中加入
User-Agent并注明联系方式。 - ✅ 控制频率(一般 ≤ 1 QPS),避开高峰。
- ❌ 不要绕过登录、付费墙、地区限制。
- ❌ 不要把抓到的内容二次分发起诉。
- ❌ 不要绕过
robots.txt(https://music.163.com/robots.txt 自行查阅)。