R / Richie全部文章 ↑

Python · 5 分钟阅读

Python:用 Selenium 爬取网易云音乐

目录


⚠️ 重要声明

  1. 网易云音乐的页面结构、接口参数经常变化,不保证示例长期可用。
  2. 抓取网易云音乐的歌曲/歌单/音频等受其 服务条款 约束,
    请勿用于商业用途 或绕过付费 / 版权限制。
  3. 大量并发抓取可能触发反爬(封 IP、要求登录、返回假数据)。
  4. 抓取前请认真阅读目标站点的 robots.txt 和相关法律法规。

本章定位为 Selenium 入门练习,示例仅演示访问首页 + 解析歌单标题的最小流程。


1. 环境准备

1.1 安装 Selenium

pip install -U selenium
# 或 uv
uv add selenium

1.2 下载 WebDriver

Selenium 需要 与浏览器版本一致 的 WebDriver。

浏览器 WebDriver
Chrome https://googlechromelabs.github.io/chrome-for-testing/
Edge https://developer.microsoft.com/microsoft-edge/tools/webdriver/
Firefox https://github.com/mozilla/geckodriver/releases
Safari 内置于 macOS(需要 safari --enable 在 Safari 中打开 Develop 菜单)

推荐做法:用 webdriver-manager 自动管理

pip install webdriver-manager
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager

driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()))

1.3 WebDriver 放在哪?

方式 路径 优缺点
全局 PATH /usr/local/bin 或 %PATH% 多个项目共享,最常用
项目目录 ./drivers/chromedriver 多版本管理方便,但部署复杂
自动管理 webdriver-manager 强烈推荐,免去手动同步

2. 最小可运行示例

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC


def make_driver() -> webdriver.Chrome:
    options = webdriver.ChromeOptions()
    options.add_argument("--no-sandbox")
    options.add_argument("--disable-dev-shm-usage")
    # 打开无头模式:服务器 / CI 环境用
    # options.add_argument("--headless=new")
    return webdriver.Chrome(
        service=Service(),                # 让 webdriver-manager 自动选 driver
        options=options,
    )


def main() -> None:
    driver = make_driver()
    try:
        driver.get("https://music.163.com/")
        print("title:", driver.title)

        # 等到 id="g_iframe" 出现(歌单内容通常在 iframe 中)
        WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.ID, "g_iframe"))
        )
        driver.switch_to.frame("g_iframe")

        # 找歌单里的链接 / 标题(选择器随时可能失效)
        titles = driver.find_elements(By.CSS_SELECTOR, "a.msk")
        for t in titles[:5]:
            print(t.get_attribute("title") or t.text)
    finally:
        driver.quit()


if __name__ == "__main__":
    main()

要点:

  • 永远用 try / finally 或 with 包住 driver,保证 driver.quit() 关闭浏览器。
  • WebDriverWait 显式等待比 time.sleep 稳定很多。
  • 网易云的很多内容在 iframe 中,记得 switch_to.frame。
  • 关闭浏览器时浏览器进程可能不退出,再加一句 driver.quit() 即可。

3. 更稳妥的写法:selenium.webdriver.remote.webdriver 通用基类

如果你的脚本会在 Chrome / Edge / Firefox 之间切换,建议用 webdriver-manager
管理多浏览器,并抽象一个工厂函数:

from selenium import webdriver
from webdriver_manager.chrome import ChromeDriverManager
from webdriver_manager.firefox import GeckoDriverManager
from selenium.webdriver.chrome.service import Service as ChromeService
from selenium.webdriver.firefox.service import Service as FirefoxService


def make_driver(browser: str = "chrome", headless: bool = True) -> webdriver.Remote:
    if browser == "chrome":
        options = webdriver.ChromeOptions()
        if headless:
            options.add_argument("--headless=new")
        return webdriver.Chrome(
            service=ChromeService(ChromeDriverManager().install()),
            options=options,
        )
    if browser == "firefox":
        options = webdriver.FirefoxOptions()
        if headless:
            options.add_argument("-headless")
        return webdriver.Firefox(
            service=FirefoxService(GeckoDriverManager().install()),
            options=options,
        )
    raise ValueError(f"unsupported browser: {browser}")

4. 常见反爬与缓解

现象 排查 / 缓解
selenium.common.exceptions.WebDriverException WebDriver 与浏览器版本不匹配。pip install -U webdriver-manager
打开空白 / 一直转圈 检测 UA / 用 WebDriverWait 等元素
“检测到自动化工具” 用 undetected-chromedriver,加 --disable-blink-features=AutomationControlled
频繁封 IP 限速 + 代理池(前提是合法)
登录后才返回真实数据 用 Selenium 自动登录,但请勿绕过付费内容
# 让 Selenium 看起来更像真实用户
options.add_experimental_option("excludeSwitches", ["enable-automation"])
options.add_experimental_option("useAutomationExtension", False)
options.add_argument(
    "user-agent=Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) "
    "AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0.0.0 Safari/537.36"
)

5. 比 Selenium 更轻的方案

当目标页面有公开 API 时,优先用 HTTP 而非浏览器:

方案 适合
requests 接口是公开的、或只是 GET/POST JSON
httpx 同步/异步都要用
playwright 浏览器自动化,但 API 更现代、速度更快
pyppeteer Chrome DevTools Protocol(已被 playwright 取代)

用浏览器爬一切是“最慢、也最容易被识别”的方案。除非页面是 重度 SPA / 强
JavaScript 渲染
,否则应优先尝试直接请求 API。


6. 调试技巧

  • 截图:driver.save_screenshot("debug.png")
  • 查看页面源码:print(driver.page_source)
  • 浏览器控制台:driver.get_log("browser")
  • 断点:在脚本里 import pdb; pdb.set_trace() 后手工 driver.get(...)。
  • 用 --user-data-dir 复用同一个浏览器配置,避开重复登录。

7. 合规清单(务必阅读)

  • ✅ 只抓取 公开 且 允许爬取 的内容。
  • ✅ 在请求中加入 User-Agent 并注明联系方式。
  • ✅ 控制频率(一般 ≤ 1 QPS),避开高峰。
  • ❌ 不要绕过登录、付费墙、地区限制。
  • ❌ 不要把抓到的内容二次分发起诉。
  • ❌ 不要绕过 robots.txt(https://music.163.com/robots.txt 自行查阅)。