Skip to Content
GuidesUse the Python SDK

Use the Lightpanda Python SDK

In this guide, you’ll scrape a JavaScript-rendered page with the lightpanda Python package and extract structured data from it with page.extract.

The package bundles the Lightpanda browser binary, so there’s no client to wire up and no separate browser download, unlike driving Lightpanda from Playwright or Selenium. It also uses an order of magnitude less memory than a Selenium and Chrome stack  for the same scrape.

Prerequisites

You’ll need Python 3.10 or newer.

Install the lightpanda package from PyPI.

pip install lightpanda

Create scrape.py

Create a scrape.py file. Import Browser and open a session:

from lightpanda import Browser with Browser() as browser, browser.new_session() as page: pass

Browser() starts the bundled Lightpanda binary and stops it when the with block exits.

Use page.goto to load quotes.toscrape.com/js , a page whose quotes are rendered by client-side JavaScript. Then call page.extract with a schema describing the fields to pull from each .quote element:

SCHEMA = { "quotes": [ { "selector": ".quote", "fields": {"text": ".text", "author": ".author", "tags": [".tag"]}, } ], } page.goto(url="https://quotes.toscrape.com/js/") data = page.extract(schema=SCHEMA) print(data["quotes"])

extract returns each .quote element as a record, with tags resolved to a list of every .tag inside it. A plain requests.get on this page returns zero quotes because the DOM is only built after the page’s JavaScript runs.

Full script

Follow the site’s own “Next” link to collect every page instead of just the first:

from collections import Counter from lightpanda import Browser SCHEMA = { "quotes": [ { "selector": ".quote", "fields": {"text": ".text", "author": ".author", "tags": [".tag"]}, } ], "next": {"selector": "li.next a", "attr": "href"}, } with Browser() as browser, browser.new_session() as page: quotes = [] url = "https://quotes.toscrape.com/js/" while url: page.goto(url=url) data = page.extract(schema=SCHEMA) quotes.extend(data["quotes"]) url = data["next"] print(f"{len(quotes)} quotes") tags = Counter(tag for q in quotes for tag in q["tags"]) print("Top tags:", tags.most_common(5))

Run it

python scrape.py
$ python scrape.py 100 quotes Top tags: [('love', 14), ('inspirational', 13), ('life', 13), ('humor', 12), ('books', 11)]

Scrape concurrently with asyncio

Use AsyncBrowser instead of Browser when you want to scrape many pages at once instead of one after another, or when your code already runs inside an async application (a web server built with FastAPI, for example) where a blocking call would freeze everything else it’s doing. Every method becomes awaitable; AsyncBrowser itself runs each call on the browser’s own thread pool, so it never blocks your event loop.

Since the site’s ten pages are numbered, fetch all of them at once with asyncio.gather instead of following the pagination link one page at a time. browser.session() opens a session scoped to one task and closes it when that task ends:

import asyncio from collections import Counter from lightpanda import AsyncBrowser SCHEMA = { "quotes": [ { "selector": ".quote", "fields": {"text": ".text", "author": ".author", "tags": [".tag"]}, } ], } URLS = [f"https://quotes.toscrape.com/js/page/{n}/" for n in range(1, 11)] async def scrape_one(browser, url): async with browser.session() as page: await page.goto(url=url) data = await page.extract(schema=SCHEMA) return data["quotes"] async def main(): async with AsyncBrowser() as browser: results = await asyncio.gather(*(scrape_one(browser, url) for url in URLS)) quotes = [quote for page_quotes in results for quote in page_quotes] print(f"{len(quotes)} quotes") tags = Counter(tag for q in quotes for tag in q["tags"]) print("Top tags:", tags.most_common(5)) asyncio.run(main())
ℹ️

Every browser action is a Session method, typed and documented in your IDE, with the action and its arguments in snake_case (wait_for_selector, backend_node_id).

Find every method’s signature and docstring in the Python SDK reference.

Replay a saved script

If you already have a PandaScript, a .js file recorded by Lightpanda Agent or the MCP save tool, replay it from Python with run_script, no LLM call needed:

from lightpanda import run_script output = run_script("scrape.js") print(output)

run_script shells out to lightpanda run <script> and returns its stdout; a non-zero exit raises ScriptError. run_script_async is the awaitable variant.

Respect robots.txt

The browser can enforce robots.txt for you (off by default). To enable it, pass browser flags through args= on any entry point:

With Browser (or AsyncBrowser):

Browser(args=["--obey-robots"])

With run_script:

run_script("scrape.js", args=["--obey-robots"])

A request the site disallows then fails rather than being sent: a Session call raises ToolError, and run_script raises ScriptError.

The same flag works on CDPServer and BiDiServer, where a blocked request surfaces as a navigation error in Playwright or Selenium rather than as a Python exception.

⚠️

The rule applies to every request, not just the page you asked for. Plenty of sites disallow the directory their own assets live in, so a page you are allowed to fetch can load with its scripts and styles blocked, and render blank or empty. That is robots.txt being honoured, not a failure. If a page comes back strangely empty under --obey-robots, read the site’s robots.txt before assuming something is broken.