Skip to Content
ReferencePython SDK

Python SDK

Every public class, method, property and exception of the lightpanda package, with the signatures and docstrings shipped in the code. See Use the Python SDK for a practical walkthrough. Every sync class has an asyncio twin with the same methods, awaitable; the async sections below list only what the twin adds.

Browser

A lightpanda browser process. Spawns the bundled binary on first use.

Not fork-inheritable: after os.fork()/multiprocessing, create a fresh Browser in the child.

class Browser( binary: str | os.PathLike | None = None, env: dict[str, str] | None = None, timeout: float = 300.0, verbose: bool = False, args: Sequence[str] = () )

Spawn the browser process and fetch its tool list.

Arguments:

  • binary: Path to a lightpanda binary. When omitted, resolved from the LIGHTPANDA_BIN environment variable, then the binary bundled in the package, then PATH.
  • env: Extra environment variables for the spawned process.
  • timeout: Seconds to wait for a response to any request before raising ProtocolError.
  • verbose: Let the browser’s own logging through to stderr.
  • args: Extra CLI flags for the spawned browser process, e.g. ["--obey-robots"] to enforce robots.txt, ["--http-cache-dir", path] or cookie flags.

Usable as a context manager (with).

Browser.tools

tools: dict[str, dict]

Property. Tool name → {description, schema, output_schema}, as reported by the browser. output_schema is None for tools that declare none.

Browser.new_session

def new_session(self) -> Session

Open a new isolated browsing context: its own page, cookies and memory. Close it with Session.close() or a with block.

Browser.close

def close(self) -> None

Stop the browser process, closing every session with it.

Session

One isolated browsing context (own page, cookies, memory).

Do not construct directly — use Browser.new_session().

Browser actions are keyword-only methods, named in snake_case after the browser’s own action names: the waitForSelector action is wait_for_selector(), and its backendNodeId argument is backend_node_id.

Where a method accepts both selector and backend_node_id, pass one of the two. selector is preferred for reproducibility and wins when both are given; backend_node_id takes the values returned by tree(), links() or find_element().

call() is the escape hatch that takes the action and argument names exactly as the browser declares them. A failed action raises ToolError.

class Session(browser: Browser, session_id: str)

Usable as a context manager (with).

Session.id

id: str

Property. The session id, as the browser knows it.

Session.call

def call(self, tool: str, **kwargs)

Invoke a browser tool by name. The generated methods route here.

Accepts the tool and argument names as the browser declares them (waitForSelector, backendNodeId) as well as their snake_case forms. Returns parsed JSON for JSON-carrying tools, bytes for image results (screenshot without path), otherwise the result text. Raises ToolError when the tool reports a failure.

Arguments:

  • tool: The tool name.
  • **kwargs: The tool’s arguments; None values are omitted.

Session.close

def close(self) -> None

Release the session’s page. Idempotent; calls made after this raise ToolError. Closing the browser closes every session.

Session.click

def click( self, *, selector: str | None = None, backend_node_id: int | None = None ) -> PageResult

Click on an interactive element. Provide either a CSS selector (preferred for reproducibility) or a backendNodeId. Returns the current page URL and title after the click.

Arguments:

  • selector: CSS selector of the element to click. Preferred over backendNodeId.
  • backend_node_id: The backend node ID of the element to click.

Returns:

PageResult: The sentence above, with these attributes (None when absent):

  • url: URL of the page the call left loaded.
  • http_status: Response status of that page’s own document. Absent when no response has arrived. 4xx/5xx means an error page, not the content.
  • title: Document title. Empty or absent when the page has none.

Session.console_logs

def console_logs(self) -> Any

Get buffered console.log/warn/error messages from the current page. Returns all messages since last call and clears the buffer.

Session.detect_forms

def detect_forms(self, *, url: str | None = None, timeout: int | None = None) -> Any

List the forms on the page as JSON: each form’s backendNodeId, action, method and fields, where each field has backendNodeId, tagName, name, inputType, required, disabled, and when present value, placeholder and select options. Use it before filling a form to see every field it expects. It returns no CSS selectors; get one per field with nodeDetails so the fill calls stay replayable. If a url is provided, it navigates there first.

Arguments:

  • url: Optional URL to navigate to before processing.
  • timeout: Optional timeout in milliseconds. Defaults to 10000.

Session.evaluate

def evaluate( self, *, script: str, url: str | None = None, timeout: int | None = None, save: str | None = None ) -> Any

Evaluate JavaScript in the current page context — an escape hatch for page-side logic the dedicated tools can’t express; prefer extract for data and click/fill/etc. for actions. It runs in the page, so it cannot see the agent script’s variables or builtins — interpolate any value into the script string. A bare trailing expression yields its value; top-level await and return are supported (the body then runs as an async function, so use return to produce a value). Objects and arrays return as JSON, so no JSON.stringify is needed. If a url is provided, it navigates there first. The globalThis.lp object exposes a Session-scoped bridge store: values written via lp.foo = ... auto-sync at end of evaluate, surviving navigation; values previously set via /extract save= or /evaluate save= appear as lp.<name>.

Arguments:

  • script: JavaScript run in the page context. A bare trailing expression, or return with top-level await, is the result.
  • url: Optional URL to navigate to before evaluating.
  • timeout: Optional timeout in milliseconds. Defaults to 10000.
  • save: Optional bridge-store key. The evaluate’s return value is stored under this name and re-exposed as lp.<name> to subsequent evaluates. Objects, arrays, and strings are serialized automatically — no JSON.stringify needed.

Session.extract

def extract(self, *, schema: str | dict | list, save: str | None = None) -> Any

Extract structured data from the current page (navigate first). schema is a JSON object (passed as a string) mapping output field names to CSS-selector specs. It is NOT a JSON Schema — no “type”/“properties” wrappers; the keys ARE your output fields. Value shapes:

"<sel>" → first match's text (trimmed; null if no match) ["<sel>"] → every match's text (string[]) {"selector":"<sel>","attr":"<name>"} → first match's attribute value (href/src resolved to absolute URLs) [{"selector":"<sel>","attr":"<name>"}] → every match's attribute (string[]) [{"selector":"<sel>","fields":{…}}] → one object per match; field selectors resolve relative to that match and accept any shape above ("" = the match's own text; nest arrays for per-item sub-lists)

Add “limit”: N inside any array’s object spec to cap matches. Every extracted value is a string or null — parse numbers downstream. An empty array is a valid result, but if ALL top-level keys miss, the call errors: inspect the page (tree/markdown) and retry with corrected selectors. Finish data tasks with extract — it is the only read recorded as a replayable extract(...) script call; answers lifted from markdown text in chat are not.

Examples (schema → result):

{"karma": "#karma"} → {"karma":"42"} {"items": [".story .title"]} → {"items":["Title 1","Title 2"]} {"top3": [{"selector":".story .title","limit":3}]} → {"top3":["A","B","C"]} {"links": [{"selector":"a.title","attr":"href"}]} → {"links":["https://site/a","https://site/b"]} {"stories": [{"selector":".athing","fields":{"title":".titleline","rank":".rank"}}]} → {"stories":[{"title":"Foo","rank":"1"}]}

Arguments:

  • schema: Extraction schema as a string: a JSON object literal mapping output field names to CSS-selector specs (see tool description). Not a JSON Schema.
  • save: Optional bridge-store key. The extracted JSON is stored under this name and exposed as lp.<name> in subsequent /evaluate calls.

Session.fill

def fill( self, *, value: str, selector: str | None = None, backend_node_id: int | None = None ) -> PageResult

Fill text into an input element. Provide either a CSS selector (preferred for reproducibility) or a backendNodeId.

Arguments:

  • value: The text to fill into the input element.
  • selector: CSS selector of the input element to fill. Preferred over backendNodeId.
  • backend_node_id: The backend node ID of the input element to fill.

Returns:

PageResult: The sentence above, with these attributes (None when absent):

  • url: URL of the page the call left loaded.
  • http_status: Response status of that page’s own document. Absent when no response has arrived. 4xx/5xx means an error page, not the content.
  • title: Document title. Empty or absent when the page has none.

Session.find_element

def find_element(self, *, role: str | None = None, name: str | None = None) -> Any

Find interactive elements by role and/or accessible name. Returns matching elements with their backend node IDs. Useful for locating specific elements without parsing the full semantic tree.

Arguments:

  • role: Optional ARIA role to match (e.g. ‘button’, ‘link’, ‘textbox’, ‘checkbox’).
  • name: Optional accessible name to match, case-insensitive: a substring, or a JavaScript regex literal such as /sign (in|up)/ (unanchored; flags i, m, s, u accepted; case-insensitive even without i, prefix (?-i) to make it case-sensitive).

Session.get_cookies

def get_cookies(self, *, url: str | None = None, all: bool | None = None) -> Any

Cookies stored in the browser. Defaults to cookies whose domain matches the current page’s host. Pass url=<URL> to filter for another host, or all=true to dump every cookie regardless of host. Useful for debugging authentication and session state.

Arguments:

  • url: Restrict output to cookies matching this URL’s host. Defaults to the current page.
  • all: If true, dump every cookie regardless of host. Overrides url.

Session.get_env

def get_env(self, *, name: str | None = None) -> Any

With name: read an LP_* env var (other namespaces report as not set) — for non-secret config only (base URLs, flags). Without name: list LP_* names that are set (no values) — safe credential discovery. For secrets, pass $LP_* placeholders in tool args; never request a credential by name (the value would land in your context).

Arguments:

  • name: Optional. If provided, must start with LP_; returns the value. If omitted, returns the list of LP_* names that are set.

Session.get_url

def get_url(self) -> Any

Current page URL. The browser may already have a page loaded (command, replayed script) not visible in this conversation — call this before assuming nothing is loaded when the user references the current page/site. Also useful to verify a navigation or detect a redirect.

Session.goto

def goto( self, *, url: str, timeout: int | None = None, wait_until: str | None = None ) -> PageResult

Navigate the current page to a URL. Returns the HTTP status once waitUntil fires (default load), or a timeout notice; a 4xx or 5xx means the page you got is an error page, not the content — check it before reading on; content rendered by post-load JavaScript may not be there yet (see waitForState). The page stays loaded for later reads and actions. To navigate and read in one call, pass url to markdown, tree or html instead; use goto when the next step is an action or extract.

Arguments:

  • url: The URL to navigate to, must be a valid URL.
  • timeout: Optional timeout in milliseconds. Defaults to 10000.
  • wait_until: Event that completes the navigation. Defaults to ‘load’. Prefer ‘domcontentloaded’ followed by waitForSelector on pages whose late scripts (ads) hold ‘load’ back. Avoid ‘done’ (full quiescence): on pages with constant background activity it is the slowest choice and can run to the timeout.

Returns:

PageResult: The sentence above, with these attributes (None when absent):

  • url: URL of the page the call left loaded.
  • http_status: Response status of that page’s own document. Absent when no response has arrived. 4xx/5xx means an error page, not the content.
  • title: Document title. Empty or absent when the page has none.

Session.hover

def hover( self, *, selector: str | None = None, backend_node_id: int | None = None ) -> PageResult

Hover over an element, triggering mouseover and mouseenter events. Provide either a CSS selector (preferred for reproducibility) or a backendNodeId. Useful for menus, tooltips, and hover states.

Arguments:

  • selector: CSS selector of the element to hover over. Preferred over backendNodeId.
  • backend_node_id: The backend node ID of the element to hover over.

Returns:

PageResult: The sentence above, with these attributes (None when absent):

  • url: URL of the page the call left loaded.
  • http_status: Response status of that page’s own document. Absent when no response has arrived. 4xx/5xx means an error page, not the content.
  • title: Document title. Empty or absent when the page has none.

Session.html

def html( self, *, selector: str | None = None, backend_node_id: int | None = None, max_bytes: int | None = None, strip: dict | None = None, url: str | None = None, timeout: int | None = None ) -> Any

Raw HTML for the document or, with selector/backendNodeId, a single node’s outerHTML. Verbose; use only when you need attributes that markdown discards.

Arguments:

  • selector: Optional CSS selector. When set, dump only that element’s outerHTML.
  • backend_node_id: Optional backend node ID. When set, dump only that node’s outerHTML. 0 is treated as omitted.
  • max_bytes: Optional soft cap on output size in bytes. Content is truncated at a UTF-8 boundary and a short ‘[truncated]’ marker is appended past the cap.
  • strip: Optional. Omit element groups from the output: js (script, noscript, script preloads), css (style, stylesheet links), ui (css plus img, picture, video, audio, svg, canvas, iframe), invisible (elements an author rule or inline style sets to display:none), shell (nav, aside, dialog, page-level header/footer and the matching landmark roles; skipped when that would drop most of the text), clutter (keep only the main content, in the manner of reader modes; includes shell and invisible, and falls back to shell when it finds too little). {“js”:true,“css”:true} keeps a page dump small.
  • url: Optional URL to navigate to before dumping.
  • timeout: Optional timeout in milliseconds. Defaults to 10000.

Session.interactive_elements

def interactive_elements(self, *, url: str | None = None, timeout: int | None = None) -> Any

List every visible interactive element on the page as a JSON array: native controls, ARIA widgets, contenteditable regions, elements with event listeners, and focusable elements. Each entry has backendNodeId, tagName, role, name, type (why it counts as interactive), tabIndex, and when present listeners, disabled, id, class, href, inputType, value, elementName and placeholder. Use it to survey what can be acted on; to locate one element by role or name, findElement is cheaper. If a url is provided, it navigates there first.

Arguments:

  • url: Optional URL to navigate to before processing.
  • timeout: Optional timeout in milliseconds. Defaults to 10000.
def links( self, *, limit: int | None = None, url: str | None = None, timeout: int | None = None ) -> Any

Extract the visible links in the opened page as JSON objects with text (anchor text, falling back to aria-label/title/image alt), href (resolved URL), and backendNodeId (pass to click/nodeDetails). One entry per href; hidden links are omitted. If a url is provided, it navigates to that url first.

Arguments:

  • limit: Optional. Return at most this many links, in document order.
  • url: Optional URL to navigate to before processing.
  • timeout: Optional timeout in milliseconds. Defaults to 10000.

Session.markdown

def markdown( self, *, selector: str | None = None, backend_node_id: int | None = None, max_bytes: int | None = None, strip: dict | None = None, url: str | None = None, timeout: int | None = None ) -> Any

Render the page (or a subtree) as markdown. Scope with selector or backendNodeId to read just the relevant region — full-page markdown is the last resort. Use maxBytes to cap long pages.

Arguments:

  • selector: Optional CSS selector. Render markdown for just that element’s subtree.
  • backend_node_id: Optional backend node ID. Render markdown for just that node’s subtree. 0 is treated as omitted.
  • max_bytes: Optional soft cap on output size in bytes. Content is truncated at a UTF-8 boundary and a short ‘[truncated]’ marker is appended past the cap.
  • strip: Optional. Omit element groups from the output; same groups as the html tool’s strip. shell (page chrome by markup) and clutter (keep only the main content, in the manner of reader modes) are the ones that matter for reading; ui also drops images.
  • url: Optional URL to navigate to before rendering.
  • timeout: Optional timeout in milliseconds. Defaults to 10000.

Session.node_details

def node_details(self, *, backend_node_id: int) -> Any

Details for a node by backendNodeId: a ready-to-use CSS selector that resolves to the node (the first match, as click/fill resolve it), plus tag, role, name, interactivity, disabled, value, input type, placeholder, href, id, class, checked, select options. The canonical way to turn a tree backendNodeId into a CSS selector.

Arguments:

  • backend_node_id: The backend node ID of the element to inspect.

Session.press

def press( self, *, key: str, selector: str | None = None, backend_node_id: int | None = None ) -> PageResult

Press a keyboard key, dispatching keydown and keyup events. Use key names like ‘Enter’, ‘Tab’, ‘Escape’, ‘ArrowDown’, ‘Backspace’, or single characters like ‘a’, ‘1’. Common shorthand is normalized: ‘enter’/‘return’ → ‘Enter’, ‘esc’ → ‘Escape’, ‘up’/‘down’/‘left’/‘right’ → ‘Arrow*’, ‘space’ → ’ ’. Pressing ‘Enter’ on a form input or submit button triggers implicit form submission.

Arguments:

  • key: The key to press (e.g. ‘Enter’, ‘Tab’, ‘a’).
  • selector: Optional CSS selector of the element to target. Preferred over backendNodeId.
  • backend_node_id: Optional backend node ID of the element to target. Defaults to the document when neither selector nor backendNodeId is provided; 0 is treated as omitted.

Returns:

PageResult: The sentence above, with these attributes (None when absent):

  • url: URL of the page the call left loaded.
  • http_status: Response status of that page’s own document. Absent when no response has arrived. 4xx/5xx means an error page, not the content.
  • title: Document title. Empty or absent when the page has none.

Session.screenshot

def screenshot( self, *, path: str | None = None, selector: str | None = None, backend_node_id: int | None = None, full_page: bool | None = None, strip: dict | None = None, url: str | None = None, timeout: int | None = None ) -> Any

Render the page, or one node, as a PNG: the text layout Lightpanda computes, not a pixel-accurate browser rendering (no images, fonts or CSS colours). With path, writes the file at full size and returns its location; without it, returns the image inline where the client can display one, at most 1280px wide and 4096px tall. Use it to see spatial layout; read content with markdown/tree.

Arguments:

  • path: Optional relative path (no ’..’ segments) to write the PNG to. Created or overwritten. Without it the image is returned inline, which needs a client that can display images.
  • selector: Optional CSS selector. When set, render only that element.
  • backend_node_id: Optional backend node ID. When set, render only that node. 0 is treated as omitted.
  • full_page: Render the whole content height instead of one viewport. Defaults to false.
  • strip: Optional. Omit element groups from the render; same groups as the html tool’s strip (js, css, ui, invisible, shell, clutter).
  • url: Optional URL to navigate to before rendering.
  • timeout: Optional timeout in milliseconds. Defaults to 10000.

Session.scroll

def scroll( self, *, selector: str | None = None, backend_node_id: int | None = None, x: int | None = None, y: int | None = None ) -> PageResult

Scroll the window, or an element’s scroll container, to an absolute position; an omitted axis keeps its current offset. Target an element with a CSS selector (preferred for reproducibility) or a backendNodeId; omit both to scroll the window. Page scripts receive a scroll event, so content that loads on scroll (infinite feeds, lazy lists) may appear: read the page again afterwards, with waitForState if it is still loading. Returns the final scroll position and the current page URL and title.

Arguments:

  • selector: Optional: CSS selector of the element to scroll. Preferred over backendNodeId. If the element is not itself a scroll container, its nearest scrollable ancestor is scrolled instead.
  • backend_node_id: Optional: The backend node ID of the element to scroll. If the element is not itself a scroll container, its nearest scrollable ancestor is scrolled instead. If neither this nor selector is given (or it is 0), scrolls the window.
  • x: Optional: The horizontal scroll offset.
  • y: Optional: The vertical scroll offset.

Returns:

PageResult: The sentence above, with these attributes (None when absent):

  • url: URL of the page the call left loaded.
  • http_status: Response status of that page’s own document. Absent when no response has arrived. 4xx/5xx means an error page, not the content.
  • title: Document title. Empty or absent when the page has none.
def search(self, *, query: str, timeout: int | None = None) -> Any

Run a web search and return results as markdown: a numbered list of {title, url, snippet}. Search tries brave, tavily, exa, then keenable in order, each when its API key (BRAVE_API_KEY, TAVILY_API_KEY, EXA_API_KEY or KEENABLE_API_KEY) is set; keenable also works without a key through its public endpoint (rate-limited per client IP). Prefer this over goto-ing google.com/search directly (Google blocks the browser on User-Agent/TLS). The browser does not navigate — to open a result, use goto with its URL.

Arguments:

  • query: The search query.
  • timeout: Optional timeout in milliseconds. Defaults to 10000.

Session.select_option

def select_option( self, *, value: str, selector: str | None = None, backend_node_id: int | None = None ) -> PageResult

Select an option in a <select> dropdown element by its value. Provide either a CSS selector (preferred for reproducibility) or a backendNodeId. Dispatches input and change events.

Arguments:

  • value: The value of the option to select.
  • selector: CSS selector of the <select> element. Preferred over backendNodeId.
  • backend_node_id: The backend node ID of the <select> element.

Returns:

PageResult: The sentence above, with these attributes (None when absent):

  • url: URL of the page the call left loaded.
  • http_status: Response status of that page’s own document. Absent when no response has arrived. 4xx/5xx means an error page, not the content.
  • title: Document title. Empty or absent when the page has none.

Session.set_checked

def set_checked( self, *, checked: bool, selector: str | None = None, backend_node_id: int | None = None ) -> PageResult

Check or uncheck a checkbox or radio button. Provide either a CSS selector (preferred for reproducibility) or a backendNodeId. Dispatches input, change, and click events.

Arguments:

  • checked: Whether to check (true) or uncheck (false) the element.
  • selector: CSS selector of the checkbox or radio input element. Preferred over backendNodeId.
  • backend_node_id: The backend node ID of the checkbox or radio input element.

Returns:

PageResult: The sentence above, with these attributes (None when absent):

  • url: URL of the page the call left loaded.
  • http_status: Response status of that page’s own document. Absent when no response has arrived. 4xx/5xx means an error page, not the content.
  • title: Document title. Empty or absent when the page has none.

Session.structured_data

def structured_data(self, *, url: str | None = None, timeout: int | None = None) -> Any

Page metadata as JSON: jsonLd (each JSON-LD block as a string), openGraph, twitterCard, meta and links (key/value lists), plus alternate (hreflang variants) and linkHeaders (relations from the HTTP Link header) when present. Empty sections come back as empty arrays. Use it for publisher-declared facts such as product price, article author or canonical URL before scraping the visible text for them. If a url is provided, it navigates there first.

Arguments:

  • url: Optional URL to navigate to before processing.
  • timeout: Optional timeout in milliseconds. Defaults to 10000.

Session.tree

def tree( self, *, url: str | None = None, timeout: int | None = None, backend_node_id: int | None = None, max_depth: int | None = None ) -> Any

Semantic outline of the page as indented text: one node per line with its role, accessible name, value and backendNodeId, plus checked state and select options with the selected one marked. The default first read of an unfamiliar page; input and select values are already here, so no nodeDetails call is needed to read them. Pass backendNodeId to scope to a subtree and maxDepth to survey structure before going deeper. Read it again after any page-changing action, since the DOM it describes may have changed; use nodeDetails to turn a backendNodeId into a CSS selector for actions.

Arguments:

  • url: Optional URL to navigate to before fetching the semantic tree.
  • timeout: Optional timeout in milliseconds. Defaults to 10000.
  • backend_node_id: Optional backend node ID to get the tree for a specific element instead of the document root. 0 is treated as omitted.
  • max_depth: Optional maximum depth of the tree to return. Useful for exploring high-level structure first.

Session.wait_for_script

def wait_for_script(self, *, script: str, timeout: int | None = None) -> Any

Wait until a JS expression returns truthy. Re-evaluates on each tick of the event loop. Use for synchronization beyond what CSS selectors can express — e.g. window.dataLoaded === true, document.readyState === 'complete', document.querySelectorAll('.row').length >= 5.

Arguments:

  • script: JS expression evaluated each tick until truthy. Must be an expression (not a statement).
  • timeout: Optional timeout in milliseconds. Defaults to 5000, or 15000 when the page has not reached ‘load’ yet.

Session.wait_for_selector

def wait_for_selector(self, *, selector: str, timeout: int | None = None) -> Any

Wait for an element matching a CSS selector to appear in the page. Returns the backend node ID of the matched element.

Arguments:

  • selector: The CSS selector to wait for.
  • timeout: Optional timeout in milliseconds. Defaults to 5000, or 15000 when the page has not reached ‘load’ yet.

Session.wait_for_state

def wait_for_state(self, *, state: str, timeout: int | None = None) -> Any

Wait for the CURRENT page to reach a load state (no navigation). After a goto, the page is returned at the fast load snapshot, so content rendered by post-load JS (XHR-loaded lists, feeds, search results) may still be missing. When a read looks incomplete — empty lists, spinners, skeletons — call this with ‘networkidle’ and re-read. Prefer ‘networkidle’; ‘done’ can be slow on sites with constant background activity (ads, polling).

Arguments:

  • state: Load state to wait for. ‘networkidle’ = network settled (the usual choice to finish a dynamic page).
  • timeout: Optional timeout in milliseconds. Defaults to 5000.

PageResult

What a navigation or action answered: a sentence that also carries fields.

goto and the action tools describe what they did in prose — which element was clicked, whether a new window took over — and separately report where that left the page. Dropping either would lose something, so this is the sentence, unchanged for printing and comparison, with the fields as attributes:

r = page.goto(url="https://example.com/missing") print(r) # Navigated successfully. HTTP 404 Not Found. r.http_status # 404

http_status is None before any response has arrived, title on a document that has none.

class PageResult

PageResult.url

url: str

PageResult.http_status

http_status: int | None

PageResult.title

def title(self, /)

Return a version of the string where each word is titlecased.

More specifically, words start with uppercased characters and all remaining cased characters have lower case.

PageResult.encode

def encode(self, /, encoding='utf-8', errors='strict')

Encode the string using the codec registered for encoding.

encoding

The encoding in which to encode the string.

errors

The error handling scheme to use for encoding errors. The default is 'strict' meaning that encoding errors raise a UnicodeEncodeError. Other possible values are 'ignore', 'replace' and 'xmlcharrefreplace' as well as any other name registered with codecs.register_error that can handle UnicodeEncodeErrors.

PageResult.replace

def replace(self, old, new, /, count=-1)

Return a copy with all occurrences of substring old replaced by new.

count Maximum number of occurrences to replace. -1 (the default value) means replace all occurrences.

If the optional argument count is given, only the first count occurrences are replaced.

PageResult.split

def split(self, /, sep=None, maxsplit=-1)

Return a list of the substrings in the string, using sep as the separator string.

sep The separator used to split the string.
When set to None (the default value), will split on any whitespace character (including \n \r \t \f and spaces) and will discard empty strings from the result. maxsplit Maximum number of splits. -1 (the default value) means no limit.

Splitting starts at the front of the string and works to the end.

Note, str.split() is mainly useful for data that has been intentionally delimited. With natural text that includes punctuation, consider using the regular expression module.

PageResult.rsplit

def rsplit(self, /, sep=None, maxsplit=-1)

Return a list of the substrings in the string, using sep as the separator string.

sep The separator used to split the string.
When set to None (the default value), will split on any whitespace character (including \n \r \t \f and spaces) and will discard empty strings from the result. maxsplit Maximum number of splits. -1 (the default value) means no limit.

Splitting starts at the end of the string and works to the front.

PageResult.join

def join(self, iterable, /)

Concatenate any number of strings.

The string whose method is called is inserted in between each given string. The result is returned as a new string.

Example: ’.‘.join([‘ab’, ‘pq’, ‘rs’]) -> ‘ab.pq.rs’

PageResult.capitalize

def capitalize(self, /)

Return a capitalized version of the string.

More specifically, make the first character have upper case and the rest lower case.

PageResult.casefold

def casefold(self, /)

Return a version of the string suitable for caseless comparisons.

PageResult.center

def center(self, width, fillchar=' ', /)

Return a centered string of length width.

Padding is done using the specified fill character (default is a space).

PageResult.count

def count(unknown)

Return the number of non-overlapping occurrences of substring sub in string S[start:end].

Optional arguments start and end are interpreted as in slice notation.

PageResult.expandtabs

def expandtabs(self, /, tabsize=8)

Return a copy where all tab characters are expanded using spaces.

If tabsize is not given, a tab size of 8 characters is assumed.

PageResult.find

def find(unknown)

Return the lowest index in S where substring sub is found, such that sub is contained within S[start:end].

Optional arguments start and end are interpreted as in slice notation. Return -1 on failure.

PageResult.partition

def partition(self, sep, /)

Partition the string into three parts using the given separator.

This will search for the separator in the string. If the separator is found, returns a 3-tuple containing the part before the separator, the separator itself, and the part after it.

If the separator is not found, returns a 3-tuple containing the original string and two empty strings.

PageResult.index

def index(unknown)

Return the lowest index in S where substring sub is found, such that sub is contained within S[start:end].

Optional arguments start and end are interpreted as in slice notation. Raises ValueError when the substring is not found.

PageResult.ljust

def ljust(self, width, fillchar=' ', /)

Return a left-justified string of length width.

Padding is done using the specified fill character (default is a space).

PageResult.lower

def lower(self, /)

Return a copy of the string converted to lowercase.

PageResult.lstrip

def lstrip(self, chars=None, /)

Return a copy of the string with leading whitespace removed.

If chars is given and not None, remove characters in chars instead.

PageResult.rfind

def rfind(unknown)

Return the highest index in S where substring sub is found, such that sub is contained within S[start:end].

Optional arguments start and end are interpreted as in slice notation. Return -1 on failure.

PageResult.rindex

def rindex(unknown)

Return the highest index in S where substring sub is found, such that sub is contained within S[start:end].

Optional arguments start and end are interpreted as in slice notation. Raises ValueError when the substring is not found.

PageResult.rjust

def rjust(self, width, fillchar=' ', /)

Return a right-justified string of length width.

Padding is done using the specified fill character (default is a space).

PageResult.rstrip

def rstrip(self, chars=None, /)

Return a copy of the string with trailing whitespace removed.

If chars is given and not None, remove characters in chars instead.

PageResult.rpartition

def rpartition(self, sep, /)

Partition the string into three parts using the given separator.

This will search for the separator in the string, starting at the end. If the separator is found, returns a 3-tuple containing the part before the separator, the separator itself, and the part after it.

If the separator is not found, returns a 3-tuple containing two empty strings and the original string.

PageResult.splitlines

def splitlines(self, /, keepends=False)

Return a list of the lines in the string, breaking at line boundaries.

Line breaks are not included in the resulting list unless keepends is given and true.

PageResult.strip

def strip(self, chars=None, /)

Return a copy of the string with leading and trailing whitespace removed.

If chars is given and not None, remove characters in chars instead.

PageResult.swapcase

def swapcase(self, /)

Convert uppercase characters to lowercase and lowercase characters to uppercase.

PageResult.translate

def translate(self, table, /)

Replace each character in the string using the given translation table.

table Translation table, which must be a mapping of Unicode ordinals to Unicode ordinals, strings, or None.

The table must implement lookup/indexing via getitem, for instance a dictionary or list. If this operation raises LookupError, the character is left untouched. Characters mapped to None are deleted.

PageResult.upper

def upper(self, /)

Return a copy of the string converted to uppercase.

PageResult.startswith

def startswith(unknown)

Return True if the string starts with the specified prefix, False otherwise.

prefix

A string or a tuple of strings to try.

start

Optional start position. Default: start of the string.

end

Optional stop position. Default: end of the string.

PageResult.endswith

def endswith(unknown)

Return True if the string ends with the specified suffix, False otherwise.

suffix

A string or a tuple of strings to try.

start

Optional start position. Default: start of the string.

end

Optional stop position. Default: end of the string.

PageResult.removeprefix

def removeprefix(self, prefix, /)

Return a str with the given prefix string removed if present.

If the string starts with the prefix string, return string[len(prefix):]. Otherwise, return a copy of the original string.

PageResult.removesuffix

def removesuffix(self, suffix, /)

Return a str with the given suffix string removed if present.

If the string ends with the suffix string and that suffix is not empty, return string[:-len(suffix)]. Otherwise, return a copy of the original string.

PageResult.isascii

def isascii(self, /)

Return True if all characters in the string are ASCII, False otherwise.

ASCII characters have code points in the range U+0000-U+007F. Empty string is ASCII too.

PageResult.islower

def islower(self, /)

Return True if the string is a lowercase string, False otherwise.

A string is lowercase if all cased characters in the string are lowercase and there is at least one cased character in the string.

PageResult.isupper

def isupper(self, /)

Return True if the string is an uppercase string, False otherwise.

A string is uppercase if all cased characters in the string are uppercase and there is at least one cased character in the string.

PageResult.istitle

def istitle(self, /)

Return True if the string is a title-cased string, False otherwise.

In a title-cased string, upper- and title-case characters may only follow uncased characters and lowercase characters only cased ones.

PageResult.isspace

def isspace(self, /)

Return True if the string is a whitespace string, False otherwise.

A string is whitespace if all characters in the string are whitespace and there is at least one character in the string.

PageResult.isdecimal

def isdecimal(self, /)

Return True if the string is a decimal string, False otherwise.

A string is a decimal string if all characters in the string are decimal and there is at least one character in the string.

PageResult.isdigit

def isdigit(self, /)

Return True if the string is a digit string, False otherwise.

A string is a digit string if all characters in the string are digits and there is at least one character in the string.

PageResult.isnumeric

def isnumeric(self, /)

Return True if the string is a numeric string, False otherwise.

A string is numeric if all characters in the string are numeric and there is at least one character in the string.

PageResult.isalpha

def isalpha(self, /)

Return True if the string is an alphabetic string, False otherwise.

A string is alphabetic if all characters in the string are alphabetic and there is at least one character in the string.

PageResult.isalnum

def isalnum(self, /)

Return True if the string is an alpha-numeric string, False otherwise.

A string is alpha-numeric if all characters in the string are alpha-numeric and there is at least one character in the string.

PageResult.isidentifier

def isidentifier(self, /)

Return True if the string is a valid Python identifier, False otherwise.

Call keyword.iskeyword(s) to test whether string s is a reserved identifier, such as “def” or “class”.

PageResult.isprintable

def isprintable(self, /)

Return True if all characters in the string are printable, False otherwise.

A character is printable if repr() may use it in its output.

PageResult.zfill

def zfill(self, width, /)

Pad a numeric string with zeros on the left, to fill a field of the given width.

The string is never truncated.

PageResult.format

def format(self, /, *args, **kwargs)

Return a formatted version of the string, using substitutions from args and kwargs. The substitutions are identified by braces (’{’ and ’}’).

PageResult.format_map

def format_map(self, mapping, /)

Return a formatted version of the string, using substitutions from mapping. The substitutions are identified by braces (’{’ and ’}’).

PageResult.maketrans

def maketrans(unknown)

Return a translation table usable for str.translate().

If there is only one argument, it must be a dictionary mapping Unicode ordinals (integers) or characters to Unicode ordinals, strings or None. Character keys will be then converted to ordinals. If there are two arguments, they must be strings of equal length, and in the resulting dictionary, each character in x will be mapped to the character at the same position in y. If there is a third argument, it must be a string, whose characters will be mapped to None in the result.

run_script

def run_script( script: str | os.PathLike, env: dict[str, str] | None = None, binary: str | os.PathLike | None = None, timeout: float | None = None, args: Sequence[str] = () ) -> str

Replay a saved lightpanda script (no LLM) and return its stdout.

env entries (e.g. LP_* placeholder values) are added to the child’s environment. args are extra CLI flags for the browser, e.g. ["--obey-robots"] to enforce robots.txt. Raises ScriptError on a non-zero exit.

AsyncBrowser

A lightpanda browser process, driven from asyncio.

The subprocess is spawned by start() — called automatically on async with entry and by new_session(). Not fork-inheritable, same as Browser.

Every public method and property of Browser exists on AsyncBrowser with the same name and signature; methods are coroutines to await. Only the members AsyncBrowser adds are listed below.

class AsyncBrowser( binary: str | os.PathLike | None = None, env: dict[str, str] | None = None, timeout: float = 300.0, verbose: bool = False, args: Sequence[str] = (), max_concurrency: int = 32 )

Store the settings; the process is spawned by start().

binary, env, timeout, verbose and args are forwarded to Browser.

Arguments:

  • max_concurrency: Caps the tool calls executing concurrently across this browser’s sessions; worker threads are created lazily.

Usable as a context manager (async with).

AsyncBrowser.wrap

@classmethod def wrap( cls, browser: Browser, max_concurrency: int = 32 ) -> AsyncBrowser

Adopt an already-running Browser — the migration path for driving existing sync setup from asyncio. close() shuts down the facade but leaves the wrapped browser running.

AsyncBrowser.start

async def start(self) -> AsyncBrowser

Spawn the browser process and fetch its tool list. Idempotent.

AsyncBrowser.session

@contextlib.asynccontextmanager async def session(self)

async with browser.session() as page: — a session scoped to the block and closed on exit. Use new_session() for the unscoped form.

AsyncSession

One isolated browsing context (own page, cookies, memory), async.

Do not construct directly — use AsyncBrowser.new_session().

Same conventions as Session — keyword-only actions in snake_case, selector winning over backend_node_id, call() as the escape hatch — with every method a coroutine to await.

Every public method and property of Session exists on AsyncSession with the same name and signature; methods are coroutines to await. AsyncSession adds no members of its own.

class AsyncSession( session: Session, executor: concurrent.futures.thread.ThreadPoolExecutor )

Usable as a context manager (async with).

run_script_async

async def run_script_async( script: str | os.PathLike, env: dict[str, str] | None = None, binary: str | os.PathLike | None = None, timeout: float | None = None, args: Sequence[str] = () ) -> str

Async variant of lightpanda.run_script() (runs in a worker thread).

CDPServer

A lightpanda process serving the Chrome DevTools Protocol on 127.0.0.1.

from lightpanda import CDPServer from playwright.sync_api import sync_playwright with CDPServer() as server, sync_playwright() as p: browser = p.chromium.connect_over_cdp(server.ws_endpoint) page = browser.new_context().new_page() page.goto("https://example.com")

Every connected client gets its own browser; up to 16 connect at once by default (args=["--cdp-max-connections", "N"] to change). The process is stopped by close() / leaving the with block, and on Linux also when the interpreter dies.

class CDPServer( binary: str | os.PathLike | None = None, env: dict[str, str] | None = None, verbose: bool = False, args: Sequence[str] = (), port: int | None = None )

Spawn the server process.

Arguments:

  • binary: Path to a lightpanda binary. When omitted, resolved from the LIGHTPANDA_BIN environment variable, then the binary bundled in the package, then PATH.
  • env: Extra environment variables for the spawned process.
  • verbose: Let the browser’s own logging through to stderr.
  • args: Extra lightpanda serve flags, e.g. ["--obey-robots"] to enforce robots.txt; pass port= rather than --port.
  • port: Pin the listening port. Defaults to a free one.

Usable as a context manager (with).

CDPServer.ws_endpoint

ws_endpoint: str

Property. The CDP WebSocket URL, ws://127.0.0.1:<port>/.

Keep it as is: the server only upgrades on path / and only accepts an IP-literal or localhost host.

CDPServer.version

def version(self) -> dict

The /json/version document (browser, protocol version, webSocketDebuggerUrl).

CDPServer.port

port: int

Property. The port the server listens on.

CDPServer.http_endpoint

http_endpoint: str

Property. http://127.0.0.1:<port>, the server’s HTTP root: what Puppeteer (browserURL) and Playwright (connect_over_cdp with an http URL) discover the CDP WebSocket from, and Selenium’s command_executor.

CDPServer.close

def close(self) -> None

Stop the server process. Idempotent.

AsyncCDPServer

CDPServer for asyncio: the process is spawned by start(), called automatically on async with entry.

async with AsyncCDPServer() as server, async_playwright() as p: browser = await p.chromium.connect_over_cdp(server.ws_endpoint)

Every public method and property of CDPServer exists on AsyncCDPServer with the same name and signature; methods are coroutines to await. Only the members AsyncCDPServer adds are listed below.

class AsyncCDPServer( binary: str | os.PathLike | None = None, env: dict[str, str] | None = None, verbose: bool = False, args: Sequence[str] = (), port: int | None = None )

Store the settings; the process is spawned by start().

Arguments are forwarded to the sync class.

Usable as a context manager (async with).

AsyncCDPServer.start

async def start(self)

Spawn the server process. Idempotent.

BiDiServer

A lightpanda process serving WebDriver BiDi on 127.0.0.1.

from lightpanda import BiDiServer from selenium import webdriver from selenium.webdriver.common.options import ArgOptions options = ArgOptions() options.web_socket_url = True # ask for a WebDriver BiDi session with BiDiServer() as server: driver = webdriver.Remote(command_executor=server.http_endpoint, options=options) context = driver.browsing_context.create(type="tab") driver.browsing_context.navigate(context=context, url="https://example.com", wait="complete") print(driver.script.execute("() => document.title", context_id=context)["value"]) driver.quit()

http_endpoint is Selenium’s command_executor. The browser serves the BiDi modules (session, browser, browsingContext, script, input) over the WebSocket plus the classic session bootstrap (GET /status, POST /session with the webSocketUrl capability, DELETE /session/<id>); other classic WebDriver commands such as Selenium’s driver.get or find_element are not served, so drive the page through driver.browsing_context and driver.script with an explicit context, created first as above. Pass args=["--protocol", "cdp"] to serve CDP on the same port as well (--protocol is additive). The process is stopped by close() / leaving the with block, and on Linux also when the interpreter dies.

class BiDiServer( binary: str | os.PathLike | None = None, env: dict[str, str] | None = None, verbose: bool = False, args: Sequence[str] = (), port: int | None = None )

Spawn the server process.

Arguments:

  • binary: Path to a lightpanda binary. When omitted, resolved from the LIGHTPANDA_BIN environment variable, then the binary bundled in the package, then PATH.
  • env: Extra environment variables for the spawned process.
  • verbose: Let the browser’s own logging through to stderr.
  • args: Extra lightpanda serve flags, e.g. ["--obey-robots"] to enforce robots.txt; pass port= rather than --port.
  • port: Pin the listening port. Defaults to a free one.

Usable as a context manager (with).

BiDiServer.bidi_endpoint

bidi_endpoint: str

Property. The session-less BiDi WebSocket URL, ws://127.0.0.1:<port>/session, for clients that speak BiDi directly (session.new over the socket). A session bootstrapped through POST /session gets its own socket at <bidi_endpoint>/<sessionId>, returned as the webSocketUrl capability.

Keep the IP literal: the WebSocket upgrade rejects any Origin header and only accepts an IP-literal or localhost host.

BiDiServer.status

def status(self) -> dict

The GET /status value, {"ready": True, "message": ""}.

BiDiServer.port

port: int

Property. The port the server listens on.

BiDiServer.http_endpoint

http_endpoint: str

Property. http://127.0.0.1:<port>, the server’s HTTP root: what Puppeteer (browserURL) and Playwright (connect_over_cdp with an http URL) discover the CDP WebSocket from, and Selenium’s command_executor.

BiDiServer.close

def close(self) -> None

Stop the server process. Idempotent.

AsyncBiDiServer

BiDiServer for asyncio: the process is spawned by start(), called automatically on async with entry.

Every public method and property of BiDiServer exists on AsyncBiDiServer with the same name and signature; methods are coroutines to await. Only the members AsyncBiDiServer adds are listed below.

class AsyncBiDiServer( binary: str | os.PathLike | None = None, env: dict[str, str] | None = None, verbose: bool = False, args: Sequence[str] = (), port: int | None = None )

Store the settings; the process is spawned by start().

Arguments are forwarded to the sync class.

Usable as a context manager (async with).

AsyncBiDiServer.start

async def start(self)

Spawn the server process. Idempotent.

Exceptions

LightpandaError

class LightpandaError(Exception)

Base error for the lightpanda package.

ProcessError

class ProcessError(LightpandaError)

The browser binary could not be found, started, or reached.

ProtocolError

class ProtocolError(LightpandaError) ProtocolError(message: str, code: int | None = None)

JSON-RPC level failure (invalid request, timeout, internal error).

  • code. The JSON-RPC error code, when the server sent one.

ScriptError

class ScriptError(LightpandaError) ScriptError(message: str, returncode: int, stdout: str = '', stderr: str = '')

A script replay (run_script) exited with a failure.

  • returncode. The process exit status, or -1 when the script file does not exist.
  • stdout. What the script wrote to stdout before failing.
  • stderr. What the script wrote to stderr.

ToolError

class ToolError(LightpandaError)

A browser tool reported failure (bad selector, JS exception, …).