Python SDK
Every public class, method, property and exception of the lightpanda package, with the signatures and docstrings shipped in the code. See Use the Python SDK for a practical walkthrough. Every sync class has an asyncio twin with the same methods, awaitable; the async sections below list only what the twin adds.
Browser
A lightpanda browser process. Spawns the bundled binary on first use.
Not fork-inheritable: after os.fork()/multiprocessing, create a
fresh Browser in the child.
class Browser(
binary: str | os.PathLike | None = None,
env: dict[str, str] | None = None,
timeout: float = 300.0,
verbose: bool = False,
args: Sequence[str] = ()
)Spawn the browser process and fetch its tool list.
Arguments:
- binary: Path to a lightpanda binary. When omitted, resolved from
the
LIGHTPANDA_BINenvironment variable, then the binary bundled in the package, thenPATH. - env: Extra environment variables for the spawned process.
- timeout: Seconds to wait for a response to any request before
raising
ProtocolError. - verbose: Let the browser’s own logging through to stderr.
- args: Extra CLI flags for the spawned browser process, e.g.
["--obey-robots"]to enforcerobots.txt,["--http-cache-dir", path]or cookie flags.
Usable as a context manager (with).
Browser.tools
tools: dict[str, dict]Property. Tool name → {description, schema, output_schema}, as reported by
the browser. output_schema is None for tools that declare none.
Browser.new_session
def new_session(self) -> SessionOpen a new isolated browsing context: its own page, cookies and
memory. Close it with Session.close() or a with block.
Browser.close
def close(self) -> NoneStop the browser process, closing every session with it.
Session
One isolated browsing context (own page, cookies, memory).
Do not construct directly — use Browser.new_session().
Browser actions are keyword-only methods, named in snake_case after the
browser’s own action names: the waitForSelector action is
wait_for_selector(), and its backendNodeId argument is
backend_node_id.
Where a method accepts both selector and backend_node_id, pass one
of the two. selector is preferred for reproducibility and wins when
both are given; backend_node_id takes the values returned by
tree(), links() or find_element().
call() is the escape hatch that takes the action and argument names
exactly as the browser declares them. A failed action raises
ToolError.
class Session(browser: Browser, session_id: str)Usable as a context manager (with).
Session.id
id: strProperty. The session id, as the browser knows it.
Session.call
def call(self, tool: str, **kwargs)Invoke a browser tool by name. The generated methods route here.
Accepts the tool and argument names as the browser declares them
(waitForSelector, backendNodeId) as well as their snake_case
forms. Returns parsed JSON for JSON-carrying tools, bytes for
image results (screenshot without path), otherwise the result
text. Raises ToolError when the tool reports a failure.
Arguments:
- tool: The tool name.
- **kwargs: The tool’s arguments;
Nonevalues are omitted.
Session.close
def close(self) -> NoneRelease the session’s page. Idempotent; calls made after this
raise ToolError. Closing the browser closes every session.
Session.click
def click(
self,
*,
selector: str | None = None,
backend_node_id: int | None = None
) -> PageResultClick on an interactive element. Provide either a CSS selector (preferred for reproducibility) or a backendNodeId. Returns the current page URL and title after the click.
Arguments:
- selector: CSS selector of the element to click. Preferred over backendNodeId.
- backend_node_id: The backend node ID of the element to click.
Returns:
PageResult: The sentence above, with these attributes (None when absent):
url: URL of the page the call left loaded.http_status: Response status of that page’s own document. Absent when no response has arrived. 4xx/5xx means an error page, not the content.title: Document title. Empty or absent when the page has none.
Session.console_logs
def console_logs(self) -> AnyGet buffered console.log/warn/error messages from the current page. Returns all messages since last call and clears the buffer.
Session.detect_forms
def detect_forms(self, *, url: str | None = None, timeout: int | None = None) -> AnyList the forms on the page as JSON: each form’s backendNodeId, action, method and fields, where each field has backendNodeId, tagName, name, inputType, required, disabled, and when present value, placeholder and select options. Use it before filling a form to see every field it expects. It returns no CSS selectors; get one per field with nodeDetails so the fill calls stay replayable. If a url is provided, it navigates there first.
Arguments:
- url: Optional URL to navigate to before processing.
- timeout: Optional timeout in milliseconds. Defaults to 10000.
Session.evaluate
def evaluate(
self,
*,
script: str,
url: str | None = None,
timeout: int | None = None,
save: str | None = None
) -> AnyEvaluate JavaScript in the current page context — an escape hatch for page-side logic the dedicated tools can’t express; prefer extract for data and click/fill/etc. for actions. It runs in the page, so it cannot see the agent script’s variables or builtins — interpolate any value into the script string. A bare trailing expression yields its value; top-level await and return are supported (the body then runs as an async function, so use return to produce a value). Objects and arrays return as JSON, so no JSON.stringify is needed. If a url is provided, it navigates there first. The globalThis.lp object exposes a Session-scoped bridge store: values written via lp.foo = ... auto-sync at end of evaluate, surviving navigation; values previously set via /extract save= or /evaluate save= appear as lp.<name>.
Arguments:
- script: JavaScript run in the page context. A bare trailing expression, or
returnwith top-levelawait, is the result. - url: Optional URL to navigate to before evaluating.
- timeout: Optional timeout in milliseconds. Defaults to 10000.
- save: Optional bridge-store key. The evaluate’s return value is stored under this name and re-exposed as
lp.<name>to subsequent evaluates. Objects, arrays, and strings are serialized automatically — no JSON.stringify needed.
Session.extract
def extract(self, *, schema: str | dict | list, save: str | None = None) -> AnyExtract structured data from the current page (navigate first). schema is a JSON object (passed as a string) mapping output field names to CSS-selector specs. It is NOT a JSON Schema — no “type”/“properties” wrappers; the keys ARE your output fields. Value shapes:
"<sel>" → first match's text (trimmed; null if no match)
["<sel>"] → every match's text (string[])
{"selector":"<sel>","attr":"<name>"} → first match's attribute value (href/src resolved to absolute URLs)
[{"selector":"<sel>","attr":"<name>"}] → every match's attribute (string[])
[{"selector":"<sel>","fields":{…}}] → one object per match; field selectors resolve relative to that match and accept any shape above ("" = the match's own text; nest arrays for per-item sub-lists)Add “limit”: N inside any array’s object spec to cap matches.
Every extracted value is a string or null — parse numbers downstream. An empty array is a valid result, but if ALL top-level keys miss, the call errors: inspect the page (tree/markdown) and retry with corrected selectors.
Finish data tasks with extract — it is the only read recorded as a replayable extract(...) script call; answers lifted from markdown text in chat are not.
Examples (schema → result):
{"karma": "#karma"} → {"karma":"42"}
{"items": [".story .title"]} → {"items":["Title 1","Title 2"]}
{"top3": [{"selector":".story .title","limit":3}]} → {"top3":["A","B","C"]}
{"links": [{"selector":"a.title","attr":"href"}]} → {"links":["https://site/a","https://site/b"]}
{"stories": [{"selector":".athing","fields":{"title":".titleline","rank":".rank"}}]} → {"stories":[{"title":"Foo","rank":"1"}]}Arguments:
- schema: Extraction schema as a string: a JSON object literal mapping output field names to CSS-selector specs (see tool description). Not a JSON Schema.
- save: Optional bridge-store key. The extracted JSON is stored under this name and exposed as
lp.<name>in subsequent /evaluate calls.
Session.fill
def fill(
self,
*,
value: str,
selector: str | None = None,
backend_node_id: int | None = None
) -> PageResultFill text into an input element. Provide either a CSS selector (preferred for reproducibility) or a backendNodeId.
Arguments:
- value: The text to fill into the input element.
- selector: CSS selector of the input element to fill. Preferred over backendNodeId.
- backend_node_id: The backend node ID of the input element to fill.
Returns:
PageResult: The sentence above, with these attributes (None when absent):
url: URL of the page the call left loaded.http_status: Response status of that page’s own document. Absent when no response has arrived. 4xx/5xx means an error page, not the content.title: Document title. Empty or absent when the page has none.
Session.find_element
def find_element(self, *, role: str | None = None, name: str | None = None) -> AnyFind interactive elements by role and/or accessible name. Returns matching elements with their backend node IDs. Useful for locating specific elements without parsing the full semantic tree.
Arguments:
- role: Optional ARIA role to match (e.g. ‘button’, ‘link’, ‘textbox’, ‘checkbox’).
- name: Optional accessible name to match, case-insensitive: a substring, or a JavaScript regex literal such as /sign (in|up)/ (unanchored; flags i, m, s, u accepted; case-insensitive even without i, prefix (?-i) to make it case-sensitive).
Session.get_cookies
def get_cookies(self, *, url: str | None = None, all: bool | None = None) -> AnyCookies stored in the browser. Defaults to cookies whose domain matches the current page’s host. Pass url=<URL> to filter for another host, or all=true to dump every cookie regardless of host. Useful for debugging authentication and session state.
Arguments:
- url: Restrict output to cookies matching this URL’s host. Defaults to the current page.
- all: If true, dump every cookie regardless of host. Overrides
url.
Session.get_env
def get_env(self, *, name: str | None = None) -> AnyWith name: read an LP_* env var (other namespaces report as not set) — for non-secret config only (base URLs, flags). Without name: list LP_* names that are set (no values) — safe credential discovery. For secrets, pass $LP_* placeholders in tool args; never request a credential by name (the value would land in your context).
Arguments:
- name: Optional. If provided, must start with LP_; returns the value. If omitted, returns the list of LP_* names that are set.
Session.get_url
def get_url(self) -> AnyCurrent page URL. The browser may already have a page loaded (command, replayed script) not visible in this conversation — call this before assuming nothing is loaded when the user references the current page/site. Also useful to verify a navigation or detect a redirect.
Session.goto
def goto(
self,
*,
url: str,
timeout: int | None = None,
wait_until: str | None = None
) -> PageResultNavigate the current page to a URL. Returns the HTTP status once waitUntil fires (default load), or a timeout notice; a 4xx or 5xx means the page you got is an error page, not the content — check it before reading on; content rendered by post-load JavaScript may not be there yet (see waitForState). The page stays loaded for later reads and actions. To navigate and read in one call, pass url to markdown, tree or html instead; use goto when the next step is an action or extract.
Arguments:
- url: The URL to navigate to, must be a valid URL.
- timeout: Optional timeout in milliseconds. Defaults to 10000.
- wait_until: Event that completes the navigation. Defaults to ‘load’. Prefer ‘domcontentloaded’ followed by waitForSelector on pages whose late scripts (ads) hold ‘load’ back. Avoid ‘done’ (full quiescence): on pages with constant background activity it is the slowest choice and can run to the timeout.
Returns:
PageResult: The sentence above, with these attributes (None when absent):
url: URL of the page the call left loaded.http_status: Response status of that page’s own document. Absent when no response has arrived. 4xx/5xx means an error page, not the content.title: Document title. Empty or absent when the page has none.
Session.hover
def hover(
self,
*,
selector: str | None = None,
backend_node_id: int | None = None
) -> PageResultHover over an element, triggering mouseover and mouseenter events. Provide either a CSS selector (preferred for reproducibility) or a backendNodeId. Useful for menus, tooltips, and hover states.
Arguments:
- selector: CSS selector of the element to hover over. Preferred over backendNodeId.
- backend_node_id: The backend node ID of the element to hover over.
Returns:
PageResult: The sentence above, with these attributes (None when absent):
url: URL of the page the call left loaded.http_status: Response status of that page’s own document. Absent when no response has arrived. 4xx/5xx means an error page, not the content.title: Document title. Empty or absent when the page has none.
Session.html
def html(
self,
*,
selector: str | None = None,
backend_node_id: int | None = None,
max_bytes: int | None = None,
strip: dict | None = None,
url: str | None = None,
timeout: int | None = None
) -> AnyRaw HTML for the document or, with selector/backendNodeId, a single node’s outerHTML. Verbose; use only when you need attributes that markdown discards.
Arguments:
- selector: Optional CSS selector. When set, dump only that element’s outerHTML.
- backend_node_id: Optional backend node ID. When set, dump only that node’s outerHTML. 0 is treated as omitted.
- max_bytes: Optional soft cap on output size in bytes. Content is truncated at a UTF-8 boundary and a short ‘[truncated]’ marker is appended past the cap.
- strip: Optional. Omit element groups from the output:
js(script, noscript, script preloads),css(style, stylesheet links),ui(css plus img, picture, video, audio, svg, canvas, iframe),invisible(elements an author rule or inline style sets to display:none),shell(nav, aside, dialog, page-level header/footer and the matching landmark roles; skipped when that would drop most of the text),clutter(keep only the main content, in the manner of reader modes; includesshellandinvisible, and falls back toshellwhen it finds too little). {“js”:true,“css”:true} keeps a page dump small. - url: Optional URL to navigate to before dumping.
- timeout: Optional timeout in milliseconds. Defaults to 10000.
Session.interactive_elements
def interactive_elements(self, *, url: str | None = None, timeout: int | None = None) -> AnyList every visible interactive element on the page as a JSON array: native controls, ARIA widgets, contenteditable regions, elements with event listeners, and focusable elements. Each entry has backendNodeId, tagName, role, name, type (why it counts as interactive), tabIndex, and when present listeners, disabled, id, class, href, inputType, value, elementName and placeholder. Use it to survey what can be acted on; to locate one element by role or name, findElement is cheaper. If a url is provided, it navigates there first.
Arguments:
- url: Optional URL to navigate to before processing.
- timeout: Optional timeout in milliseconds. Defaults to 10000.
Session.links
def links(
self,
*,
limit: int | None = None,
url: str | None = None,
timeout: int | None = None
) -> AnyExtract the visible links in the opened page as JSON objects with text (anchor text, falling back to aria-label/title/image alt), href (resolved URL), and backendNodeId (pass to click/nodeDetails). One entry per href; hidden links are omitted. If a url is provided, it navigates to that url first.
Arguments:
- limit: Optional. Return at most this many links, in document order.
- url: Optional URL to navigate to before processing.
- timeout: Optional timeout in milliseconds. Defaults to 10000.
Session.markdown
def markdown(
self,
*,
selector: str | None = None,
backend_node_id: int | None = None,
max_bytes: int | None = None,
strip: dict | None = None,
url: str | None = None,
timeout: int | None = None
) -> AnyRender the page (or a subtree) as markdown. Scope with selector or backendNodeId to read just the relevant region — full-page markdown is the last resort. Use maxBytes to cap long pages.
Arguments:
- selector: Optional CSS selector. Render markdown for just that element’s subtree.
- backend_node_id: Optional backend node ID. Render markdown for just that node’s subtree. 0 is treated as omitted.
- max_bytes: Optional soft cap on output size in bytes. Content is truncated at a UTF-8 boundary and a short ‘[truncated]’ marker is appended past the cap.
- strip: Optional. Omit element groups from the output; same groups as the html tool’s strip.
shell(page chrome by markup) andclutter(keep only the main content, in the manner of reader modes) are the ones that matter for reading;uialso drops images. - url: Optional URL to navigate to before rendering.
- timeout: Optional timeout in milliseconds. Defaults to 10000.
Session.node_details
def node_details(self, *, backend_node_id: int) -> AnyDetails for a node by backendNodeId: a ready-to-use CSS selector that resolves to the node (the first match, as click/fill resolve it), plus tag, role, name, interactivity, disabled, value, input type, placeholder, href, id, class, checked, select options. The canonical way to turn a tree backendNodeId into a CSS selector.
Arguments:
- backend_node_id: The backend node ID of the element to inspect.
Session.press
def press(
self,
*,
key: str,
selector: str | None = None,
backend_node_id: int | None = None
) -> PageResultPress a keyboard key, dispatching keydown and keyup events. Use key names like ‘Enter’, ‘Tab’, ‘Escape’, ‘ArrowDown’, ‘Backspace’, or single characters like ‘a’, ‘1’. Common shorthand is normalized: ‘enter’/‘return’ → ‘Enter’, ‘esc’ → ‘Escape’, ‘up’/‘down’/‘left’/‘right’ → ‘Arrow*’, ‘space’ → ’ ’. Pressing ‘Enter’ on a form input or submit button triggers implicit form submission.
Arguments:
- key: The key to press (e.g. ‘Enter’, ‘Tab’, ‘a’).
- selector: Optional CSS selector of the element to target. Preferred over backendNodeId.
- backend_node_id: Optional backend node ID of the element to target. Defaults to the document when neither selector nor backendNodeId is provided; 0 is treated as omitted.
Returns:
PageResult: The sentence above, with these attributes (None when absent):
url: URL of the page the call left loaded.http_status: Response status of that page’s own document. Absent when no response has arrived. 4xx/5xx means an error page, not the content.title: Document title. Empty or absent when the page has none.
Session.screenshot
def screenshot(
self,
*,
path: str | None = None,
selector: str | None = None,
backend_node_id: int | None = None,
full_page: bool | None = None,
strip: dict | None = None,
url: str | None = None,
timeout: int | None = None
) -> AnyRender the page, or one node, as a PNG: the text layout Lightpanda computes, not a pixel-accurate browser rendering (no images, fonts or CSS colours). With path, writes the file at full size and returns its location; without it, returns the image inline where the client can display one, at most 1280px wide and 4096px tall. Use it to see spatial layout; read content with markdown/tree.
Arguments:
- path: Optional relative path (no ’..’ segments) to write the PNG to. Created or overwritten. Without it the image is returned inline, which needs a client that can display images.
- selector: Optional CSS selector. When set, render only that element.
- backend_node_id: Optional backend node ID. When set, render only that node. 0 is treated as omitted.
- full_page: Render the whole content height instead of one viewport. Defaults to false.
- strip: Optional. Omit element groups from the render; same groups as the html tool’s strip (
js,css,ui,invisible,shell,clutter). - url: Optional URL to navigate to before rendering.
- timeout: Optional timeout in milliseconds. Defaults to 10000.
Session.scroll
def scroll(
self,
*,
selector: str | None = None,
backend_node_id: int | None = None,
x: int | None = None,
y: int | None = None
) -> PageResultScroll the window, or an element’s scroll container, to an absolute position; an omitted axis keeps its current offset. Target an element with a CSS selector (preferred for reproducibility) or a backendNodeId; omit both to scroll the window. Page scripts receive a scroll event, so content that loads on scroll (infinite feeds, lazy lists) may appear: read the page again afterwards, with waitForState if it is still loading. Returns the final scroll position and the current page URL and title.
Arguments:
- selector: Optional: CSS selector of the element to scroll. Preferred over backendNodeId. If the element is not itself a scroll container, its nearest scrollable ancestor is scrolled instead.
- backend_node_id: Optional: The backend node ID of the element to scroll. If the element is not itself a scroll container, its nearest scrollable ancestor is scrolled instead. If neither this nor selector is given (or it is 0), scrolls the window.
- x: Optional: The horizontal scroll offset.
- y: Optional: The vertical scroll offset.
Returns:
PageResult: The sentence above, with these attributes (None when absent):
url: URL of the page the call left loaded.http_status: Response status of that page’s own document. Absent when no response has arrived. 4xx/5xx means an error page, not the content.title: Document title. Empty or absent when the page has none.
Session.search
def search(self, *, query: str, timeout: int | None = None) -> AnyRun a web search and return results as markdown: a numbered list of {title, url, snippet}. Search tries brave, tavily, exa, then keenable in order, each when its API key (BRAVE_API_KEY, TAVILY_API_KEY, EXA_API_KEY or KEENABLE_API_KEY) is set; keenable also works without a key through its public endpoint (rate-limited per client IP). Prefer this over goto-ing google.com/search directly (Google blocks the browser on User-Agent/TLS). The browser does not navigate — to open a result, use goto with its URL.
Arguments:
- query: The search query.
- timeout: Optional timeout in milliseconds. Defaults to 10000.
Session.select_option
def select_option(
self,
*,
value: str,
selector: str | None = None,
backend_node_id: int | None = None
) -> PageResultSelect an option in a <select> dropdown element by its value. Provide either a CSS selector (preferred for reproducibility) or a backendNodeId. Dispatches input and change events.
Arguments:
- value: The value of the option to select.
- selector: CSS selector of the <select> element. Preferred over backendNodeId.
- backend_node_id: The backend node ID of the <select> element.
Returns:
PageResult: The sentence above, with these attributes (None when absent):
url: URL of the page the call left loaded.http_status: Response status of that page’s own document. Absent when no response has arrived. 4xx/5xx means an error page, not the content.title: Document title. Empty or absent when the page has none.
Session.set_checked
def set_checked(
self,
*,
checked: bool,
selector: str | None = None,
backend_node_id: int | None = None
) -> PageResultCheck or uncheck a checkbox or radio button. Provide either a CSS selector (preferred for reproducibility) or a backendNodeId. Dispatches input, change, and click events.
Arguments:
- checked: Whether to check (true) or uncheck (false) the element.
- selector: CSS selector of the checkbox or radio input element. Preferred over backendNodeId.
- backend_node_id: The backend node ID of the checkbox or radio input element.
Returns:
PageResult: The sentence above, with these attributes (None when absent):
url: URL of the page the call left loaded.http_status: Response status of that page’s own document. Absent when no response has arrived. 4xx/5xx means an error page, not the content.title: Document title. Empty or absent when the page has none.
Session.structured_data
def structured_data(self, *, url: str | None = None, timeout: int | None = None) -> AnyPage metadata as JSON: jsonLd (each JSON-LD block as a string), openGraph, twitterCard, meta and links (key/value lists), plus alternate (hreflang variants) and linkHeaders (relations from the HTTP Link header) when present. Empty sections come back as empty arrays. Use it for publisher-declared facts such as product price, article author or canonical URL before scraping the visible text for them. If a url is provided, it navigates there first.
Arguments:
- url: Optional URL to navigate to before processing.
- timeout: Optional timeout in milliseconds. Defaults to 10000.
Session.tree
def tree(
self,
*,
url: str | None = None,
timeout: int | None = None,
backend_node_id: int | None = None,
max_depth: int | None = None
) -> AnySemantic outline of the page as indented text: one node per line with its role, accessible name, value and backendNodeId, plus checked state and select options with the selected one marked. The default first read of an unfamiliar page; input and select values are already here, so no nodeDetails call is needed to read them. Pass backendNodeId to scope to a subtree and maxDepth to survey structure before going deeper. Read it again after any page-changing action, since the DOM it describes may have changed; use nodeDetails to turn a backendNodeId into a CSS selector for actions.
Arguments:
- url: Optional URL to navigate to before fetching the semantic tree.
- timeout: Optional timeout in milliseconds. Defaults to 10000.
- backend_node_id: Optional backend node ID to get the tree for a specific element instead of the document root. 0 is treated as omitted.
- max_depth: Optional maximum depth of the tree to return. Useful for exploring high-level structure first.
Session.wait_for_script
def wait_for_script(self, *, script: str, timeout: int | None = None) -> AnyWait until a JS expression returns truthy. Re-evaluates on each tick of the event loop. Use for synchronization beyond what CSS selectors can express — e.g. window.dataLoaded === true, document.readyState === 'complete', document.querySelectorAll('.row').length >= 5.
Arguments:
- script: JS expression evaluated each tick until truthy. Must be an expression (not a statement).
- timeout: Optional timeout in milliseconds. Defaults to 5000, or 15000 when the page has not reached ‘load’ yet.
Session.wait_for_selector
def wait_for_selector(self, *, selector: str, timeout: int | None = None) -> AnyWait for an element matching a CSS selector to appear in the page. Returns the backend node ID of the matched element.
Arguments:
- selector: The CSS selector to wait for.
- timeout: Optional timeout in milliseconds. Defaults to 5000, or 15000 when the page has not reached ‘load’ yet.
Session.wait_for_state
def wait_for_state(self, *, state: str, timeout: int | None = None) -> AnyWait for the CURRENT page to reach a load state (no navigation). After a goto, the page is returned at the fast load snapshot, so content rendered by post-load JS (XHR-loaded lists, feeds, search results) may still be missing. When a read looks incomplete — empty lists, spinners, skeletons — call this with ‘networkidle’ and re-read. Prefer ‘networkidle’; ‘done’ can be slow on sites with constant background activity (ads, polling).
Arguments:
- state: Load state to wait for. ‘networkidle’ = network settled (the usual choice to finish a dynamic page).
- timeout: Optional timeout in milliseconds. Defaults to 5000.
PageResult
What a navigation or action answered: a sentence that also carries fields.
goto and the action tools describe what they did in prose — which
element was clicked, whether a new window took over — and separately
report where that left the page. Dropping either would lose something, so
this is the sentence, unchanged for printing and comparison, with the
fields as attributes:
r = page.goto(url="https://example.com/missing")
print(r) # Navigated successfully. HTTP 404 Not Found.
r.http_status # 404http_status is None before any response has arrived, title on a
document that has none.
class PageResultPageResult.url
url: strPageResult.http_status
http_status: int | NonePageResult.title
def title(self, /)Return a version of the string where each word is titlecased.
More specifically, words start with uppercased characters and all remaining cased characters have lower case.
PageResult.encode
def encode(self, /, encoding='utf-8', errors='strict')Encode the string using the codec registered for encoding.
encoding
The encoding in which to encode the string.errors
The error handling scheme to use for encoding errors.
The default is 'strict' meaning that encoding errors raise a
UnicodeEncodeError. Other possible values are 'ignore', 'replace'
and 'xmlcharrefreplace' as well as any other name registered with
codecs.register_error that can handle UnicodeEncodeErrors.PageResult.replace
def replace(self, old, new, /, count=-1)Return a copy with all occurrences of substring old replaced by new.
count
Maximum number of occurrences to replace.
-1 (the default value) means replace all occurrences.If the optional argument count is given, only the first count occurrences are replaced.
PageResult.split
def split(self, /, sep=None, maxsplit=-1)Return a list of the substrings in the string, using sep as the separator string.
sep
The separator used to split the string. When set to None (the default value), will split on any
whitespace character (including \n \r \t \f and spaces) and
will discard empty strings from the result.
maxsplit
Maximum number of splits.
-1 (the default value) means no limit.Splitting starts at the front of the string and works to the end.
Note, str.split() is mainly useful for data that has been intentionally delimited. With natural text that includes punctuation, consider using the regular expression module.
PageResult.rsplit
def rsplit(self, /, sep=None, maxsplit=-1)Return a list of the substrings in the string, using sep as the separator string.
sep
The separator used to split the string. When set to None (the default value), will split on any
whitespace character (including \n \r \t \f and spaces) and
will discard empty strings from the result.
maxsplit
Maximum number of splits.
-1 (the default value) means no limit.Splitting starts at the end of the string and works to the front.
PageResult.join
def join(self, iterable, /)Concatenate any number of strings.
The string whose method is called is inserted in between each given string. The result is returned as a new string.
Example: ’.‘.join([‘ab’, ‘pq’, ‘rs’]) -> ‘ab.pq.rs’
PageResult.capitalize
def capitalize(self, /)Return a capitalized version of the string.
More specifically, make the first character have upper case and the rest lower case.
PageResult.casefold
def casefold(self, /)Return a version of the string suitable for caseless comparisons.
PageResult.center
def center(self, width, fillchar=' ', /)Return a centered string of length width.
Padding is done using the specified fill character (default is a space).
PageResult.count
def count(unknown)Return the number of non-overlapping occurrences of substring sub in string S[start:end].
Optional arguments start and end are interpreted as in slice notation.
PageResult.expandtabs
def expandtabs(self, /, tabsize=8)Return a copy where all tab characters are expanded using spaces.
If tabsize is not given, a tab size of 8 characters is assumed.
PageResult.find
def find(unknown)Return the lowest index in S where substring sub is found, such that sub is contained within S[start:end].
Optional arguments start and end are interpreted as in slice notation. Return -1 on failure.
PageResult.partition
def partition(self, sep, /)Partition the string into three parts using the given separator.
This will search for the separator in the string. If the separator is found, returns a 3-tuple containing the part before the separator, the separator itself, and the part after it.
If the separator is not found, returns a 3-tuple containing the original string and two empty strings.
PageResult.index
def index(unknown)Return the lowest index in S where substring sub is found, such that sub is contained within S[start:end].
Optional arguments start and end are interpreted as in slice notation. Raises ValueError when the substring is not found.
PageResult.ljust
def ljust(self, width, fillchar=' ', /)Return a left-justified string of length width.
Padding is done using the specified fill character (default is a space).
PageResult.lower
def lower(self, /)Return a copy of the string converted to lowercase.
PageResult.lstrip
def lstrip(self, chars=None, /)Return a copy of the string with leading whitespace removed.
If chars is given and not None, remove characters in chars instead.
PageResult.rfind
def rfind(unknown)Return the highest index in S where substring sub is found, such that sub is contained within S[start:end].
Optional arguments start and end are interpreted as in slice notation. Return -1 on failure.
PageResult.rindex
def rindex(unknown)Return the highest index in S where substring sub is found, such that sub is contained within S[start:end].
Optional arguments start and end are interpreted as in slice notation. Raises ValueError when the substring is not found.
PageResult.rjust
def rjust(self, width, fillchar=' ', /)Return a right-justified string of length width.
Padding is done using the specified fill character (default is a space).
PageResult.rstrip
def rstrip(self, chars=None, /)Return a copy of the string with trailing whitespace removed.
If chars is given and not None, remove characters in chars instead.
PageResult.rpartition
def rpartition(self, sep, /)Partition the string into three parts using the given separator.
This will search for the separator in the string, starting at the end. If the separator is found, returns a 3-tuple containing the part before the separator, the separator itself, and the part after it.
If the separator is not found, returns a 3-tuple containing two empty strings and the original string.
PageResult.splitlines
def splitlines(self, /, keepends=False)Return a list of the lines in the string, breaking at line boundaries.
Line breaks are not included in the resulting list unless keepends is given and true.
PageResult.strip
def strip(self, chars=None, /)Return a copy of the string with leading and trailing whitespace removed.
If chars is given and not None, remove characters in chars instead.
PageResult.swapcase
def swapcase(self, /)Convert uppercase characters to lowercase and lowercase characters to uppercase.
PageResult.translate
def translate(self, table, /)Replace each character in the string using the given translation table.
table
Translation table, which must be a mapping of Unicode ordinals
to Unicode ordinals, strings, or None.The table must implement lookup/indexing via getitem, for instance a dictionary or list. If this operation raises LookupError, the character is left untouched. Characters mapped to None are deleted.
PageResult.upper
def upper(self, /)Return a copy of the string converted to uppercase.
PageResult.startswith
def startswith(unknown)Return True if the string starts with the specified prefix, False otherwise.
prefix
A string or a tuple of strings to try.start
Optional start position. Default: start of the string.end
Optional stop position. Default: end of the string.PageResult.endswith
def endswith(unknown)Return True if the string ends with the specified suffix, False otherwise.
suffix
A string or a tuple of strings to try.start
Optional start position. Default: start of the string.end
Optional stop position. Default: end of the string.PageResult.removeprefix
def removeprefix(self, prefix, /)Return a str with the given prefix string removed if present.
If the string starts with the prefix string, return string[len(prefix):]. Otherwise, return a copy of the original string.
PageResult.removesuffix
def removesuffix(self, suffix, /)Return a str with the given suffix string removed if present.
If the string ends with the suffix string and that suffix is not empty, return string[:-len(suffix)]. Otherwise, return a copy of the original string.
PageResult.isascii
def isascii(self, /)Return True if all characters in the string are ASCII, False otherwise.
ASCII characters have code points in the range U+0000-U+007F. Empty string is ASCII too.
PageResult.islower
def islower(self, /)Return True if the string is a lowercase string, False otherwise.
A string is lowercase if all cased characters in the string are lowercase and there is at least one cased character in the string.
PageResult.isupper
def isupper(self, /)Return True if the string is an uppercase string, False otherwise.
A string is uppercase if all cased characters in the string are uppercase and there is at least one cased character in the string.
PageResult.istitle
def istitle(self, /)Return True if the string is a title-cased string, False otherwise.
In a title-cased string, upper- and title-case characters may only follow uncased characters and lowercase characters only cased ones.
PageResult.isspace
def isspace(self, /)Return True if the string is a whitespace string, False otherwise.
A string is whitespace if all characters in the string are whitespace and there is at least one character in the string.
PageResult.isdecimal
def isdecimal(self, /)Return True if the string is a decimal string, False otherwise.
A string is a decimal string if all characters in the string are decimal and there is at least one character in the string.
PageResult.isdigit
def isdigit(self, /)Return True if the string is a digit string, False otherwise.
A string is a digit string if all characters in the string are digits and there is at least one character in the string.
PageResult.isnumeric
def isnumeric(self, /)Return True if the string is a numeric string, False otherwise.
A string is numeric if all characters in the string are numeric and there is at least one character in the string.
PageResult.isalpha
def isalpha(self, /)Return True if the string is an alphabetic string, False otherwise.
A string is alphabetic if all characters in the string are alphabetic and there is at least one character in the string.
PageResult.isalnum
def isalnum(self, /)Return True if the string is an alpha-numeric string, False otherwise.
A string is alpha-numeric if all characters in the string are alpha-numeric and there is at least one character in the string.
PageResult.isidentifier
def isidentifier(self, /)Return True if the string is a valid Python identifier, False otherwise.
Call keyword.iskeyword(s) to test whether string s is a reserved identifier, such as “def” or “class”.
PageResult.isprintable
def isprintable(self, /)Return True if all characters in the string are printable, False otherwise.
A character is printable if repr() may use it in its output.
PageResult.zfill
def zfill(self, width, /)Pad a numeric string with zeros on the left, to fill a field of the given width.
The string is never truncated.
PageResult.format
def format(self, /, *args, **kwargs)Return a formatted version of the string, using substitutions from args and kwargs. The substitutions are identified by braces (’{’ and ’}’).
PageResult.format_map
def format_map(self, mapping, /)Return a formatted version of the string, using substitutions from mapping. The substitutions are identified by braces (’{’ and ’}’).
PageResult.maketrans
def maketrans(unknown)Return a translation table usable for str.translate().
If there is only one argument, it must be a dictionary mapping Unicode ordinals (integers) or characters to Unicode ordinals, strings or None. Character keys will be then converted to ordinals. If there are two arguments, they must be strings of equal length, and in the resulting dictionary, each character in x will be mapped to the character at the same position in y. If there is a third argument, it must be a string, whose characters will be mapped to None in the result.
run_script
def run_script(
script: str | os.PathLike,
env: dict[str, str] | None = None,
binary: str | os.PathLike | None = None,
timeout: float | None = None,
args: Sequence[str] = ()
) -> strReplay a saved lightpanda script (no LLM) and return its stdout.
env entries (e.g. LP_* placeholder values) are added to the
child’s environment. args are extra CLI flags for the browser, e.g.
["--obey-robots"] to enforce robots.txt. Raises
ScriptError on a non-zero exit.
AsyncBrowser
A lightpanda browser process, driven from asyncio.
The subprocess is spawned by start() — called automatically on
async with entry and by new_session(). Not fork-inheritable,
same as Browser.
Every public method and property of Browser exists on AsyncBrowser with the same name and signature; methods are coroutines to await. Only the members AsyncBrowser adds are listed below.
class AsyncBrowser(
binary: str | os.PathLike | None = None,
env: dict[str, str] | None = None,
timeout: float = 300.0,
verbose: bool = False,
args: Sequence[str] = (),
max_concurrency: int = 32
)Store the settings; the process is spawned by start().
binary, env, timeout, verbose and args are
forwarded to Browser.
Arguments:
- max_concurrency: Caps the tool calls executing concurrently across this browser’s sessions; worker threads are created lazily.
Usable as a context manager (async with).
AsyncBrowser.wrap
@classmethod
def wrap(
cls,
browser: Browser,
max_concurrency: int = 32
) -> AsyncBrowserAdopt an already-running Browser — the migration path for
driving existing sync setup from asyncio. close() shuts down
the facade but leaves the wrapped browser running.
AsyncBrowser.start
async def start(self) -> AsyncBrowserSpawn the browser process and fetch its tool list. Idempotent.
AsyncBrowser.session
@contextlib.asynccontextmanager
async def session(self)async with browser.session() as page: — a session scoped to
the block and closed on exit. Use new_session() for the
unscoped form.
AsyncSession
One isolated browsing context (own page, cookies, memory), async.
Do not construct directly — use AsyncBrowser.new_session().
Same conventions as Session — keyword-only actions in snake_case,
selector winning over backend_node_id, call() as the escape
hatch — with every method a coroutine to await.
Every public method and property of Session exists on AsyncSession with the same name and signature; methods are coroutines to await. AsyncSession adds no members of its own.
class AsyncSession(
session: Session,
executor: concurrent.futures.thread.ThreadPoolExecutor
)Usable as a context manager (async with).
run_script_async
async def run_script_async(
script: str | os.PathLike,
env: dict[str, str] | None = None,
binary: str | os.PathLike | None = None,
timeout: float | None = None,
args: Sequence[str] = ()
) -> strAsync variant of lightpanda.run_script() (runs in a worker thread).
CDPServer
A lightpanda process serving the Chrome DevTools Protocol on 127.0.0.1.
from lightpanda import CDPServer
from playwright.sync_api import sync_playwright
with CDPServer() as server, sync_playwright() as p:
browser = p.chromium.connect_over_cdp(server.ws_endpoint)
page = browser.new_context().new_page()
page.goto("https://example.com")Every connected client gets its own browser; up to 16 connect at once
by default (args=["--cdp-max-connections", "N"] to change). The
process is stopped by close() / leaving the with block, and on
Linux also when the interpreter dies.
class CDPServer(
binary: str | os.PathLike | None = None,
env: dict[str, str] | None = None,
verbose: bool = False,
args: Sequence[str] = (),
port: int | None = None
)Spawn the server process.
Arguments:
- binary: Path to a lightpanda binary. When omitted, resolved from
the
LIGHTPANDA_BINenvironment variable, then the binary bundled in the package, thenPATH. - env: Extra environment variables for the spawned process.
- verbose: Let the browser’s own logging through to stderr.
- args: Extra
lightpanda serveflags, e.g.["--obey-robots"]to enforcerobots.txt; passport=rather than--port. - port: Pin the listening port. Defaults to a free one.
Usable as a context manager (with).
CDPServer.ws_endpoint
ws_endpoint: strProperty. The CDP WebSocket URL, ws://127.0.0.1:<port>/.
Keep it as is: the server only upgrades on path / and only
accepts an IP-literal or localhost host.
CDPServer.version
def version(self) -> dictThe /json/version document (browser, protocol version,
webSocketDebuggerUrl).
CDPServer.port
port: intProperty. The port the server listens on.
CDPServer.http_endpoint
http_endpoint: strProperty. http://127.0.0.1:<port>, the server’s HTTP root: what Puppeteer
(browserURL) and Playwright (connect_over_cdp with an http URL)
discover the CDP WebSocket from, and Selenium’s command_executor.
CDPServer.close
def close(self) -> NoneStop the server process. Idempotent.
AsyncCDPServer
CDPServer for asyncio: the process is spawned by
start(), called automatically on async with entry.
async with AsyncCDPServer() as server, async_playwright() as p:
browser = await p.chromium.connect_over_cdp(server.ws_endpoint)Every public method and property of CDPServer exists on AsyncCDPServer with the same name and signature; methods are coroutines to await. Only the members AsyncCDPServer adds are listed below.
class AsyncCDPServer(
binary: str | os.PathLike | None = None,
env: dict[str, str] | None = None,
verbose: bool = False,
args: Sequence[str] = (),
port: int | None = None
)Store the settings; the process is spawned by start().
Arguments are forwarded to the sync class.
Usable as a context manager (async with).
AsyncCDPServer.start
async def start(self)Spawn the server process. Idempotent.
BiDiServer
A lightpanda process serving WebDriver BiDi on 127.0.0.1.
from lightpanda import BiDiServer
from selenium import webdriver
from selenium.webdriver.common.options import ArgOptions
options = ArgOptions()
options.web_socket_url = True # ask for a WebDriver BiDi session
with BiDiServer() as server:
driver = webdriver.Remote(command_executor=server.http_endpoint, options=options)
context = driver.browsing_context.create(type="tab")
driver.browsing_context.navigate(context=context, url="https://example.com", wait="complete")
print(driver.script.execute("() => document.title", context_id=context)["value"])
driver.quit()http_endpoint is Selenium’s command_executor. The browser
serves the BiDi modules (session, browser, browsingContext,
script, input) over the WebSocket plus the classic session
bootstrap (GET /status, POST /session with the webSocketUrl
capability, DELETE /session/<id>); other classic WebDriver commands
such as Selenium’s driver.get or find_element are not served, so
drive the page through driver.browsing_context and driver.script
with an explicit context, created first as above. Pass
args=["--protocol", "cdp"] to serve CDP on the same port as well
(--protocol is additive). The process is stopped by close() /
leaving the with block, and on Linux also when the interpreter dies.
class BiDiServer(
binary: str | os.PathLike | None = None,
env: dict[str, str] | None = None,
verbose: bool = False,
args: Sequence[str] = (),
port: int | None = None
)Spawn the server process.
Arguments:
- binary: Path to a lightpanda binary. When omitted, resolved from
the
LIGHTPANDA_BINenvironment variable, then the binary bundled in the package, thenPATH. - env: Extra environment variables for the spawned process.
- verbose: Let the browser’s own logging through to stderr.
- args: Extra
lightpanda serveflags, e.g.["--obey-robots"]to enforcerobots.txt; passport=rather than--port. - port: Pin the listening port. Defaults to a free one.
Usable as a context manager (with).
BiDiServer.bidi_endpoint
bidi_endpoint: strProperty. The session-less BiDi WebSocket URL, ws://127.0.0.1:<port>/session,
for clients that speak BiDi directly (session.new over the socket).
A session bootstrapped through POST /session gets its own socket at
<bidi_endpoint>/<sessionId>, returned as the webSocketUrl
capability.
Keep the IP literal: the WebSocket upgrade rejects any Origin
header and only accepts an IP-literal or localhost host.
BiDiServer.status
def status(self) -> dictThe GET /status value, {"ready": True, "message": ""}.
BiDiServer.port
port: intProperty. The port the server listens on.
BiDiServer.http_endpoint
http_endpoint: strProperty. http://127.0.0.1:<port>, the server’s HTTP root: what Puppeteer
(browserURL) and Playwright (connect_over_cdp with an http URL)
discover the CDP WebSocket from, and Selenium’s command_executor.
BiDiServer.close
def close(self) -> NoneStop the server process. Idempotent.
AsyncBiDiServer
BiDiServer for asyncio: the process is spawned by
start(), called automatically on async with entry.
Every public method and property of BiDiServer exists on AsyncBiDiServer with the same name and signature; methods are coroutines to await. Only the members AsyncBiDiServer adds are listed below.
class AsyncBiDiServer(
binary: str | os.PathLike | None = None,
env: dict[str, str] | None = None,
verbose: bool = False,
args: Sequence[str] = (),
port: int | None = None
)Store the settings; the process is spawned by start().
Arguments are forwarded to the sync class.
Usable as a context manager (async with).
AsyncBiDiServer.start
async def start(self)Spawn the server process. Idempotent.
Exceptions
LightpandaError
class LightpandaError(Exception)Base error for the lightpanda package.
ProcessError
class ProcessError(LightpandaError)The browser binary could not be found, started, or reached.
ProtocolError
class ProtocolError(LightpandaError)
ProtocolError(message: str, code: int | None = None)JSON-RPC level failure (invalid request, timeout, internal error).
code. The JSON-RPC error code, when the server sent one.
ScriptError
class ScriptError(LightpandaError)
ScriptError(message: str, returncode: int, stdout: str = '', stderr: str = '')A script replay (run_script) exited with a failure.
returncode. The process exit status, or-1when the script file does not exist.stdout. What the script wrote to stdout before failing.stderr. What the script wrote to stderr.
ToolError
class ToolError(LightpandaError)A browser tool reported failure (bad selector, JS exception, …).