Weave documentation
Rust referenceweave-browser-engine

weave-browser-engine · parse

Source declarations, signatures and documentation for parse.

Reviewed implementation boundary: The serve path warns but still binds when WEAVE_API_KEY is absent; the authentication check then allows requests. Configure and verify authorization before exposing a server. See /libraries/weave-browser-engine/research-server.

Source: sigil/weave/libs/weave-browser-engine/src/parse.rs. SHA-256: 5b87ef8b1f9cc0fd6ace8ad2407a861aaeff12f68bf45e8133917107d2a7097f.

This reference follows declared source modules, retains conditional attributes, and includes public declarations and implementation methods. Private-module re-exports and trait resolution require the compiler; this is a source reference, not a claim that every listed item is a root import. Function bodies and constant values are omitted.

parse::extract_text

Extract visible text from an HTML document, stripping scripts and styles.

Walks the <body> element (or the root if no body exists) and collects all text nodes that are not inside <script> or <style> elements.

pub fn extract_text(doc: &Html) -> String;

Source line: 14.

parse::extract_title

Extract the <title> element's text content.

pub fn extract_title(doc: &Html) -> String;

Source line: 62.

parse::extract_description

Extract the meta description from <meta name="description">.

pub fn extract_description(doc: &Html) -> String;

Source line: 71.

Extract all <a href="..."> links from the document.

Resolves relative URLs against base_url. Each result is a JSON object with "href" and "text" fields.

pub fn extract_links(doc: &Html, base_url: &str) -> Vec<Value>;

Source line: 84.

parse::extract_images

Extract all <img src="..."> images from the document.

Resolves relative src URLs against base_url. Each result is a JSON object with "src" and "alt" fields.

pub fn extract_images(doc: &Html, base_url: &str) -> Vec<Value>;

Source line: 112.

parse::extract_by_selector

Extract elements matching a CSS selector.

Returns a JSON array. Each element has "text" and "html" fields. If the selector is invalid, returns an array with a single error object.

pub fn extract_by_selector(doc: &Html, selector: &str) -> Vec<Value>;

Source line: 140.

parse::extract_structured_data

Extract structured data (JSON-LD, Open Graph) from the document.

Returns a vector of JSON objects, each with "type" and "data" fields.

pub fn extract_structured_data(doc: &Html) -> Vec<Value>;

Source line: 160.

On this page