Use before comparing or indexing Unicode text. Applies exactly NFC, NFD, NFKC, or NFKD and reports UTF-16 lengths. Compatibility forms can change meaning; this is not transliteration, case folding, or confusable detection.
Choose for: Apply an explicit Unicode normalization form.
Outside this profile: Transliteration, confusable detection, or case folding.
Use to split a document into bounded Unicode-code-point windows with configurable overlap. Exact source slices use UTF-16 offsets, so adjacent windows can be verified without losing source characters. Not tokenization or grapheme segmentation. At most 2,000 chunks and 500,000 returned UTF-16 units.
Choose for: Create bounded code-point windows with verifiable UTF-16 source offsets.
Outside this profile: Tokenizer-accurate token budgets or semantic passage segmentation.
Use for deterministic document sizing and basic QA. Counts UTF-16 units, Unicode code points, UTF-8 bytes, whitespace, logical lines, blank-line-separated paragraphs, and word-like spans. Words use Unicode letters/numbers with internal apostrophes; this is not linguistic segmentation, token counting, or a readability score. Empty text has zero lines; a final newline adds an empty line.
Choose for: Measure deterministic Unicode and byte sizes with basic word-like counts.
Outside this profile: Model token counts, language-specific segmentation, or readability scoring.
Use to remove repeated records from newline-delimited text while keeping the first original spelling and order. Comparison optionally trims whitespace and uses locale-independent lowercase. Output separators become LF; no fuzzy matching or Unicode normalization. Empty text has zero lines; a trailing newline represents a final blank line.
Choose for: Keep the first spelling and order with optional trim/lowercase comparison.
Outside this profile: Fuzzy entity resolution or Unicode normalization.
Use to produce a small, reproducible frequency index for a document. Counts Unicode word-like spans containing letters, lowercases without stemming, optionally filters a fixed English stop list plus your stop words, and sorts by frequency then first occurrence. No semantic ranking, language detection, phrase extraction, or AI. First offsets refer to original UTF-16 text.
Choose for: Count word-like spans with optional stopword filtering.
Outside this profile: Semantic keywords, sentiment, topic modeling, or language detection.
Use for auditable changes between small documents. Computes a minimal insertion/deletion script by bounded longest-common-subsequence, with one-based old/new line numbers; ties delete first. Maximum 300 lines and 30,000 UTF-16 units per document. Compares line content exactly but ignores CR/LF separator style; a final newline is an additional empty line. Not a byte diff, patch file, or fuzzy diff.
Choose for: Inspect exact bounded line changes with source line numbers.
Outside this profile: Large files, fuzzy matching, byte comparison, or executable patch output.
Use to build a lightweight navigation index or section slicer. Finds ATX headings (# through ######) with up to three leading spaces, ignores fenced code, and supplies parent indexes and exact section offsets. Heading text remains inline Markdown. Does not parse Setext headings, indented code, block quotes, HTML blocks, or a full CommonMark AST. At most 5,000 headings.
Choose for: Find supported ATX headings and exact section offsets.
Outside this profile: A complete CommonMark AST or Setext/HTML heading support.
Use to inventory explicit Markdown links without fetching them. Supports single-line inline links/images with non-nested labels, bare or angle-bracket destinations, balanced destination parentheses to depth 8, quoted titles, full/collapsed references with single-line definitions, and http(s) autolinks. Ignores fenced/inline code. No shortcut references, nested labels, HTML links, indented code, or full CommonMark parsing. URLs are untrusted, unvalidated strings; entities stay literal. Maximum 5,000 links, definitions, and unresolved references each.
Choose for: Collect supported explicit links and unresolved references without fetching.
Outside this profile: Link availability checks, full CommonMark, or URL safety certification.
Use to flatten supplied HTML snippets into readable text. Uses parse5 HTML parsing, decodes entities, excludes script/style/template/noscript, optionally includes image alt text, and collapses source whitespace. Block elements and br yield line breaks or spaces. No network, CSS/layout/visibility computation, sanitization guarantee, or browser rendering; hidden elements can remain. Input is parsed as an HTML fragment. Conservative limits: 1,000 less-than characters, 64 lexically open tags (explicitly close optional-end tags), and 4,096 visited parser nodes.
Choose for: Parse provided HTML fragments locally and omit script/style/template/noscript text.
Outside this profile: URL scraping, browser-rendered visibility, or HTML security sanitization.
Use to fill deterministic text templates without evaluating code. Replaces only exact {{name}} placeholders, where names use ASCII letters/digits/underscore/dot/hyphen and start with a letter or underscore. Values are literal strings and are never re-expanded or escaped for HTML/SQL. All recognized variables must exist; other brace syntax stays literal. At most 1,000 variables, 5,000 replacements, 200,000 value units, and 500,000 output units.
Choose for: Substitute literal named values without evaluating expressions.
Outside this profile: HTML/SQL escaping, recursive templates, or executable template languages.
Use as a review aid before inspecting logs or sample text. Replaces selected ASCII email, formatted US phone, XXX-XX-XXXX SSN-shaped, and valid dotted IPv4 patterns with fixed labels; returns original-source offsets without match values. Phone patterns require separators; SSN patterns do not validate issuance. Non-exhaustive heuristics can miss PII and produce false positives. Never a security, compliance, or anonymization guarantee. Maximum 10,000 matches.
Choose for: Aid review by masking a bounded set of common ASCII patterns.
Outside this profile: Complete anonymization, legal compliance, or exhaustive PII detection.
Use for predictable whitespace cleanup before diffing or storing text. Normalizes CR/LF separators to LF or CRLF, optionally collapses ASCII spaces/tabs, trims selected line edges, limits consecutive blank lines, and removes outer blank lines. Blank lines contain only spaces/tabs; other Unicode whitespace is preserved. Set maxBlankLines to null to preserve every blank line. This can change indentation-sensitive documents.
Choose for: Apply explicit line separator and ASCII whitespace cleanup.
Outside this profile: Preserving all indentation-sensitive semantics automatically.
Produce an atomic source-ordered localization bundle after exact key coverage and repeated {{name}} placeholder checks. Explicit source fallback provenance; malformed double braces or mismatches block the bundle. No translation, language verification, rendering or escaping.
Choose for: Build a source-ordered target localization dictionary with explicit fallback provenance Check literal translated placeholder names and repeated counts before releasing a bundle
Outside this profile: Generate or evaluate translations, detect language or verify semantic quality ICU/MessageFormat plural rules, printf, HTML/Markdown interpretation, sanitization or locale fallback chains Live execution, network lookups, payments, reservation, or inferring factual correctness of supplied data
Rank supplied documents with fixed-parameter lexical BM25 and exact original UTF-16 token spans. Scores are binary64 logarithm approximations ordered at 12 decimal places, with exact ID ties. No embeddings, persistent index, web search or semantic relevance guarantee.
Choose for: Rank my supplied documents by lexical BM25 Retrieve matching passages from caller-supplied text with exact token offsets
Outside this profile: Web search, semantic/vector search, linguistic segmentation, relevance certification or verifying whether a passage supports a claim Network calls, external actions, supplied-evidence authenticity or factual correctness certification
Render bounded literal/argument/select/plural message trees for explicit cases using fixed CLDR47 integer cardinal rules for en/de/ja/ar/pt-BR/pt-PT. Exact branches precede categories; other is required. Returns selected branches and finite case coverage. No translation, escaping, ICU syntax or runtime Intl dependency.
Choose for: Render supplied plural/select messages into concrete localized preview strings Preview fixed integer cardinal branches and finite case coverage for supplied message trees
Outside this profile: Translate or verify language quality; parse full ICU MessageFormat; decimal/ordinal plurals; format dates/numbers; send messages; escape HTML or SQL Network calls, external actions, supplied-evidence authenticity or factual correctness certification
Convert a strict plain SRT/WebVTT caption file with caller-supplied rational time scaling and offset, returning canonical caption text and an exact cue timing map.
Choose for: Supplied literal caption text with explicit hour timestamps and a known affine millisecond transform
Outside this profile: Transcription, translation, speaker identification, media access, inferred sync anchors or guaranteed audiovisual alignment Rich captions, markup, entities, whitespace-only cue lines, cue settings, STYLE, REGION, NOTE or arbitrary WebVTT identifiers Fetching, uploading, sending, writing files, accessing media or changing external state
Write a real fixed-width UTF-8 import artifact from caller-declared byte columns and string records, with explicit ASCII padding, strict overflow rejection and byte offset manifest.
Choose for: Explicit UTF-8 byte column layout, ASCII pad characters and exact-key string records
Outside this profile: Terminal display width, glyph alignment, packed decimals, spreadsheet formatting or named bank/government formats Truncation, numeric coercion, parsing fixed-width input or CSV comparison Blind transcoding after layout; changing encoding can invalidate UTF-8 byte widths Fetching, uploading, sending, writing files, accessing media or changing external state
Create source-preserving UTF-8-byte-bounded text windows with literal double-LF, LF and ASCII-space cut priority, scalar-safe overlap, exact UTF-16/UTF-8 spans and an aggregate duplication budget.
Choose for: Creating scalar-safe UTF-8 payload windows from supplied Unicode text Favoring literal double-LF, LF and ASCII-space boundaries with exact core reconstruction Bounding aggregate source-byte duplication independently of per-chunk overlap
Outside this profile: Model tokenizer limits, token costs, embedding or semantic retrieval Sentence, grapheme, word, CommonMark or Unicode line segmentation Source authenticity, linguistic understanding, normalization or privacy redaction Fetching sources, writing files, model calls, target execution or changing external state