Text with the parts matching a search query highlighted, built so that neither the text nor the query can break it — and both of them routinely arrive from a URL query string. Reach for it wherever somebody has just typed something and needs to see where it landed: search results and their snippets, a filtered list or a table under an active filter, in-page and in-document find, a docs or knowledge-base search page, log and diff viewers, autocomplete and combobox options, a tag or user picker, admin record lookup, and the "showing 12 results for …" line above any of them. Common asks it answers: "highlight search term react", "highlight matching text react", "react-highlight-words alternative", "highlight-words-core", "react highlighter component", "shadcn highlight text", "highlight search results react", "mark tag react", "wrap matches in mark react", "highlight substring in string react", "highlight multiple words react", "case insensitive highlight react", "highlight text without dangerouslySetInnerHTML", "highlight search term xss", "escape regex special characters search", "new RegExp from user input invalid regular expression", "search query with special characters breaks", "highlight accent insensitive search", "remove diacritics search javascript", "normalize NFD strip combining marks", "highlight overlapping matches", "find in page highlight react", "scroll to current match react", "検索ワード ハイライト react", "該当箇所を光らせる", "検索結果 マーカー 表示", "全角 正規表現 エスケープ 検索". The reason to install one rather than write it is that the three-line version is a security bug. `text.replace(query, '<mark>' + query + '</mark>')` into dangerouslySetInnerHTML is what everybody writes first, and it injects HTML from two directions at once — the body text and the search term, which on a results page is usually `?q=` straight off the address bar. This never builds a string of HTML at all. It cuts the text into runs and React renders them, so a `<script>` in either input is characters on screen and nothing else. The query is never compiled to a regular expression either, which is the other half of the same problem and the one that shows up in the bug tracker rather than the security report. `C++`, `a.b`, `$100`, `(draft)` and `[WIP]` are all ordinary things to type into a search box and all of them are also regex syntax: `new RegExp(q)` throws `Invalid regular expression` on some and, worse, quietly matches the wrong text on others — search `a.b` and watch `axb` light up. The usual patch is an escaping helper copied from a gist. Here there is nothing to escape, because matching is `indexOf` over a folded copy of the string. That folded copy is where the real work is. Matching has to ignore case and accents — `café` should be found by `cafe` — but the highlight has to be drawn on the original text, and the two strings do not have the same length. Stripping a combining mark shortens it; `İ` lowercases to two characters and lengthens it; `Σ` and its word-final form `ς` are the same letter written two ways, so `ΕΛΛΑΣ` is invisible to anyone typing `ελλάς`. An offset found in one string and used in the other is wrong by a few characters, and the damage lands next to the highlight rather than in it — a duplicated or eaten letter somewhere along the line, where nobody is looking. So no offset is ever converted: the text is split into grapheme clusters with `Intl.Segmenter`, each is folded on its own, and every match is carried back through a recorded boundary. Concatenating the output always reproduces the input exactly, whatever the query. Working in graphemes is also what stops it tearing a character in half. A flag is two regional indicators, a thumbs-up with a skin tone is a base plus a modifier, and a `<mark>` drawn around the first half of either one splits it on screen — the flag falls apart into two letters, the modifier is orphaned beside the thumb. A match that covers only part of a character is passed over and the search carries on from the next position. Multiple terms are handled the way a reader expects rather than the way the loop falls out. A string query is split on whitespace, because that is what the search backend did with it and a phrase that never occurs verbatim would otherwise highlight nothing at all; pass an array to match exact phrases instead. Terms that overlap — `ab` and `bc` both land on `abc` — are merged into one highlight, since a `<mark>` cannot be nested inside another and emitting the shared letter twice corrupts the text. Matches that merely touch are deliberately left as two, so a find bar still counts what the reader can see. It is a real `<mark>` element rather than a styled span, which matters on somebody else's machine: in Windows high contrast mode the browser replaces author colours with system ones and knows to give a `<mark>` the system's own highlight pair, while a span painted to look identical is handed the ordinary page colours and every highlight on the page silently disappears for the readers who turned high contrast on in order to see things. The user-agent `color: black` that ships with `<mark>` is overridden so the text keeps the colour it already had instead of turning black in dark mode, and the highlight carries no horizontal padding, which would otherwise re-space the line as the reader types. The text itself is never cut down or rewritten, so screen readers, find-in-page, selection and copy all see exactly the string that was passed in. For a find bar with next and previous buttons, `activeIndex` marks one match as the current one: the marks carry `data-match-index`, the active one carries `data-active="true"` and `aria-current`, and bringing it into view is `container.querySelector('[data-active="true"]')?.scrollIntoView()`. `splitHighlight(text, query, options)` is exported on its own — a pure function from a string and a query to a list of runs — for counting matches, for highlighting into a canvas or a PDF, or for testing your own rendering. Official shadcn/ui has nothing for this, and the measurement is not close: fetching all sixty-three registry entries today (sixty-two are fetchable, questionnaire alone 404s, and nine of the newer ones are served only on the new-york-v4 style track) and concatenating the component sources gives 211,787 bytes, in which `<mark`, `Highlight`, `escapeRegExp`, `Intl.Segmenter`, `normalize("NFD")`, `searchWords`, `caseSensitive` and `matchDiacritics` are every one of them zero hits. The two occurrences of the word are `data-highlighted`, the Radix menu-item state, which is a different thing entirely. Within pulld it is the display half of search: search-input is where the query is typed, command-palette highlights the fuzzy subsequence it matched inside its own option list, and this is the one you point at arbitrary body text. It composes with read-more and middle-truncate on the same line of a result, and sits naturally in a table cell, a tree-view label or a diff-view row. No dependencies, and no hooks — so it renders inside a React server component with no "use client" of its own and ships no client JavaScript, which is the common case, because search results have usually just been fetched on the server.
pnpm dlx shadcn@latest add "https://pulld.pages.dev/r/highlight-text.json"