From urlextract¶
urlextract pulls the URLs out of a run of plain text. It works from IANA’s
top-level-domain list. A regular expression locates every TLD in the text, then each hit grows left and right to the
stop characters that bound a URL and the result has to pass a host check. urlextract downloads the list on first use and
caches it on disk, and update() or update_when_older(days) refresh that cache. find_urls returns the matched
strings, only_unique de-duplicates them, get_indices pairs each with its offsets, and has_urls answers the
yes/no question. Its scope stops at locating URLs in text: it has no opinion about HTML.
turbohtml covers the same ground with LinkDetector, whose find hands back a
LinkSpan per match with the offsets, the matched text and a normalized url. turbohtml
compiles the TLD table into the extension instead of fetching one at runtime, and runs the scan in C, so there is no
cache directory, no first-call network round trip and no per-process warm-up. linkify() goes one
step further and rewrites HTML into <a> tags, which urlextract does not attempt.
turbohtml vs urlextract¶
Dimension |
turbohtml |
urlextract |
|---|---|---|
Scope |
Locate links in text and rewrite HTML into |
Locate URLs in text only |
Result |
A |
The matched string, with its offsets when you ask for them |
TLD list |
IANA’s, compiled into the extension and versioned with the release; |
IANA’s, downloaded on first use and cached on disk; |
Network and state |
None; the same input gives the same answer on every machine |
Downloads and caches the TLD list; the answer follows whatever the cache holds |
Phone numbers |
|
None |
Performance |
C scan (see below) |
Pure-Python regex plus per-match validation |
Typing |
Fully annotated, |
Typed public surface |
Dependencies |
Single package, C extension bundled |
|
Maintenance |
Active, part of the turbohtml project |
Active |
Feature overlap¶
The detection surface ports one-to-one:
URLExtract().find_urls(text)->LinkDetector().find(text), returning a list of spans rather than a list of strings.find_urls(text, only_unique=True)->find(text, unique=True).find_urls(text, get_indices=True)-> nothing to ask for: everyLinkSpancarriesstartandend.find_urls(text, with_schema_only=True)->LinkDetector(bare_domains=False), which detects only URLs whose scheme is written.has_urls(text)->has_link().extract_email-> theemailsargument, on by default.gen_urls(text)->find()returns a list; iterate it. The C scan is a single linear pass with no per-match allocation, so there is nothing for a generator to defer.
What turbohtml adds¶
linkify()rewrites HTML, leaving URLs that already sit inside an<a>, inside<script>/<style>, or inside caller-named skip tags untouched. urlextract reports strings; the anchor construction, the entity escaping and the “is this already a link” walk would be yours to write.A normalized
urlon every span:mailto:for an address,http://for a bare domain, the text itself for a URL that carries its own scheme.LinkSpan.textkeeps the original substring.Phone numbers as
tel:links throughPhoneNumbers, with the parsedPhoneNumberon the span.A
Linkifyconfiguration object with callbacks that can adjust or veto each link.No runtime download and no cache: the TLD table is part of the build, so a sandboxed or offline process behaves the same as a connected one, and two machines running the same release agree.
What urlextract has that turbohtml does not¶
check_dns=True, which resolves each candidate host before reporting it. turbohtml never touches the network. Filter the spans yourself if you need this.Bare IP addresses. urlextract reports
192.168.1.1/abc; turbohtml links an IP only when its scheme is written, as inhttp://192.168.1.1/abc. The same holds for a single-label host:http://localhost:8000/links,localhost:8000does not, because without a scheme there is nothing to tell the name apart from a word.ignore_listandpermit_list. Filter the returned spans onspan.url, which is the set membership test urlextract runs over its own results.A configurable result cap.
_limitandURLExtractErrorguard against a pathological input; the C scan is a single linear pass, so there is no blow-up to cap.Tunable boundaries:
set_stop_chars_left/set_stop_chars_right,set_after_tld_charsandadd_enclosure. turbohtml’s rules are fixed – brackets, parentheses and braces balance, and trailing sentence punctuation is trimmed – and are not configurable.update()andupdate_when_older(). The table moves when you upgrade turbohtml. Passtlds=for a suffix IANA does not list.A command-line entry point. urlextract ships a
urlextractcommand; turbohtml’s CLI has no link-extraction subcommand.
Performance¶
operation |
turbohtml |
urlextract |
|---|---|---|
find comment (1 link, 1 email) |
784 ns |
231 µs (295x ±2%) |
find prose (1 KiB) |
10.4 µs |
3.69 ms (356x ±5%) |
has_link comment |
155 ns |
224 µs (1446x ±8%) |
has_link prose (1 KiB) |
156 ns |
4.46 ms (28591x ±96%) |
has_link early (220 KiB tail) |
135 ns |
785 ms (5795802x ±41%) |
find numeric prose, phones off (1 KiB) |
401 ns |
3.63 ms (9058x ±3%) |
find ucs2 prose, phones off (100 KiB) |
109 µs |
362 ms (3334x ±42%) |
find ucs4 prose, phones off (100 KiB) |
78.1 µs |
273 ms (3502x ±32%) |
urlextract only locates URLs in plain text, so the comparison is on detection alone. The gap on the presence test is the
widest: has_link() stops at the first match, while has_urls drives the generator
that scans for every top-level domain in the input, so a link near the start of a long document costs turbohtml the
bytes up to it and urlextract the whole document.
How to migrate¶
Swap the import and build a reusable LinkDetector instead of a URLExtract:
# urlextract
from urlextract import URLExtract
urls = URLExtract().find_urls("see https://example.com")
# -> a list of str, ["https://example.com"]
turbohtml |
|
|---|---|
|
|
|
|
|
|
|
|
|
|
|
the |
|
nothing: the table is compiled in; |
(rewrite HTML yourself) |
from turbohtml.clean import LinkDetector
detector = LinkDetector()
print([span.text for span in detector.find("see https://example.com")])
print([span.url for span in detector.find("a.com, b.com, a.com", unique=True)])
['https://example.com']
['http://a.com', 'http://b.com']
To rewrite HTML rather than list matches, reach for linkify(), which has no urlextract
counterpart:
from turbohtml.clean import linkify
print(linkify("visit example.com for more"))
visit <a href="http://example.com" rel="nofollow">example.com</a> for more
Gotchas and pitfalls¶
find_urlsreturns strings;find()returns spans. Takespan.textfor the substring as written, orspan.urlfor the href-ready form – they differ for a bare domain and for an address.De-duplication keys differ.
only_uniquecompares the matched strings, soexample.comandhttp://example.comare two results;unique=Truecompares the normalizedurl, so they are one.urlextract reports an address as a URL only with
extract_email=True; turbohtml detects addresses by default. Passemails=Falseto turn them off.Build the detector once and reuse it. urlextract compiles its regex per instance and turbohtml compiles its configuration per instance, so constructing one per call wastes the same work in both.
linkify()parses its input as HTML, so<and&are markup, not literal characters. For plain-text-only work useLinkDetector.