The package is built to be cheap to call: one pass over the tokens, constant-time word lookups, and lexicons stored in a form PHP can share between requests.
Benchmarks
Measured with benchmarks/run.php from the repository, with OPcache enabled. These are single runs on a laptop (Docker on WSL2), so treat them as a rough guide: repeated runs vary by 10% or more.
| Measure | English | Indonesian |
|---|---|---|
| PHP version | 8.5.11 | 8.5.11 |
| Short texts (a tweet) per second | 80,000 | 66,000 |
| Reviews (about 100 words) per second | 14,000 | 14,000 |
| Long texts (about 2,000 words) per second | 770 | 770 |
The peak memory of the whole benchmark process, with both languages loaded, was about 2.2 MB.
Run the benchmark on your own hardware to see numbers for your setup:
php -d opcache.enable_cli=1 benchmarks/run.php
Why OPcache matters
The lexicons ship as plain PHP files that return an array literal. With OPcache enabled, PHP compiles such a file once and keeps the array as an immutable structure in shared memory. Every request then reads it with no parsing and no copying.
Without OPcache the file is parsed once per process and kept in a static cache, so a long-running worker pays the cost once, and a classic request-per-process setup pays it on every request.
OPcache is on by default in most PHP installs, but not for the CLI. In production with PHP-FPM, check that it is active:
opcache.enable=1
opcache.validate_timestamps=0 ; deploy-time cache reset, saves a stat() per file
For CLI jobs such as queue workers or imports, enable it explicitly:
php -d opcache.enable_cli=1 artisan queue:work
Get the most out of it
- Reuse the analyzer.
Sentiment::analyze()already shares one default analyzer per language, built lazily. If you customize withwithWords()orwithThreshold(), build the analyzer once and reuse it. - Loop freely. Each call is independent and keeps no state between texts, so one process can analyze thousands of texts.
- Keep texts short. The work is linear in the number of tokens. For long documents, analyze per sentence for better quality at the same total cost.
- Plain ASCII text is the fastest path. Non-ASCII text (accents, emoji, other scripts) takes extra checks per token.
Memory
Lexicons load once per process and per language. Analyzing a text allocates only its token list, so memory use stays flat however many texts you analyze. Without OPcache, loading a lexicon costs about 1 MB for English and 0.5 MB for Indonesian per process, and the whole benchmark process peaks at about 3 MB. With OPcache the lexicons live in shared memory and the per-process cost is a few KB.