umbot performance test resultsThis page is machine-translated from the Russian original. If something reads oddly, the Russian version is the source of truth — open an issue.
All tests were run on a clean installation without network calls, external APIs or user logic.
Only the execution time of the framework's internal logic was measured: from receiving the request to building the response object.
Below are the results of measuring the execution time of only the framework's internal logic,
excluding the execution time of action. The tests were run on a system with an AMD Ryzen 5 5600G, 32 GB RAM, SSD.
Note: the initial upload time of new images or audio files (when the file's
tokenis not yet cached in the database) can significantly exceed 1 second and is listed separately. After the first upload and caching, subsequent access to the same files is fast.
Special attention was paid to optimizing regular expressions (RegExp). With isPattern: true,
umbot caches compiled RegExp objects.
This means that the first run() call with commands using new patterns **compiles the RegExp
**, which takes longer.
On subsequent calls with the same patterns, the compiled RegExp are taken from the cache, which **significantly speeds up
** execution.
The framework also groups regular expressions from different commands.
That is, when many commands with regular expressions are defined, these regular expressions are combined into groups to
reduce the number of calls.
The benchmark/comparison/ bench compares umbot with real competitor
frameworks per platform. The bench has its own
benchmark/comparison/package.json (the competitor dependencies do not get into
the repository's package.json). Each bench is a separate file: the scenarios, lockstep,
metrics and printing are shared (runner.js), and the platform file defines the input and connects
the competitors. To run:
npm i --prefix benchmark/comparison
npm run compare # Telegram: umbot vs grammy, telegraf
npm run compare:alisa # Alisa: umbot vs yandex-dialogs-sdk
npm run compare:vk # VK: umbot vs vk-io (+ @vk-io/hear)
npm run compare:viber # Viber: umbot vs viber-bot
npm run compare:max # MAX: umbot vs @maxhub/max-bot-api
The bench files are structured the same way (telegram.js, alisa.js, vk.js,
viber.js, max.js): makeUpdate — the platform update, umbot — the adapter
and botType, competitors() — the offline path of each competitor (handling via
the standard handleUpdate/handleRequest/handleWebhookUpdate, no network).
Adding a new competitor: the package goes into the dependencies of
benchmark/comparison/package.json, the implementation into competitors() of the
platform file following the existing ones.
Also in benchmark/comparison/:
stress.js (run with npm run stress:compare:vk and so on,
or node stress.js <file> [sec] [window]) — behavior under a continuous
stream: a window of 200 in-flight requests, 30 s, a traffic mix
(exact/partial/regex/fallback = 40/25/25/10), rotation of 1000
user_ids, 1000 commands. Metrics: aggregate RPS under the stream (not to be confused
with 1000/p50 of the greenhouse bench — a different mode), p50/p95/p99, the share
of time in GC, the heap floor trend (a garbage sawtooth is normal, a rising floor is
a problem), the retained delta after the final GC (a leak). Each
participant runs in an isolated child process: if one fails on the heap
limit, only its row is lost.baseline.js for a CI night job:
npm run baseline:update — record umbot's p50 on the marker scenario
"50 commands, exact match" of all five benches;
npm run baseline:check — compare against it (a +15% JIT noise tolerance; exit 1 on
a regression). It catches "someone put a heavy RegExp compilation on the hot path"
automatically rather than by a manual run. The baseline file is created on the same
machine where check runs.Fairness of the bench (the same for all platforms):
handleUpdate in grammy/telegraf/max-bot-api,
handleRequest in yandex-dialogs-sdk, handleWebhookUpdate in vk-io,
_handleEventReceived in viber-bot).
botInfo is set for the competitors explicitly (a documented option/property),
otherwise their handleUpdate goes to the network for getMe.ctx.skipAutoReply = true
(without it, its getContent would go to the platform API — the internet would be measured);
the competitors do not call ctx.reply/ctx.send. VkAdapter in the bench
runs with vk_load_user_info: false — otherwise the adapter by default goes
to the VK API for the user name on every message.hears(new RegExp(...))/an equivalent (a shared asRegExp wrapper in runner.js;
without this alignment their hears(string) matches only an exact match —
verified in the sources and by a behavioral test on each library:
grammy — txt === t, telegraf and max-bot-api — new RegExp('^…$'),
@vk-io/hear — text === condition, viber-bot has no string triggers).
The RegExp variant is their fair price for the same semantics as an umbot slot.delay is not set for its participant): the deferred build of
the search indexes does not get a chance to run in the bench's synchronous loop, and the first
requests go through a regular scan until the plan is built on the request path (after
64 lookups). This tail ends up in umbot's cold start and in the beginning of the warm-up — not in
umbot's favor.The results below are a run of umbot 3.1.4 from September 2026 (methodology v4). Environment: AMD Ryzen 5 5600G
(6 cores / 12 threads), 32 GB RAM, Windows 10, Node.js 24.21.0, without re2. Competitor versions
are pinned in benchmark/comparison/package-lock.json: grammy 1.46.0, telegraf 4.16.3,
yandex-dialogs-sdk 2.3.0, vk-io 4.10.1 (+ @vk-io/hear 1.1.1), viber-bot 1.0.18,
@maxhub/max-bot-api 0.3.0.
1000 / p50 for a single core, a theoretical ceiling
of routing, not server throughput. "1 000 000 RPS" means "routing takes 1 µs",
not "the server will handle a million requests per second".bot.hears(...), equivalents for the others). In grammY/telegraf each such handler is a separate
middleware layer, and the update goes through the layers one by one. That is why their cost grows linearly with
the number of commands. In umbot an exact match is looked up in a hash table, a substring — in a substring index, and a regex runs
only if the text contains its mandatory part: the time barely depends on the number of commands. If you write a single
shared handler with your own Map in grammY, the gap on exact matches disappears — the bench compares built-in routing, not
"the best possible code on each framework".npm run build
npm i --prefix benchmark/comparison
npm run compare # Telegram (grammy, telegraf); also: compare:alisa / compare:vk / compare:viber / compare:max
npm run stress:compare # stress under a stream; also: stress:compare:alisa / :vk / :viber / :max
node benchmark/comparison/stress.js telegram.js 10 2000 # a window of 2000 concurrent requests
Absolute numbers will differ on another machine; the ratios and the shape of the curve ("does the time grow with the number of commands") should match. If you get something different, that is a reason to open an issue with the bench output.
| Scenario (the matched command is in the middle of the list) | Implementation | p50 time, µs | p95, µs | RPS (1000/p50) | Memory, KB/req |
|---|---|---|---|---|---|
| 6 commands, exact match | umbot | 0.9 | 1.8 | 1 111 111 | 2.2 |
| grammy | 3.0 | 7.5 | 333 333 | 13.2 | |
| telegraf | 3.9 | 5.8 | 256 410 | 15.9 | |
| 6 commands, regex | umbot | 1.1 | 2.9 | 909 091 | 2.2 |
| grammy | 3.1 | 5.3 | 322 581 | 13.2 | |
| telegraf | 3.8 | 5.2 | 263 158 | 15.7 | |
| 50 commands, exact match | umbot | 1.1 | 2.2 | 909 091 | 2.2 |
| grammy | 14.5 | 25.2 | 68 966 | 73.4 | |
| telegraf | 18.8 | 30.1 | 53 191 | 97.2 | |
| 50 commands, partial match | umbot | 1.4 | 3.3 | 714 286 | 2.3 |
| grammy | 14.7 | 22.9 | 68 027 | 73.5 | |
| telegraf | 18.9 | 32.7 | 52 910 | 97.2 | |
| 50 commands, regex | umbot | 1.5 | 3.8 | 666 667 | 2.3 |
| grammy | 14.5 | 23.4 | 68 966 | 73.5 | |
| telegraf | 19.5 | 30.9 | 51 282 | 97.2 | |
| 50 commands, fallback | umbot | 1.7 | 3.3 | 588 235 | 2.2 |
| grammy | 24.1 | 38.0 | 41 494 | 115.7 | |
| telegraf | 31.0 | 48.7 | 32 258 | 27.4 | |
| 500 commands, exact match | umbot | 0.9 | 1.5 | 1 111 111 | 2.2 |
| grammy | 139.7 | 226.9 | 7 158 | 40.6 | |
| telegraf | 180.7 | 265.2 | 5 534 | 126.0 | |
| 500 commands, partial match | umbot | 1.6 | 3.1 | 625 000 | 2.3 |
| grammy | 141.6 | 204.8 | 7 062 | 39.2 | |
| telegraf | 177.2 | 257.6 | 5 643 | 126.0 | |
| 500 commands, regex | umbot | 1.5 | 2.3 | 666 667 | 2.3 |
| grammy | 160.5 | 265.9 | 6 231 | 39.2 | |
| telegraf | 195.1 | 342.5 | 5 126 | 126.0 | |
| 1000 commands, exact match | umbot | 1.0 | 2.0 | 1 000 000 | 2.2 |
| grammy | 301.0 | 567.0 | 3 322 | 73.9 | |
| telegraf | 375.8 | 691.4 | 2 661 | 118.6 | |
| 1000 commands, regex | umbot | 3.5 | 5.7 | 285 714 | 5.0 |
| grammy | 354.4 | 855.9 | 2 822 | 73.9 | |
| telegraf | 433.7 | 975.9 | 2 306 | 118.6 | |
| 1000 commands, partial match | umbot | 1.5 | 3.9 | 666 667 | 2.3 |
| grammy | 320.2 | 595.8 | 3 123 | 73.9 | |
| telegraf | 442.5 | 1061.7 | 2 260 | 118.6 | |
| 1000 commands, fallback | umbot | 1.9 | 7.6 | 526 316 | 2.2 |
| grammy | 611.2 | 1823.1 | 1 636 | 30.2 | |
| telegraf | 762.9 | 1722.0 | 1 311 | 100.7 |
How to read it (including the inconvenient parts):
node benchmark/comparison/stress.js telegram.js 10 2000,
1000 commands, 10 s): umbot — 598 272 RPS, p99 3.8 ms, heap 18 MB; grammy — 422 RPS, p99 8.0 s,
heap ~1.2 GB; telegraf — 239 RPS, p99 14.7 s, heap ~3.2 GB (close to Node's default heap limit
of ~4 GB).The same scenarios and rules (methodology v4). The key numbers are the p50 time per update; each bench prints the full output:
| Competitor | umbot's score | 6 commands, exact | 1000 commands, exact | 1000 commands, regex | Memory per update |
|---|---|---|---|---|---|
| yandex-dialogs-sdk | 13 wins of 13 | 1.6 vs 3.0 µs | 1.3 vs 148.1 µs | 4.2 vs 223.2 | umbot is lower everywhere (3.3–6.3 vs 7.5–112) |
| vk-io | 11 wins, 2 parities | 1.0 vs 1.2 µs | 0.8 vs 26.1 µs | 3.5 vs 52.5 | umbot is lower everywhere (2.1–4.8 vs 3.3–100) |
| viber-bot | 13 wins of 13 | 1.9 vs 3.7 µs | 2.1 vs 231.4 µs | 6.2 vs 269.6 | 6 commands: viber-bot is lower (1.5–1.6 vs 2.1); from 50 commands umbot is lower |
| max-bot-api | 12 wins, 1 parity | 1.4 vs 1.5 µs | 1.3 vs 207.3 µs | 5.2 vs 226.2 | umbot is lower everywhere (2.2–4.8 vs 6.3–117) |
How to read it (including the inconvenient parts):
Promise.all over
all commands on every request to find "similar" phrases (verified in the sources), hence
the growth with the number of commands. The SDK has not been updated since 2022.The stress bench (stress.js) measures not the latency of a single request but behavior under a stream:
a window of 200 in-flight requests, 30 s, a 40/25/25/10 mix (exact/partial/regex/fallback),
1000 commands (an intentionally heavy scenario), rotation of 1000 different user_ids. Each participant runs in
a separate process:
| Bench | Participant | Requests | RPS | p50 | p99 | GC, % of time | Leak |
|---|---|---|---|---|---|---|---|
| VK | umbot | 21 473 753 | 715 481 | 133 µs | 434 µs | 1.0 | no |
| vk-io | 755 903 | 25 170 | 3.9 ms | 11 ms | 2.9 | no | |
| Telegram | umbot | 20 369 920 | 678 320 | 138 µs | 444 µs | 1.0 | no |
| grammy | 9 262 | 307 | 608 ms | 1 135 ms | 63.9 | no | |
| telegraf | 6 266 | 207 | 852 ms | 1 836 ms | 66.2 | no | |
| Alisa | umbot | 17 500 009 | 583 037 | 185 µs | 554 µs | 1.3 | no |
| dialogs-sdk | 80 716 | 2 687 | 68 ms | 103 ms | 17.3 | no | |
| Viber | umbot | 22 731 200 | 757 386 | 127 µs | 403 µs | 1.0 | no |
| viber-bot | 193 107 | 6 435 | 16 ms | 37 ms | 0.3 | no | |
| MAX | umbot | 19 879 695 | 662 382 | 150 µs | 435 µs | 1.0 | no |
| max-bot-api | 19 512 | 649 | 287 ms | 426 ms | 66.4 | no |
Leak is the retained delta after the final GC: 0.0 KB per 1000 requests for all participants, the heap does not grow from the beginning to the end of the test.
A methodology fix (3.1.0). In the previous version of the bench, retained memory was measured while the worker's own latency array (8 bytes per request) was still alive. The "leak" came out proportional to the number of requests and penalized the fastest participant. Quantiles are now computed before the measurement and the array is released; after the fix, no participant leaks.
How to read it (including the inconvenient parts):
* fallback command, umbot's base controller replies to the user on a miss, that is,
makes a network call to the platform API. The bench registers * with skipAutoReply to measure
routing, not the internet. In production a miss without a handler is a real request to the platform.The time of the first 200 requests after registering commands, without warm-up (important for Cloud Functions: the first request after an instance comes up). A cold JIT makes the first request tens of times more expensive than the plateau; the participants start in the same conditions:
| Platform | Scenario | umbot: 1st / avg of 200, µs | Competitor: 1st / avg of 200, µs |
|---|---|---|---|
| Telegram | 6 commands, exact | 1 303 / 14.1 | grammy 720 / 14.8 |
| Telegram | 1000 commands, exact | 124 / 4.2 | grammy 1 413 / 374.7 |
| Telegram | 1000 commands, regex | 1 687 / 46.4 | grammy 3 490 / 423.9 |
| Alisa | 6 commands, exact | 1 675 / 17.4 | dialogs-sdk 760 / 13.0 |
| VK | 6 commands, exact | 3 087 / 35.5 | vk-io 978 / 11.3 |
| VK | 1000 commands, exact | 146 / 5.4 | vk-io 501 / 54.8 |
| Viber | 6 commands, exact | 2 786 / 35.9 | viber-bot 1 594 / 18.2 |
| MAX | 6 commands, exact | 1 444 / 18.4 | max-bot-api 611 / 15.1 |
How to read it: on small command sets umbot's first request is 1.7–3 times slower than the competitors'
(1.3–3.1 ms vs 0.6–1.6 ms) — umbot has more code that V8 compiles on the first call
(lazy function compilation; Node's compile cache, NODE_COMPILE_CACHE, does not remove this —
verified). For serverless this is 1–2 ms per cold instance start. The very first scenario of the process
(VK, Viber) additionally pays for loading the modules. As the number of commands grows, umbot is ahead in the cold
start too. Search indexes for large command sets are built after commands are registered, and the first
requests before the build go through a regular scan — the first request does not pay for building the indexes.
| Competitor | Greenhouse bench (13 scenarios) | Stress, 1000 commands (RPS) | Memory per update |
|---|---|---|---|
| grammy (Telegram) | umbot is faster in all 13 | 678 320 vs 307 (2 210×) | umbot is lower everywhere (6–53 times) |
| telegraf | umbot is faster in all 13 | 678 320 vs 207 (3 277×) | umbot is lower everywhere (7–57 times) |
| yandex-dialogs-sdk | umbot is faster in all 13 | 583 037 vs 2 687 (217×) | umbot is lower everywhere |
| vk-io | 11 wins, 2 parities (6 commands) | 715 481 vs 25 170 (28×) | umbot is lower everywhere |
| viber-bot | umbot is faster in all 13 | 757 386 vs 6 435 (118×) | at 6 commands viber-bot is lower, then umbot |
| max-bot-api | 12 wins, 1 parity (6 commands, exact) | 662 382 vs 649 (1 021×) | umbot is lower everywhere |
Where umbot falls behind or is on par:
Where umbot is consistently better:
The practical takeaway: on small bots (up to ~50 commands) speed is not what matters — a difference of fractions of a microsecond is invisible against the platform's network round trip of 50–300 ms — functionality is (multi-platform support, steps, NLU, state). On bots with hundreds of commands and a high stream, umbot has a multiple headroom in CPU and memory.
An honest assessment of the tool's boundaries (not just its strengths):
umbot is built for:
addStep), questionnaire forms (addForm);Scenarios where you will need additional logic on your side:
bot.addEvent('photo' | 'voice' | 'callback' | 'inline' | ...), 17
universal types, platform adapters declare their support via
supportedEvents (validation and warnings when connecting).
Extra code is needed only where deep platform-specific event details are required
(for example, business logic based on fields that make no sense to
unify) — read controller.eventType, controller.payload
and requestObject right in the event handler;bot.addEvent(...) +
the cross-platform API facade ctx.api?.sendPhoto/sendDocument/ sendAudio/sendVideo/answerCallback (with a per-platform support matrix
and can(method), see api-reference.md). However, full Bot API coverage
(web apps, payments, games, business mode and the rest of the long tail of methods)
stays with TelegramRequest — the calls are built manually via
the API client;/(да|нет)/) are checked one by one, and indexes for tens of thousands of commands take
memory and take noticeable time to build after registration. For such volumes the framework
recommends parameterized commands/NLU/an external API (see the bench output), or
enable re2 for regular expressions.The key idea: umbot is not a "universal framework for every case" but a tool for the "dialog bot/skill on many platforms with predictable resources" scenario. If your product is dialogs, commands and steps, you get everything out of the box and performance above the competitors. If your product is media stream processing or fine-grained Telegram specifics, weigh it: additional code on umbot versus giving up 6 other platforms and resource efficiency.
The bench and stress benchmarks check whether there is enough memory before starting.
The check takes three real limits into account: the V8 heap limit (--max-old-space-size or ~4 GB
by default — even with 16 GB of RAM Node does not allocate more), the cgroup limit (docker/k8s) and
physical memory including the page cache (the kernel drops the cache under pressure — os.freemem()
on unix understates what is available, which made the test falsely refuse to start). The consumption estimate
is based on actual measurements: ~466 B per string command, ~0.8 KB per isPattern command,
~1.6 KB including V8 heap fragmentation at extreme counts.
| Command | Description |
|---|---|
npm run bench |
Command processing time with different regex complexity (9 levels from 50 to 1 000 000 commands) |
npm run stress |
Stress test: full load (1003 commands, concurrent requests, burst tests, maximum RPS over 15 s) |
npm run stress:lite |
A lightweight stress test (8 commands instead of 1003). Suitable for a quick check on weak hardware |
npm run stress:long |
A long-running 48-hour test. Checks memory leaks and performance stability |
npm run comparison |
A fair comparison of umbot against a "pure" router implementation (lockstep, medians; your own router — via UM_COMPARISON_ROUTER) |
npm run compare |
A comparison with real frameworks: compare — Telegram (grammy, telegraf); compare:alisa — Alisa (yandex-dialogs-sdk); compare:vk — VK (vk-io); compare:viber — Viber (viber-bot); compare:max — MAX (max-bot-api). Identical input, p50/p95/IQR/RPS/memory + cold start (see above) |
npm run stress:compare:* |
A stress bench against competitors: stress:compare — Telegram; stress:compare:alisa/vk/viber/max — other platforms. 30 s of a continuous stream (a window of 200 in-flight, a 40/25/25/10 mix, 1000 commands, 1000 user_ids), aggregate RPS, GC share, heap trend, retained leak. Each participant in an isolated process |
npm run baseline:update / baseline:check |
Regression control for a CI night job: umbot's p50 on the marker scenario of all 5 benches against baseline.json (a +15% tolerance); check returns exit 1 on a regression |
Full reference — API v-3.1 · all versions.