umbot
    Preparing search index...

    Performance: what to expect and when to worry

    This page is machine-translated from the Russian original. If something reads oddly, the Russian version is the source of truth — open an issue.

    umbot is designed with high performance and predictable response times in mind. This is especially important given the strict time limits of voice platforms (the framework logs a warning when processing a request takes longer than 2000 ms and an error when it takes longer than 2900 ms).

    This document describes:

    • how much time the framework itself spends processing a request,
    • in which scenarios delays can occur,
    • how its performance compares with solutions tailored to a single platform.

    All data was obtained under controlled conditions and is reproducible — you can run the tests yourself.

    In typical scenarios (up to 1 000 commands) umbot's internal processing takes less than 30 ms on a cold start. This includes:

    • parsing the incoming request,
    • detecting the platform
    • reading/writing user data,
    • finding the matching command,
    • building the response.

    In the vast majority of scenarios umbot stays within tens of milliseconds. However, there are two cases when the processing time can grow — and for each of them the framework offers a ready-made solution.

    1. The first upload of media files
      When a voice skill or chatbot sends an image or audio for the first time, the framework uploads the file to the platform's server and saves the received token in the database.
      The delay depends on the file size, the connection speed to the platform API and the queue on the platform side.
      The solution is the built-in Preload class: it lets you upload all media files in advance, at application startup, and cache their tokens. Then sending during a real dialog is instant — the framework simply uses the ready token.

    2. A cold start with a large number of complex regular expressions
      If your application uses 10 000+ regular expressions, the framework has to compile all of them on the first request.
      Solutions:

      • Install the re2 library — it speeds up regular expression processing 2–15 times and reduces memory usage 3–7 times.
      • Use RegExp caching (the framework does this automatically).
      • "Warm up" the cache at startup by sending a few test requests.
      • Enable the strict mode (strict_prod) to reject potentially dangerous ReDoS expressions as early as command registration.

    All the situations described are solved with the framework's standard tools. In a typical project (up to 1000 commands, moderate use of regexes) you will not run into delays — they occur only in extreme scenarios and have simple optimization paths.

    Important: the execution time of your own logic (your code in addCommand or action) is not included in these numbers. If your handler runs for 2.5 s, that is your responsibility.


    umbot is a multi-platform framework, whereas Telegraf, vk-io, alice-kit, etc. are specialized SDKs.

    A direct performance comparison is not correct for several reasons:

    1. Different architectural goals — specialized libraries focus on deep integration with one platform (for example, polls in Telegram), whereas umbot provides a single API for all platforms
    2. Different measurement points — umbot benchmarks measure pure command routing time, while other libraries often include network interaction with platform APIs
    3. Different priorities — umbot is optimized for scenarios where supporting several platforms with one codebase matters

    The key takeaway: in real projects the bottleneck is almost always either the platform API (RPS limits) or your business logic (database calls, external services) — but not command routing.

    Even under peak load, the umbot core (16 000+ RPS in benchmarks) processes requests faster than:

    • platforms deliver them to the bot (API limits)
    • typical application logic runs
    • network requests to databases complete

    Performance depends heavily on the server configuration and the load. Here are reference points for different conditions:

    Conditions RPS
    A loaded production server (2 cores / 4 GB RAM, background load: 3 websites + 10 skills + MariaDB)¹ 16 000+
    A clean server (1 core / 1 GB RAM, benchmark only)¹ 27 000–35 000
    An isolated benchmark (AMD Ryzen 5 5600G, 32 GB RAM): the full cycle (input → normalization → logic → response) 68 000
    An isolated benchmark, sequential scenario (maximum single-thread speed) ~80 000
    The comparison stress bench: 1000 commands, 200 in-flight requests, an exact/partial/regex/fallback mix 580 000–760 000

    ¹ Measured on versions before 3.1.4. On the same machine, npm run stress on 3.1.4 vs 3.1.3: full cycle 60 000 → 68 000 RPS, sequential scenario ~74 500 → ~80 000 RPS, burst tests 42 000 → 45 000 RPS; the comparison stress bench (npm run stress:compare*) — 69 000–78 000 → 580 000–760 000 RPS (details in BENCHMARKS).

    Burst tests (thousands of concurrent requests) on a loaded server complete without errors in less than 1 s.

    ⚠️ Important: RPS depends heavily on the environment. The numbers above are reference points for comparing scenarios, not guarantees. A less loaded system achieves a higher RPS.

    📊 Context: popular platforms have different request rate limits — from tens to several thousand per second, depending on the plan and the request type. umbot's numbers are at or above these limits, which means the framework will not become the bottleneck of your application.

    Memory usage (incremental, the difference before/after initialization):

    • Base initialization: < 1 MB
    • 1000 commands: < 1 MB
    • 10 000 regexes: +8 MB

    💡 The absolute footprint of the Node.js process (~45–60 MB) is not counted — it is the cost of the runtime, not of the framework.


    Long-running testing (48 hours) revealed no memory leaks or performance degradation: the average throughput in the sequential scenario stayed at ~78 000 RPS, and memory usage was stable. The 3.1.4 comparison stress bench (30 s under a stream) shows no leaks either: the heap does not grow from the beginning to the end of the test, the retained delta after GC is 0.

    Key test metrics:

    • Requests processed: 13.53 billion in 172 800 seconds
    • Throughput: 78 314 RPS (average, no dips)
    • Memory (heap): start 11 MB → average ~90 MB → end 89 MB
    • Memory (RSS): start 2179 MB → average ~350 MB → end 391 MB (Δ: -16 MB)

    What this means:

    • No leaks. Memory usage does not grow linearly but fluctuates within a corridor. The negative RSS delta (-16 MB) shows that the garbage collector releases resources correctly. The application will not "bloat" and will not crash after a week of running due to a lack of memory.
    • Performance does not degrade. The results of the 3-second (84 024 RPS), 15-second (79 566 RPS) and 48-hour (78 314 RPS) tests differ by no more than 1.6% — within environment noise, with no one-directional drift. There is no "warm-up" effect, no dips due to memory fragmentation or accumulated state.
    • Predictability. The framework behaves the same over short and long runs. This lets you plan resources precisely: if the core delivered 78k RPS in the test, you will get comparable core performance in production (adjusted for the network and the database).

    💡 Note: the test was run in an isolated environment (no network, no database). In a real project the total memory usage will depend on your business logic, but the framework core will not become a source of leaks or degradation. The source code of the benchmarks is open — you can run them yourself and check the results on your own infrastructure.

    If the DB adapter could not connect, the request is processed without the database (userData is neither loaded nor saved), and the next connection attempt is made only after a pause: 5 s after the first failure, then the pause doubles up to 60 s and is reset after a successful connection. Every failure is logged with the time of the next attempt. This way a database outage does not turn into bot unavailability: only the request that happens to trigger the attempt has to wait for the connection (for MongoAdapter this is a few seconds), not every request in a row — otherwise replies would be late beyond the platform's limit (3 s for Alice).

    Requests from one user (platform + userId) run strictly one after another, requests from different users run in parallel. So userData does not lose changes when a user presses a button twice or the platform sends several updates in a row. If the user's previous request takes longer than 10 seconds, the next one starts without waiting for it: a stuck handler does not block the user forever. Alice, SmartApp and Marusia need a reply within ~3 seconds, so there a request waits for the previous one no longer than half of the remaining time. The guarantee holds within one process.

    A redelivery of an already accepted delivery (Telegram, VK, MAX, Viber) is acknowledged with 200 ok without running the logic again: the user does not get the reply twice. A redelivery that arrives while the original request is being processed waits for its outcome (up to 30 seconds) and is processed again if the original failed with 500: the update is not lost. Details and exceptions are in "Supported platforms" → "Receiving updates".

    Without a custom logger, errors and warnings are written to error.log and warn.log (the error_log folder). A file larger than 10 MB is renamed to <name>.1 (the previous copy is replaced), so logs take no more than ~40 MB. This matters because junk webhook requests from anyone also end up in the log. File writes are queued, so rotation under high load does not overwrite the .1 archive.

    All the tests are open and live in the repository:

    # Installation
    git clone https://github.com/max36895/umbot.git
    npm install
    npm run build

    # Performance test by the number of commands
    npm run bench

    # Stress test for concurrent requests
    npm run stress

    The full results (including memory, different regex types, the "first vs repeated run" comparison) are in the src/docs/BENCHMARKS.md file (also online: https://www.maxim-m.ru/docs/umbot/v-3.1/guides/BENCHMARKS).

    You can also run the tests in the lite and long modes. In the lite mode 5 commands are registered instead of 1000 (8 vs 1003 in total with the fixed start/help/fallback). The long mode runs a 48-hour test.

    # Stress test in the lite mode
    npm run stress:lite

    # Stress test in the long mode
    npm run stress:long

    ✅ It fits if:

    • You are building a single application for several platforms.
    • You expect a high load.
    • Predictability and ease of maintenance matter to you.
    • You are building a voice assistant, a reference service, a text game.

    umbot handles dialog scenarios very well, but if your task is deep integration with specific features of one platform (for example, Telegram polls, which require calling the API directly), you may need to write a small wrapper. Even then, 90% of the logic (command handling, state, buttons) stays the same for all platforms.

    By default the command registered first wins — the order of addCommand calls matters, just as on other platforms. To keep a large command set from turning into a linear scan, the lookup works like this:

    1. An exact match is checked first (a hash table, O(1)).
    2. String slots as substrings (from 16 commands) are looked up in a substring index (an Aho–Corasick automaton): the time depends on the length of the utterance, not on the number of commands.
    3. A regex (from 16 of them) runs only if the utterance contains its mandatory part: for /order_\d+/ it is "order_". Regexes without such a part (/(yes|no)/) are always checked.
    4. The indexes are built once after commands are registered (with a ~40 ms delay), not on every request. Requests that arrive earlier (serverless: the first request right after startup) use a regular scan: building the indexes for 1000 commands in a cold process costs a few milliseconds, while a scan for a single request is cheaper.

    If you need a different selection logic (fuzzy search, your own priority), plug in your own lookup algorithm:

    import { Bot } from 'umbot';

    const bot = new Bot();
    bot.setCustomCommandResolver((userCommand, commands) => {
    // Example: return a command by hash (your own rules)
    for (const [name, cmd] of commands) {
    // slots is optional and can contain RegExp — handle both cases
    const found = cmd.slots?.some((slot) =>
    typeof slot === 'string' ? userCommand.includes(slot) : slot.test(userCommand),
    );
    if (found) {
    return name;
    }
    }
    return null;
    });

    💡 Recommendations:

    Keep the lookup order if it is critical for your logic. Use caching (Map<string, string>) for frequent phrases. For fuzzy search, consider fuse.js or natural. When using regexes, do not forget about ReDoS protection.

    If your application uses regular expressions to find commands, consider switching to re2: processing time can drop significantly, and memory usage drops too. In an internal test with a large number of regular expressions (>10 000), re2 reduced the processing time 5–7 times (the overall stated speedup range is 2–15 times depending on the patterns) and reduced memory usage.

    Slow command processing can be caused by:

    1. A large number of commands. In this case, reconsider the approach and rewrite the commands themselves, merging them if needed.
    2. Suboptimal code in the command handler. The command code may contain heavy operations and take a significant part of the processing time. Make sure your code has no heavy operations, and if they are necessary, try splitting a large operation into commands, or run the long processing asynchronously, outside the user request.

    If your application has many images or sounds, upload them in advance; otherwise uploading a resource happens while the user request is being processed, which can slow down the response significantly. So prepare a list of resources to upload in advance and pass it to Preload.

    umbot is a high-performance multi-platform framework that:

    • Performs on par with specialized solutions.
    • Handles tens of thousands of requests per second on a single VPS.
    • Lets you maintain a single codebase for several platforms while keeping acceptable performance for most use cases. It does not solve every task — but for dialog and interactive scenarios it provides a stable, predictable and scalable foundation.

    If your task is a dialog with the user rather than chat management, umbot gives you a single, predictable and scalable foundation.