Skip to content
Usef Farahmand

Handling 1 Million Leaderboard Records in Unity with the Job System and Burst

How a Unity leaderboard loads, sorts, and searches 1,000,000 CSV records without blocking the main thread, with the design trade-offs and profiler results.

  • Unity
  • C#
  • Job System
  • Burst
  • Performance
  • UI Virtualization
Handling 1 Million Leaderboard Records in Unity with the Job System and Burst

A leaderboard system that reads 1,000,000 records from a CSV file, sorts them, and keeps them searchable, without blocking the main thread for even a single frame. That is the whole point of this project: to show how the Unity Job System and Burst can handle a large amount of data without frame-rate drops or unnecessary memory pressure.

Unity Profiler and the Game View of the leaderboard during a search

Project structure

The scripts are organized into these folders:

  • Application: Manager.cs, which coordinates the whole flow (Load → Sort → Search → UI)
  • Core: LoadService, SortService, SearchService
  • Data: FileReader, LeaderboardEntry, FindLineOffsetsJob, ParseLineJob
  • Search: IdSearchJob, UsernameSearchJob
  • Sorting: SortJob
  • UI: LeaderboardUIManager, ItemScrollView, ItemContainer, ItemSlot, SearchInput
  • Test: debugging and performance measurement tools, including LeaderboardProfilerReport

I kept the data and processing logic (Core, Data, Search, Sorting) completely separate from the UI. The two layers only meet through Manager, so each one can be tested on its own.

1. Loading and parsing the data

  1. FileReader.ReadAsync reads the file asynchronously with File.ReadAllBytesAsync, then copies the bytes into a NativeArray<byte>.
  2. FindLineOffsetsJob (IJob, Burst) scans the whole byte buffer once and finds the start and length of every line (both \n and \r\n are supported).
  3. ParseLineJob (IJobParallelFor, Burst, batch size 64) parses the lines in parallel, directly on the bytes and without string.Split, so there are no GC allocations and no string-conversion overhead. The result goes into a NativeArray<LeaderboardEntry> allocated with Allocator.Persistent.

Parsing is relatively heavy, so I wait for both jobs with the same pattern I use for Sort and Search (explained below). A helper called WaitForJobAsync checks JobHandle.IsCompleted over several frames with Awaitable.NextFrameAsync instead of calling Complete() directly. That way the main thread never waits for a job to finish.

Since this wait can last longer than one frame, the intermediate buffers (lineStartOffsets and lineLengths) are created with Allocator.Persistent, not Allocator.TempJob (which is only valid for a few frames). WaitForJobAsync also completes the job inside a finally block, so even if the wait is interrupted midway (for example when leaving Play Mode), the following Dispose runs without errors.

2. Sorting

SortJob (IJob, Burst) uses the built-in NativeArray<T>.Sort(IComparer<T>) (Introsort in Unity.Collections) and sorts by Score in descending order. It runs on a worker thread, and its completion is checked by polling IsCompleted in SortService.Update().

3. Search and filtering

  • User input is controlled by a debounce (SearchInput, 0.15 seconds), so a keystroke does not create a new job every time.
  • A numeric query runs IdSearchJob (prefix match on the ID). Anything else runs UsernameSearchJob (case-insensitive prefix match on FixedString64Bytes).
  • Both are IJobParallelFor jobs that scan all 1 million records in parallel. Results are collected with NativeList<int>.ParallelWriter, with no resizing, because the capacity is reserved up front for the total number of records.
  • SearchService polls IsCompleted just like SortService. If the user sends a new query while a search is running, only the latest query is kept and it runs after the current search finishes.

4. Display and scrolling (UI)

  • Object pooling: ItemContainer creates only as many objects as fit in the viewport (plus a buffer), not one per record.
  • Virtualization: ItemScrollView calculates the index of the first visible item from the scroll position and only repositions and repopulates the existing pooled items.
  • Filtered results go through the same path (SetResults), so the behavior is identical for all records and for search results.

Design FAQ

Why Awaitable instead of Task or UniTask?

Awaitable is native to the Unity engine (2023.1+), integrates directly with the PlayerLoop, and needs no external package. Unlike Task, it does not depend on the thread pool or the usual .NET SynchronizationContext; it produces fewer allocations and is lighter for per-frame and per-operation scenarios in Unity. Automatic cancellation when the object is destroyed is also built in. Its only limitation is that it is available only on Unity 2023.1+, and its ecosystem is not yet as mature as UniTask's.

Why parse with IJobParallelFor and a batch size of 64?

Parsing each line is independent of every other line (embarrassingly parallel), so parallelization is the natural choice. A batch size of 64 balances per-batch scheduling overhead against worker-thread utilization:

Batch sizeScheduling overheadLoad balance across threadsResult
32 or lessHigh: more batches mean more dispatch overheadGoodRejected: the scheduling overhead eats the gain from parallelism
64LowGoodChosen
128 or moreVery lowPoor: fewer batches than available cores, so some threads sit idleRejected: threads end up unevenly loaded

Why are line boundaries found with a single-threaded job (IJob) and not a parallel one?

Finding line boundaries is a simple sequential scan over bytes. Even single-threaded and Burst-compiled, it takes a few milliseconds for tens of megabytes of data. Parallelizing it would require merging results between chunks, which is extra complexity with no noticeable gain at this scale.

Why sort with NativeArray<T>.Sort (Introsort) and not a hand-written algorithm?

The built-in Introsort in Unity.Collections gives acceptable O(n log n) performance for 1 million items and runs on a worker thread without blocking the main thread. A Radix Sort on Score (potentially O(n), since it is an integer) or a parallel Merge Sort would be faster, but sorting happens only once, at load time (not on every search), so the extra complexity was not worth the gain. In the real profile (table below), sorting 1 million records took 114.51 ms spread over 3 frames, never exceeding 59.89 ms in the worst frame, so the built-in Introsort is enough for this data size.

Why a parallel linear scan for search instead of a hash map or trie?

Username search needs a prefix match. A regular hash map only makes exact matches O(1); for prefixes you would have to build a trie, which costs more memory and complexity. And since IDs are parsed in input order (not sorted by ID), binary search would require maintaining an extra index. Building and maintaining an extra index (more memory, invalidation complexity) was rejected given the gain: with parallelism across all CPU cores, a linear scan over 1 million records took on average 12.64 ms for ID and 15.55 ms for Username in the real profile (401 sample queries, table below), regardless of whether a query had zero matches or 111,112. That is good enough for this scale.

Why debounce the search input, and why 0.15 seconds?

Without debounce, every keystroke creates a parallel job over 1 million records, which wastes resources and creates a race between the results of consecutive searches. 0.15 seconds is short enough for the UI to feel immediate, but long enough to stop a job from being created for every typed character.

Why are Sort and Search polled in Update() instead of calling Complete() directly?

Calling JobHandle.Complete() right after Schedule() is equivalent to blocking the main thread until the job finishes. By checking IsCompleted every frame inside Update(), the main thread never waits and the frame rate does not drop; the result is consumed only once the job has actually finished.

Why doesn't the 1-million-item list freeze the UI?

Because of the combination of object pooling (only visible items are created) and virtualization (the position of each pooled item is recalculated from the scroll offset, instead of re-rendering the whole list). Rendering cost is independent of the total number of records and depends only on the number of items inside the viewport.

Limitations and trade-offs

  • FixedString64Bytes has limited capacity for usernames (about 61 bytes of UTF-8); longer usernames are truncated or produce an error.
  • Reading the file currently creates two copies of the data in memory (a managed byte[] plus the NativeArray<byte>). For much larger files, reading directly into a native buffer would remove this cost.
  • Each search performs a new full scan over all the data (no index). For scales far beyond 1 million records, a secondary index may become necessary.
  • The very large Content height of the ScrollRect (proportional to 1 million items) has not been precisely tested for float precision at the end of the list; it is worth checking with a fast scroll to the very end.
  • The classes inside Test/ are manual debugging and measurement tools and are not part of the main product flow.

Profiler and performance results

The LeaderboardProfilerReport script (inside Test/) ran the Read → FindLines → Parse → Sort → Search stages once in the Editor and measured each stage's wall time, the number of frames it spanned, and the worst frame time. It printed a complete Markdown table to the Console and saved it to leaderboard_profiler_report.md, the file attached to the repository, which includes the details of each of the 401 sample queries (ID and Username, each with real, invalid, and partial data). The table below summarizes that file.

Scrolling was not measured with this tool because it needs real touch or drag simulation. The scroll figure comes from the screenshot and the demo video below (Unity Profiler, PlayerLoop, inside the Editor, during real scrolling of the list).

Unity Profiler and the Game View of the leaderboard during a search

The live Profiler + Search demo video shows the Profiler's frame time while typing in the search field. No noticeable CPU spike appears at the moment of the search, because the parallel job runs on worker threads, not on the main thread.

Live Profiler + Search demo
StageWall time (ms)Frames elapsedWorst single frame (ms)Notes
File read (I/O)44.97144.3837,137,815 bytes
Line-offset scan112.331156.961,000,001 lines
Parse (1M records)72.66196.151,000,000 records
Sort (1M records)114.51359.89Descending by Score
Search by ID (401 queries)avg 12.64 / max 112.041 per queryavg 12.95 / max 119.180 to 111,112 results per query
Search by Username (401 queries)avg 15.55 / max 44.811 per queryavg 15.87 / max 45.260 to 14,270 results per query
Scrolling, steady staten/an/a10.62Worst frame during a fast scroll to the end of the list (the PlayerLoop row in the Profiler screenshot above)

These figures were taken inside the Editor and do not include the EditorLoop cost, because it is reported separately in the hierarchy and does not exist at all in a real build. On a development build these numbers should be lower, not higher.

Search comparison: ID vs Username

MetricID searchUsername search
Average wall time12.64 ms15.55 ms
Maximum wall time112.04 ms44.81 ms
Average worst frame12.95 ms15.87 ms
Maximum worst frame119.18 ms45.26 ms
Match count range0 to 111,1120 to 14,270

Username search is on average about 3 ms slower than ID search, because FixedString64Bytes needs a byte-by-byte case-insensitive comparison, whereas ID search is a simple numeric comparison. The higher maximum for ID search (112.04 ms vs 44.81 ms) comes from a single outlier query rather than a stable pattern: the second and third highest ID search values fall in the same 20 to 32 ms range as Username search. The queries with higher overhead (above 20 ms) mostly occurred in the second half of the run. That comes from more than 800 accumulated Debug.Log lines in the diagnostic script itself, not from the search job; in the real UI search, without the extra logging, this increase does not appear.

The peak native memory for the main array was 83.92 MB (1,000,000 records × 88 bytes).

The managed memory reported in this run was a drop of −517,682 KB. A negative number indicates a garbage-collection pass during the run, not a memory leak. This figure is unrelated to the main pipeline: Profiler.GetTotalAllocatedMemoryLong() returns the cumulative allocation of the whole session rather than current usage, and a large part of the fluctuation comes from the diagnostic script itself (800+ formatted log lines). For an accurate figure of the pipeline's real usage, take a separate Memory Profiler snapshot from a normal run, without this test tool.

Overall conclusion from the profile: the cost of each search is practically independent of the number of matches (whether 0 records or 111,112, the run time stays in the same few-millisecond range), which is exactly what a parallel scan over the whole array should do. Sort, at 114.51 ms spread over 3 frames, never lets a single frame exceed 60 ms, so it produces no noticeable hitch. Scrolling took only 10.62 ms at its worst, which means object pooling and virtualization work as expected. None of the heavy operations (load, sort, search, scroll) noticeably lowers the frame rate.

How to run and test

  1. Set the CSV file path in the path field on Manager (or on TestParse, SortJobTests, or LeaderboardProfilerReport for a separate test).
  2. Press Play. The order of execution is: Read → Find Lines → Parse → Sort → initial display → search preparation.
  3. To test search, type a number (ID) or part of a username into the UI.

Source Code

Links