EN

Android Performance

Focus on Android Performance

loading
SmartPerfetto v1.12.0 Release Notes

The previous update ended with v1.7.0 on August 21, 2026. From v1.7.0 on August 21 to v1.12.0 on September 17, SmartPerfetto shipped nine new versions and merged 141 non-merge commits across 28 calendar days.

The period added an arbitrary dual-Trace workspace, source analysis without a required prebuilt index, explicit Auto/Fast/Full modes, durable investigation checkpoints, and claim verification that reaches the Web UI, CLI, reports, and snapshots. Across these changes, a run now keeps the request, tool execution, evidence, claims, and delivery status available for inspection.

Project: github.com/Gracker/SmartPerfetto.

Half a Year in the Making: Open-Sourcing an Android Internals Ebook

There is no shortage of material on Android internals. AOSP, official docs, blog posts, papers, trace discussions — it is all findable. The annoying part is that every piece speaks for a different version: this post concludes something about Android 12, that one does not say which version it tested, a third describes what happened on one particular device.

That gap is what AIW is for. The full name is Android Internal Wiki; the Chinese book title is 《Android 技术内幕:系统机制、性能优化与工具实战》 — five parts, twenty-six chapters, with the content pinned to Android 17. The repository is about to be open-sourced:

https://github.com/Gracker/android-internals-wiki

What the repository gives you is twenty-six chapters of content, a full table of contents, and an EPUB rebuilt every week.

Those twenty-six chapters are now far enough along to put in front of people. If you read a passage that does not match the source, an issue or a pull request is welcome.

SmartPerfetto v1.7.0 Release Notes

The previous update ended with v1.1.1 on July 17, 2026. By then, SmartPerfetto already had a dual-trace workspace, Quick Mode, private analysis context, a unified Agent runtime architecture, and dedicated Camera, Heap, and GPU analysis capabilities.

Over the next five weeks, from July 18 through August 21, the repository advanced to v1.7.0. New features continued to land, but much of the work shifted toward knowledge provenance, distribution quality, access isolation, result completeness, and reversible improvement.

SmartPerfetto is moving from “complete one analysis” toward “complete analyses reliably and traceably across more environments.”

This post uses the publicly released v1.7.0 source as its endpoint and reviews the major capabilities added after v1.1.1.

Project and previous post:

SmartPerfetto v1.0.28 Release Notes

SmartPerfetto 2026-05-17 to 2026-06-04 update cover

The previous SmartPerfetto update was published on May 17. At that point the project had already moved from “an AI Assistant inside Perfetto UI” to “a reusable trace analysis platform.” By June 4, the new work is concentrated in five areas: Smart Mode, selected-range quick analysis, CLI capture, stronger evidence rules for Power / ANR / Input / IO / Network, and four Agent runtimes.

This article is based on the SmartPerfetto main branch on June 4, 2026. The latest public release at this point is v1.0.28. The goal is simple: list what changed after May 17, explain how runtimes and evidence sources are handled, and show what information makes a useful bug report.

Project links:

SmartPerfetto v1.0.7 Release Notes

SmartPerfetto Update Cover

When I wrote the SmartPerfetto open-source introduction on April 29, the headline was still “put an AI Assistant that can run SQL, invoke Skills, and generate reports inside the Perfetto UI.” Two weeks later the repository has moved a long way. The feature surface expanded from single-trace Q&A to reusable analysis results, multi-trace comparison, dual Claude/OpenAI runtimes, SQL guardrails, evidence-source indexing, no-install packages, rendering-pipeline teaching, and a more complete Provider diagnostic flow.

This article is based on the SmartPerfetto repository state as of May 17, 2026, and adds a new feature description on top of the previous post. Readers should walk away knowing three things: what was added in the past two weeks, where the current full feature boundary sits, and what information a good bug report should include.

Project links:

Android Perfetto Series 18: Response Latency in Practice, from Input Events to First-Frame Feedback

This article addresses a topic easily overshadowed by startup speed and smoothness: first-frame interaction response. It asks how long the screen takes to provide the first frame of feedback after a user action. Complete page loading and stability later in an animation require other metrics.

After tapping a button, how long until its pressed state appears? After tapping a WeChat conversation, how long until the first frame of the conversation page starts moving? After swiping a list, how long until its content follows the finger? These questions relate to slow startup and dropped frames, but need their own measurement definitions.

We use Perfetto to break down this interval: from InputReader read_time, through the App receiving the event, to an associated frame’s present. This gives a candidate internal delay, estimated_input_to_present_ms. Only after confirming the target layer and that the business state visibly changed on that frame should a report call it input_to_present_ms. Neither value covers touch IC latency, delays before driver reporting, display scanning, or panel response. By the end, you should be able to define a start, endpoint, capture configuration, SQL metrics, and supplementary App markers for a tap or scroll scenario.

The article examines response latency through four independent measurements—read/dispatch, handling, ACK→first frame, and scroll responsiveness. Each comes with capture configuration, the first tracks to inspect in the UI, and SQL. It then covers five common attribution paths and the respective boundaries of high-speed cameras and Perfetto.

Android Perfetto Series 17: Domain Automation, Issue Checklists, and Platform Infrastructure

For slow camera opens, audio underruns, a stalled WebView first screen, or dropped frames in Flutter and games, the evidence is scattered across the app, system_server, cameraserver/audioserver, HAL, SurfaceFlinger, and scheduling tracks. Simply inspecting a few more tracks in the UI can easily leave critical timestamps unnoticed.

Using Camera and Audio as examples, this article describes a reusable method for domain-specific analysis: capture the right data, identify the objects, calculate stages and blocking, and preserve timestamps in the report so that reviewers can return to the UI. Camera and Audio are examples; the core is a domain schema and collaboration with platform tracing. By the end, you should at least be able to break a slow camera open into reviewable stages and an audio underrun into cycle anomalies, rather than delivering only a table of the longest slices.

Android Perfetto Series 16: GPU, Power Counters, and Hardware Bottleneck Analysis

FrameTimeline marks a frame as late, yet the main thread, RenderThread, and SurfaceFlinger all seem to stay within their budgets. It is tempting to write “possibly a GPU bottleneck.” That is a risky claim: GPU frequency, GPU counters, battery current, and power rails are not tools for attributing an individual App frame to a root cause.

Parts 07 and 08 covered the rendering path from the App to SurfaceFlinger; Parts 06 and 09 covered refresh rates and CPU scheduling/frequency. This article adds an often-skipped step without repeating those topics: when a problem appears to have reached the hardware resource layer, how should GPU, devfreq, battery counters, and power rails enter the analysis of the same trace interval?

These counters are difficult because not every device provides them, and different GPU vendors use different counter names. A conclusion cannot simply say “a high GPU value means a GPU bottleneck.” FrameTimeline, RenderThread, SurfaceFlinger, the GPU timeline, frequency, and power data must be examined over the same interval.

Android Perfetto Series 15: Boot Traces, Long-Running Field Traces, and Capturing Intermittent Problems

Many performance problems cannot be reproduced with a single tap. Slow boot, intermittent jank, problems after turning the screen on or off, audio interruptions after several hours, periodic stutters in an automotive system, and performance degradation as temperature rises are all poor fits for a manually recorded 10-second trace.

This article covers long-running field traces: defining the observation window in advance, using triggers to preserve the incident, ensuring the file finishes writing safely, and determining whether the collected data is trustworthy. Boot tracing is covered below as a special observation window. Part 02 already covered basic interactive capture methods—the command line, official scripts, Developer options, and the web interface. Here, we focus on what changes in the field and during long-running recording.

Android Perfetto Series 14: heapprofd and Memory Profiling

The hardest part of a memory problem is that a rising number does not immediately tell you what is growing. dumpsys meminfo provides RSS/PSS, and a Java heap dump reveals object retention relationships. But for JNI/C++, CPU-side Skia/Bitmap allocations, and native wrappers in media players, you often also need to know which call stack issued the malloc/new request.

This article covers heapprofd. Its value is bringing native allocations, frees, and call stacks into a Perfetto trace, where you can inspect them alongside application events, thread scheduling, and Binder on the same timeline.

Use heapprofd after confirming the trend. I would split the growth between native heap, graphics buffers, and other mappings with meminfo or smaps before opening a flamegraph. Otherwise, it is easy to chase a malloc stack when the growth came from GraphicBuffer. This article does not address Java object references or graphics backing memory.

We will cover when to use heapprofd, how to capture a profile, how to read it in the UI and SQL, how to interpret sampling intervals and symbolization, and finally what to check before publishing a report.

Android Perfetto Series 13: Perfetto SDK, Track Event, and App Field Tracing

A system trace can tell you when a thread ran, which CPU core it ran on, and which FrameTimeline frame missed its deadline. With Binder/ftrace/sched evidence enabled, it can also reconstruct Binder calls and clues about waiting. But it does not know which frame ID a player is decoding, which scene a game engine is loading, or how long an application task has spent in a queue.

This article answers one question: how do application semantics enter a trace? The system can tell you that the main thread was slow; Track Event tells you which frame was being decoded, which queue was blocked, and which request moved between threads. Combining the two lets you narrow “the main thread was slow by 18 ms” to a candidate texture-upload interval around frame 1082, then continue validating it with RenderThread/HWUI, GPU, and FrameTimeline evidence.

The article follows the practical integration order: decide whether instrumentation is needed and at which layer, define an application-phase dictionary and backend, then build the minimal integration, category controls, combined system capture, and cross-thread flows. Finally, control write volume and define the app-side field tracing protocol.

Android Perfetto Series 12: Trace Dataflow and Data Loss Troubleshooting

One of the most troublesome situations when opening a trace is finding that the file appears to contain data, but some of it was lost along the way. The UI still shows tracks, and SQL still returns rows, yet part of a thread’s state history is missing, process names do not line up, or some Track Events have disappeared. The further you analyze it, the more your conclusions start to resemble guesses.

SQL can only analyze evidence that still exists in the trace. I would not discard an entire trace just because it reports ftrace loss; first identify whether scheduling, app markers, or frame data were affected.

Part 02 covered capture, and Part 11 covered SQL queries. This article focuses on one question: where did the trace lose data, which conclusions does that affect, and how should the next capture configuration change?

We will follow the data from the kernel to the trace file, locating loss at each stage: ftrace, producer shared memory, the central buffer, incremental state, and flush. Each stage loses data differently and needs a different remedy. At the end, there is a template for explaining how much of a trace remains trustworthy in an analysis report.

Android Perfetto Series 11: PerfettoSQL, Trace Processor, and Regression Detection

Selecting a window in Perfetto UI, inspecting the main thread, RenderThread, and CPU states, then taking screenshots works well for one investigation. The harder question is whether another trace or device supports the same conclusion. I care more about being able to recompute that conclusion next time than about making one screenshot look convincing.

Part 11 covers PerfettoSQL and Trace Processor. The aim is straightforward: turn judgments made in the UI into queries that others can check, then use Python to run them across multiple traces. You do not need to become a SQL expert first. Treat a trace as a set of timestamped tables, and translate your visual assessment into a query you can run repeatedly.

The article follows one path: understand Trace Processor, express an assessment as a query over a defined time window, run it from the command line, process multiple traces with Python, and finally turn the results into stable metrics and Trace Summary output. Each step includes SQL you can reuse.

SmartPerfetto Is Open Source: A Perfetto AI Assistant for Android Trace Analysis

SmartPerfetto is now fully open source. Open the repository and you can see the runnable mainline project as it exists today: the Perfetto UI fork, the agentv3 backend, MCP tools, YAML Skills, scene strategies, scripts, and documentation. There is no private core module kept outside the repository, and this is not just a thin demo shell.

The project comes from a very concrete daily workflow: you have a trace in hand, Perfetto has already exposed the facts, but moving from facts to judgment still means jumping through tables, writing SQL, matching threads, checking FrameTimeline, finding the Binder peer, and then returning to the timeline to confirm everything again. SmartPerfetto tries to turn those repeated actions into tools, so performance engineers can spend more time on judgment.

It is still in development. I am releasing it now because trace analysis grows from real samples: real devices, real vendor differences, real product traces, and real PRs all change how Skills and strategies should be written. Waiting until every capability is stable before publishing it one way would miss the stage where samples matter most.

If you often open Perfetto to inspect scrolling jank, startup, ANR, Binder, CPU scheduling, or rendering pipelines, SmartPerfetto provides a Perfetto UI with an AI Assistant. After loading a trace, you ask questions in natural language. The backend queries trace_processor_shell, invokes YAML Skills, organizes evidence, and streams conclusions plus data tables back into the browser.

Project links:

For normal trial use, you only need the main repository. Gracker/perfetto is the frontend fork used by the perfetto/ submodule. It mainly matters to developers who want to modify the AI Assistant plugin UI.

The previous two technical articles are better for readers who want the engineering details:

Those two articles go deep into the internal architecture. Once the source is public, readers usually care more about what is actually in the repository, whether it can run, and which parts are not stable yet. This article focuses on the open-source release itself: what is open, what works today, how the internal pieces are divided, how to run it locally, and where collaboration is most useful.