A local Instruments run may show that the first screen appears in “about one second,” while the merged build performs noticeably worse in remote testing. The tool is rarely the problem. More often, the measurement has no clear boundary: network waits, JSON parsing, database writes, and UI commits are bundled into one total duration. Even a small change in machine state then makes the results impossible to compare. A more reliable approach is to mark business-level intervals with os_signpost, then use xctrace on a cloud Mac to record the same fixed scenario repeatedly.
Define a Repeatable Critical Path First
Do not begin by asking, “How fast is the entire app?” Start with a user action that has stable start and end points. For example, after opening an order page, measure from the beginning of a local cache read until the first batch of list cells has finished binding its data. Split that path into three intervals:
| Interval | Start | End | Suitable for a gate |
|---|---|---|---|
load_cache |
Begin reading the cache | Return structured data | Yes |
decode_payload |
Begin decoding | Produce domain objects | Yes |
render_first_batch |
Submit the first batch of models | Finish laying out the first batch of views | Yes |
wait_remote_api |
Send an external request | Receive the response | Diagnostic only |
External network time depends on routing and server conditions, so it should not directly block a merge. Performance gates should prioritize work that runs within the process, uses fixed inputs, and can be triggered repeatedly.
One interval should answer one question. If the same marker includes networking, disk access, and rendering, exceeding the limit still will not tell you what needs to be fixed.
Add Stable, Correlatable Signposts
Use a fixed subsystem and category, and do not append user data to interval names. When measuring concurrent work, create a separate OSSignpostID for each business operation so that adjacent tasks cannot be paired with the wrong start or end event.
import os
enum PerfTrace {
static let log = OSLog(
subsystem: "com.example.mobile",
category: "CriticalPath"
)
static func begin(_ name: StaticString) -> OSSignpostID {
let id = OSSignpostID(log: log)
os_signpost(.begin, log: log, name: name, signpostID: id)
return id
}
static func end(_ name: StaticString, id: OSSignpostID) {
os_signpost(.end, log: log, name: name, signpostID: id)
}
}
Call sites must use defer to close the interval. Otherwise, an error path can leave an interval with a start event but no matching end event:
let traceID = PerfTrace.begin("decode_payload")
defer { PerfTrace.end("decode_payload", id: traceID) }
let models = try decoder.decode([Item].self, from: fixtureData)
Do not include tokens, file contents, or user identifiers in signpost text. Performance evidence becomes part of the build artifacts, so fields should remain low-cardinality and safe to archive.
Keep xctrace Recording Conditions Fixed
First verify in the graphical interface that the template name is correct and the intervals are visible. Then move to the command line. The test app should use a fixed data fixture, a fixed simulator model, and the same build configuration. Terminate any old process before recording so that state left by a previous run does not affect the result.
set -euo pipefail
OUT="$PWD/perf-artifacts"
mkdir -p "$OUT"
xcrun xctrace record \
--template "Points of Interest" \
--time-limit 45s \
--output "$OUT/critical-path.trace" \
--launch -- "$APP_PATH" \
-PerfScenario first-batch \
-PerfFixture "$PWD/Fixtures/items.json"
xcrun xctrace export \
--input "$OUT/critical-path.trace" \
--toc > "$OUT/toc.xml"
The XPath and table structure produced by xctrace export may change between tool versions. The parsing script should therefore verify that the target schema exists before reading intervals. Do not treat an awk command that depends on column positions as a stable long-term interface. Prefer parsing the exported XML, and retain one minimal trace as a contract test when upgrading versions.
When running in a LemonVM remote environment, first record xcodebuild -version, the operating system version, the build configuration, and the test fixture hash. If the node or toolchain changes, establish a new baseline instead of directly comparing figures from the old and new environments.
Build Thresholds from Multiple Samples
A single run is easily affected by first-load work, background indexing, and thermal state. Run one warm-up first, followed by at least five valid recordings. Compare medians for the gate, while also setting a more lenient upper limit for the slowest sample.
If the baseline median is 420 ms, the warning threshold could be set to 1.15 times the baseline, or 483 ms. Do not apply that ratio to every interval, however. A small function that takes tens of milliseconds and a data-preparation task that takes several seconds do not have the same variability model.
The evaluation script should distinguish at least three outcomes:
- The trace cannot be generated: this is an infrastructure failure, not a performance regression.
- The target interval is missing: the scenario or instrumentation failed, and the test needs to be fixed.
- Valid samples exceed the threshold: the performance gate fails, and the evidence must be archived.
The baseline file should be reviewed alongside the code. It should include the interval name, sample count, median, permitted upper limit, toolchain version, and fixture hash. Any change that raises a threshold should explain why in the commit description, preventing limits from only ever moving upward.
Eliminate Common False Regressions
If the first run is slow but later runs are stable, caching or dynamic-linking preparation is usually responsible. Separate cold-start and warm-path measurements into two groups. If every run is slow, then inspect the code changes. If only one run is abnormal, first check whether indexing, backups, an additional simulator, or a parallel build was active at the time.
Also verify the following:
- Do not mix Release and Debug data;
- Maintain separate baselines for different simulator OS versions;
- Restore the same business state before every test run;
- Do not run performance jobs concurrently with large build jobs;
- Archive the trace, parsed results, and standard error output together;
- Never write
0 msby default when an interval cannot be parsed; - After a timeout, terminate the app under test and any leftover recording processes.
If the test depends on a graphical session, add a visibility preflight before the job begins. If the test only covers parsing or the storage layer, use a separate test target whenever possible to reduce noise from UI state.
Make Failures Immediately Actionable
A gate should report more than “performance failed.” At minimum, print the interval name, current median, baseline, percentage change, number of valid samples, and trace path. For example:
FAIL render_first_batch
median_ms=512
baseline_ms=420
change=21.9%
samples=7
trace=perf-artifacts/run-06.trace
After seeing the result, a reviewer should be able to download the corresponding trace, locate the exact interval, and determine whether the issue is a regression in business logic or an abnormal measurement environment. A stable performance gate does not require every result to be identical. Its purpose is to make changes under the same inputs, toolchain, and runtime conditions explainable and reproducible. Start with one high-value path, observe its variability over time, and then add intervals gradually. That is more effective than introducing dozens of unreliable metrics at once.
Frequently asked questions
Should an os_signpost performance gate use the mean or median?
Use the median of several valid runs and pair it with a high percentile or maximum allowance. A mean is easily distorted by one system-level outlier, while a single run is not a defensible gate.
Should a failed xctrace recording count as a performance regression?
No. Recording failure, missing intervals, and threshold violations need separate exit codes. Report a performance regression only after the job has produced a valid measurement set.
Should every os_signpost interval run in continuous integration?
No. Gate only critical paths with stable boundaries and reproducible inputs. Keep intervals that depend heavily on manual interaction or live external services as diagnostic signals.
Choose a LemonVM Cloud Mac plan for your project
Lemon M4 and Lemon M4 Pro are available in Singapore, Tokyo, Seoul, Hong Kong, and the US West, with daily, weekly, monthly, and quarterly billing.