Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 4 additions & 2 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,7 +84,9 @@ src/lib.rs All implementation + map definitions:
dwarf_unwind_step_impl() — tail-call target (5 frames per call)
dwarf_finalize_stack() — write completed stack to maps
handle_process_exit() — PID exit detection
src/pt_regs.rs pt_regs struct definition for x86_64
(register access is arch-neutral via RawRegs +
reg_ip/reg_sp/reg_fp, gated on bpf_target_arch;
build.rs emits that cfg — x86_64 + aarch64)
```

## Data Flow — One Sample
Expand Down Expand Up @@ -306,7 +308,7 @@ sudo tests/run_e2e.sh --filter dwarf
| Entries per shard | 65,536 (`MAX_SHARD_ENTRIES`) | Very large binaries truncated |
| CFA registers | RSP, RBP only | Other registers skipped |
| DWARF expressions | Unsupported | Except PLT-stub and signal-frame patterns |
| Architecture | x86_64 only | Hardcoded register rules |
| Architecture | x86_64 and aarch64 | Per-arch register rules (`SP_REG`/`FP_REG`/`RA_REG`); aarch64 adds an RA column (`UnwindEntry::ra_offset`) since the return address is in LR, not always at CFA-8 |

## Key Dependencies

Expand Down
10 changes: 10 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,15 @@
# Changelog

## v0.3.22

### New Features

- **aarch64 support for frame-pointer unwinding and Java/HotSpot interpreter naming** — the eBPF register access was hardcoded to x86_64 (`pt_regs.rip/rbp/rsp/orig_rax`), so the kernel programs could not even compile for an aarch64 BPF target. Register access is now arch-neutral: a `RawRegs` alias plus `reg_ip/reg_sp/reg_fp/reg_syscall_nr` accessors gated on `bpf_target_arch` (emitted by a new `profile-bee-ebpf/build.rs` mirroring aya-ebpf's logic). On aarch64 the register file is read as `user_pt_regs` (FP = x29, IP = pc, SP = sp). Frame-pointer stack walking, kernel-stack symbolization, and HotSpot interpreter-frame `Method*` naming (`interpreter_frame_method_offset = -3` words = -24 bytes, identical on x86_64 and aarch64) all work. Validated end-to-end on Corretto 17 aarch64 (`-Xint`): leaf interpreter frame resolves to `Burn.fib`, and the FP fixtures resolve full `_start → main → …` chains.
- A per-arch prebuilt eBPF object (`ebpf-bin/profile-bee.<arch>.bpf.o`, e.g. `profile-bee.aarch64.bpf.o`) is now selected by `build.rs` from `CARGO_CFG_TARGET_ARCH`, falling back to `profile-bee.bpf.o` (x86_64), so stable `cargo install` works on aarch64 without a nightly eBPF rebuild.
- Removed the dead, misleading `profile-bee-ebpf/src/pt_regs.rs` (an unused x86_64 bindgen dump — the real `pt_regs` comes from `aya_ebpf::bindings`).
- **aarch64 DWARF `.eh_frame` unwinding** — `--dwarf` now works on aarch64, not just x86_64. The generator emits per-arch DWARF register rules (`SP_REG`/`FP_REG`/`RA_REG`), and `UnwindEntry` gains a return-address column (`ra_offset`, reusing former padding so the struct stays 12 bytes and the x86_64 prebuilt stays byte-compatible): unlike x86_64 where the RA is always at CFA-8, on aarch64 it lives in LR (x30) and may not be on the stack, so leaf frames are recovered from the sampled link register (`DwarfUnwindState.lr`, occupying formerly-reserved bytes). All new eBPF DWARF logic is `bpf_target_arch = "aarch64"`-gated, so x86_64 codegen is unchanged. Validated on aarch64 (kernel 6.12): the full e2e DWARF suite passes (no-FP callstack, O2, deep recursion, 50-level deep stacks, cross-`.so`, PIE, Rust, off-CPU) via the deep tail-call path.
- **Fix: kernel/user address boundary on aarch64** — `__START_KERNEL_MAP` was the x86_64 constant `0xffffffff80000000` on all arches, but aarch64 kernel addresses (`0xffff_0000_…`) are numerically *below* it, so kernel pointers were wrongly accepted as valid user frames. This surfaced as bogus `[unknown]` frames when unwinding kernel register state (off-CPU profiling). Now `2^48` (top of the 48-bit user VA range) on aarch64.

## v0.3.21

### Bug Fixes
Expand Down
10 changes: 5 additions & 5 deletions Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -335,7 +335,7 @@ sudo probee -o output.svg -t 5000 -- ./my-fp-binary
- Rust: Add `-g` flag when compiling
- C/C++: Compile with debug symbols (`-g` flag)

**Limitations:** Max 8 executable mappings per process, 131K unwind table entries per binary (up to 64 binaries), up to 165 frame depth (via tail-call chaining; legacy fallback: 21 frames). x86_64 only. Libraries loaded via dlopen are detected within ~1 second.
**Limitations:** Max 8 executable mappings per process, 131K unwind table entries per binary (up to 64 binaries), up to 165 frame depth (via tail-call chaining; legacy fallback: 21 frames). x86_64 and aarch64 are supported. Libraries loaded via dlopen are detected within ~1 second.

See [docs/dwarf_unwinding_design.md](docs/dwarf_unwinding_design.md) for architecture details, and [Polar Signals' article on profiling without frame pointers](https://www.polarsignals.com/blog/posts/2022/11/29/profiling-without-frame-pointers) for background.

Expand Down Expand Up @@ -454,7 +454,7 @@ let session = ProfilingSession::new(config).await?;
## Limitations

- Linux only (requires eBPF support)
- DWARF unwinding: x86_64 only, see limits above
- Architecture: x86_64 and aarch64 supported (frame-pointer + DWARF unwinding, Java/HotSpot `Method*` naming)
- JIT-compiled code (HotSpot Java, V8/Node) is symbolized via `/tmp/perf-<pid>.map`; interpreter and inlined frames are not yet reconstructed. `Compiler.perfmap` auto-dump requires JDK 17+ (on JDK 8/11 use perf-map-agent/async-profiler)
- [VDSO](https://man7.org/linux/man-pages/man7/vdso.7.html) `.eh_frame` parsed for DWARF unwinding; VDSO symbolization not yet supported

Expand Down
2 changes: 1 addition & 1 deletion profile-bee-common/Cargo.toml
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
[package]
name = "profile-bee-common"
version = "0.3.21"
version = "0.3.22"
edition = "2021"
license = "MIT"
description = "Shared types between profile-bee userspace and eBPF programs"
Expand Down
49 changes: 36 additions & 13 deletions profile-bee-common/src/lib.rs
Original file line number Diff line number Diff line change
Expand Up @@ -225,11 +225,21 @@ pub const REG_RULE_UNDEFINED: u8 = 2;
pub const REG_RULE_REGISTER: u8 = 3;
pub const REG_RULE_EXPRESSION: u8 = 4;

/// Compact unwind table entry for eBPF-side stack unwinding.
/// Compact unwind table entry for eBPF-side stack unwinding (12 bytes).
///
/// On x86_64, the return address is always at CFA-8, so we don't store RA
/// rule/offset. Using u32 for PC (file-relative addresses fit in 4GB) and
/// i16 offsets gives us 12 bytes per entry.
/// `rbp_offset`/`rbp_type` describe the frame-pointer register (RBP on x86_64,
/// x29 on aarch64). `ra_offset` describes where the return address lives:
///
/// - On x86_64 the return address is always at CFA-8, so the eBPF x86_64 path
/// ignores `ra_offset` and hardcodes CFA-8. Generation still records -8 for
/// documentation.
/// - On aarch64 the return address is in LR (x30) — not always on the stack —
/// so `ra_offset` is load-bearing: a byte offset from CFA where RA is saved,
/// or the sentinel [`RA_OFFSET_IN_LR`] meaning "RA is still in the LR
/// register" (leaf frames; recoverable only for the sampled frame).
///
/// `u32` PC (file-relative addresses fit in 4GB) and `i16` offsets keep this at
/// 12 bytes per entry.
#[derive(Copy, Clone, Debug, Eq, PartialEq)]
#[repr(C)]
pub struct UnwindEntry {
Expand All @@ -238,13 +248,26 @@ pub struct UnwindEntry {
pub rbp_offset: i16,
pub cfa_type: u8,
pub rbp_type: u8,
pub _pad: [u8; 2],
/// Return-address recovery (see struct docs). Occupies the two bytes that
/// were previously padding, so both the size (12 bytes) and the offsets of
/// every other field are unchanged — an existing x86_64 prebuilt (which
/// ignores this field) stays byte-compatible.
pub ra_offset: i16,
}

impl UnwindEntry {
pub const STRUCT_SIZE: usize = size_of::<UnwindEntry>();
}

/// `ra_offset` sentinel: the return address is still in the LR register (x30),
/// not saved to the stack (leaf frames / function prologue on aarch64). Only
/// recoverable for the sampled (leaf) frame, where the LR register value is
/// known; deeper frames with this rule terminate unwinding.
pub const RA_OFFSET_IN_LR: i16 = i16::MAX;
/// `ra_offset` sentinel: the return address is undefined (top of stack /
/// outermost frame) — terminate unwinding.
pub const RA_OFFSET_UNDEFINED: i16 = i16::MIN;

/// Maximum frames to unwind per tail-call iteration
pub const FRAMES_PER_TAIL_CALL: usize = 5;
/// Maximum tail-call depth (kernel limit is 33)
Expand Down Expand Up @@ -411,14 +434,14 @@ pub struct DwarfUnwindState {
/// Saved CPU ID (for finalization)
pub cpu: u32,
pub _pad2: u32,
/// Reserved. Originally intended to thread the normalized execution context
/// through the tail-call chain, but the finalizers read `preempt_count`
/// directly instead (threading it from `collect_trace` forked verifier state
/// through the FP-walk loop and blew the instruction limit). Kept for layout
/// stability; tail calls preserve interrupt context so the direct read is
/// equivalent.
pub context: u32,
pub _pad3: u32,
/// Link register (x30) captured from the sampled register file. Used by the
/// aarch64 DWARF unwinder to recover the leaf frame's return address when it
/// is still in LR (not yet spilled to the stack). Zero / unused on x86_64.
///
/// Occupies the 8 bytes formerly split between a reserved `context` field
/// (the finalizers read `preempt_count` directly, so it was unused) and
/// `_pad3`, so the struct size is unchanged.
pub lr: u64,
/// V8 SharedFunctionInfo tagged pointers extracted during FP+V8 tail-call
/// walking. Parallel to `pointers[0..MAX_V8_FRAMES]`. Zero means "not a
/// V8 frame" or "beyond V8 extraction limit".
Expand Down
2 changes: 1 addition & 1 deletion profile-bee-ebpf/Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

23 changes: 23 additions & 0 deletions profile-bee-ebpf/build.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,23 @@
use std::env;

/// Emit the `bpf_target_arch` cfg for this crate.
///
/// aya-ebpf sets this cfg in its own build script, but `cargo:rustc-cfg` only
/// applies to the crate that emits it — it does not propagate to dependents.
/// We therefore mirror aya-ebpf's logic here so our `#[cfg(bpf_target_arch =
/// ...)]` gates select the same architecture aya used to generate `pt_regs`:
/// prefer `CARGO_CFG_BPF_TARGET_ARCH` (set for cross-builds), else derive it
/// from the host triple (the BPF `TARGET` is `bpf`, which is useless here).
fn main() {
println!("cargo:rerun-if-env-changed=CARGO_CFG_BPF_TARGET_ARCH");
let arch = env::var("CARGO_CFG_BPF_TARGET_ARCH").unwrap_or_else(|_| {
let host = env::var("HOST").expect("HOST not set");
host.split_once('-')
.map(|(arch, _)| arch.to_owned())
.unwrap_or(host)
});
println!("cargo:rustc-cfg=bpf_target_arch=\"{arch}\"");
println!(
"cargo::rustc-check-cfg=cfg(bpf_target_arch, values(\"x86_64\",\"arm\",\"aarch64\",\"riscv64\",\"powerpc64\",\"s390x\"))"
);
}
Loading
Loading