Skip to content
1 change: 1 addition & 0 deletions .github/copilot-instructions.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
@../AGENTS.md
128 changes: 128 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,128 @@
# AGENTS.md

This file provides guidance to AI Agents when working with code in this repository.

Comment thread
mmichel11 marked this conversation as resolved.
Comment thread
mmichel11 marked this conversation as resolved.
## Project Overview

oneDPL (oneAPI Data Parallel Library) is a header-only C++17 library implementing the [oneAPI specification](https://github.com/uxlfoundation/oneAPI-spec/tree/main/source/elements/oneDPL) for parallel algorithms. It provides C++ standard library-like parallel algorithms that work across heterogeneous devices (CPUs, GPUs) using different parallel backends (TBB, DPCPP/SYCL, OpenMP, serial).

**Key Resources:**
- [IMPLEMENTATION_DETAILS.md](IMPLEMENTATION_DETAILS.md) - architecture, roadmap and development patterns
- [cmake/README.md](cmake/README.md) - Complete build system documentation
- [CONTRIBUTING.md](CONTRIBUTING.md) - Testing requirements and contribution guide
- [Official Documentation](https://uxlfoundation.github.io/oneDPL)

## Quick Start: Building and Testing

### Configure Build
```bash
cmake -DCMAKE_CXX_COMPILER=icpx \
-DCMAKE_CXX_STANDARD=17 \
-DONEDPL_BACKEND=dpcpp \
-DCMAKE_BUILD_TYPE=Release \
-B build
```

**Common CMake Variables:**
- `ONEDPL_BACKEND`: tbb (default), dpcpp, dpcpp_only, omp, serial
- `CMAKE_BUILD_TYPE`: Debug, Release, RelWithDebInfo, RelWithAsserts

**Note:** `RelWithAsserts` is Release without `-DNDEBUG`, useful for testing assert-heavy code.
**Device selection:** For SYCL/DPC++ device selection, use `ONEAPI_DEVICE_SELECTOR` as documented in the [Device Selection](#device-selection) section below.

### Build and Run Tests

```bash
# Build specific test
cmake --build build --target sort.pass

# Build all tests
cmake --build build --target build-onedpl-tests

# Build by category
cmake --build build --target build-onedpl-algorithm-tests

# Run specific test
cd build && ctest -R ^sort.pass$

# Run tests by category
ctest -L ^algorithm$

# Run test executable directly
./build/test/sort.pass
```
Comment thread
danhoeflinger marked this conversation as resolved.

**See `cmake/README.md` for complete testing options.**

Comment thread
danhoeflinger marked this conversation as resolved.
### Build Documentation

```bash
# Setup documentation
cd documentation
python3 -m venv docdpl
source docdpl/bin/activate
pip install -r _auxiliary/requirements.txt

# Generate HTML in build/html
make html
```

## Architecture Quick Reference

oneDPL uses a three-tier architecture (see `IMPLEMENTATION_DETAILS.md` for details):
Comment thread
danhoeflinger marked this conversation as resolved.

1. **Public API Layer** (`include/oneapi/dpl/`) - Standard-like algorithm headers
2. **Pattern Layer** (`include/oneapi/dpl/pstl/`, `include/oneapi/dpl/internal/`) - `__pattern_*()` implementations and extensions
3. **Backend Layer** (`include/oneapi/dpl/pstl/`) - Parallel execution primitives

**Key Pattern:** Algorithm → Pattern Function → Backend Primitive

Example: `std::any_of()` → `__pattern_any_of()` → `__parallel_or()`

## Test Organization

- `test/parallel_api/` - Parallel algorithm tests (algorithm, numeric, memory, ranges)
- `test/xpu_api/` - Device-side API tests
- `test/general/` - General functionality
- `test/kt/` - Kernel templates (experimental)

Tests use `.pass.cpp` suffix.

# Code styling
- Code should be self-documenting.
- Only use comments when necessary to explain something that is unclear.
- Do not create short 1-2 line functions only used once unless it improves understandability.

## Code Formatting

**clang-format is required** for all code except tests:
```bash
clang-format -i <file>
```
Contributors can override clang-format suggestions for readability in exceptional cases.

When touching existing code, migrate any `::std::` usages to `std::` within the functions touched.

Comment thread
danhoeflinger marked this conversation as resolved.
## Device Selection

SYCL/DPC++ device selection via environment variables:

```bash
# Intel LLVM Compiler >= 2023.1
# Uncomment exactly one of the following selectors:
export ONEAPI_DEVICE_SELECTOR=level_zero:gpu
# export ONEAPI_DEVICE_SELECTOR=opencl:gpu
# export ONEAPI_DEVICE_SELECTOR=*:cpu
```

## Code Review Guidelines

- **Formatting-only changes are frowned upon.** PRs should not include changes that are purely cosmetic (whitespace, brace style, etc.) without substantive functional changes. If a file is being modified for functional reasons, incidental formatting fixes in the same area are acceptable, but reformatting unrelated to the PR's purpose should be flagged in reviews and requested to be removed.

Comment thread
mmichel11 marked this conversation as resolved.
## Important Notes

- **Header-only library** - no binary artifacts
- **C++17 minimum** required; parallel ranges API requires **C++20**
- **Backend selection is compile-time only** - zero runtime overhead
- **Two-phase headers** - `*_defs.h` (declarations), `*_impl.h` (implementations)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we want a requirement specifying that all AI-authored commits should be disclosed in the commit message? I think if we list it here, the agent should respect it.

It brings up another question about if we should adjust the contribution guide to give user guidance for AI generated code (e.g. all commit should be human reviewed prior to opening a PR, but I think that can be addressed separately).

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Its a good question. I think ultimately the committing user has to be responsible for the code whether or not it was partially or fully generated by an agent. In this case, I'm not so sure that adding something to the commit message is important, but its worth discussing and perhaps making a policy.


1 change: 1 addition & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
@AGENTS.md
6 changes: 6 additions & 0 deletions GEMINI.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,6 @@
Use your tool to write directly to files rather than using cat or other workarounds.

Unless asked to fix or implement something, don't jump to editing code or fixing a problem,
first explain your plan and check with the user before going forward.

@AGENTS.md
153 changes: 153 additions & 0 deletions IMPLEMENTATION_DETAILS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,153 @@
# oneDPL Implementation Details

This document describes the high-level architecture and development patterns in oneDPL.

oneDPL is a header-only library, and has a C++17 minimum requirement.

## Three-Tier Design Pattern

oneDPL uses a layered architecture to provide a single API across multiple backends:

### 1. Glue/Public API Layer
**Location:** `include/oneapi/dpl/`

Public headers matching C++ standard library naming:
- `algorithm`, `execution`, `numeric`, `memory`, `ranges`, `iterator`

Implementation in `glue_*.h` files:
- Thin wrappers accepting execution policies
- Dispatch to pattern layer using `__select_backend()`
- Files split: `*_defs.h` (declarations), `*_impl.h` (implementations)

### 2. Pattern Implementation Layer
**Location:** `include/oneapi/dpl/pstl/` (e.g., `algorithm_impl.h`, `numeric_impl.h`)

Core algorithm logic in `__pattern_*()` functions:
- Polymorphic overloads based on:
- Iterator type (forward, random-access)
- Execution policy (serial, parallel, hetero)
- Vectorization capability (`_IsVector` template parameter)
- Enable single algorithm to serve multiple execution contexts

Note: `include/oneapi/dpl/internal/` contains some extension implementations (scan-by-segment, binary search, dynamic selection, etc.).

Example flow: `std::any_of()` → `__internal::__pattern_any_of()` → backend primitives

### 3. Backend Implementation Layer
**Location:** `include/oneapi/dpl/pstl/`

Backend-specific parallelism primitives:

**Host Backends:**
- `parallel_backend_tbb.h` - Intel Threading Building Blocks
- `parallel_backend_omp.h` - OpenMP pragmas
- `parallel_backend_serial.h` - Sequential fallback

**Heterogeneous Backend:**
- `hetero/dpcpp/parallel_backend_sycl*.h` - SYCL/DPC++ for GPUs/accelerators

Each backend implements common interface:
- `__parallel_for()`, `__parallel_reduce()`, `__parallel_scan()`, etc.
- Compile-time selection via namespace aliasing (`__par_backend`)

## Execution Policies

Policies control algorithm execution strategy:

| Policy | Execution | Use Case |
|--------|-----------|----------|
| `sequenced_policy` (seq) | Sequential | Single-threaded CPU |
Comment thread
danhoeflinger marked this conversation as resolved.
| `parallel_policy` (par) | Parallel threads | Multi-core CPU with TBB/OpenMP |
| `parallel_unsequenced_policy` (par_unseq) | Parallel + SIMD | Multi-core CPU with vectorization |
| `unsequenced_policy` (unseq) | SIMD only | Single-threaded with vectorization |
| Device policies | GPU/accelerator | SYCL/DPC++ execution |

Policies map to execution tags:
- `__serial_tag` - Sequential execution
- `__parallel_tag<_IsVector>` - Parallel execution with optional vectorization
- `__hetero_tag<_BackendTag>` - Device execution

## Backend Architecture

### Backend Selection

Backends selected at **compile-time only** via CMake `ONEDPL_BACKEND`:

```cmake
-DONEDPL_BACKEND=tbb # TBB (default for most platforms)
-DONEDPL_BACKEND=dpcpp # SYCL/DPC++ with TBB for host
-DONEDPL_BACKEND=dpcpp_only # SYCL/DPC++ without TBB
-DONEDPL_BACKEND=omp # OpenMP
-DONEDPL_BACKEND=serial # Sequential
```

Selection logic in `CMakeLists.txt`:
- When `ONEDPL_BACKEND` is not set, auto-detects SYCL support: defaults to `dpcpp` if SYCL is available, otherwise `tbb`
- When `ONEDPL_BACKEND` is explicitly set, auto-detection is skipped; the specified backend's dependencies must be available
- Namespace `__par_backend` points to active backend

### SYCL/DPC++ Backend

**Specialized Kernel Implementations:**

| File | Purpose |
|------|---------|
| `parallel_backend_sycl_for.h` | Kernel launching primitives |
Comment thread
danhoeflinger marked this conversation as resolved.
| `parallel_backend_sycl_reduce.h` | Reduction operations |
| `parallel_backend_sycl_*scan*.h` | Scan/prefix-sum operations |
| `parallel_backend_sycl_merge*.h` | Merge algorithms |
| `parallel_backend_sycl_radix_sort*.h` | Optimized radix sort |
| `parallel_backend_sycl_reduce_by_segment.h` | Segmented operations |

## Development Patterns

### Pattern-Based Implementation

Standard pattern for implementing algorithms:

1. **Public API** (`glue_algorithm_impl.h`)
- Accepts execution policy + algorithm parameters
- Dispatches to pattern function

2. **Pattern Function** (`algorithm_impl.h`)
- `__pattern_*()` with overloads for different contexts
- Selects appropriate brick/backend based on:
- Iterator category
- Execution policy
- Vectorization support

3. **Backend Primitive**
- Actual parallel execution via backend
- Example: `__par_backend::__parallel_or()`

### Brick Functions

`__brick_*()` functions provide elementary implementations:
- Used within work units by patterns
- Enable code reuse between serial/parallel variants
- Example: `__brick_any_of()` used by both serial and parallel `__pattern_any_of()`


## Testing Architecture

Tests organized by execution model:

**`test/parallel_api/`** - Tests using host execution policies
- `algorithm/` - Parallel algorithm tests
- `numeric/` - Reduction, scan operations
- `memory/` - Memory algorithms
- `ranges/` - Range-based API

**`test/xpu_api/`** - Tests for device-side API
- C++ standard APIs callable from SYCL kernels
- Validates device-side iterator support

**`test/general/`** - General functionality
- SYCL iterator tests
- Policy behavior validation

**`test/kt/`** - Kernel templates (experimental)
- Hardware-specific optimizations
- Requires specific device capabilities

Test naming: `*.pass.cpp` suffix for tests expected to pass.
Loading