Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 12 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,8 +4,20 @@ All notable changes to capcut-cli are documented here. The format follows [Keep

## [Unreleased]

## [0.26.0] — 2026-09-25

### Added

- `lint --frame-grid` validates every segment start and end against the draft's fps. `lint --frame-grid --fix` snaps the two boundaries and derives the duration from them, avoiding the one-microsecond overlaps that independent start/duration rounding can create; an unchanged 1x source range follows the repaired duration.
- `caption --script` now aligns Chinese and Japanese at character granularity, including when Whisper returns several characters as one timed token. `--min-script-match <0..1>` turns the alignment report into a write gate: a mismatched script is refused before the draft changes.
- `caption --audio-stream <n>` selects a zero-based audio stream from a multi-stream input. FFmpeg extracts that stream to a temporary mono 16 kHz WAV before Whisper runs; `--ffmpeg-cmd` selects the binary and the result reports `audio_stream` while keeping `source_audio` pointed at the original container.
- Boundary-safe ripple editing: `remove --ripple` closes the removed span across every track and refuses before writing when another segment crosses it; `shift-all --from <time>` moves only segments at or after an exact boundary and likewise refuses a boundary that cuts through a segment.
- Large proxy renders automatically pass filter chains longer than 8 KiB through FFmpeg's filter-script input instead of the process command line, then remove the temporary script after the render. Dry-run plans expose the script path and content without writing it.
- `import-timeline` flattens nested OTIO `Timeline.1`, `Stack.1`, and `Track.1` sequences into editable CapCut tracks while preserving parent gaps, offsets, durations, and nested caption markers. Unsupported effects and items remain explicitly reported.
- `render --crf <0..51>` controls constant quality (default 28); `render --video-bitrate <rate>` selects a target bitrate such as `2500k` or `4M`. The two modes are mutually exclusive and the render plan reports which one it uses.
- `caption --word-reveal` writes progressive, word-timed caption prefixes (`one` → `one two` → `one two three`) without flattening them into pixels. It is mutually exclusive with karaoke.
- `restyle <project> --preset <file> [--track-name <name>]` applies one text-style preset atomically to every text segment or to one caption track. Explicit style flags override the preset and the result reports affected track/segment counts.

- `doctor` reports what each draft store holds — for every default CapCut/JianYing project directory it finds (or the one folder named with the new `--drafts <dir>`), a `draft-store` check counts the projects as readable, markerless, encrypted or unreadable. A JianYing 6.0+ store, where every project the app wrote is an encrypted payload, is now named once and up front (warn) with what still works — `init`, `quickstart` and `compile` build plaintext drafts from the bundled template — instead of being discovered one failed command at a time. Same classification as the `template.store` report of `init`/`quickstart`/`compile`. `capcut doctor --drafts <dir>` also makes the check usable on a machine without the app, and in CI.
- `examples/short-video-narration.md` (+ zh-CN) — silent clip → 9:16 draft with a TTS voiceover and script-accurate captions, as four commands (`quickstart --ratio 9:16` → `tts --text-file` → `caption --from-segment --script` → `lint`) and as one script, `examples/scripts/narrate-short.sh`. `examples/scripts/edge-tts-wav.sh` bridges edge-tts (MP3 only) to the WAV `tts` expects at `{out}`; any other engine plugs in through `--tts-cmd`. The vision-model step that writes the script is optional and stays outside the CLI: the script is a text file.
- `python/` — a thin Python client, published to PyPI as `capcut` (`pip install capcut`). `capcut.run(cmd, *args, **flags)` spawns the CLI once without a shell and returns the JSON it prints; keyword arguments become flags (`font_size=16` → `--font-size 16`), positionals pass through as single argv tokens, a non-zero exit raises `CommandError` with `status` and the CLI's JSON. `capcut.serve(jobs)` feeds the stateless JSONL queue and returns one result per job. `capcut.describe()`, `capcut.doctor()`, `capcut.version()`. Pure Python, no dependencies, Python ≥ 3.9; the binary is found on PATH or through `CAPCUT_CLI`. Not part of the npm tarball.
Expand Down
2 changes: 2 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -105,6 +105,8 @@ The host reads a draft and passes its JSON as tool input. The component itself h

## Release notes

> **New in v0.26.0:** exact frame-grid lint/fix; character-level Chinese/Japanese script alignment with an optional match gate; explicit caption audio-stream selection; safe ripple delete and boundary shifts; scalable FFmpeg filter scripts; nested OTIO import; CRF/bitrate proxy controls; progressive word-reveal captions; and atomic whole-track `restyle`. Full details in the [changelog](./CHANGELOG.md).

> **New in v0.25.0:** `caption` follows the transcript's script. Whisper's "words" for Chinese and Japanese are single characters or short tokens, so the Latin defaults (four words per cue, joined with spaces) produced fragments with spaces between the characters; cues are now joined without spaces and bounded by characters alone, at the width `lint` holds captions to (zh 16, ja 13, ko 16), and the result reports `caption_script`. An explicit `--max-words` / `--max-chars` still wins. Full details in the [changelog](./CHANGELOG.md).

> **New in v0.24.0:** captions in Chinese, Japanese and Korean are held to their own limits — `lint` flags a 32-character Chinese line and a 15 chars/s cue that the Latin defaults (42, 20) let through, and `--fix` re-wraps between characters (zh 16/9, ja 13/4, ko 16/12; an explicit `--max-chars` / `--max-cps` still applies everywhere). On a JianYing 6.0+ drafts folder, where every app-written project is encrypted, `init` / `quickstart` / `compile` now say that none could seed the new draft (`template.store`, a WARNING) and `lint` reports `template-unverified-store` instead of nothing. Plus a one-command agent install: `npx skills add renezander030/capcut-cli`. Full details in the [changelog](./CHANGELOG.md).
Expand Down
2 changes: 2 additions & 0 deletions README.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -82,6 +82,8 @@ Claude Code 也可以把它作为插件加载:

## 发布说明

> **v0.26.0 新增:** 精确帧网格检查/修复;中日文按字对齐脚本并可设置匹配率门槛;字幕音轨选择;安全波纹删除和边界平移;大型 FFmpeg 滤镜脚本;嵌套 OTIO 导入;CRF/码率预览控制;逐词显现字幕;以及整轨原子化 `restyle`。完整说明见[更新日志](./CHANGELOG.md)。

> **v0.25.0 新增:** `caption` 按转写文本的文字来分句。Whisper 对中文、日文给出的"词"是单个字或很短的片段,按拉丁默认(每句 4 词、用空格连接)会生成字与字之间带空格的碎片;现在中日文按字直接连接、只按字数上限分句,上限就是 `lint` 对字幕的行宽(zh 16、ja 13、ko 16),结果里会报告 `caption_script`。显式传入的 `--max-words` / `--max-chars` 仍然优先。完整说明见[更新日志](./CHANGELOG.md)。

> **v0.24.0 新增:** 中文、日文、韩文字幕按各自的规范检查 —— `lint` 会指出 32 字的中文单行和每秒 15 字的字幕(拉丁默认的 42 字 / 每秒 20 字会放过它们),`--fix` 按字重新折行(zh 16/9、ja 13/4、ko 16/12;显式传入 `--max-chars` / `--max-cps` 仍对所有文字生效)。在剪映 6.0+ 的草稿目录里(应用写出的项目全部加密),`init` / `quickstart` / `compile` 现在会明确说明没有任何项目可作为种子(`template.store` 与 WARNING),`lint` 会报告 `template-unverified-store` 而不是沉默。另外,一条命令即可把它装进 Agent:`npx skills add renezander030/capcut-cli`。完整说明见[更新日志](./CHANGELOG.md)。
Expand Down
Loading
Loading