diff --git a/CHANGELOG.md b/CHANGELOG.md index 3ccd6c8..956147d 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -4,6 +4,18 @@ All notable changes to capcut-cli are documented here. The format follows [Keep ## [Unreleased] +## [0.25.0] — 2026-09-18 + +### Added + +- `caption` follows the transcript's script. Whisper's "words" for Chinese and Japanese are single characters or short tokens, so the Latin defaults (four words per cue, joined with spaces) produced fragments with spaces between the characters. Cues in Chinese and Japanese are now joined without spaces and bounded by characters alone, at the width `lint` holds captions to (16 zh, 13 ja; Korean keeps spaces and four words at 16); karaoke ranges follow the new offsets. The result reports `caption_script` (`latin` | `zh` | `ja` | `ko`). An explicit `--max-words` / `--max-chars` applies as given. Library: `groupWords` takes a `separator`; `wordSeparator`, `groupingDefaults` and `GroupingDefaults` are exported; the script detection moved to `script.ts` (still exported from the lint entry points). +- `caption --max-chars` is documented in the command reference alongside `--max-words`. +- The Chinese README notes the feature in 科技爱好者周刊 issue 413. + +### Unchanged + +- `--script` alignment still tokenizes the script file on whitespace; a Chinese script line is one token to it. + ## [0.24.0] — 2026-09-18 ### Added diff --git a/README.md b/README.md index eb32dd4..9ea67e3 100644 --- a/README.md +++ b/README.md @@ -103,9 +103,9 @@ The host reads a draft and passes its JSON as tool input. The component itself h ## Release notes -> **New in v0.24.0:** captions in Chinese, Japanese and Korean are held to their own limits — `lint` flags a 32-character Chinese line and a 15 chars/s cue that the Latin defaults (42, 20) let through, and `--fix` re-wraps between characters (zh 16/9, ja 13/4, ko 16/12; an explicit `--max-chars` / `--max-cps` still applies everywhere). On a JianYing 6.0+ drafts folder, where every app-written project is encrypted, `init` / `quickstart` / `compile` now say that none could seed the new draft (`template.store`, a WARNING) and `lint` reports `template-unverified-store` instead of nothing. Plus a one-command agent install: `npx skills add renezander030/capcut-cli`. Full details in the [changelog](./CHANGELOG.md). +> **New in v0.25.0:** `caption` follows the transcript's script. Whisper's "words" for Chinese and Japanese are single characters or short tokens, so the Latin defaults (four words per cue, joined with spaces) produced fragments with spaces between the characters; cues are now joined without spaces and bounded by characters alone, at the width `lint` holds captions to (zh 16, ja 13, ko 16), and the result reports `caption_script`. An explicit `--max-words` / `--max-chars` still wins. Full details in the [changelog](./CHANGELOG.md). -> **New in v0.23.0:** drafts that open on the CapCut you actually have. A draft built from the bundled 6.5.0 template is refused by CapCut 8.4+, 8.7 Windows and 9.3 as "from an unusual path" ([#67](https://github.com/renezander030/capcut-cli/issues/67), [#111](https://github.com/renezander030/capcut-cli/issues/111) — the real 8.7 Windows round-trip, negative with the bundled template and positive with one captured from the installed app). `init`, `quickstart` and `compile` now seed new drafts from the newest app-authored project in your drafts folder by default (its version markers and settings, none of its content, never its `Timelines/` mirrors); `migrate --from-store` restamps drafts built earlier, and `lint` reports the stale signature as `template-stale`. Media gets its `local_material_id` link to `draft_materials` at add time — the key JianYing 5.9+ and CapCut 9.3 resolve local clips by ([JmsLdrn/capcut-mcp#1](https://github.com/JmsLdrn/capcut-mcp/issues/1)) — and `lint --fix` writes it for existing drafts (`media-unlinked`). Plus `source-range-exceeds-material`, a `compile --check` that names flat `text-style` keys ([#110](https://github.com/renezander030/capcut-cli/issues/110)), the macOS permission hint on `media-outside-draft`, and `init` stamping both timeline mirrors so `register` accepts its own drafts. No command was removed and no existing output changed shape. Full details in the [changelog](./CHANGELOG.md). +> **New in v0.24.0:** captions in Chinese, Japanese and Korean are held to their own limits — `lint` flags a 32-character Chinese line and a 15 chars/s cue that the Latin defaults (42, 20) let through, and `--fix` re-wraps between characters (zh 16/9, ja 13/4, ko 16/12; an explicit `--max-chars` / `--max-cps` still applies everywhere). On a JianYing 6.0+ drafts folder, where every app-written project is encrypted, `init` / `quickstart` / `compile` now say that none could seed the new draft (`template.store`, a WARNING) and `lint` reports `template-unverified-store` instead of nothing. Plus a one-command agent install: `npx skills add renezander030/capcut-cli`. Full details in the [changelog](./CHANGELOG.md). ## Built with capcut-cli diff --git a/README.zh-CN.md b/README.zh-CN.md index 4d0011b..ebc24c6 100644 --- a/README.zh-CN.md +++ b/README.zh-CN.md @@ -40,6 +40,8 @@ capcut info ./my-first/ -H 有用的话,[给 capcut-cli 加个 Star](https://github.com/renezander030/capcut-cli),帮助更多剪辑师和 Agent 开发者发现它。 +入选[《科技爱好者周刊》第 413 期](https://github.com/ruanyf/weekly/blob/master/docs/issue-413.md)。 + 想了解更多实用的 AI Agent 工具,从视频自动化到交付前的检查,[在 GitHub 上关注 René](https://github.com/renezander030)。 也可以从源码构建:`git clone https://github.com/renezander030/capcut-cli && cd capcut-cli && npm install && npm run build`(然后用 `npm link` 暴露出 `capcut`)。或者不安装,直接运行任意命令:`npx capcut-cli `。 @@ -78,9 +80,9 @@ Claude Code 也可以把它作为插件加载: ## 发布说明 -> **v0.24.0 新增:** 中文、日文、韩文字幕按各自的规范检查 —— `lint` 会指出 32 字的中文单行和每秒 15 字的字幕(拉丁默认的 42 字 / 每秒 20 字会放过它们),`--fix` 按字重新折行(zh 16/9、ja 13/4、ko 16/12;显式传入 `--max-chars` / `--max-cps` 仍对所有文字生效)。在剪映 6.0+ 的草稿目录里(应用写出的项目全部加密),`init` / `quickstart` / `compile` 现在会明确说明没有任何项目可作为种子(`template.store` 与 WARNING),`lint` 会报告 `template-unverified-store` 而不是沉默。另外,一条命令即可把它装进 Agent:`npx skills add renezander030/capcut-cli`。完整说明见[更新日志](./CHANGELOG.md)。 +> **v0.25.0 新增:** `caption` 按转写文本的文字来分句。Whisper 对中文、日文给出的"词"是单个字或很短的片段,按拉丁默认(每句 4 词、用空格连接)会生成字与字之间带空格的碎片;现在中日文按字直接连接、只按字数上限分句,上限就是 `lint` 对字幕的行宽(zh 16、ja 13、ko 16),结果里会报告 `caption_script`。显式传入的 `--max-words` / `--max-chars` 仍然优先。完整说明见[更新日志](./CHANGELOG.md)。 -> **v0.23.0 新增:** 生成的草稿能在你实际安装的 CapCut 里打开。用内置 6.5.0 模板生成的草稿会被 CapCut 8.4+、8.7 Windows 和 9.3 以"项目来自异常路径"拒绝([#67](https://github.com/renezander030/capcut-cli/issues/67)、[#111](https://github.com/renezander030/capcut-cli/issues/111)——这是等待已久的 8.7 Windows 真机验证:内置模板失败,从已安装应用捕获的模板成功)。`init`、`quickstart` 与 `compile` 现在默认以草稿目录中最新的应用生成项目为种子(保留其版本标记与设置,不带任何内容,绝不复制其 `Timelines/` 镜像);`migrate --from-store` 为旧版本生成的草稿重新盖上标记,`lint` 以 `template-stale` 报告过期签名。素材在添加时即写入 `draft_materials` 并回填 `local_material_id`——剪映 5.9+ 与 CapCut 9.3 正是靠这个键定位本地素材([JmsLdrn/capcut-mcp#1](https://github.com/JmsLdrn/capcut-mcp/issues/1)),已有草稿可用 `lint --fix` 补链(`media-unlinked`)。另有 `source-range-exceeds-material` 检查、能指出扁平 `text-style` 键的 `compile --check`([#110](https://github.com/renezander030/capcut-cli/issues/110))、`media-outside-draft` 的 macOS 权限提示,以及 `init` 同时盖章两份时间线镜像,使 `register` 接受自己生成的草稿。没有删除任何命令,现有输出结构均未改变。详见[更新日志](./CHANGELOG.md)。 +> **v0.24.0 新增:** 中文、日文、韩文字幕按各自的规范检查 —— `lint` 会指出 32 字的中文单行和每秒 15 字的字幕(拉丁默认的 42 字 / 每秒 20 字会放过它们),`--fix` 按字重新折行(zh 16/9、ja 13/4、ko 16/12;显式传入 `--max-chars` / `--max-cps` 仍对所有文字生效)。在剪映 6.0+ 的草稿目录里(应用写出的项目全部加密),`init` / `quickstart` / `compile` 现在会明确说明没有任何项目可作为种子(`template.store` 与 WARNING),`lint` 会报告 `template-unverified-store` 而不是沉默。另外,一条命令即可把它装进 Agent:`npx skills add renezander030/capcut-cli`。完整说明见[更新日志](./CHANGELOG.md)。 ## 常用命令 diff --git a/docs/command-reference.json b/docs/command-reference.json index 00da6aa..06b64d6 100644 --- a/docs/command-reference.json +++ b/docs/command-reference.json @@ -1,6 +1,6 @@ { "name": "capcut-cli", - "version": "0.24.0", + "version": "0.25.0", "schema_version": 2, "description": "Edit CapCut/JianYing draft_content.json directly. JSON in, JSON out.", "global_flags": [ @@ -3659,7 +3659,16 @@ ], "type": "number", "required": false, - "description": "Maximum words per cue." + "description": "Maximum words per karaoke cue. Unset: 4; for Chinese and Japanese transcripts unlimited (the character cap bounds the cue)." + }, + { + "name": "max_chars", + "flags": [ + "--max-chars" + ], + "type": "number", + "required": false, + "description": "Maximum characters per cue (karaoke) or per script line. Unset: 28 / 42; Chinese 16 / 16, Japanese 13 / 13, Korean 16 / 16." }, { "name": "max_gap_ms", diff --git a/package-lock.json b/package-lock.json index 89997b2..5e8f7c8 100644 --- a/package-lock.json +++ b/package-lock.json @@ -1,12 +1,12 @@ { "name": "capcut-cli", - "version": "0.24.0", + "version": "0.25.0", "lockfileVersion": 3, "requires": true, "packages": { "": { "name": "capcut-cli", - "version": "0.24.0", + "version": "0.25.0", "license": "MIT", "bin": { "capcut": "dist/index.js", diff --git a/package.json b/package.json index ba2bcbd..786f50a 100644 --- a/package.json +++ b/package.json @@ -1,6 +1,6 @@ { "name": "capcut-cli", - "version": "0.24.0", + "version": "0.25.0", "description": "Independent, unofficial CLI to create and edit CapCut projects — build drafts from scratch, add video/audio/text, subtitles, timing, speed, volume, templates, cut long-form to shorts. No API needed. Not affiliated with ByteDance.", "type": "module", "bin": { @@ -25,6 +25,7 @@ "dist/draft.d.ts", "dist/lint.d.ts", "dist/runner.d.ts", + "dist/script.d.ts", "dist/store.d.ts", "dist/text-offsets.d.ts", "dist/time.d.ts", diff --git a/src/caption.ts b/src/caption.ts index f63228e..fba99c3 100644 --- a/src/caption.ts +++ b/src/caption.ts @@ -14,6 +14,7 @@ import { import type { Draft, Segment, Track } from "./draft.js"; import { findSegment } from "./draft.js"; import { captionStyleFromPreset, type TextStylePreset } from "./preset.js"; +import { type CaptionScript, captionScript, groupingDefaults, wordSeparator } from "./script.js"; import { parseSrt } from "./srt.js"; import { storedTextLength } from "./text-offsets.js"; @@ -72,6 +73,9 @@ export interface CaptionResult { color_cycle?: number; /** --script alignment quality (only when a script was given). */ script?: AlignmentReport; + /** The script the transcript is written in, which chose the word separator + * and the --max-words / --max-chars defaults (see script.ts). */ + caption_script: CaptionScript; } /** @@ -95,6 +99,15 @@ export function captionDraft(draft: Draft, opts: CaptionOptions): CaptionResult const audio = resolveAudio(draft, opts); const transcription = runWhisper(audio, opts); const recognizedWords = transcription.words.length > 0 ? transcription.words : wordsFromCues(transcription.cues); + // The transcript's script decides how words join into a cue (no space inside + // Chinese or Japanese) and how many of them make one when the caller set no + // --max-words / --max-chars; the script file's wording when one was given, + // since that is the text the cues will carry. + const textScript = captionScript( + opts.scriptText !== undefined ? opts.scriptText : recognizedWords.map((word) => word.word).join(""), + ); + const separator = wordSeparator(textScript); + const grouping = groupingDefaults(textScript); let scriptReport: AlignmentReport | undefined; let cues: CaptionCue[]; if (opts.scriptText !== undefined) { @@ -106,15 +119,33 @@ export function captionDraft(draft: Draft, opts: CaptionOptions): CaptionResult const aligned = alignScript(lines, recognizedWords); scriptReport = aligned.report; cues = opts.karaoke - ? groupWords(aligned.words, opts.maxWords ?? 4, opts.maxChars ?? 28, (opts.maxGapMs ?? 500) * 1000) + ? groupWords( + aligned.words, + opts.maxWords ?? grouping.karaokeMaxWords, + opts.maxChars ?? grouping.karaokeMaxChars, + (opts.maxGapMs ?? 500) * 1000, + separator, + ) : // One cue per script line — the author's chunking — split only when a // line outgrows --max-chars (words and gaps never split a line). aligned.lines.flatMap((line) => - groupWords(line, Number.POSITIVE_INFINITY, opts.maxChars ?? 42, Number.POSITIVE_INFINITY), + groupWords( + line, + Number.POSITIVE_INFINITY, + opts.maxChars ?? grouping.lineMaxChars, + Number.POSITIVE_INFINITY, + separator, + ), ); } else { cues = opts.karaoke - ? groupWords(recognizedWords, opts.maxWords ?? 4, opts.maxChars ?? 28, (opts.maxGapMs ?? 500) * 1000) + ? groupWords( + recognizedWords, + opts.maxWords ?? grouping.karaokeMaxWords, + opts.maxChars ?? grouping.karaokeMaxChars, + (opts.maxGapMs ?? 500) * 1000, + separator, + ) : transcription.cues; } if (cues.length === 0) { @@ -166,7 +197,7 @@ export function captionDraft(draft: Draft, opts: CaptionOptions): CaptionResult presetRanges, }); if (opts.karaoke && cue.words && cue.words.length > 0) { - const fullText = cue.words.map((word) => word.word).join(" "); + const fullText = cue.words.map((word) => word.word).join(separator); let cursor = 0; let cueMatches = 0; for (const word of cue.words) { @@ -188,7 +219,7 @@ export function captionDraft(draft: Draft, opts: CaptionOptions): CaptionResult } else { setTextRanges(draft, segmentId, [karaokeRange]); } - cursor = end + 1; + cursor = end + separator.length; created++; } keywordMatches += cueMatches; @@ -212,6 +243,7 @@ export function captionDraft(draft: Draft, opts: CaptionOptions): CaptionResult source_audio: audio, engine: opts.whisperCmd ? "shell" : "whisper-cli", engine_name: transcription.engine, + caption_script: textScript, words: transcription.words.length, karaoke: opts.karaoke ?? false, // undefined when the flags are off, so JSON output stays byte-identical. @@ -432,7 +464,21 @@ export function wordsFromCues(cues: CaptionCue[]): CaptionWord[] { return words; } -export function groupWords(words: CaptionWord[], maxWords = 4, maxChars = 28, maxGapUs = 500_000): CaptionCue[] { +/** + * Group timed words into cues: a cue closes when the next word would exceed + * maxWords, would push the joined text past maxChars, or starts after a gap + * longer than maxGapUs. `separator` is what joins the words into the cue's + * text — a space for Latin and Korean, nothing for Chinese and Japanese + * (wordSeparator in script.ts) — and counts towards maxChars like any other + * character. + */ +export function groupWords( + words: CaptionWord[], + maxWords = 4, + maxChars = 28, + maxGapUs = 500_000, + separator = " ", +): CaptionCue[] { const cues: CaptionCue[] = []; let group: CaptionWord[] = []; const flush = () => { @@ -440,7 +486,7 @@ export function groupWords(words: CaptionWord[], maxWords = 4, maxChars = 28, ma cues.push({ startUs: group[0].startUs, endUs: group[group.length - 1].endUs, - text: group.map((word) => word.word).join(" "), + text: group.map((word) => word.word).join(separator), words: group, }); group = []; @@ -450,7 +496,9 @@ export function groupWords(words: CaptionWord[], maxWords = 4, maxChars = 28, ma const gap = group.length === 0 ? 0 : word.startUs - group[group.length - 1].endUs; if ( group.length > 0 && - (candidate.length > maxWords || candidate.map((item) => item.word).join(" ").length > maxChars || gap > maxGapUs) + (candidate.length > maxWords || + candidate.map((item) => item.word).join(separator).length > maxChars || + gap > maxGapUs) ) { flush(); } diff --git a/src/command-specs.ts b/src/command-specs.ts index ac47d8e..8fbeb23 100644 --- a/src/command-specs.ts +++ b/src/command-specs.ts @@ -538,7 +538,18 @@ const optionsByCommand: Record = { option("whisper_model", ["--whisper-model"], "string", "Whisper model."), option("language", ["--language"], "string", "Language code."), option("karaoke", ["--karaoke"], "boolean", "Create word-highlight caption ranges."), - option("max_words", ["--max-words"], "number", "Maximum words per cue."), + option( + "max_words", + ["--max-words"], + "number", + "Maximum words per karaoke cue. Unset: 4; for Chinese and Japanese transcripts unlimited (the character cap bounds the cue).", + ), + option( + "max_chars", + ["--max-chars"], + "number", + "Maximum characters per cue (karaoke) or per script line. Unset: 28 / 42; Chinese 16 / 16, Japanese 13 / 13, Korean 16 / 16.", + ), option("max_gap_ms", ["--max-gap-ms"], "number", "Maximum gap inside a karaoke cue."), TRACK_NAME, STYLE_REF, diff --git a/src/lib.ts b/src/lib.ts index 8654a9b..8a40033 100644 --- a/src/lib.ts +++ b/src/lib.ts @@ -53,6 +53,8 @@ export { } from "./lint.js"; export type { RunCommandRequest, RunCommandResult } from "./runner.js"; export { runCommand } from "./runner.js"; +export type { GroupingDefaults } from "./script.js"; +export { groupingDefaults, wordSeparator } from "./script.js"; export { fromStoredOffset, rangesLookDoubled, diff --git a/src/lint.ts b/src/lint.ts index de2964b..09fa848 100644 --- a/src/lint.ts +++ b/src/lint.ts @@ -7,11 +7,17 @@ import { type Category, listEnum, type Namespace } from "./enums.js"; import { copyAssetDeduped, effectCatalogue, filterCatalogue, storeOutgrowsTemplate } from "./factory.js"; import { linkLocalMaterialIds, readSidecar, unlinkedMaterials } from "./materials-register.js"; import { ffprobeAvailable, isVfr, probeMedia } from "./probe.js"; +import { type CaptionScript, CJK_SCRIPT_LIMITS, captionScript, type ScriptLimit, type ScriptLimits } from "./script.js"; import { assessMediaRegistrationAt } from "./store.js"; import { rangesLookDoubled, repairDoubledRanges } from "./text-offsets.js"; import { allUserEnumIds } from "./user-enums.js"; import { atLeast } from "./version.js"; +export type { CaptionScript, ScriptLimit, ScriptLimits } from "./script.js"; +// The script detection and its limit table live in script.ts (caption shares +// them); re-exported here so lint stays the one import for lint callers. +export { CJK_SCRIPT_LIMITS, captionScript } from "./script.js"; + export type Severity = "error" | "warning" | "info"; export interface LintIssue { @@ -64,79 +70,6 @@ const FIXABLE_CODES = new Set([ // fixable:false instead. export const MIN_CAPTION_DURATION_US = 100_000; -/** The script a caption is written in, as far as the line-length and - * reading-speed rules care: Latin (and everything else), or one of the three - * CJK scripts whose subtitling conventions differ from Latin ones. */ -export type CaptionScript = "latin" | "zh" | "ja" | "ko"; - -export interface ScriptLimit { - maxCharsPerLine?: number; - maxCharsPerSecond?: number; -} - -/** Caption limits that replace `maxCharsPerLine` / `maxCharsPerSecond` for a - * cue written in the given script. An absent key falls back to the Latin - * value for that rule. */ -export type ScriptLimits = Partial, ScriptLimit>>; - -/** - * Where the Latin defaults (42 characters per line, 20 per second) come from - * a Latin alphabet, a CJK character carries a syllable or a word, so a line - * a third as long is already full and a third the speed is already fast: - * the streaming style guides sit at 16 characters per line and 9 per second - * for Simplified Chinese, 13 and 4 for Japanese, 16 and 12 for Korean. A - * 30-character Chinese line passing a 42-character check is the failure this - * table exists for. - */ -export const CJK_SCRIPT_LIMITS: ScriptLimits = { - zh: { maxCharsPerLine: 16, maxCharsPerSecond: 9 }, - ja: { maxCharsPerLine: 13, maxCharsPerSecond: 4 }, - ko: { maxCharsPerLine: 16, maxCharsPerSecond: 12 }, -}; - -const KANA = /[\u3040-\u30ff]/; -const HANGUL = /[\u1100-\u11ff\u3130-\u318f\uac00-\ud7af]/; -const HAN = /[\u3400-\u4dbf\u4e00-\u9fff\uf900-\ufaff]/; -// Full-width punctuation and symbols travel with the CJK scripts and count -// towards the share, without deciding which script it is. -const CJK_ANY = - /[\u3000-\u303f\u3040-\u30ff\u3130-\u318f\u3400-\u4dbf\u4e00-\u9fff\uac00-\ud7af\uf900-\ufaff\uff00-\uffef]/; - -/** - * The script a caption's limits should follow: CJK when at least half of its - * visible characters are CJK, then Japanese if any kana is present, Korean if - * any hangul, else Chinese. A mixed caption below that share (a Latin line - * with one CJK name) keeps the Latin limits. - */ -export function captionScript(text: string): CaptionScript { - let total = 0; - let cjk = 0; - let kana = 0; - let hangul = 0; - let han = 0; - for (const ch of text) { - if (/\s/.test(ch)) continue; - total++; - if (KANA.test(ch)) { - kana++; - cjk++; - } else if (HANGUL.test(ch)) { - hangul++; - cjk++; - } else if (HAN.test(ch)) { - han++; - cjk++; - } else if (CJK_ANY.test(ch)) { - cjk++; - } - } - if (total === 0 || cjk * 2 < total) return "latin"; - if (kana > 0) return "ja"; - if (hangul > 0) return "ko"; - if (han > 0) return "zh"; - return "latin"; -} - /** The line-length and reading-speed limits for one caption's text, with the * suffix its messages carry when a script-specific default applied. */ export function captionLimits( diff --git a/src/script.ts b/src/script.ts new file mode 100644 index 0000000..eb92785 --- /dev/null +++ b/src/script.ts @@ -0,0 +1,118 @@ +/** + * Which script a caption is written in, and what that changes: the limits a + * line and a reading speed are held to (lint), and how whisper's words are + * joined and grouped into cues (caption). Latin and everything else share one + * bucket; Chinese, Japanese and Korean each get their own because their + * subtitling conventions differ from Latin ones and from each other. + */ + +/** The script a caption is written in, as far as the line-length and + * reading-speed rules care: Latin (and everything else), or one of the three + * CJK scripts whose subtitling conventions differ from Latin ones. */ +export type CaptionScript = "latin" | "zh" | "ja" | "ko"; + +export interface ScriptLimit { + maxCharsPerLine?: number; + maxCharsPerSecond?: number; +} + +/** Caption limits that replace `maxCharsPerLine` / `maxCharsPerSecond` for a + * cue written in the given script. An absent key falls back to the Latin + * value for that rule. */ +export type ScriptLimits = Partial, ScriptLimit>>; + +/** + * Where the Latin defaults (42 characters per line, 20 per second) come from + * a Latin alphabet, a CJK character carries a syllable or a word, so a line + * a third as long is already full and a third the speed is already fast: + * the streaming style guides sit at 16 characters per line and 9 per second + * for Simplified Chinese, 13 and 4 for Japanese, 16 and 12 for Korean. A + * 30-character Chinese line passing a 42-character check is the failure this + * table exists for. + */ +export const CJK_SCRIPT_LIMITS: ScriptLimits = { + zh: { maxCharsPerLine: 16, maxCharsPerSecond: 9 }, + ja: { maxCharsPerLine: 13, maxCharsPerSecond: 4 }, + ko: { maxCharsPerLine: 16, maxCharsPerSecond: 12 }, +}; + +const KANA = /[\u3040-\u30ff]/; +const HANGUL = /[\u1100-\u11ff\u3130-\u318f\uac00-\ud7af]/; +const HAN = /[\u3400-\u4dbf\u4e00-\u9fff\uf900-\ufaff]/; +// Full-width punctuation and symbols travel with the CJK scripts and count +// towards the share, without deciding which script it is. +const CJK_ANY = + /[\u3000-\u303f\u3040-\u30ff\u3130-\u318f\u3400-\u4dbf\u4e00-\u9fff\uac00-\ud7af\uf900-\ufaff\uff00-\uffef]/; + +/** + * The script a caption's limits should follow: CJK when at least half of its + * visible characters are CJK, then Japanese if any kana is present, Korean if + * any hangul, else Chinese. A mixed caption below that share (a Latin line + * with one CJK name) keeps the Latin limits. + */ +export function captionScript(text: string): CaptionScript { + let total = 0; + let cjk = 0; + let kana = 0; + let hangul = 0; + let han = 0; + for (const ch of text) { + if (/\s/.test(ch)) continue; + total++; + if (KANA.test(ch)) { + kana++; + cjk++; + } else if (HANGUL.test(ch)) { + hangul++; + cjk++; + } else if (HAN.test(ch)) { + han++; + cjk++; + } else if (CJK_ANY.test(ch)) { + cjk++; + } + } + if (total === 0 || cjk * 2 < total) return "latin"; + if (kana > 0) return "ja"; + if (hangul > 0) return "ko"; + if (han > 0) return "zh"; + return "latin"; +} + +/** + * How a cue's words are joined into its text: Chinese and Japanese are written + * without spaces, so a space between two of whisper's tokens would land in the + * caption; Korean and Latin scripts are written with them. + */ +export function wordSeparator(script: CaptionScript): string { + return script === "zh" || script === "ja" ? "" : " "; +} + +export interface GroupingDefaults { + /** --max-words default for karaoke cues. */ + karaokeMaxWords: number; + /** --max-chars default for karaoke cues. */ + karaokeMaxChars: number; + /** --max-chars default when a script line outgrows one cue. */ + lineMaxChars: number; +} + +/** + * The defaults `caption` groups words with when the caller sets nothing. A + * "word" whisper emits for Chinese or Japanese is a character or a short + * token, so the Latin four-words-per-cue rule would cut a sentence into + * fragments; there the character cap is the only bound, and it is the line + * width the lint limits already hold captions to (16 zh, 13 ja, 16 ko). + */ +export function groupingDefaults(script: CaptionScript): GroupingDefaults { + switch (script) { + case "zh": + return { karaokeMaxWords: Number.POSITIVE_INFINITY, karaokeMaxChars: 16, lineMaxChars: 16 }; + case "ja": + return { karaokeMaxWords: Number.POSITIVE_INFINITY, karaokeMaxChars: 13, lineMaxChars: 13 }; + case "ko": + return { karaokeMaxWords: 4, karaokeMaxChars: 16, lineMaxChars: 16 }; + default: + return { karaokeMaxWords: 4, karaokeMaxChars: 28, lineMaxChars: 42 }; + } +} diff --git a/test/caption-cjk.test.mjs b/test/caption-cjk.test.mjs new file mode 100644 index 0000000..c8e4a7e --- /dev/null +++ b/test/caption-cjk.test.mjs @@ -0,0 +1,157 @@ +import assert from "node:assert/strict"; +import { mkdtempSync, readFileSync, rmSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { dirname, join } from "node:path"; +import { describe, it } from "node:test"; +import { fileURLToPath } from "node:url"; +import { groupWords } from "../dist/caption.js"; +import { captionScript, groupingDefaults, wordSeparator } from "../dist/script.js"; +import { spawnCli } from "./helpers/spawn-cli.mjs"; +import { tmpDraft } from "./helpers/tmp-draft.mjs"; + +// Whisper's "words" for Chinese or Japanese are characters or short tokens. +// Joined with spaces and cut four to a cue, a Chinese sentence came out as +// "今 天 我 们" fragments; now the script chooses the separator (none for zh / +// ja) and the grouping defaults (the character cap alone, at the line width +// lint holds captions to). + +const __dirname = dirname(fileURLToPath(import.meta.url)); +const FAKE_WHISPER = join(__dirname, "helpers", "fake-whisper.mjs"); +const isWindows = process.platform === "win32"; + +const ZH = "今天我们来聊一聊剪映草稿的自动化处理"; // 18 characters +const timed = (text, stepS = 0.2) => + [...text].map((word, i) => ({ + word, + startUs: Math.round(i * stepS * 1e6), + endUs: Math.round((i + 1) * stepS * 1e6), + })); + +describe("script detection and defaults (script.ts)", () => { + it("names the script from the visible characters", () => { + assert.equal(captionScript(ZH), "zh"); + assert.equal(captionScript("きょうはじまる物語"), "ja", "kana decides even beside kanji"); + assert.equal(captionScript("오늘은 새로운 이야기"), "ko"); + assert.equal(captionScript("This is a Latin line"), "latin"); + assert.equal(captionScript("Rene 说 hello world today again"), "latin", "one CJK word in a Latin line"); + assert.equal(captionScript(""), "latin"); + }); + + it("joins without a space in Chinese and Japanese, with one elsewhere", () => { + assert.equal(wordSeparator("zh"), ""); + assert.equal(wordSeparator("ja"), ""); + assert.equal(wordSeparator("ko"), " "); + assert.equal(wordSeparator("latin"), " "); + }); + + it("bounds Chinese and Japanese cues by characters only, at the lint line width", () => { + assert.deepEqual(groupingDefaults("latin"), { karaokeMaxWords: 4, karaokeMaxChars: 28, lineMaxChars: 42 }); + assert.deepEqual(groupingDefaults("zh"), { + karaokeMaxWords: Number.POSITIVE_INFINITY, + karaokeMaxChars: 16, + lineMaxChars: 16, + }); + assert.deepEqual(groupingDefaults("ja"), { + karaokeMaxWords: Number.POSITIVE_INFINITY, + karaokeMaxChars: 13, + lineMaxChars: 13, + }); + assert.deepEqual(groupingDefaults("ko"), { karaokeMaxWords: 4, karaokeMaxChars: 16, lineMaxChars: 16 }); + }); +}); + +describe("groupWords with a separator", () => { + it("joins Chinese characters without spaces and splits at the character cap", () => { + const cues = groupWords(timed(ZH), Number.POSITIVE_INFINITY, 16, Number.POSITIVE_INFINITY, ""); + assert.deepEqual( + cues.map((c) => c.text), + ["今天我们来聊一聊剪映草稿的自动化", "处理"], + ); + assert.equal(cues[0].words.length, 16); + }); + + it("keeps the Latin behaviour when no separator is given", () => { + const words = ["one", "two", "three", "four", "five"].map((word, i) => ({ + word, + startUs: i * 200_000, + endUs: (i + 1) * 200_000, + })); + const cues = groupWords(words, 3, 20); + assert.deepEqual( + cues.map((c) => c.text), + ["one two three", "four five"], + ); + }); +}); + +function trackTexts(draftPath, trackName) { + const draft = JSON.parse(readFileSync(draftPath, "utf-8")); + const track = draft.tracks.find((t) => t.type === "text" && t.name === trackName); + assert.ok(track, `track ${trackName} present`); + return track.segments.map((s) => { + const mat = draft.materials.texts.find((m) => m.id === s.material_id); + return JSON.parse(mat.content); + }); +} + +describe("capcut caption --karaoke on a Chinese transcript (fake whisper)", { skip: isWindows }, () => { + it("writes space-free cues of at most 16 characters and reports caption_script", (t) => { + const fix = tmpDraft(); + t.after(() => fix.cleanup()); + const dir = mkdtempSync(join(tmpdir(), "capcut-caption-cjk-")); + t.after(() => rmSync(dir, { recursive: true, force: true })); + const audio = join(dir, "voice.wav"); + writeFileSync(audio, "stub"); + const heard = join(dir, "heard.json"); + writeFileSync( + heard, + JSON.stringify({ + segments: [{ words: [...ZH].map((word, i) => ({ word, start: i * 0.2, end: (i + 1) * 0.2 })) }], + }), + ); + + const r = spawnCli( + ["caption", fix.path, "--audio", audio, "--whisper-cmd", FAKE_WHISPER, "--karaoke", "--track-name", "zh"], + { env: { FAKE_WHISPER_SOURCE: heard } }, + ); + assert.equal(r.status, 0, `stderr: ${r.stderr}`); + assert.equal(r.json.caption_script, "zh"); + assert.equal(r.json.words, 18); + assert.equal(r.json.first_cue.text, "今天我们来聊一聊剪映草稿的自动化"); + assert.equal(r.json.last_cue.text, "处理"); + + // Karaoke writes one segment per word carrying the whole cue's text. + const texts = trackTexts(fix.path, "zh"); + assert.equal(texts.length, 18); + for (const content of texts) { + assert.doesNotMatch(content.text, / /, "no space between Chinese characters"); + assert.ok(content.text.length <= 16, `cue within 16 characters: ${content.text}`); + } + // The third word's highlight covers exactly its own character. + const third = texts[2].styles.find((s) => s.range && s.range[0] === 2); + assert.ok(third, `third word range present: ${JSON.stringify(texts[2].styles)}`); + assert.deepEqual(third.range, [2, 3]); + }); + + it("still honours an explicit --max-words / --max-chars", (t) => { + const fix = tmpDraft(); + t.after(() => fix.cleanup()); + const dir = mkdtempSync(join(tmpdir(), "capcut-caption-cjk-")); + t.after(() => rmSync(dir, { recursive: true, force: true })); + const audio = join(dir, "voice.wav"); + writeFileSync(audio, "stub"); + const heard = join(dir, "heard.json"); + writeFileSync( + heard, + JSON.stringify({ + segments: [{ words: [...ZH].map((word, i) => ({ word, start: i * 0.2, end: (i + 1) * 0.2 })) }], + }), + ); + const r = spawnCli( + ["caption", fix.path, "--audio", audio, "--whisper-cmd", FAKE_WHISPER, "--karaoke", "--max-words", "6"], + { env: { FAKE_WHISPER_SOURCE: heard } }, + ); + assert.equal(r.status, 0, `stderr: ${r.stderr}`); + assert.equal(r.json.first_cue.text, "今天我们来聊"); + }); +});