Audit de fidélité powerpoint vers html
/SKILLVérifier une exportation Python-PPTX par rapport à la présentation HTML d'origine, repérer les écarts de mise en page ou de contenu (pied de page qui
name: pptx-html-fidelity-audit
description: Audit a python-pptx export against its source HTML deck, identify layout/content drift (footer overflow, cropped content, missing italic/em, lost styling, off-rhythm spacing), and re-export with strict footer-rail + cursor-flow layout discipline. Use this skill whenever the user has a .pptx that was generated from an HTML slide deck and asks to compare/audit/verify/fix the export — including phrases like "compare ppt with html", "fidelity audit", "fix the pptx", "ppt is cut off", "footer overlap", "italic missing in pptx", "re-export the deck", "pptx-html-fidelity-audit", or any case where a python-pptx → HTML round-trip needs verification or repair. Also trigger when the user shows you a deck.html and a deck.pptx side by side and is debugging visual differences.
triggers:
- "pptx fidelity"
- "pptx audit"
- "ppt 跑掉"
- "字型不對"
- "footer overlap"
- "verify pptx"
- "html to pptx"
od:
mode: utility
scenario: engineering
PPTX ↔ HTML Fidelity Audit
A repeatable workflow for catching the ways a python-pptx export silently drifts from its HTML source — and fixing them with a layout discipline that prevents the same regressions on the next pass.
When this skill applies
The user has:
- A source HTML slide deck (typically a single-file deck with
<section class="slide">blocks):
<section class="slide light">
<div class="chrome">2026 · Q2 review</div>
<span class="kicker">Pillar 03</span>
<h2 class="h-xl">Shipping <em>velocity</em> doubled</h2>
<p class="lead">…</p>
<div class="foot">page 5 / 14</div>
</section>- A PPTX file generated from that deck via python-pptx (or similar).
- A suspicion (or visible evidence) that the PPTX doesn't match the HTML — text bleeding into the footer, italic words gone flat, hero slides not centered, sections cropped, tag styling lost.
If the user only has one of those two artifacts, this skill doesn't apply yet — first generate the missing one, or ask the user to provide it.
Why this is hard (and why a skill helps)
PPTX is a fixed-canvas, absolute-positioned medium. HTML is a fluid, flow-based medium. A naive python-pptx export pins each block at hand-picked (top, left) coordinates, which works for the first slide it was tested on and silently fails for every other slide whose content has different intrinsic height. The result is the most common drift modes:
- Footer overflow — content's
top + heightcrosses into the footer row. - Off-canvas content — bottom of last block exceeds
7.5"(16:9 canvas). - Italic loss —
<em>in HTML never getsrun.font.italic = True. - Hero slides not centered — vertical-stack slides use
MARGIN_TOPinstead of computing center. - Box bounds intruding — the text fits, but the shape's bounding box is oversized and visually crosses the rail.
- Tag/styling loss — colored chrome rows, kicker uppercase tracking, mono-vs-serif assignments quietly fall back to defaults.
Every one of these is a layout discipline problem, not a content problem. Once you adopt the discipline, they stop happening.
Workflow
The audit is five steps. Don't skip any of them — the discipline only works if the audit produces a real list of issues to drive the re-export. A fix-without-audit pass tends to leave half the issues alive.
Step 1 — Extract ground truth from the PPTX
Run scripts/extract_pptx.py <path-to.pptx> > pptx_dump.json. The script walks every shape on every slide and dumps text, position (top / left), size (width / height), and per-run typography (font name, size pt, bold, italic, color). This is the actual state of the export — don't trust the export script's intent, trust the dump.
For 14-slide decks, the dump is ~30–60 KB and human-readable.
Step 2 — Walk the HTML structure
Read the source HTML and enumerate <section class="slide"> blocks. For each, note:
- The slide's theme (
light/dark/hero light/hero dark). - The
chromerow text (top metadata). - The
kicker(small uppercase eyebrow above the headline). - The headline (h-hero / h-xl / etc.) and any sub-head.
- The body copy and any structured blocks (pipeline steps, cards, pillars, observation cards).
- The
footrow (bottom metadata). - Any
<em>or italic-styled spans — italic is the silent regression.
Map each HTML slide to a PPTX slide index. For decks following the convention "slide 1 = cover, slide N = closing", the mapping is positional.
Step 3 — Build the audit table
For each slide, walk shapes from the dump and check against expected layout rules. Use this exact table format — the severity column is what drives the fix priority:
| Slide | Issue | Severity |
|---|---|---|
| 1 cover | meta-row 底端 6.95" 蓋過 footer (6.7") | 🔴 |
| 5 checklist | row B 步驟描述底端 7.2" 切到 footer | 🔴 |
| 8 3E | 收束段落直接坐在 footer 起點 | 🔴 |
| 9 on-day | step 描述底端剛好碰 footer,無安全距 | 🟠 |
| 多處 | em (Play