Skip to content

Fix (file names): Keep duplicate page postfixes and avoid broken attachment and image links - #193

Open
klangme1ster wants to merge 1 commit into
theohbrothers:masterfrom
klangme1ster:fix/file-names-keep-duplicate-postfix-and-avoid-broken-links
Open

klangme1ster wants to merge 1 commit into
theohbrothers:masterfrom
klangme1ster:fix/file-names-keep-duplicate-postfix-and-avoid-broken-links

Conversation

@klangme1ster

Copy link
Copy Markdown
  • Duplicate page names: the '-1', '-2' postfix was appended before truncating to mdFileNameAndFolderNameMaxLength, so for long names it was cut off again, and the pages overwrote each other. Duplicates are now detected on the truncated name (which also catches long names that only differ after the cut-off), and the postfix survives truncation (new Truncate-PathFileName -KeepSuffix).
  • Inserted attachments: references were replaced with a plain search, so with attachments '11-36002.pdf' and 'Datenblatt_11-36002.pdf' on one page, the first name also matched inside the second, nesting one link inside the other. Only whole file names are matched now.
  • Images: image file names are prefixed with the page path, which can contain characters like '#' and "'" (e.g. 'c't-expert-community--Newsletter-Errors all over the place #10-image1.jpg'). A '#' in a link is read as a URL fragment, breaking the image. Image names now go through Remove-InvalidFileNameCharsInsertedFiles, like inserted attachments.

…chment and image links

- Duplicate page names: the '-1', '-2' postfix was appended before truncating to mdFileNameAndFolderNameMaxLength, so for long names it was cut off again, and the pages overwrote each other. Duplicates are now detected on the truncated name (which also catches long names that only differ after the cut-off), and the postfix survives truncation (new Truncate-PathFileName -KeepSuffix).
- Inserted attachments: references were replaced with a plain search, so with attachments '11-36-00002.pdf' and 'Datenblatt_11-36-00002.pdf' on one page, the first name also matched inside the second, nesting one link inside the other. Only whole file names are matched now.
- Images: image file names are prefixed with the page path, which can contain characters like '#' and "'" (e.g. 'c't-expert-community--Newsletter-theohbrothers#10-image1.jpg'). A '#' in a link is read as a URL fragment, breaking the image. Image names now go through Remove-InvalidFileNameCharsInsertedFiles, like inserted attachments.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant