Markdown File Converter
Paste the following prompt into your AI chat to install this skill:
Please install @user_e8ecaff7/test1-new according to https://skillhub.cn/install/skillhub.md.
About this skill
Problem
When working with source materials, text in PDFs, Office files, HTML, CSV, JSON, and images is often trapped in binary layouts or markup. Copying it by hand can discard headings, tables, lists, and links, while raw files are awkward to feed into search indexes, knowledge bases, or model pipelines. This skill turns those common inputs into Markdown, a text-friendly intermediate format for reading, indexing, and further processing.
How it works
The skill runs uvx markitdown, so the main converter does not require a separate installation. It supports Documents such as PDF, Word (.docx), PowerPoint (.pptx), and Excel (.xlsx / .xls); Web/Data such as HTML, CSV, JSON, and XML; Media such as image EXIF extraction and OCR, plus audio EXIF and transcription; and Other inputs including ZIP contents, YouTube URLs, and EPub files. The output preserves document structure, including headings, tables, lists, and links. The first run caches dependencies, making later runs faster. For complex PDFs or files with poor extraction, the notes suggest using the -d option with Azure Document Intelligence. It fits cases where scattered files should become versionable plain text, but not DRM-protected files, very poor scans, or tasks requiring exact original layout.
Use Cases
- Convert multiple PDF, Word, and PPT reports into Markdown for unified search and model summarization.
- Export HTML, CSV, JSON, and XML samples to Markdown for archiving API notes.
- Extract image EXIF/OCR and audio transcription into editable Markdown notes.
- Iterate over a ZIP archive's files and convert each to Markdown for bulk offline documentation cleanup.
Best For
- Knowledge-base engineers who need PDFs, Office docs, and web pages as searchable Markdown.
- Data engineers archiving interface samples, CSV, JSON, and HTML docs.
- Research assistants extracting OCR text and audio transcriptions from media files.
- Documentation maintainers batch-processing ZIP archives into plain-text records.
Related Skills
Reads conversation-trace files to generate an animal- or mythology-based soul mirror card and Johari Window insights.
A Deling knowledge-base research workflow that clarifies intent, runs broad and vertical searches, supplements with web sources, validates diversity, and traces key claims.
A SiYuan knowledge-base management skill for double-link parent indexes, MOCs, numbered documents, tags, repo sync, and WeChat import workflows.
A structured workflow for academic literature reviews, covering multi-database search, screening, thematic synthesis, citation validation, and PDF output.