MongoDB Data Modeling
Paste the following prompt into your AI chat to install this skill:
Install @kunlungrowth/mongodb-modeling using the official guide at https://skillhub.cn/install/skillhub.md.
About this skill
Problem
MongoDB schemas often fail when they normalize too early or embed too aggressively. Lists may run repeated $lookup, detail pages may unwind large arrays, and oversized documents degrade writes. This skill helps developers choose a schema from query paths, write frequency, and document size before production pressure forces a rewrite.
Core workflow
The skill first checks whether MongoDB is the right tool: flexible documents, read-heavy workloads, JSON-shaped data, and horizontal scaling are good fits. Strict cross-document transactions, complex many-table joins, and fixed-schema reporting should usually favor PostgreSQL. It then applies the “read and write together” rule: one-to-one relations are often embedded; one-to-many relations are usually separate collections with references; low-cardinality references may use Extended Reference, such as caching author_name, while metrics like comment_count may use a Computed Pattern. Finally, it proposes indexes and pipelines aligned to the access path, for example articles: { created_at: -1, status: 1 } and comments: { article_id: 1, created_at: -1 }, and prefers $lookup plus $project or two targeted queries over $lookup + $unwind when that is cheaper.
Boundaries
Use it for schema selection, slow collection diagnosis, sharding strategy, and aggregation pipeline review. Avoid relying on it for strict ACID financial flows, complex OLAP, or multi-table reporting. It reads like an architecture checklist, not a substitute for workload profiling and capacity testing.
Use Cases
- Choose collection boundaries for a blog's articles, comments, and authors, and select indexes plus query paths.
- Reduce repeated author lookups in time-ordered list pages by tuning aggregation pipelines and reference denormalization.
- Evaluate moving large comment sets to a separate collection with indexes and pagination strategies.
- Design shard key, cluster topology, and routing for TB-scale business data.
Best For
- Backend developers who need to choose between embedded and referenced models for JSON-heavy data without post-launch rewrites.
- Content platform engineers who need to optimize article list, detail, and comment aggregation query paths.
- Data architecture owners who need to assess whether MongoDB fits the workload and define a sharding strategy.
- Developers inheriting legacy systems who need to diagnose slow collections, add indexes, or restructure document relationships.
Related Skills
Automatically indexes Gradle-cached AAR/JAR dependency classes and returns library coordinates, versions, and public APIs by fully qualified name, using only the Python standard library.
Codifies AMT and YourMT3 training conventions, script patterns, hyperparameters, precision, checkpoints, and NaN safeguards.
Retrieve relevant chunks from a customer-managed PKM dataset by dataset_id and return concise, source-annotated answers.
Convert PRDs, user stories, or functional specs into prioritized test-point checklists covering functional, business-rule, boundary, exception, and non-functional dimensions.