Ai Research skills for AI agents
12 practitioner-grade ai research skills, each a focused Markdown document your agent loads into context on demand. Search them from Claude Desktop, Cursor or any MCP client, or pull one with the CLI.
All 12 skills
- AI Ethics Responsible
Triggers when users need help with AI ethics, fairness, or responsible AI development. Activate for questions about fairness metrics (demographic parity, equalized odds, calibration), bias detection and mitigation (pre-processing, in-processing, post-processing), model cards, datasheets for datasets, AI impact assessments, transparency and explainability requirements, dual-use considerations, and environmental impact of training large models.
125 lines - AI Grant Funding
Triggers when users need help writing AI/ML research grant proposals or planning funded research. Activate for questions about NSF, DARPA, NIH, or industry research grants, positioning AI research for funding, budget planning for compute-heavy research, collaboration letters, broader impacts statements, research timeline planning for ML projects, and industry research partnerships.
124 lines - AI Peer Review
Triggers when users need help reviewing ML papers or understanding the peer review process. Activate for questions about NeurIPS, ICML, or ICLR review guidelines, writing constructive criticism, evaluating experimental rigor, spotting common statistical errors in ML papers, assessing novelty versus incrementalism, reviewer calibration, area chair responsibilities, ethical review considerations, and meta-review writing.
135 lines - AI Research Methodology
Triggers when users need help designing ML experiments, formulating research hypotheses, or planning ablation studies. Activate for questions about baseline selection, controlled experiments in machine learning, statistical significance testing for model comparisons, reproducibility practices (random seeds, environment pinning), compute budgeting for research, and refining research questions in AI/ML.
103 lines - AI Safety Alignment
Triggers when users need help with AI safety, alignment research, or responsible AI development. Activate for questions about RLHF (reward modeling, PPO, DPO, KTO), constitutional AI, scalable oversight, interpretability methods (mechanistic interpretability, probing, attention analysis), red teaming, jailbreak robustness, alignment tax, corrigibility, the value alignment problem, and AI governance frameworks.
119 lines - Experiment Tracking Research
Triggers when users need help with experiment management and tracking for ML research. Activate for questions about experiment management systems (Weights & Biases, MLflow, Neptune, Aim), hyperparameter logging, artifact versioning, experiment comparison and visualization, collaborative research workflows, compute cost tracking, experiment metadata standards, and reproducing experiments from logs.
145 lines - Literature Survey AI
Triggers when users need help conducting systematic literature reviews in AI/ML, constructing taxonomies of research areas, or writing survey papers. Activate for questions about identifying research trends and gaps, keeping up with the arXiv flood, daily paper filtering, conference proceeding analysis, citation network analysis, and research community mapping.
131 lines - ML Benchmarking
Triggers when users need help with ML benchmark design, dataset curation for evaluation, or evaluation protocol construction. Activate for questions about leaderboard management, benchmark saturation detection, cross-dataset generalization, statistical testing for model comparison (McNemar, paired bootstrap), avoiding benchmark gaming, and standard benchmark suites like GLUE, SuperGLUE, ImageNet, MMLU, HELM, and BIG-Bench.
121 lines - ML Paper Writing
Triggers when users need help writing ML research papers for top venues. Activate for questions about writing for NeurIPS, ICML, ICLR, CVPR, paper structure conventions, abstract writing, related work positioning, experiment section best practices, figure and table design, rebuttal writing, camera-ready preparation, supplementary material organization, and arXiv preprint strategy.
128 lines - Open Source ML
Triggers when users need help with open-source ML practices, model release, or community engagement. Activate for questions about Hugging Face Hub (model cards, datasets, spaces), open-weight vs open-source licensing (Apache 2.0, MIT, Llama license), documentation for ML projects, community building around ML tools, contributing to open-source ML frameworks like PyTorch or JAX, and responsible model release practices.
137 lines - Paper Reading Reproduction
Triggers when users need help reading ML papers efficiently, critically evaluating research claims, or reproducing results from published work. Activate for questions about the three-pass paper reading method, understanding mathematical notation in ML papers, identifying key contributions versus incremental work, building literature maps, and troubleshooting common reproducibility pitfalls.
125 lines - Research Engineering
Triggers when users need help with ML infrastructure, GPU cluster management, or distributed training. Activate for questions about SLURM, Kubernetes for ML workloads, distributed training setup (data parallel, model parallel, pipeline parallel, FSDP, DeepSpeed ZeRO), experiment management at scale, Weights & Biases or MLflow for research, debugging distributed training, profiling GPU utilization, and memory optimization.
125 lines