Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Build private, on-device AI features on iPhone, iPad, and Mac with Foundation Models, Core ML, MLX Swift, or llama.cpp. Use when choosing an Apple-local model runtime, building an Apple Intelligence chatbot or tool-calling feature, running an LLM on Apple Silicon, converting or compressing a Python model for Core ML, or comparing on-device inference backends. For Swift Core ML loading and prediction code, use the coreml skill.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 57% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 88% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 97% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 284% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 146% | 0% |
Guide for selecting, deploying, and optimizing on-device ML models. Covers Apple Foundation Models, Core ML, MLX Swift, and llama.cpp.
Use this decision tree to pick the right framework for your use case.
When to use: Text generation, summarization, entity extraction, structured output, and short dialog on iOS 26+ / macOS 26+ devices with Apple Intelligence enabled. No app-managed API key, network round trip, or model hosting; still handle system model asset readiness.
Best for:
@Generable typesTool protocolNot suited for: Complex math, code generation, factual accuracy tasks, or apps targeting pre-iOS 26 devices.
When to use: Deploying custom trained models (vision, NLP, audio) across all Apple platforms. Converting models from PyTorch, TensorFlow, or scikit-learn with coremltools.
Best for:
When to use: Running specific open-source LLMs (Llama, Mistral, Qwen, Gemma) on Apple Silicon with maximum throughput. Research and prototyping.
Best for:
mlx-communityWhen to use: Cross-platform LLM inference using GGUF model format. Production deployments needing broad device support.
Best for:
| Scenario | Framework | |---|---| | Text generation on Apple Intelligence devices (iOS 26+) | Foundation Models | | Structured output from on-device LLM | Foundation Models (@Generable) | | Image classification, object detection | Core ML | | Custom model from PyTorch/TensorFlow | Core ML + coremltools | | Running specific open-source LLMs | MLX Swift or llama.cpp | | Maximum throughput on Apple Silicon | MLX Swift | | Cross-platform LLM inference | llama.cpp | | OCR and text recognition | Vision framework | | Sentiment analysis, NER, tokenization | Natural Language framework | | Training custom classifiers on device | Create ML |
Use the system language model for short generation, summarization, tagging, structured output, and tool-augmented tasks on Apple Intelligence devices. Gate every entry point before creating a session:
swiftimport FoundationModels switch SystemLanguageModel.default.availability { case .available: guard SystemLanguageModel.default.supportsLocale(Locale.current) else { // Use locale fallback before generating break } // Proceed with model usage case .unavailable(.appleIntelligenceNotEnabled): // Guide user to enable Apple Intelligence in Settings case .unavailable(.modelNotReady): // System model assets are not ready; show loading state case .unavailable(.deviceNotEligible): // Device cannot run Apple Intelligence; use fallback case .unavailable(let reason): // Unknown or future unavailable reason; use fallback and log reason }
Then create a session and keep its shared context budget small:
swiftlet session = LanguageModelSession { "You are a helpful cooking assistant." } session.prewarm() let response = try await session.respond(to: "Suggest a quick pasta recipe")
Required guardrails:
check isResponding before issuing another response.
context window. Register only necessary tools and keep schemas compact.
supportsLocale(_:); do not raw-match language lists.remain active, so handle refusal and other generation errors with fallback UI.
Load the Foundation Models reference when the task needs @Generable, @Guide, streaming, tool definitions, transcripts, generation options, custom adapters, prompt design, or detailed error handling.
Apple's framework for deploying trained models. Automatically dispatches to the optimal compute unit (CPU, GPU, or Neural Engine).
| Format | Extension | When to Use | |---|---|---| | .mlpackage | Directory (mlprogram) | All new models (iOS 15+) | | .mlmodel | Single file (neuralnetwork) | Legacy only (iOS 11-14) | | .mlmodelc | Compiled | Pre-compiled for faster loading |
Always use mlprogram (.mlpackage) for new work.
pythonimport coremltools as ct # PyTorch conversion (torch.jit.trace) model.eval() # CRITICAL: always call eval() before tracing traced = torch.jit.trace(model, example_input) mlmodel = ct.convert( traced, inputs=[ct.TensorType(shape=(1, 3, 224, 224), name="image")], minimum_deployment_target=ct.target.iOS18, convert_to='mlprogram', ) mlmodel.save("Model.mlpackage")
tolerances before conversion.
precision, and preprocessing; fix the conversion and rerun the fixtures.
each compression change and undo or tune changes that miss the threshold.
correctness, latency, memory, and package-size targets all pass.
coremlThis skill owns Python-side conversion, compression, profiling, and framework selection. Use the sibling coreml skill for Swift app integration, prediction APIs, runtime configuration, Vision request wiring, and detailed model loading.
> See references/coreml-conversion.md for the > full conversion pipeline and references/coreml-optimization.md > for optimization techniques.
Apple's ML framework for Swift. Highest sustained generation throughput on Apple Silicon via unified memory architecture.
swiftimport MLX import MLXLLM import MLXLMCommon import MLXLMHFAPI let container = try await LLMModelFactory.shared.loadContainer( from: HubClient.default, using: TokenizersLoader(), configuration: .init(id: "mlx-community/Qwen3-4B-4bit") ) let session = ChatSession(container) print(try await session.respond(to: "Hello"))
| Device | RAM | Recommended Model | RAM Usage | |---|---|---|---| | iPhone 12-14 | 4-6 GB | SmolLM2-135M or Qwen 2.5 0.5B | ~0.3 GB | | iPhone 15 Pro+ | 8 GB | Gemma 3n E4B 4-bit | ~3.5 GB | | Mac 8 GB | 8 GB | Llama 3.2 3B 4-bit | ~3 GB | | Mac 16 GB+ | 16 GB+ | Mistral 7B 4-bit | ~6 GB |
Memory.cacheLimit = 512 * 1024 * 1024also call Memory.clearCache() after generation-heavy phases
exercise Metal-dependent inference, memory, or performance
> See references/mlx-swift.md for full MLX Swift > patterns and llama.cpp integration.
When an app needs multiple AI backends (e.g., Foundation Models + MLX fallback):
swiftfunc respond(to prompt: String) async throws -> String { if SystemLanguageModel.default.isAvailable { return try await foundationModelsRespond(prompt) } else if canLoadMLXModel() { return try await mlxRespond(prompt) } else { throw AIError.noBackendAvailable } }
Serialize all model access through a coordinator actor to prevent contention:
swiftactor ModelCoordinator { func withExclusiveAccess<T>(_ work: () async throws -> T) async rethrows -> T { try await work() } }
For custom Core ML models, name only the conversion/optimization handoff here: send Swift app integration, model loading, Vision wiring, and prediction lifecycle to coreml. Keep private user content, such as journals, on device unless product explicitly opts into a nonlocal fallback.
"Debug Executable")
session.prewarm() for Foundation Models before user interaction.mlmodelc for faster loadingintegration to the sibling framework skills
SystemLanguageModel.default.availability leaves unsupported devices with failures instead of fallback UI.
see nothing. Always provide a graceful degradation path.
Monitor usage via tokenCount(for:) and summarize when needed.
LanguageModelSession supports onerequest at a time. Check session.isResponding or serialize access.
parameter bypasses guardrail boundaries. Keep user content in the prompt.
on fixed fixtures, then fix and reconvert before compressing or shipping.
model.eval() before Core ML tracing. PyTorch models must bein eval mode before torch.jit.trace. Training-mode artifacts corrupt output.
mlprogram (.mlpackage) for newCore ML models. The legacy neuralnetwork format is deprecated.
physical devices; Simulator is only a UI/control-flow smoke test.
Memory.clearCache().@Generable properties in logical generation ordercontextSize)Sendable-conformant or @MainActor-isolated@Generable, tool calling, prompt designOther measured skills in the registry, with their headline benchmark lift.