Install any skill in seconds. Free to start, no credit card required.
Get Started Free →デバイス上基盤モデルの実装パターン、量子化、最適化、およびプライバシーを考慮した推論。
.claude/skills/affaan-m-foundation-models-on-device/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 32% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 40% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 50% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -7% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 42% | 0% |
使用 FoundationModels 框架将苹果的设备端语言模型集成到应用中的模式。涵盖文本生成、使用 @Generable 的结构化输出、自定义工具调用以及快照流式传输——全部在设备端运行,以保护隐私并支持离线使用。
在创建会话之前,始终检查模型可用性:
swiftstruct GenerativeView: View { private var model = SystemLanguageModel.default var body: some View { switch model.availability { case .available: ContentView() case .unavailable(.deviceNotEligible): Text("Device not eligible for Apple Intelligence") case .unavailable(.appleIntelligenceNotEnabled): Text("Please enable Apple Intelligence in Settings") case .unavailable(.modelNotReady): Text("Model is downloading or not ready") case .unavailable(let other): Text("Model unavailable: \(other)") } } }
swift// Single-turn: create a new session each time let session = LanguageModelSession() let response = try await session.respond(to: "What's a good month to visit Paris?") print(response.content) // Multi-turn: reuse session for conversation context let session = LanguageModelSession(instructions: """ You are a cooking assistant. Provide recipe suggestions based on ingredients. Keep suggestions brief and practical. """) let first = try await session.respond(to: "I have chicken and rice") let followUp = try await session.respond(to: "What about a vegetarian option?")
指令的关键点:
生成结构化的 Swift 类型,而不是原始字符串:
swift@Generable(description: "Basic profile information about a cat") struct CatProfile { var name: String @Guide(description: "The age of the cat", .range(0...20)) var age: Int @Guide(description: "A one sentence profile about the cat's personality") var profile: String }
swiftlet response = try await session.respond( to: "Generate a cute rescue cat", generating: CatProfile.self ) // Access structured fields directly print("Name: \(response.content.name)") print("Age: \(response.content.age)") print("Profile: \(response.content.profile)")
.range(0...20) — 数值范围.count(3) — 数组元素数量description: — 生成的语义引导让模型调用自定义代码以执行特定领域的任务:
swiftstruct RecipeSearchTool: Tool { let name = "recipe_search" let description = "Search for recipes matching a given term and return a list of results." @Generable struct Arguments { var searchTerm: String var numberOfResults: Int } func call(arguments: Arguments) async throws -> ToolOutput { let recipes = await searchRecipes( term: arguments.searchTerm, limit: arguments.numberOfResults ) return .string(recipes.map { "- \($0.name): \($0.description)" }.joined(separator: "\n")) } }
swiftlet session = LanguageModelSession(tools: [RecipeSearchTool()]) let response = try await session.respond(to: "Find me some pasta recipes")
swiftdo { let answer = try await session.respond(to: "Find a recipe for tomato soup.") } catch let error as LanguageModelSession.ToolCallError { print(error.tool.name) if case .databaseIsEmpty = error.underlyingError as? RecipeSearchToolError { // Handle specific tool error } }
使用 PartiallyGenerated 类型为实时 UI 流式传输结构化响应:
swift@Generable struct TripIdeas { @Guide(description: "Ideas for upcoming trips") var ideas: [String] } let stream = session.streamResponse( to: "What are some exciting trip ideas?", generating: TripIdeas.self ) for try await partial in stream { // partial: TripIdeas.PartiallyGenerated (all properties Optional) print(partial) }
swift@State private var partialResult: TripIdeas.PartiallyGenerated? @State private var errorMessage: String? var body: some View { List { ForEach(partialResult?.ideas ?? [], id: \.self) { idea in Text(idea) } } .overlay { if let errorMessage { Text(errorMessage).foregroundStyle(.red) } } .task { do { let stream = session.streamResponse(to: prompt, generating: TripIdeas.self) for try await partial in stream { partialResult = partial } } catch { errorMessage = error.localizedDescription } } }
| 决策 | 理由 | |----------|-----------| | 设备端执行 | 隐私性——数据不离开设备;支持离线工作 | | 4,096 个令牌限制 | 设备端模型约束;跨会话分块处理大数据 | | 快照流式传输(非增量) | 对结构化输出友好;每个快照都是一个完整的部分状态 | | @Generable 宏 | 为结构化生成提供编译时安全性;自动生成 PartiallyGenerated 类型 | | 每个会话单次请求 | isResponding 防止并发请求;如有需要,创建多个会话 | | response.content(而非 .output) | 正确的 API——始终通过 .content 属性访问结果 |
model.availability——处理所有不可用的情况instructions 来引导模型行为——它们的优先级高于提示词isResponding——会话一次处理一个请求response.content 访问结果——而不是 .output@Generable——比解析原始字符串提供更强的保证GenerationOptions(temperature:) 来调整创造力(值越高越有创意)model.availability 就创建会话.output 而不是 .content 来访问响应数据@Generable 结构化输出可行时,却去解析原始字符串响应| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-02 | fail→pass | 16,344 | 10,825 | -34% | 1 | 1 | 0% | 3,151 | 4,163 | +32% | 0 | 0 | — |
case-01 | fail→pass | 23,319 | 13,978 | -40% | 1 | 1 | 0% | 3,437 | 4,802 | +40% | 0 | 0 | — |
case-03 | fail→pass | 19,221 | 17,269 | -10% | 1 | 1 | 0% | 3,685 | 5,526 | +50% | 0 | 0 | — |
case-04 | fail→pass | 20,636 | 7,730 | -63% | 1 | 1 | 0% | 3,610 | 3,351 | -7% | 0 | 0 | — |
case-05 | pass→pass | 14,588 | 11,140 | -24% | 1 | 1 | 0% | 2,388 | 3,941 | +65% | 0 | 0 | — |
case-06 | pass→pass | 9,027 | 3,197 | -65% | 1 | 1 | 0% | 1,436 | 2,490 | +73% | 0 | 0 | — |
case-07 | fail→pass | 11,133 | 4,720 | -58% | 1 | 1 | 0% | 1,999 | 2,833 | +42% | 0 | 0 | — |
case-08 | fail→pass | 14,715 | 7,010 | -52% | 1 | 1 | 0% | 2,426 | 3,227 | +33% | 0 | 0 | — |
case-09 | fail→pass | 12,315 | 7,963 | -35% | 1 | 1 | 0% | 2,015 | 3,473 | +72% | 0 | 0 | — |
case-10 | pass→pass | 11,652 | 10,390 | -11% | 1 | 1 | 0% | 2,075 | 3,938 | +90% | 0 | 0 | — |
case-11 | fail→pass | 12,531 | 5,479 | -56% | 1 | 1 | 0% | 2,142 | 2,897 | +35% | 0 | 0 | — |
case-12 | fail→pass | 13,088 | 4,976 | -62% | 1 | 1 | 0% | 2,246 | 2,917 | +30% | 0 | 0 | — |
case-13 | pass→pass | 10,176 | 4,351 | -57% | 1 | 1 | 0% | 1,447 | 2,648 | +83% | 0 | 0 | — |
case-14 | fail→pass | 12,743 | 4,454 | -65% | 1 | 1 | 0% | 1,971 | 2,639 | +34% | 0 | 0 | — |
case-15 | fail→pass | 24,312 | 5,087 | -79% | 1 | 1 | 0% | 2,545 | 2,735 | +7% | 0 | 0 | — |
case-16 | pass→pass | 13,379 | 10,226 | -24% | 1 | 1 | 0% | 1,995 | 3,629 | +82% | 0 | 0 | — |
case-17 | pass→pass | 7,839 | 7,483 | -5% | 1 | 1 | 0% | 1,342 | 3,269 | +144% | 0 | 0 | — |
case-18 | fail→fail | 8,193 | 6,568 | -20% | 1 | 1 | 0% | 1,406 | 2,916 | +107% | 0 | 0 | — |
case-19 | pass→pass | 13,532 | 12,423 | -8% | 1 | 1 | 0% | 2,359 | 4,093 | +74% | 0 | 0 | — |
case-20 | pass→pass | 8,749 | 6,220 | -29% | 1 | 1 | 0% | 1,645 | 2,927 | +78% | 0 | 0 | — |
case-21 | pass→pass | 7,902 | 7,721 | -2% | 1 | 1 | 0% | 1,488 | 3,475 | +134% | 0 | 0 | — |
case-22 | pass→pass | 12,185 | 11,702 | -4% | 1 | 1 | 0% | 2,250 | 3,979 | +77% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +50 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.