LLM Server

LLM Server

Private AI Inference on Your Device设备端私密 AI 推理デバイス上のプライベート AI 推論

Turn your iPhone or iPad into a fully functional AI inference server. Run large language models entirely on-device with no cloud dependency. Your conversations and data never leave your phone.将您的 iPhone 或 iPad 变成一个功能完备的 AI 推理服务器。完全在设备上运行大型语言模型,无需云端依赖。您的对话和数据永远不会离开您的手机。iPhone や iPad を本格的な AI 推論サーバーに変えましょう。大規模言語モデルをデバイス上で完全に実行し、クラウドに依存しません。会話やデータが端末から外に出ることはありません。

iOS AppiOS 应用iOS アプリ Fully Offline完全离线完全オフライン OpenAI Compatible兼容 OpenAIOpenAI 互換 Built-in Chat内置聊天チャット内蔵

Key Features主要功能主な機能

Fully Offline Inference完全离线推理完全オフライン推論

Once a model is downloaded, everything runs locally. No internet connection required for inference. Your data never leaves your device, and no subscription is needed.模型下载完成后,所有推理均在本地运行。推理过程无需互联网连接。您的数据永远不会离开设备,也不需要任何订阅。モデルをダウンロードすれば、すべてローカルで動作します。推論にインターネット接続は不要です。データが端末から外に出ることはなく、サブスクリプションも不要です。

Built-in Chat Interface内置聊天界面内蔵チャットインターフェース

Start chatting immediately with the integrated chat UI. No external app needed. The built-in interface connects directly to your local server for instant, private conversations.通过内置聊天界面即刻开始对话,无需外部应用。内置界面直接连接本地服务器,实现即时、私密的对话。内蔵チャット UI ですぐに会話を開始できます。外部アプリは不要です。ローカルサーバーに直接接続し、即座にプライベートな会話ができます。

OpenAI-Compatible API兼容 OpenAI 的 APIOpenAI 互換 API

Drop-in replacement for the OpenAI API. Supports chat completions, text completions, streaming (SSE), and model listing. Also compatible with Ollama CLI commands.可直接替代 OpenAI API。支持聊天补全、文本补全、流式传输 (SSE) 和模型列表。同时兼容 Ollama CLI 命令。OpenAI API のドロップイン代替。チャット補完、テキスト補完、ストリーミング (SSE)、モデル一覧に対応。Ollama CLI コマンドとも互換性があります。

TLS & AuthenticationTLS 与身份认证TLS と認証

Secure your server with HTTPS/TLS encryption. Generate self-signed certificates or import your own. Protect access with Bearer token API key authentication and manage multiple keys.使用 HTTPS/TLS 加密保护您的服务器。生成自签名证书或导入您自己的证书。通过 Bearer Token API 密钥认证保护访问,并管理多个密钥。HTTPS/TLS 暗号化でサーバーを保護します。自己署名証明書の生成や独自の証明書のインポートが可能です。Bearer トークン API キー認証でアクセスを保護し、複数のキーを管理できます。

Broad Model Support广泛的模型支持幅広いモデル対応

Supports any GGUF-format model powered by llama.cpp. Run LLaMA, Mistral, Phi, Gemma, Qwen, DeepSeek, and many more. Browse and download directly from Hugging Face with built-in search.支持 llama.cpp 驱动的任何 GGUF 格式模型。可运行 LLaMA、Mistral、Phi、Gemma、Qwen、DeepSeek 等众多模型。内置搜索功能,直接从 Hugging Face 浏览和下载。llama.cpp による GGUF 形式モデルに対応。LLaMA、Mistral、Phi、Gemma、Qwen、DeepSeek など多数のモデルを実行可能。Hugging Face からの検索・ダウンロード機能を内蔵しています。

Metal GPU AccelerationMetal GPU 加速Metal GPU アクセラレーション

Hardware-accelerated inference using Apple Metal. Configure GPU layer offloading, thread count, context sizes up to 32K tokens, and fine-tune temperature, top-p, top-k, and penalties.使用 Apple Metal 进行硬件加速推理。配置 GPU 层卸载、线程数、高达 32K Token 的上下文大小,并精细调节温度、top-p、top-k 和惩罚参数。Apple Metal によるハードウェアアクセラレーション推論。GPU レイヤーオフロード、スレッド数、最大 32K トークンのコンテキストサイズを設定し、温度、top-p、top-k、ペナルティを細かく調整できます。

Network Server网络服务器ネットワークサーバー

Expose your server on the local network for use with any OpenAI-compatible client. Bind to all interfaces, localhost only, or a specific network interface. Configurable port and CORS support.在局域网上公开服务器,供任何兼容 OpenAI 的客户端使用。可绑定到所有接口、仅本地主机或指定网络接口。支持自定义端口和 CORS。ローカルネットワーク上でサーバーを公開し、OpenAI 互換クライアントから利用できます。全インターフェース、ローカルホストのみ、または特定のネットワークインターフェースにバインド可能。ポート設定と CORS に対応。

Smart Resource Management智能资源管理スマートリソース管理

Real-time thermal and memory monitoring. Automatic thread reduction under heat, request rejection at critical temperatures, and conservative memory budgeting to keep your device responsive.实时温度和内存监控。高温时自动减少线程,临界温度时拒绝请求,保守的内存预算策略保持设备流畅响应。リアルタイムの温度・メモリ監視。高温時の自動スレッド削減、臨界温度でのリクエスト拒否、保守的なメモリ管理でデバイスの快適さを維持します。

Developer Tools开发者工具開発者ツール

Live API documentation with curl examples, structured logging with configurable levels, built-in self-test diagnostics, token throughput monitoring (tokens/sec), and request queue management.提供带有 curl 示例的实时 API 文档、可配置级别的结构化日志、内置自检诊断、令牌吞吐量监控(tokens/sec)以及请求队列管理。curl サンプル付きのライブ API ドキュメント、レベル設定可能な構造化ログ、内蔵セルフテスト診断、トークンスループット監視(tokens/sec)、リクエストキュー管理を備えています。

Privacy Policy隐私政策プライバシーポリシー

Last updated: September 2026最后更新:2026 年 9 月最終更新:2026 年 9 月

No Data Collection不收集数据データ収集なし

LLM Server does not collect, store, or transmit any personal data to Linosec or to any analytics service. All AI inference runs entirely on your device. No analytics, no telemetry, no tracking of any kind.LLM Server 不会向 Linosec 或任何分析服务收集、存储或传输任何个人数据。所有 AI 推理完全在您的设备上运行。没有任何分析、遥测或追踪。LLM Server は、Linosec や分析サービスに個人データを収集・保存・送信することは一切ありません。すべての AI 推論はデバイス上で完全に実行されます。分析、テレメトリ、トラッキングは一切ありません。

On-Device Processing设备端处理デバイス上で処理

All AI model inference occurs locally on your device using llama.cpp. Your conversations, prompts, and generated text stay on your device and are not sent to any external server, unless you explicitly enable the optional web search tool described below.所有 AI 模型推理均通过 llama.cpp 在您的设备本地完成。除非您主动启用下文所述的可选网络搜索工具,否则您的对话、提示和生成的文本都会保留在设备上,不会发送至任何外部服务器。すべての AI モデル推論は llama.cpp を使用してデバイス上でローカルに実行されます。下記のオプションのウェブ検索ツールをお客様が明示的に有効にしない限り、会話、プロンプト、生成テキストはデバイス内に留まり、外部サーバーに送信されることはありません。

Internet Usage互联网使用インターネットの使用

Network activity is limited to downloading AI models from Hugging Face and, if you turn it on, the optional web search tool — both initiated solely by you. With web search left off and a model already downloaded, the app operates fully offline. No background network requests are made.网络活动仅限于从 Hugging Face 下载 AI 模型,以及您启用后的可选网络搜索工具——两者均完全由您主动发起。在网络搜索保持关闭且模型已下载的情况下,应用完全离线运行,不会发起任何后台网络请求。ネットワーク通信は、Hugging Face からの AI モデルのダウンロードと、お客様が有効にした場合のオプションのウェブ検索ツールに限られ、いずれもお客様の操作によってのみ開始されます。ウェブ検索を無効のままにし、モデルをダウンロード済みであれば、アプリは完全にオフラインで動作し、バックグラウンドでのネットワークリクエストは行いません。

Optional Web Search可选的网络搜索オプションのウェブ検索

Web search is disabled by default. If you enable a web search tool, the parts of your chat needed to build a query — and, depending on the provider, the content retrieved in response — are transmitted over a TLS-encrypted connection to the third-party search provider you select: Brave Search, Tavily, Ollama Search or Perplexity. Linosec does not operate these services and has no control over how they handle the data they receive. Please refer to the privacy policy of the provider you choose before enabling this feature.网络搜索默认处于关闭状态。如果您启用网络搜索工具,构建查询所需的聊天内容(视提供商而定,还包括返回的搜索结果内容)将通过 TLS 加密连接传输至您所选择的第三方搜索提供商:Brave Search、Tavily、Ollama Search 或 Perplexity。Linosec 不运营这些服务,也无法控制其如何处理所接收的数据。启用此功能前,请查阅您所选提供商的隐私政策。ウェブ検索は既定で無効です。ウェブ検索ツールを有効にすると、クエリの作成に必要なチャットの内容(プロバイダーによっては取得した検索結果の内容も含む)が、お客様が選択した第三者の検索プロバイダー(Brave Search、Tavily、Ollama Search、Perplexity)へ TLS 暗号化通信で送信されます。Linosec はこれらのサービスを運営しておらず、受信したデータの取り扱いを管理することはできません。本機能を有効にする前に、選択したプロバイダーのプライバシーポリシーをご確認ください。

Local Storage本地存储ローカル保存

Models, settings, and API keys are stored locally on your device only. API keys are encrypted using the system keychain. You have complete control over your data and can delete everything at any time.模型、设置和 API 密钥仅存储在您的设备本地。API 密钥通过系统钥匙串加密存储。您可以完全控制自己的数据,随时删除所有内容。モデル、設定、API キーはお使いのデバイス上にのみ保存されます。API キーはシステムキーチェーンで暗号化されます。データを完全に管理でき、いつでもすべてを削除できます。

Permissions权限権限

The app requests access to local files for model import and local network access for serving the API. No camera, microphone, contacts, or location permissions are required.本应用请求访问本地文件以导入模型,以及本地网络访问以提供 API 服务。不需要相机、麦克风、通讯录或位置权限。アプリはモデルのインポートのためのファイルアクセスと、API 提供のためのローカルネットワークアクセスを要求します。カメラ、マイク、連絡先、位置情報の権限は不要です。

Contact联系我们お問い合わせ

If you have any questions or concerns about this Privacy Policy, please contact us:如果您对本隐私政策有任何疑问或顾虑,请联系我们:本プライバシーポリシーに関するご質問やご懸念がありましたら、お気軽にお問い合わせください:

  • info@linosec.lu
  • 96 Avenue Gaston Diderich, L-1420 Luxembourg
  • (+352) 661 140 732

Your Privacy Matters您的隐私至关重要あなたのプライバシーが大切です

LLM Server was designed with privacy as a core principle. All AI inference happens on your device. Internet access is limited to downloading models from Hugging Face and to the optional web search tool, and both are always initiated by you. No data is collected and no analytics are tracked; nothing is shared with third parties other than the search queries you deliberately send to the search provider you selected. You are in complete control.LLM Server 将隐私作为核心设计原则。所有 AI 推理均在您的设备上进行。互联网访问仅限于从 Hugging Face 下载模型以及可选的网络搜索工具,且两者始终由您主动发起。不收集任何数据,不追踪任何分析;除您有意发送至所选搜索提供商的搜索查询外,不与第三方共享任何信息。一切尽在您的掌控之中。LLM Server はプライバシーを基本原則として設計されています。すべての AI 推論はデバイス上で行われます。インターネットへのアクセスは Hugging Face からのモデルダウンロードとオプションのウェブ検索ツールに限られ、いずれも常にお客様の操作によって開始されます。データの収集や分析の追跡は一切なく、お客様が意図的に選択した検索プロバイダーへ送信する検索クエリを除き、第三者と情報を共有することはありません。すべてお客様の管理下にあります。