Sednium Oorty.
One Interface.
Every AI Model.

Run native quantized GGUF models directly on your phone with real-time token telemetry. Auto-save every conversation into an open Obsidian-compatible Markdown vault with YAML frontmatter, execute autonomous MCP tools, and orchestrate top cloud models seamlessly.

100% Offline GGUF
384-D Vector Recall
8+ Cloud AI Providers
.md Obsidian Vault
17:02
5G 100%
Logo
Oorty LOCAL_GGUF • Quick
🟢 RAM: 420MB
Explain StateFlow in Kotlin Jetpack Compose
Qwen 2.5 (0.5B) 28.4 tok/s
val state by viewModel.uiState
    .collectAsStateWithLifecycle()

Text("Result: ${state.data}")
⚡ 0.78s TTFT • 420MB / Safe Vault ✓
---
id: "a1b2c3d4-e5f6"
title: "StateFlow Compose"
model: "Qwen2.5-0.5B"
tags: ["kotlin", "compose"]
---

# StateFlow Compose
## 🤖 Oorty — 17:02 | ~520 tok
StateFlow is an observable holder...
$ mcp::recall_from_vault
query: "Room Database KSP"
[SUCCESS] Score 0.892
🤖 Injected 480 tokens into context.
Gemini 2.5 Flash Google AI Studio
Claude 3.7 Sonnet Anthropic BYOK
GPT-4o OpenAI Direct
DeepSeek R1 Groq 400 tok/s
📎
Ask local model anything...
🎙️

Unifying On-Device & Cloud AI Ecosystems

⚡ Native GGUF (llama.cpp)
Google LiteRT
Google Gemini
Anthropic Claude
OpenAI GPT-4o
Groq LPU
NVIDIA NIM
Sednium Rosette

Built for Modern AI Engineers & Note-Takers

From offline GGUF execution to Markdown databases, every layer is engineered for performance and privacy.

On-Device GGUF Engine

Run open-weights models (Qwen 2.5, Llama 3.2, Gemma 2, Phi-3 Mini) with zero internet connection, real-time tok/s metrics, and SAF storage persistence.

Obsidian Markdown Vault

Every chat is dual-written as standard .md files into Documents/Oorty/chats/ with structured YAML frontmatter, headers, tags, and speed logs.

Dynamic RAM Watchdog

Active hardware memory safety: checks real-time available RAM via ActivityManager.MemoryInfo to prevent system out-of-memory crashes.

Hybrid Semantic AI Recall

384-dimensional vector embeddings rank past conversations using Cosine similarity to automatically inject relevant memory into fresh prompt turns.

MCP Agentic Framework

Connect to Model Context Protocol (MCP) servers and execute multi-step tool calls even on compact on-device models with JSON schema injection.

Biometric App Lock & BYOK

All API keys, model presets, and chat vaults are stored locally on your device with optional hardware biometric fingerprint and face unlock.

RAM & Model Fit Calculator

Select your phone's memory to find the perfect on-device GGUF model that runs smoothly without crashing.

Llama 3.2 (1B-Instruct)

Q4_K_M • ~850 MB VRAM
Comfortable 🟢

Excellent sweet spot for 6GB devices. Delivers 25–35 tok/s decode speed with fast responses and low battery consumption.

Expected Speed ~30 tok/s
Safe Context 2048 Tokens
Agentic Rating ⚡ Limited (Fast)

Your Chats as an Open Markdown Database

Unlike proprietary chat apps that lock your conversations into closed cloud databases, Oorty writes standard .md files straight to your device's Documents/Oorty/chats/ folder.

  • Obsidian Mobile Ready: Point Obsidian to Documents/Oorty/ to search, tag, and link your AI discussions instantly.
  • Structured YAML Frontmatter: Every note contains title, timestamp, tokens, model, and TF-IDF generated topic tags.
  • Zero Vendor Lock-In: Move, backup, sync via Git or Syncthing, or edit in any markdown editor.
📱 Obsidian Mobile Vault View
📁 Documents/Oorty/
📁 chats/
📄 kotlin-stateflow-4fa4.md
📄 room-database-setup-328a.md
📄 system-architecture-9661.md
📁 models/

Frequently Asked Questions

How does on-device GGUF inference work in Oorty?

Oorty integrates a native llama.cpp engine via Kotlin Coroutines. You can download quantized GGUF models directly from HuggingFace within the app or choose custom files from your device storage. Generation streams token-by-token directly in RAM without sending any data to the cloud.

Where are my chats saved and can I view them in Obsidian?

Yes! Every chat session is saved in public storage at Documents/Oorty/chats/ as standard .md Markdown notes with YAML frontmatter. Open the Obsidian mobile app and tap "Open folder as vault" on Documents/Oorty/ to immediately explore all your chats.

What happens if I try to load a model too large for my phone?

Oorty's built-in HardwareChecker evaluates your device's currently available RAM. If the model requires more memory than is safely available, Oorty displays a warning dialog detailing the RAM deficit and suggests lighter 4-bit quantizations to prevent system crashes.

Do I need to pay or buy a subscription to use Oorty?

No! Oorty is 100% free and open-source under the MIT license. For local GGUF models, it runs completely free offline. For cloud models (Gemini, Claude, GPT-4o, Groq), you simply supply your own API keys (BYOK) and pay direct provider rates with zero markup.

Take Control of Your AI Intelligence

Run offline GGUF models, store your knowledge in Obsidian, and orchestrate top cloud LLMs in one unified workspace.