Fork Google's Gemini CLI to create a local-first AI workflow tool that uses GGUF models (Llama, Phi, Qwen, etc.) instead of cloud APIs, maintaining all the sophisticated features while ensuring complete privacy and offline capability.
Current: Gemini API calls New: Native Node.js GGUF model inference using node-llama-cpp
import { getLlama, LlamaChatSession } from 'node-llama-cpp';
interface TrustModelClient {
generateText(prompt: string, options?: GenerationOptions): Promise<string>;
generateStream(
prompt: string,
options?: GenerationOptions,
): AsyncIterable<string>;
loadModel(modelPath: string): Promise<void>;
unloadModel(): Promise<void>;
createChatSession(): Promise<LlamaChatSession>;
}
class TrustNodeLlamaClient implements TrustModelClient {
private llama: any;
private model: any = null;
private context: any = null;
private currentModelPath: string | null = null;
async loadModel(modelPath: string) {
if (this.currentModelPath === modelPath && this.model) {
return; // Model already loaded
}
this.llama = await getLlama();
this.model = await this.llama.loadModel({
modelPath,
...this.getOptimalSettings(),
});
this.context = await this.model.createContext();
this.currentModelPath = modelPath;
}
async createChatSession(): Promise<LlamaChatSession> {
if (!this.context) {
throw new Error('Model not loaded');
}
return new LlamaChatSession({
contextSequence: this.context.getSequence(),
});
}
async generateText(
prompt: string,
options?: GenerationOptions,
): Promise<string> {
const session = await this.createChatSession();
return session.prompt(prompt, {
temperature: options?.temperature ?? 0.7,
maxTokens: options?.maxTokens ?? 2048,
...options,
});
}
async *generateStream(
prompt: string,
options?: GenerationOptions,
): AsyncIterable<string> {
const session = await this.createChatSession();
const response = session.promptWithMeta(prompt, {
temperature: options?.temperature ?? 0.7,
maxTokens: options?.maxTokens ?? 2048,
onTextChunk: (chunk: string) => ({ chunk }),
...options,
});
for await (const result of response) {
if (result.chunk) {
yield result.chunk;
}
}
}
private getOptimalSettings() {
// Auto-detect optimal settings based on system capabilities
const totalRAM = process.memoryUsage().heapTotal;
return {
threads: Math.min(8, require('os').cpus().length),
contextSize: 4096,
// Add GPU detection logic here
};
}
}interface TrustModelConfig {
name: string;
path: string;
type: 'llama' | 'phi' | 'qwen' | 'mistral';
quantization: 'Q4_K_M' | 'Q8_0' | 'FP16';
contextSize: number;
ramRequirement: string;
description: string;
trustScore?: number; // TrustOS feature: community trust rating
verificationHash?: string; // TrustOS feature: integrity verification
}
class TrustModelManager {
listAvailableModels(): TrustModelConfig[];
downloadModel(modelId: string): Promise<void>;
verifyModel(modelPath: string): Promise<boolean>; // TrustOS integrity check
switchModel(modelName: string): Promise<void>;
getCurrentModel(): TrustModelConfig | null;
getTrustRating(modelId: string): Promise<number>; // Community trust score
}# ~/.trustcli/config.yaml
trust:
privacy_mode: strict # TrustOS: no external calls
audit_logging: optional
model_verification: enabled
models:
default: 'llama-3.2-3b-instruct-q4'
directory: '~/.trustcli/models'
auto_verify: true
model_configs:
llama-3.2-3b-instruct-q4:
path: 'models/llama-3.2-3b-instruct-q4_k_m.gguf'
type: 'llama'
context_size: 4096
threads: 4
trust_score: 9.2
phi-3.5-mini-instruct-q8:
path: 'models/phi-3.5-mini-instruct-q8_0.gguf'
type: 'phi'
context_size: 4096
threads: 4
trust_score: 9.5
inference:
temperature: 0.7
top_p: 0.9
max_tokens: 2048
stream: true
transparency:
log_prompts: false # TrustOS: user controls logging
log_responses: false
show_model_info: true
show_performance_metrics: true- Interactive Chat Mode: Keep the conversational interface
- Codebase Analysis: File reading and context building
- Project Workflows: Git integration, PR analysis, etc.
- MCP Server Support: Plugin architecture for tools
- Multimodal Capabilities: Use vision models when available
- Model Switching:
gemini model switch phi-3.5-mini - Resource Monitoring: Show RAM usage, inference speed
- Offline Mode: Complete functionality without internet
- Model Recommendations: Suggest models based on task and hardware
- Performance Profiling: Benchmark different models/quantizations
- Fork Gemini CLI repository
- Replace Google auth with local model management
- Install and configure node-llama-cpp
- Replace API calls with local inference using node-llama-cpp
- Basic model loading and text generation
- Simple configuration system
- Model download/management system using Hugging Face Hub
- Support multiple model formats (Llama, Phi, Qwen, Mistral)
- Model switching functionality (
gemini model switch phi-3.5-mini) - Performance monitoring and resource usage display
- Auto-detection of optimal model settings
- Restore all Gemini CLI features with local models
- Codebase analysis and long context handling
- Git integration and workflow automation
- Streaming responses and chat history
- MCP server compatibility
- Advanced model management UI in terminal
- Hardware optimization recommendations
- Offline documentation and help system
- Performance benchmarking tools
- JSON schema enforcement for structured outputs
Problem: Local models are slower than cloud APIs Solution:
- Implement smart caching for repeated queries
- Use smaller models for simple tasks, larger for complex ones
- Parallel processing for batch operations
Problem: Large models consume significant RAM Solution:
- Model swapping based on task requirements
- Quantization recommendations based on available RAM
- Clear memory management with model unloading
Problem: Different models have different chat formats Solution:
- Abstraction layer for chat templates
- Auto-detection of model type and format
- Fallback to universal formats
# Install TrustCLI alongside Gemini CLI
npm install -g @trustos/cli
# Different binary names - no conflicts
gemini --help # Google's Gemini CLI
trust --help # TrustCLI
# Can use in same project
cd my-project/
gemini "Analyze this codebase" # Uses cloud Gemini
trust "Analyze this codebase" # Uses local model (private/trusted)
# Hybrid workflows - best of both worlds
gemini "Generate test plan" > plan.md
trust "Review this test plan" --file plan.md --private~/.gemini/ # Google Gemini CLI config
~/.trustcli/ # TrustCLI config directory
├── config.yaml # Main configuration
├── models/ # Downloaded models
├── cache/ # Model cache
├── logs/ # Audit logs (optional)
└── trust/ # Trust scores and verification datalocal-gemini setup
# Guides through:
# 1. Model directory setup
# 2. Hardware detection and recommendations
# 3. First model download (suggests appropriate size)
# 4. Basic configuration# Basic usage - same patterns as Gemini CLI
trust "Explain this error in main.py"
# Model management
trust model list
trust model download microsoft/Phi-3.5-mini-instruct-gguf
trust model switch phi-3.5-mini
trust model verify phi-3.5-mini # TrustOS: verify model integrity
trust model benchmark
# Advanced features
trust --model llama-3.2-3b "Code review this PR"
trust --offline "Help with this script"
trust --format json "Extract data from this log"
trust --temperature 0.1 "Write unit tests"
trust --private "Analyze sensitive code" # Explicit privacy mode
# Server mode for integrations
trust serve --port 8080 # OpenAI-compatible API
trust serve --unix-socket /tmp/trust.sock
# Performance monitoring & transparency
trust stats
trust benchmark --all-models
trust recommend-model --task coding --ram 16gb
trust audit --show-logs # TrustOS: transparency features
trust trust-score llama-3.2-3b # Community trust rating- Phi-3.5-mini-instruct (3.8B params, Q4_K_M): Fast coding assistance
- Qwen2.5-1.5B-Instruct (1.5B params, Q8_0): Quick questions
- Llama-3.2-3B-Instruct (3B params, Q8_0): Balanced performance
- CodeQwen1.5-7B-Chat (7B params, Q4_K_M): Code-specialized
- Llama-3.1-8B-Instruct (8B params, Q8_0): High-quality responses
- DeepSeek-Coder-7B-Instruct (7B params, Q8_0): Advanced coding
- Model sharing: Community-curated model recommendations
- Plugin marketplace: Extensible tool integrations
- Educational content: Guides for different use cases
- Enterprise adoption: Clear migration paths from cloud CLIs
- Better than Ollama CLI: More sophisticated workflows and integrations
- Better than cloud CLIs: Privacy, cost, and offline capabilities
- Better than basic LLM tools: Rich feature set and professional tooling
- Unique hybrid approach: Best of both local and cloud workflows
- Zero-config startup: Works out of the box with sensible defaults
- Graceful degradation: Fallback strategies when resources are limited
- Error recovery: Robust handling of model loading failures
- Observability: Rich logging and debugging capabilities
- Model agnostic: Support for multiple inference backends
- Format flexibility: Adapt to new model formats as they emerge
- Cloud bridge: Optional hybrid mode with cloud APIs as fallback
- Enterprise support: Clear path to commercial support if needed
- TypeScript/Node.js: Pure Node.js stack for CLI framework and model inference
- node-llama-cpp: Native Node.js bindings for GGUF model inference (no Python)
- Comprehensive testing: Multi-model testing across different hardware configurations
- Documentation: Model compatibility guides and TrustOS integration docs
- Performance benchmarking: Automated benchmarking suite for model comparison
- TrustOS standards: Follow TrustOS framework guidelines for transparency and trust
- GitHub Stars: 10K+ within first year
- Weekly Downloads: 50K+ npm downloads
- Community Engagement: 500+ contributors, active Discord/forum
- Enterprise Adoption: 100+ companies using in production
- Feature Parity: 95%+ of Gemini CLI features working locally
- Performance: <10s response time for typical queries
- Resource Efficiency: Runs smoothly on 8GB+ laptops
- Setup Experience: Complete setup in <5 minutes
- Model Ecosystem: Support for 25+ popular models
- Reliability: 99.9% uptime for local inference
- User Satisfaction: 4.5+ stars on npm/GitHub
- Documentation: Comprehensive guides for all use cases
- Support Response: <24h response time for issues