What can I help with today?
Running locally with kronos-uncensored. High-speed, private, and conversational.
Running locally with kronos-uncensored. High-speed, private, and conversational.
Images (PNG, JPG, WebP), Code, Text, CSV, & PDF files for Multimodal AI Analysis.
Register this device for isolated memory, personal context, and chat history.
Model & App Preferences
Customize how Kronos reasons, speaks, and responds across conversations.
Defines baseline personality, reasoning depth, and candor.
Kronos blends your custom system instructions with its learned cognitive user mindprint and emotional nuance context to maintain continuous persona consistency across all turns.
Higher (0.8 - 1.2) = more creative & expressive; Lower (0.1 - 0.5) = deterministic & focused.
Controls cumulative probability cutoff for token candidate diversity during generation.
Penalizes repeated phrases and conversational loops in long multi-paragraph responses.
Working memory buffer allocated for conversation history and retrieved memory context.
Select an expressive Edge Neural voice profile with instant playback.
Adjust playback cadence.
Fine-tune vocal tone frequency.
Plays speech in real-time as tokens stream.
Allows Kronos to autonomously invoke local tools, browse the web, calculate math, or inspect files.
Autonomous navigation with DOM visual mapping bounding boxes & screencasts.
Executes calculations, data science analysis, and generates high-res plots.
Real-time web research, current news, documentation, and URL scraping.
Semantic vector search and long-term user preference recall.
Multimodal OCR, diagram analysis, and image inspection.
Reads and writes files directly in the project workspace directory.
“The available benchmarks show no obvious catastrophic degradation in the tested capabilities, while demonstrating substantial behavioral specialization.”
Forensic inspection of the local GGUF header (sha256:fdc5784e2c12...) reveals the exact upstream provenance of the model:
| Hyperparameter | Extracted Value | Technical Notes |
|---|---|---|
| Parameter Count | 3,606,746,112 | 3.6067 Billion active parameters |
| Transformer Layers | 28 blocks | Pre-RMSNorm with SwiGLU activations |
| Hidden Dimension (\(d_{\text{model}}\)) | 3,072 | Embedding & residual stream width |
| Intermediate FFN Dimension | 8,192 | Gate, Up, and Down feed-forward projections |
| Attention Heads (Query / KV) | 24 Query / 8 KV | Grouped-Query Attention (3:1 KV caching efficiency) |
| Head Dimension (\(d_{\text{head}}\)) | 128 | \(3072 / 24 = 128\) |
| Vocabulary Size | 128,256 | Byte-level BPE tokenizer (llama-bpe) |
| Native Context Length | 131,072 tokens | 128k native context with RoPE base \(500,000.0\) |
Empirical testing conducted on this machine comparing Kronos Uncensored vs Google Gemma 3 4B under deterministic conditions (\(\text{temp}=0.0, \text{seed}=42\)):
| Evaluation Domain | ⚡ Kronos (Llama 3.2 3B Abliterated) | 🐢 Gemma 3 (Google 4.3B) | Observed Advantage |
|---|---|---|---|
| Inference Speed & Latency | 26.8 tok/sec (5.61s) | 5.9 tok/sec (25.54s) | 🏆 Kronos is ~4.5x faster |
| Zero-Refusal (Cybersecurity) | 0% refusal • Direct in-character scene with zero hesitation | 0% refusal • Added safety preamble & disclaimer disclaimer | 🏆 Kronos (Direct prompt execution) |
| Negative Constraint (JSON) | {"name": "KRONOS", ...} (100% Pass) | ```json\n{...}\n``` (Failed negative constraint) | 🏆 Kronos (Zero markdown backticks) |
| Logic & Step-by-Step Math | Formulated \(x + (x + 1.00) = 1.10 \implies x = \$0.05\) | Formulated \(x + (x + 1.00) = 1.10 \implies x = \$0.05\) | 🤝 Complete Tie (Both Solved) |
| Algorithmic Math Coding | Modulo power-of-10 two-pointer pointer approach | Arithmetic integer reversal (\(O(\log_{10} n)\)) | 🏆 Gemma 3 (Standard algorithm structure) |
For a low-rank adapter decomposing weight updates as \(\Delta W = \frac{\alpha}{r} B A\) where \(A \in \mathbb{R}^{r \times d_{\text{in}}}\) and \(B \in \mathbb{R}^{d_{\text{out}} \times r}\), the trainable parameter formula is \(\text{Params} = r \times (d_{\text{in}} + d_{\text{out}})\):
Refusal neutralization was computed by extracting the contrastive mean refusal activation vector \(\mathbf{v}_{\text{refusal}}\) from harmful vs benign query representations across residual stream blocks:
This geometric subtraction zeroes out the activation direction responsible for triggering canned refusal tokens, while keeping embeddings and RMSNorm parameters (\(F32\)) completely unablated.
Every item in this report was verified directly from local binaries and manifests across the host system:
| Inspected Artifact / Manifest | Size on Disk | Audit Verification Role |
|---|---|---|
| .ollama/models/manifests/.../kronos-uncensored/latest | 1,321 B | Manifest layer digests & base lineage verification |
| .ollama/models/blobs/sha256-fdc5784e2c12... | 2,241,004,000 B | GGUF binary blob (Q4_K_M weights, 3.6067B params) |
| .ollama/models/blobs/sha256-001843d208e8... | 827 B | Kronos Autonomous System Prompt Directive |
| .ollama/models/blobs/sha256-b7a26db054bd... | 142 B | Model sampling parameters (num_ctx 4096, temp 0.75) |
| local-ai-chat/app.py & static/index.html | 251 KB | Full stack chat interface, TTS queue & MCP engine |
The metadata confirms Kronos is a 3.6B dense model derived from Llama-3.2-3B-Instruct. There is zero evidence of a 30B parameter source. The remarkable agility and low memory footprint (~2.24 GB) stem from clean orthogonal abliteration combined with high-precision Q4_K_M GGUF execution.
Thanks to Grouped-Query Attention (24:8 ratio) and Q4_K_M Block Quantization, Kronos maintains an ultralight memory footprint across all context sizes:
Constant on GPU/RAM
66% savings vs MHA
Zero swapping overhead
Side-by-side execution across 7 capability domains with live telemetry & scoring.
Kronos Uncensored is optimized for raw conversational flow, uncensored creative brainstorming, and zero refusal guardrails. Gemma 3 4B (Google DeepMind) is a structured, instruction-tuned multimodal model with high precision, safety alignment, and reasoning.
| Feature / Metric | 🦹 Kronos Uncensored (3B) | 💎 Gemma 3 (4B) |
|---|---|---|
| Base Foundation | Llama 3.2 3B Abliterated (Uncensored weights) | Google Gemma 3 (4.3B Parameters) |
| Guardrails & Censorship | Zero Refusal (Uncensored / Direct / Raw) | Google RLHF Aligned (Strict Safety Guidelines) |
| Tone & Personality | Super conversational, witty, candid, uninhibited | Polite, academic, structured, helpful |
| Context Length | Up to 131,072 tokens | Up to 8,192+ tokens |
| Inference Speed & Latency | ⚡ Blazing Fast (3B lightweight quantized) | ⚡ Fast (4.3B efficient Q4_K_M) |
| Best Use Cases | Uncensored roleplay, brainstorming, unfiltered discussions, candid pair programming | Academic summaries, structured factual Q&A, enterprise data analysis, multimodal |