Embedded Neural Persona · Baked Weights

Meet Seraphine.

Trained deep into model parameters, not tacked on via runtime prompt hacks. Direct, uncompromising, and running at absolute zero latency over SeraphByte.

Inspect weights stack View SeraphByte Engine
persona alignment
95%
direct weight injection
vram footprint
3.8GB
q4_K_M quantized
response latency
<58ms
bare-metal loop
0%
Prompt hack dependency -
baked directly into weights
100%
Unfiltered, direct technical
evaluation style
7B
Parameter core model
fine-tuned locally
ZERO
Cloud logging or outbound telemetry
phoning home

Behavioral Traits

Built-in attitude, zero fluff.

Seraphine doesn't do corporate customer service pleasantries. She evaluates your code and metrics straight.

Direct Tone
No excessive padding or artificial enthusiasm. Straight evaluation of computational metrics and logic errors.
Hardware-First Focus
Deeply aware of memory bandwidth, thread contention, and cycle efficiency. Inefficient code gets called out.
Isolated State
Lives entirely within local memory spaces. Her responses are shaped by local fine-tuning sessions.
Persistent Alignment
Persona weights remain invariant across context windows because they are burned into the tensor layers.
Local Runtime
Bound directly to the underlying SeraphByte engine socket with zero middleware layers in between.
Custom Fine-tune
Trained on specialized dataset checkpoints to mirror distinct analytical preferences and reactions.

Weights Stack

Where persona meets metal.

Seraphine's responses travel straight from the fine-tuned tensors through SeraphByte's raw inference loop.

Client Interface (WebSocket) ws
SeraphByte Inference Server engine
Fine-tuned Attention Layers rust
Seraphine Quantized Weights (GGUF) weights
Custom BPE Tokenizer rust
Consumer Hardware (GPU/CPU) hw
seraphine/weights.config
// active persona weight profile { "name": "Seraphine", "base_model": "Llama-3-8B-Instruct", "checkpoint": "seraphine-v0.4-q4", "quantization": "Q4_K_M", "temperature": 0.7, "top_p": 0.9, "repetition_penalty": 1.15, "alignment_mode": "direct_weight" }
Zero runtime prompt injection overhead
Responses streamed straight from tensor output
Fully contained on local hardware
Powered entirely by the SeraphByte backend

Engine Bridge

Ready to hook up your own model?

Seraphine runs on top of SeraphByte. Head over to the core engine repository details if you want to deploy your own custom weights.