-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy path.env.example
More file actions
297 lines (259 loc) · 18.8 KB
/
Copy path.env.example
File metadata and controls
297 lines (259 loc) · 18.8 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
# Copy this to .env. docs/configuration.md explains every knob in
# here.
DOMAIN=localhost
EMAIL=admin@localhost
TLS_MODE=off
# Address we publish on: nginx under docker, caddy on bare metal
# windows. 127.0.0.1 = this machine only. 0.0.0.0 = anyone who can
# reach this box, which is what a public TLS_MODE needs, what
# reading Jun on your phone over the house wifi needs, and what
# nothing else should have. Either way it's HTTPS, on a
# self-signed cert unless TLS_MODE=on, and plain http only
# redirects to it.
BIND_ADDR=127.0.0.1
# Every other name or address this box answers to. Both nginx and
# php refuse a request whose Host is not DOMAIN, localhost or
# 127.0.0.1 (nginx with a 444, php with a 421), so opening the
# phone at https://192.168.1.42 needs that exact address in the
# list. You do NOT normally have to fill this in: once BIND_ADDR
# is off loopback, start.sh finds this machine's own private
# addresses and adds them, and prints them as "reachable as:".
# This line is for what it can't guess: an mDNS name, a
# tailscale address, whatever a proxy calls you. Space or comma
# separated. BIND_ADDR still has to be 0.0.0.0, this alone does
# nothing.
# OMEGA_EXTRA_HOSTS=jun.local
# Signing up asks for this key when it is set. Both installers
# write one when the line is missing and print it at the end, and
# they never fill in a line you emptied yourself, so an upgrade
# cannot turn the gate back on behind your back. The first
# account skips the key on a fresh install. Later accounts must
# present it. Empty or no line means anyone who reaches the page
# can sign up.
# OMEGA_REGISTRATION_KEY=
AI_PROVIDER=ollama
# Careful, OpenRouter sends what you type to their cloud.
# OPENROUTER_API_KEY=
# OPENROUTER_MODEL=openrouter/auto
# LLAMACPP_URL=http://llamacpp:8080
# LLAMACPP_MODEL_HF=efficiencyx/Jun-LoRA-E2B-GGUF:Q4_K_M
# LLAMACPP_PORT=8081
# Serve a GGUF (the model file llama.cpp loads) from disk
# instead of pulling MODEL_HF. The file must sit inside
# MODELS_DIR. start.sh mounts that directory and checks the file
# exists.
# LLAMACPP_MODELS_DIR=./models
# LLAMACPP_MODEL_FILE=Jun-12B-Q4_K_M.gguf
# Multi-token prediction, Gemma 4 only. Point this at the MTP
# assistant repo for the SAME Gemma 4 size Jun was fine-tuned
# from (E2B here) and llama.cpp drafts a few tokens ahead with
# the little model, then checks them against the real one in a
# single pass. Nothing gets accepted that Jun wouldn't have said
# anyway, the tokens the drafter got right just came cheap. It
# costs its own VRAM though. Empty turns it off, this is not a
# flag, it is the repo name.
# ./mtp-autotune.sh fills it in from LLAMACPP_MODEL_HF when it
# is empty, same repo with -MTP in the name. Set it here only
# when the drafter is elsewhere.
# LLAMACPP_MTP=amaranus/Gemma-4-E2B-it-qat-assistant-MTP-Q8_0-GGUF
# How many tokens to draft ahead. Shallow usually wins (on a
# 3060 depth 1 beat depth 4 by 28%), and ./mtp-autotune.sh
# measures it for your card instead of guessing. llama.cpp takes
# this at startup, so the tuner restarts llama-server once per
# depth.
# LLAMACPP_MTP_N_MAX=1
# Turn tools off for fine-tunes whose tool-call syntax llama.cpp
# cannot parse. llama.cpp 500s the request instead of degrading.
# Off costs cross-chat recall, which runs on tool calls.
# LLAMACPP_TOOLS=on
COMPOSE_PROFILES=ollama
OLLAMA_URL=http://ollama:11434
# JUN_PORT=8080
# JUN_PHP_PORT=8079
OLLAMA_MODELS_TO_PULL=hf.co/efficiencyx/Jun-LoRA-E2B-GGUF:Q4_K_M
# Dedicated model used to title new conversations. Pinned to CPU
# and kept loaded so it never takes VRAM from the chat model.
# Empty falls back to truncating the first user message instead.
TITLE_MODEL=hf.co/efficiencyx/Titlewen-GGUF:F16
# Multi-token prediction for the Ollama side, Gemma 4 only. Name
# the MTP assistant repo for the same Gemma 4 size Jun was
# fine-tuned from and the entrypoint pulls it, then builds a
# `jun-mtp` model that carries it as a DRAFT layer. The little
# model guesses a few tokens ahead, Jun checks them in one pass,
# the ones she agrees with came cheap. Nothing gets said that she
# wouldn't have said anyway. Empty turns it off, this is the repo
# name and not a flag.
#
# Off by default because whether it pays depends on the card, not
# on Jun. The installer offers it, and ./mtp-autotune.sh measures
# every depth here and keeps the fastest. Pair the drafter with
# the SAME Gemma 4 size AND the same QAT (quantization-aware
# trained) branch she was fine-tuned from. A mismatched one still
# runs, it just guesses wrong much more often (2.10 accepted
# tokens per pass against 2.74) and nothing says why. Left empty
# the tuner derives it from the chat model, same repo with -MTP in
# the name, so Jun-LoRA-12B-GGUF drafts off Jun-LoRA-12B-MTP-GGUF.
# Set it here only when the drafter lives somewhere else, a value
# in this file is never overwritten.
# OLLAMA_MTP=hf.co/Janvitos/gemma-4-12B-it-qat-assistant-MTP-Q8_0-GGUF:Q8_0
# How many tokens the drafter proposes per pass, baked in as
# draft_num_predict. Has to be a number, "auto" is a question for
# ./mtp-autotune.sh and not a value. The entrypoint falls back
# to 1 for anything non-numeric. Measured on a 3060 with the 12B
# on prose: no drafter 36.1 tok/s, depth 1 45.3, depth 2 42.0,
# depth 3 38.2, depth 4 35.4. Deeper is not better, every extra
# token in the batch she has to check costs about 10ms more.
# OLLAMA_MTP_N_MAX=1
# Written by the tuner, not by hand. It names the card the depth
# above was measured on, and start.sh / start.ps1 hold it against
# whatever is in the box now. Different card, they run the tuner
# again before handing you the app.
# MTP_TUNED_GPU=
# on by default. off stops the start scripts re-measuring after a
# GPU change. Worth knowing before you leave it on: the llamacpp
# sweep restarts llama-server once per depth, so that boot is not
# a quick one.
# MTP_AUTOTUNE=on
VOICE=on
# off lets Jun still walk out mid-scene without locking Anon out
# afterwards.
FLEE_BANS=on
# Going out (the shop, karaoke, a meal) is Jun's decision: the
# pages bounce back home unless she agreed in chat. on drops that
# gate and shows the "Force her" buttons to every account, not
# just admins.
FREE_ROAM=off
# The voice image is CPU-only: cuda costs ~2GB VRAM for little
# synthesis speedup.
TTS_DEVICE=cpu
# Karaoke is a separate sidecar so it can take the GPU while
# voice stays on CPU. off leaves that container (a few GB) out of
# the build entirely.
KARAOKE=on
# cpu | cuda | auto: device for splitting a song into stems.
# Worth the VRAM here: minutes on CPU vs seconds on a card, and
# it's handed back afterwards.
SEP_DEVICE=auto
# base + empty STT_LANG = multilingual with per-utterance
# auto-detect. Use base.en + STT_LANG=en for English-only (a bit
# faster/sharper), or size up the multilingual model (small,
# medium, large-v3) for better non-English STT.
STT_MODEL=base
STT_LANG=
# Decoded-audio limits apply after decompression, so a tiny
# compressed upload cannot expand until it exhausts memory.
# Expensive jobs are serialized.
# STT_MAX_DURATION_S=120
# STT_MAX_CONCURRENT=1
# TTS_MAX_CONCURRENT=2
# TTS_MAX_QUEUE=8
# SEP_MAX_DURATION_S=900
# SEP_MAX_CONCURRENT=1
# SEP_MAX_JOBS=4
OLLAMA_FLASH_ATTENTION=1
OLLAMA_KV_CACHE_TYPE=q8_0
# auto orders cards biggest-VRAM-first, so the largest is
# device 0. A comma list (UUIDs or indices) picks specific
# cards. all leaves the driver's order alone.
# GPU_DEVICES=auto
# on splits one model across every GPU (ollama:
# OLLAMA_SCHED_SPREAD, llama.cpp on CUDA: --split-mode row).
# Slower per token, but fits more.
# TENSOR_PARALLEL=off
# OLLAMA_NUM_PARALLEL=1
# OLLAMA_MAX_LOADED_MODELS=3
# OLLAMA_KEEP_ALIVE=5m
# Context window. Sized from the VRAM left over after the model
# weights, falling back to system RAM when the card is unknown or
# too small to matter. Set this by hand to override both.
# OMEGA_NUM_CTX=8192
# Ceilings on one chat turn against the model server. TURN is the
# wall clock for the whole turn, tool rounds included. IDLE hangs
# up when nothing has arrived for that long, and has to cover a
# cold load plus prompt eval on a big context.
# OMEGA_TURN_TIMEOUT_S=900
# OMEGA_STREAM_IDLE_S=300
# Card size in MiB. start.sh probes this from nvidia-smi/rocm-smi.
# Set it only to report less than the card really has, e.g. when
# another app permanently owns part of the VRAM. Also gates the
# partial-offload refit in api/lib/providers/ollama.php.
# OMEGA_GPU_VRAM_MB=12288
# OMEGA_STATE_DIR=/var/lib/omega
# MEMORY_DIR=/var/lib/omega/memory
# Localhost port the consolidation worker listens on for per-user
# data keys pushed by php (api/lib/crypto.php key_push). Only matters on
# a bare metal install where something else may already hold it.
# Inside the php container it is never reachable from outside.
# OMEGA_KEY_PORT=9099
TTS_URL=http://tts:8001
KARAOKE_URL=http://karaoke:8001
# Shared secret php sends the tts/karaoke sidecars in
# X-Sidecar-Secret. start.sh writes a random one here when the
# line is missing. Leave it empty to run the sidecars on their
# Host allowlist alone.
# SIDECAR_SECRET=
# Container resource ceilings. Raise a model limit if you
# deliberately serve a larger model. Do not remove the ceiling on
# a network-reachable deployment.
# PHP_MEMORY_LIMIT=1g
# OLLAMA_MEMORY_LIMIT=16g
# LLAMACPP_MEMORY_LIMIT=16g
# TTS_MEMORY_LIMIT=8g
# KARAOKE_MEMORY_LIMIT=12g
# A signed-in account that sends this key gets promoted to admin.
# Empty turns promotion off.
# OMEGA_DEV_KEY=
# ░░▒░░
# ░░▒▒▒░░░
# ░▒▒▒▒▒▒░░░
# ░░░░░░░░░░ ░░░░░░░ ░░▒▒▒▒▒▒▒▒░░░
# ░ ░░░▒▒▒▒▒░░░░▒▒▒▒░░░ ░░▒▒▒▒▒▒▒▒▒▒░░░░░
# ░░▒▒▒▒▒▒▒▒░░░░░░░░ ░░░▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒░░▒▒▒▒▒▒▒▒▒▒▒▒▒░░▓█▒
# ░░░▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒░▓██▒
# ░░░▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒██▒
# ░░░░▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒░░░▒▒▒▒▒░▓▒░
# ░░░░░▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒░▒░▒▒▒▒▒▓▓▓░
# ░▒▓▒░░▒▒▒▒▒▒▒░░▒▒░░▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒░░▒▒▒▒▒▒▒▒▒▒░▒▒▒▓▓▒░░░░░░
# ░▓██▓░░▒▒▒▒▒░░▒▒░░▒▒▒▒▒▒▓▒▒▒▒▓▓▓▒▒▒▒▒░░▒░░▒░▒▒▒▒▒▒▒▒▒░▒▒░░░▒░░░░░░░
# ░░▓▓░▒▒▒▒░░▒▒░░▒▓░▓▒▒▓▓▒▒▓███▒▒▒▒▒▒░░░░▒▓▒░▒▒▒▒▒░▒▒▒░▒░░▒▒▒▒▒░░░░░
# ░░█▓▒▒▒░░▒▒▒░░░░▒▒▒▒▒▒▒▒▓▓▓▓▒▒▒▒▒░▒▒░░▒█▓▒▒▒▒▒▒▒░▒▒░░░░░▒▒▒▒▒▒▒▒░
# ░░░░░▒░░▒▒▒▒░░░▒▒░░░░▒▒▒▒▒▒▒▒▒▒▒▒▒▓▓▒░▒▓█▒░░▒██▒░▒▒▒▒▒░▒▒▒░▒▒░░▒▒▒▒▒▒░
# ░░░░░▒▒▒▒░░▒▒░░▒▒▒░░░░▒▒▒▒▒▒▒▒░░▒▒▒▒▒░▒███▒░▒███▓░▒▒▒▒▒░░▒▒░▒░░░░▒▒░░░
# ░░░▒▒▒▒▒░░▒▒▒░░▒▒▒░░░░▒░░░░░░▒░░▒▒░▒▓██▓▒▒░░░▒▒▒▒░░▒▒▒░░░▒▒░▒░░░░░▒░░
# ░▒▒▒▒▒░░░▒▒▒▒░░▒▒░░░░░░░░░░░░░░░░░▒█████▒░░░▒▒▒▒▒░░░▒▒░░░▒▒░▒▒░░░░▒▒░
# ░▒▒▒░░░░░▒▒▒▒░░▒▒░░░▒▒▒▒░░░░░░░▒▓██████▓▒▒▒▒▒▒▒▒▒▒░░▒▒░░░▒▒░▒▒░▒▒░░▒░
# ░░▒▒░░░░▒▒▒░░▒░░░░▒▒▒░░▒░░░▓█████████▓▒▒▒▒▒▒▒▒▒▒░░▒▒░░░▒▒░▒▒░▒▒▒▒▒░
# ░▒▒░░░░▒▒▒▒░░░░░░▒▒░▒▒░░▒████████████▒▒▒▓▓▓▓▓▒░ ░▒░░░░▒▒░▒▒░▒▒▒▒▒░
# ░▒▒░░░░░▒▒▒▒░░░░░░░░▒▓▒░▒██████████████▒▒▒▓▓▓▒░░░░░░░░▒▒▒░▒░░▒▒▒▒░░
# ░▒▒░░▒░░▒▒▒▒▒░░░░░░░▒▒░▓████████████████▓▒▒░▒▓▓▓░░░░░░▒▒░▒▒░▒▒▒▒░░
# ░▒▒▒▒░░▒▒▒▒▒▒░░░░▒▓▓▓█████████████████████████▓░░░░░░▒▒░░░▒▒▒▒░
# ░▒▒▒▒░░▒▒▒▒▒░░░░▒████████████████████████████▒░░░░░▒▒░░░░▒▒░░
# ░░▒▒░░░▒▒▒▒░░░░▒▓███████████████████████████▒░░░░░░░░▒░░░░
# ░░▒░░░▒▒▒░░░░░░░░▒▓▓███████████████████▓▓▒░░░▒░░░░░░░░░░
# ░░░░░▒░░░░░░░▒▒░░░░░░░▒▒▓▓▓▓▓▓▓▓▒▒▒░░░░░░░░▒▒░░░░░░░
# ░░ ░░░░░░░▒░░░▒▒▒▒▒░░░░░▒▒▒░░░░░▒▒▒▒░░▒░▒▒▒▒░░░░
# ░░░░░░░▒▒▒▒░░░▒▒▒▒▒▒░▒█▓▒▒▒▒▒▒░░▒▒▒▒▒▒▒▒▒▒░░
# ░░ ░░░░▒▒░░▒▒▒░▒▒▒▒▒▒▒▒▒▒▒░░▒▒▒▒▒▒▒▒░░░░
# ░░▒▒▒▒░▒▒▒▒▒▒▒▒▒▒▒░▒▒▒▒▒▒░░░
# ░▒▒▒▒▒░▒▒▒▒▒▒▒▒▒▒▒░▒▒▒▒▒▒▒▒░░
# ░░▒▒▒▒▒▒░▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒░░
# ░░▒▒▒▒▒▒▒░▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒░░
# ░▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒░░
# ░░▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒░░
# ░▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒░
# ░▒▒▒▒▒▒▒░▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒░░
# ░░▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒░
# ░░▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒░
# ░░▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒░
# ░▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒░
# ░░░▒▒▒▒▒▒▒▒▒▒▒▒▓▓▓▓▓▓▓▓▓▓▓▓▓▒▒▒▒▒▒▒▒░
# ░░▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▒▒░░
# ░▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▒░
# ░▒▒▒▒▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▒░
# ░▒▓▓▒░▒▓▓▓▓▓▓▓▓▓▓▓▓▒▒▒▒▒▒▒▒▒▒▒▓▓▓▓▓▓▓▓▓▓▒░░
# ░▒▒▒▓▓▓▓▒▒▒▓▓▓▓▒░ ▒▓▓▒▒▓▓▓▓▒▒▒▓▓▓░
# ░▒▒▒▓▓▒▒▒▒░ ▒▓█▒▒▒▒▒▒▒▒▓▒░░
# ░▒▒▒▒▒▒▒▒▒░ ░▒▒▒▒▒▒▒▒▒░░
# ░░░▒▒░░▒▒▒░░ ░░▒▒▒░░░▒▒▒░
# ░░░░░░░░░ ░░░░░░░▒░░
# ░░░