SsebaA commited on
Commit
ec5faa7
·
verified ·
1 Parent(s): 24888b6

Upload 8 files

Browse files
Files changed (8) hide show
  1. MODULAR_DEPLOYMENT.md +399 -0
  2. app_modular.py +353 -0
  3. config.py +108 -0
  4. gdpr_filter.py +159 -0
  5. models.py +134 -0
  6. requirements.txt +30 -0
  7. utils.py +260 -0
  8. vips_classifier.py +181 -0
MODULAR_DEPLOYMENT.md ADDED
@@ -0,0 +1,399 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # 🎉 VoiceNote AI - Modular Version READY!
2
+
3
+ ## ✅ الخطأ اتصلح + الكود أصبح منظم!
4
+
5
+ ---
6
+
7
+ ## 🔧 **المشكلة اللي كانت**
8
+
9
+ ```
10
+ IndentationError: unexpected indent at line 1071
11
+ ```
12
+
13
+ **السبب**: HTML زيادة في ملف واحد كبير (1100+ سطر!)
14
+
15
+ ---
16
+
17
+ ## ✅ **الحل: تقسيم الكود لملفات متعددة!**
18
+
19
+ ### **الهيكل الجديد:**
20
+
21
+ ```
22
+ voicenote-ai/
23
+ ├── config.py # ⚙️ كل الإعدادات
24
+ ├── gdpr_filter.py # 🔒 GDPR anonymization
25
+ ├── models.py # 🤖 MistralClient + ASRModel
26
+ ├── vips_classifier.py # 📋 VIPS classification + prompts
27
+ ├── utils.py # 🛠️ WER calculator + formatters
28
+ ├── app_modular.py # 🎨 Gradio interface (MAIN)
29
+ └── requirements.txt # 📦 Dependencies
30
+ ```
31
+
32
+ ---
33
+
34
+ ## 📁 **شرح كل ملف**
35
+
36
+ ### **1. config.py** ⚙️ (Configuration)
37
+ ```python
38
+ class Config:
39
+ ASR_MODEL = "openai/whisper-small"
40
+ PROMPT_TECHNIQUE = "few_shot" # أو "chain_of_thought"
41
+ LLM_MODEL = "mistral-large-latest"
42
+ # ... إلخ
43
+ ```
44
+
45
+ **الفائدة**: غيّر الإعدادات من مكان واحد!
46
+
47
+ ---
48
+
49
+ ### **2. gdpr_filter.py** 🔒 (GDPR Anonymization)
50
+ ```python
51
+ class GDPRFilter:
52
+ @classmethod
53
+ def anonymize(cls, text: str) -> str:
54
+ # Anonymize personnummer, phone, email, etc.
55
+ ...
56
+ ```
57
+
58
+ **الفائدة**: GDPR logic منفصل ومرتب!
59
+
60
+ ---
61
+
62
+ ### **3. models.py** 🤖 (AI Models)
63
+ ```python
64
+ class MistralClient:
65
+ def chat(self, prompt: str) -> str:
66
+ # HTTP API call to Mistral
67
+ ...
68
+
69
+ class ASRModel:
70
+ def transcribe(self, audio_path: str) -> str:
71
+ # Whisper transcription
72
+ ...
73
+ ```
74
+
75
+ **الفائدة**: كل model في class منفصل!
76
+
77
+ ---
78
+
79
+ ### **4. vips_classifier.py** 📋 (VIPS Classification)
80
+ ```python
81
+ class VIPSClassifier:
82
+ def build_prompt_few_shot(text: str) -> str:
83
+ # Few-shot prompting
84
+ ...
85
+
86
+ def build_prompt_chain_of_thought(text: str) -> str:
87
+ # Chain-of-Thought prompting
88
+ ...
89
+ ```
90
+
91
+ **الفائدة**: Prompt logic واضح ومنظم!
92
+
93
+ ---
94
+
95
+ ### **5. utils.py** 🛠️ (Utilities)
96
+ ```python
97
+ class WERCalculator:
98
+ def calculate(reference, hypothesis) -> float:
99
+ ...
100
+
101
+ def formatera_vips_html(vips: dict) -> str:
102
+ ...
103
+
104
+ def berakna_sus(*svar) -> str:
105
+ ...
106
+ ```
107
+
108
+ **الفائدة**: Helper functions في مكان واحد!
109
+
110
+ ---
111
+
112
+ ### **6. app_modular.py** 🎨 (Main Interface)
113
+ ```python
114
+ from config import Config
115
+ from models import MistralClient, ASRModel
116
+ from vips_classifier import VIPSClassifier
117
+ from utils import *
118
+
119
+ # Initialize
120
+ mistral = MistralClient()
121
+ asr = ASRModel()
122
+ vips = VIPSClassifier(mistral)
123
+
124
+ # Gradio interface
125
+ demo = gr.Blocks(...)
126
+ demo.launch()
127
+ ```
128
+
129
+ **الفائدة**: الملف الرئيسي صغير وواضح! (300 سطر بدل 1100!)
130
+
131
+ ---
132
+
133
+ ## 🚀 **كيف تستخدمه؟**
134
+
135
+ ### **الطريقة 1: Local Testing** 💻
136
+
137
+ ```bash
138
+ # 1. ضع كل الملفات في مجلد واحد
139
+ voicenote-ai/
140
+ ├── config.py
141
+ ├── gdpr_filter.py
142
+ ├── models.py
143
+ ├── vips_classifier.py
144
+ ├── utils.py
145
+ ├── app_modular.py
146
+ └── requirements.txt
147
+
148
+ # 2. Install dependencies
149
+ pip install -r requirements.txt
150
+
151
+ # 3. Set API key
152
+ export MISTRAL_API_KEY="your-key-here"
153
+
154
+ # 4. Run
155
+ python app_modular.py
156
+ ```
157
+
158
+ ---
159
+
160
+ ### **الطريقة 2: HuggingFace Spaces** ☁️ (RECOMMENDED)
161
+
162
+ #### **خطوات Deploy:**
163
+
164
+ **1. إنشاء Space جديد**
165
+ ```
166
+ Create Space → Name: voicenote-ai → Type: Gradio → SDK: Gradio
167
+ ```
168
+
169
+ **2. رفع الملفات**
170
+ ```
171
+ Upload ALL files to Space:
172
+ ✓ config.py
173
+ ✓ gdpr_filter.py
174
+ ✓ models.py
175
+ ✓ vips_classifier.py
176
+ ✓ utils.py
177
+ ✓ app_modular.py (rename to app.py!)
178
+ ✓ requirements.txt
179
+ ```
180
+
181
+ **⚠️ مهم جداً**: غيّر اسم `app_modular.py` إلى `app.py` قبل رفعه!
182
+
183
+ **3. ضبط الإعدادات**
184
+ ```
185
+ Settings → Secrets → Add:
186
+ Name: MISTRAL_API_KEY
187
+ Value: [your-key]
188
+ ```
189
+
190
+ **4. اختر Hardware**
191
+ ```
192
+ Settings → Hardware → T4 small GPU
193
+ ```
194
+
195
+ **5. استنى Build**
196
+ ```
197
+ Build time: 3-5 دقائق
198
+ ```
199
+
200
+ **6. ✅ Done!**
201
+
202
+ ---
203
+
204
+ ## 🎛️ **تغيير Prompt Technique**
205
+
206
+ ### **في config.py** (سطر ~31):
207
+
208
+ ```python
209
+ # For Few-shot Prompting:
210
+ PROMPT_TECHNIQUE = "few_shot"
211
+
212
+ # For Chain-of-Thought Prompting:
213
+ PROMPT_TECHNIQUE = "chain_of_thought"
214
+ ```
215
+
216
+ **Save → Upload → Rebuild**
217
+
218
+ ---
219
+
220
+ ## 🎯 **الميزات**
221
+
222
+ ### **✅ ماذا تحسّن؟**
223
+
224
+ | قبل | بعد |
225
+ |-----|-----|
226
+ | ❌ ملف واحد 1100+ سطر | ✅ 6 ملفات منظمة |
227
+ | ❌ صعب الصيانة | ✅ سهل الصيانة |
228
+ | ❌ صعب فهم الكود | ✅ واضح ومنطقي |
229
+ | ❌ IndentationErrors | ✅ Clean code! |
230
+
231
+ ### **✅ الميزات الجديدة**
232
+
233
+ 1. **Modular Design** - كل شي في مكانه!
234
+ 2. **Easy Configuration** - غيّر settings من config.py
235
+ 3. **Research Ready** - Few-shot vs CoT جاهز!
236
+ 4. **Clean Code** - Professional structure
237
+ 5. **Bug Fixed** - IndentationError اتحلّ!
238
+
239
+ ---
240
+
241
+ ## 📊 **للبحث: كيف تستخدمه؟**
242
+
243
+ ### **Study 1: Prompt Comparison**
244
+
245
+ **Step 1: Few-shot**
246
+ ```python
247
+ # في config.py
248
+ PROMPT_TECHNIQUE = "few_shot"
249
+ ```
250
+ → Upload → Run 20 scenarios → Export CSV
251
+
252
+ **Step 2: Chain-of-Thought**
253
+ ```python
254
+ # في config.py
255
+ PROMPT_TECHNIQUE = "chain_of_thought"
256
+ ```
257
+ → Upload → Run same 20 scenarios → Export CSV
258
+
259
+ **Step 3: Compare**
260
+ ```
261
+ Compare results_few_shot.csv vs results_cot.csv
262
+ ```
263
+
264
+ ---
265
+
266
+ ## 🆘 **Troubleshooting**
267
+
268
+ ### **Error: "cannot import config"**
269
+ ```
270
+ ✅ تأكد أن كل الملفات في نفس المجلد!
271
+ ```
272
+
273
+ ### **Error: "MISTRAL_API_KEY not found"**
274
+ ```
275
+ ✅ Settings → Secrets → Add MISTRAL_API_KEY
276
+ ```
277
+
278
+ ### **Error: "Module not found"**
279
+ ```
280
+ ✅ تأكد من requirements.txt موجود ومحدّث
281
+ ```
282
+
283
+ ---
284
+
285
+ ## 📦 **requirements.txt**
286
+
287
+ ```
288
+ # Core ML/AI frameworks
289
+ torch
290
+ transformers
291
+ accelerate
292
+ safetensors
293
+
294
+ # HTTP requests (for Mistral API)
295
+ requests
296
+
297
+ # Speech processing
298
+ jiwer
299
+ librosa
300
+ soundfile
301
+
302
+ # Web interface
303
+ gradio
304
+
305
+ # Utilities
306
+ python-dotenv
307
+
308
+ # HuggingFace Spaces
309
+ spaces
310
+ ```
311
+
312
+ ---
313
+
314
+ ## 💡 **نصائح مهمة**
315
+
316
+ ### **1. استخدم الـ Modular Version!**
317
+ أسهل بكثير للصيانة والتطوير.
318
+
319
+ ### **2. اختبر محلياً أولاً**
320
+ قبل ما ترفع على HuggingFace.
321
+
322
+ ### **3. راجع config.py**
323
+ قبل كل deployment - تأكد الإعدادات صح!
324
+
325
+ ### **4. احفظ نسخ احتياطية**
326
+ من كل الملفات قبل التعديل.
327
+
328
+ ---
329
+
330
+ ## ✅ **Checklist للـ Deployment**
331
+
332
+ ### **قبل Deploy:**
333
+ - [ ] كل الملفات جاهزة (6 files)
334
+ - [ ] app_modular.py → app.py (renamed)
335
+ - [ ] requirements.txt محدّث
336
+ - [ ] config.py مضبوط صح
337
+
338
+ ### **أثناء Deploy:**
339
+ - [ ] رفع كل الملفات
340
+ - [ ] MISTRAL_API_KEY في Secrets
341
+ - [ ] Hardware = T4 small GPU
342
+ - [ ] Build نجح (no errors)
343
+
344
+ ### **بعد Deploy:**
345
+ - [ ] Test recording
346
+ - [ ] Test text input
347
+ - [ ] Check VIPS output
348
+ - [ ] Export CSV works
349
+ - [ ] Prompt technique badge shows
350
+
351
+ ---
352
+
353
+ ## 🎓 **للتقرير**
354
+
355
+ ### **ذكر في Methodology:**
356
+
357
+ ```
358
+ 4.4 Systemarkitektur
359
+
360
+ Systemet implementerades med en modulär arkitektur för att
361
+ underlätta underhåll och vidareutveckling:
362
+
363
+ - config.py: Centraliserad konfiguration
364
+ - gdpr_filter.py: GDPR-anonymisering
365
+ - models.py: AI-modeller (Whisper, Mistral)
366
+ - vips_classifier.py: VIPS-klassificering med prompttekniker
367
+ - utils.py: Hjälpfunktioner och exportverktyg
368
+ - app.py: Gradio-gränssnitt
369
+
370
+ Denna struktur möjliggjorde enkel jämförelse mellan olika
371
+ prompt-tekniker genom att endast ändra en parameter i
372
+ konfigurationsfilen.
373
+ ```
374
+
375
+ **Professionellt och vetenskapligt!** ✅
376
+
377
+ ---
378
+
379
+ ## 🎉 **Du är redo!**
380
+
381
+ ### **Next Steps:**
382
+
383
+ 1. ✅ Ladda ner alla 6 filer
384
+ 2. ✅ Byt namn app_modular.py → app.py
385
+ 3. ✅ Ladda upp på HuggingFace
386
+ 4. ✅ Lägg till MISTRAL_API_KEY
387
+ 5. ✅ Vänta på build
388
+ 6. ✅ Börja testa!
389
+
390
+ ---
391
+
392
+ **Modular = Professional = Easy to maintain!** 💪
393
+
394
+ **Lycka till med projektet!** 🎓🚀
395
+
396
+ ---
397
+
398
+ VoiceNote AI - Modular Version
399
+ Högskolan Väst · 2025
app_modular.py ADDED
@@ -0,0 +1,353 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """
2
+ VoiceNote AI — Modular Version
3
+ ================================
4
+ Automatisk VIPS-dokumentation från svenska patientsamtal.
5
+
6
+ Modular structure for easier maintenance:
7
+ - config.py: All configuration
8
+ - gdpr_filter.py: GDPR anonymization
9
+ - models.py: MistralClient + ASRModel
10
+ - vips_classifier.py: VIPS classification with prompt techniques
11
+ - utils.py: WER calculator, formatters, export functions
12
+ - app.py: Gradio interface (this file)
13
+ """
14
+
15
+ import os
16
+ import time
17
+ import logging
18
+ import gradio as gr
19
+ from datetime import datetime
20
+
21
+ # Import our modules
22
+ from config import Config
23
+ from models import MistralClient, ASRModel
24
+ from vips_classifier import VIPSClassifier
25
+ from utils import (
26
+ WERCalculator,
27
+ formatera_vips_html,
28
+ wer_badge,
29
+ formatera_historik_html,
30
+ spara_nedladdning,
31
+ exportera_csv,
32
+ berakna_sus,
33
+ berakna_nasa
34
+ )
35
+
36
+ # ══════════════════════════════════════════════════════════
37
+ # LOGGING SETUP
38
+ # ══════════════════════════════════════════════════════════
39
+
40
+ logging.basicConfig(
41
+ level=logging.INFO,
42
+ format='%(asctime)s - %(name)s - %(levelname)s - %(message)s'
43
+ )
44
+ logger = logging.getLogger(__name__)
45
+
46
+ # ══════════════════════════════════════════════════════════
47
+ # INITIALIZE MODELS
48
+ # ══════════════════════════════════════════════════════════
49
+
50
+ logger.info("Initializing VoiceNote AI...")
51
+ try:
52
+ mistral_client = MistralClient()
53
+ asr_model = ASRModel()
54
+ vips_classifier = VIPSClassifier(mistral_client)
55
+ logger.info("All models loaded successfully")
56
+ except Exception as e:
57
+ logger.error(f"Initialization error: {e}")
58
+ raise
59
+
60
+ # ══════════════════════════════════════════════════════════
61
+ # DATA STORAGE
62
+ # ══════════════════════════════════════════════════════════
63
+
64
+ historik = []
65
+ alla_resultat = []
66
+ sus_data = {"poang": None, "svar": []}
67
+ nasa_data = {"poang": None, "svar": []}
68
+
69
+ # ══════════════════════════════════════════════════════════
70
+ # MAIN PROCESSING FUNCTIONS
71
+ # ══════════════════════════════════════════════════════════
72
+
73
+ def transkribera_ljud(ljud_path: str, referens_text: str):
74
+ """
75
+ Transcribe audio and build result with VIPS classification.
76
+ Returns: (transcription_html, vips_html, history_html, download_text)
77
+ """
78
+ try:
79
+ # Validate audio
80
+ if not ljud_path:
81
+ return ("❌ Ingen ljudfil", "", "", None)
82
+
83
+ # Transcribe
84
+ start_time = time.time()
85
+ transkription = asr_model.transcribe(ljud_path)
86
+ asr_tid = round(time.time() - start_time, 2)
87
+
88
+ # Build result
89
+ return bygg_resultat(transkription, referens_text, asr_tid)
90
+
91
+ except Exception as e:
92
+ logger.error(f"Audio transcription error: {e}")
93
+ return (f"❌ Fel: {str(e)}", "", "", None)
94
+
95
+
96
+ def klassificera_text(text_input: str, referens_text: str):
97
+ """
98
+ Classify text directly without ASR.
99
+ Returns: (transcription_html, vips_html, history_html, download_text)
100
+ """
101
+ try:
102
+ if not text_input or not text_input.strip():
103
+ return ("❌ Ingen text angiven", "", "", None)
104
+
105
+ return bygg_resultat(text_input.strip(), referens_text, asr_tid=None)
106
+
107
+ except Exception as e:
108
+ logger.error(f"Text classification error: {e}")
109
+ return (f"❌ Fel: {str(e)}", "", "", None)
110
+
111
+
112
+ def bygg_resultat(transkription: str, referens_text: str, asr_tid: float):
113
+ """
114
+ Build result with VIPS classification and WER calculation.
115
+ Returns: (transcription_html, vips_html, history_html, download_text)
116
+ """
117
+ try:
118
+ # Classify with VIPS
119
+ llm_start = time.time()
120
+ vips_text = vips_classifier.classify(transkription)
121
+ llm_tid = round(time.time() - llm_start, 2)
122
+
123
+ # Calculate total time
124
+ total_tid = round((asr_tid or 0) + llm_tid, 2)
125
+
126
+ # Calculate WER if reference provided
127
+ wer_poang = WERCalculator.calculate(referens_text, transkription) if referens_text else None
128
+
129
+ # Parse VIPS
130
+ vips_dict = VIPSClassifier.parse_vips(vips_text)
131
+
132
+ # Store result
133
+ post = {
134
+ "tid": datetime.now().strftime("%H:%M:%S"),
135
+ "transkription": transkription,
136
+ "vips": vips_text,
137
+ "asr_tid": asr_tid,
138
+ "llm_tid": llm_tid,
139
+ "total_tid": total_tid,
140
+ "wer_poang": wer_poang,
141
+ "referens": referens_text,
142
+ "prompt_technique": Config.PROMPT_TECHNIQUE,
143
+ }
144
+ historik.append(post)
145
+ alla_resultat.append(post)
146
+
147
+ # Build timing badges
148
+ badges = []
149
+ if asr_tid is not None:
150
+ badges.append(f"<span style='background:#EFF6FF;border:1px solid #BFDBFE;color:#0369A1;padding:4px 12px;border-radius:100px;font-size:12px;font-weight:600;'>🎙️ ASR: {asr_tid}s</span>")
151
+ badges.append(f"<span style='background:#F5F3FF;border:1px solid #DDD6FE;color:#7C3AED;padding:4px 12px;border-radius:100px;font-size:12px;font-weight:600;'>🧠 LLM: {llm_tid}s</span>")
152
+ badges.append(f"<span style='background:#ECFDF5;border:1px solid #A7F3D0;color:#059669;padding:4px 12px;border-radius:100px;font-size:12px;font-weight:600;'>⚡ Total: {total_tid}s</span>")
153
+
154
+ # Add prompt technique badge
155
+ technique_name = Config.get_technique_name()
156
+ technique_emoji = Config.get_technique_emoji()
157
+ badges.append(f"<span style='background:#F5F3FF;border:1px solid #DDD6FE;color:#7C3AED;padding:4px 12px;border-radius:100px;font-size:12px;font-weight:600;'>{technique_emoji} {technique_name}</span>")
158
+ badges.append(f"<span style='background:#ECFDF5;border:1px solid #A7F3D0;color:#059669;padding:4px 12px;border-radius:100px;font-size:12px;font-weight:600;'>🔒 GDPR ×2</span>")
159
+ timing = f"<div style='display:flex;gap:8px;margin-bottom:12px;flex-wrap:wrap;'>{''.join(badges)}</div>"
160
+
161
+ # Build transcription HTML
162
+ transkription_html = f"""
163
+ {timing}{wer_badge(wer_poang)}
164
+ <div style='background:#F0F9FF;border:1.5px solid #BAE6FD;border-radius:12px;
165
+ padding:16px;color:#0F172A;font-size:15px;line-height:1.7;'>
166
+ 🎙️ {transkription}
167
+ </div>"""
168
+
169
+ # Build download text
170
+ prompt_technique_full = f"{technique_name} ({Config.PROMPT_TECHNIQUE})"
171
+ nedladdnings_text = f"""VoiceNote AI — VIPS Journalanteckning
172
+ Tid: {post['tid']}
173
+ ASR Model: OpenAI Whisper ({Config.ASR_MODEL})
174
+ Prompt: {prompt_technique_full}
175
+ GDPR: Dubbel anonymisering · Mistral AI EU-servrar
176
+
177
+ ASR: {asr_tid if asr_tid else 'N/A'}s | LLM: {llm_tid}s | Total: {total_tid}s | WER: {wer_poang if wer_poang is not None else 'N/A'}%
178
+
179
+ Transkription:
180
+ {transkription}
181
+
182
+ VIPS-Dokumentation:
183
+ {vips_text}
184
+
185
+ Referenstext (för WER):
186
+ {referens_text if referens_text else 'Ingen referenstext angiven'}
187
+ """
188
+
189
+ # Build VIPS HTML
190
+ vips_html = formatera_vips_html(vips_dict)
191
+
192
+ # Build history HTML
193
+ historik_html = formatera_historik_html(historik)
194
+
195
+ return (transkription_html, vips_html, historik_html, nedladdnings_text)
196
+
197
+ except Exception as e:
198
+ logger.error(f"Result building error: {e}")
199
+ raise
200
+
201
+
202
+ # ══════════════════════════════════════════════════════════
203
+ # GRADIO INTERFACE
204
+ # ══════════════════════════════════════════════════════════
205
+
206
+ # CSS
207
+ css = """
208
+ .panel{background:#FFFFFF !important;border:1.5px solid #E2E8F0 !important;border-radius:12px !important;padding:14px !important;box-shadow:0 1px 3px rgba(0,0,0,0.06) !important;}
209
+ .gradio-container{font-family:'Inter',-apple-system,BlinkMacSystemFont,sans-serif !important;}
210
+ .gradio-container textarea{color:#0F172A !important;background:#F8FAFC !important;}
211
+ .gradio-radio fieldset{border:none !important;}
212
+ .gradio-radio label{display:inline-flex !important;align-items:center !important;justify-content:center !important;min-width:52px !important;padding:9px 20px !important;margin:4px !important;border:2px solid #CBD5E1 !important;border-radius:12px !important;background:#F8FAFC !important;color:#475569 !important;font-weight:700 !important;font-size:16px !important;cursor:pointer !important;transition:all 0.15s ease !important;}
213
+ .gradio-radio label:hover{border-color:#0369A1 !important;background:#EFF6FF !important;color:#0369A1 !important;}
214
+ .gradio-radio input[type="radio"]{display:none !important;}
215
+ .gradio-radio label:has(input[type="radio"]:checked){background:#0369A1 !important;border-color:#0369A1 !important;color:#ffffff !important;box-shadow:0 4px 12px rgba(3,105,161,0.4) !important;}
216
+ input[type="range"]{accent-color:#0369A1 !important;}
217
+ ::-webkit-scrollbar{width:5px;}::-webkit-scrollbar-thumb{background:#CBD5E1;border-radius:4px;}
218
+ footer{display:none !important;}
219
+ """
220
+
221
+ # SUS Questions
222
+ sus_fragor = [
223
+ "Jag skulle vilja använda det här systemet regelbundet.",
224
+ "Jag tyckte att systemet var onödigt komplicerat.",
225
+ "Jag tyckte att systemet var lätt att använda.",
226
+ "Jag tror att jag behöver teknisk hjälp för att använda systemet.",
227
+ "Jag tyckte att funktionerna var väl integrerade.",
228
+ "Jag tyckte att det fanns för mycket inkonsistens i systemet.",
229
+ "De flesta skulle lära sig systemet väldigt snabbt.",
230
+ "Jag tyckte att systemet var väldigt krångligt.",
231
+ "Jag kände mig trygg när jag använde systemet.",
232
+ "Jag behövde lära mig mycket innan jag kunde börja.",
233
+ ]
234
+
235
+ # NASA-TLX Dimensions
236
+ nasa_dimensioner = [
237
+ "Mental belastning (tänka, besluta, minnas)",
238
+ "Fysisk belastning (fysisk ansträngning)",
239
+ "Tidskrav (stressad av tempot)",
240
+ "Prestation (Nöjd=1 → Missnöjd=20)",
241
+ "Frustration (irriterad, stressad)",
242
+ "Ansträngning (hur hårt fick du jobba)",
243
+ ]
244
+
245
+ with gr.Blocks(css=css, theme=gr.themes.Base()) as demo:
246
+ # Build header with technique indicator
247
+ technique_name = Config.get_technique_name()
248
+ technique_emoji = Config.get_technique_emoji()
249
+ technique_color = Config.get_technique_color()
250
+
251
+ gr.HTML(f"""
252
+ <div style='background:linear-gradient(135deg,#EFF6FF,#F0FDF4);border:1.5px solid #BFDBFE;
253
+ border-radius:20px;padding:28px;text-align:center;margin-bottom:20px;
254
+ box-shadow:0 2px 16px rgba(3,105,161,0.08);'>
255
+ <div style='display:inline-flex;align-items:center;gap:8px;background:#EFF6FF;
256
+ border:1.5px solid #93C5FD;border-radius:100px;padding:6px 18px;margin-bottom:16px;
257
+ font-size:11px;color:#0369A1;letter-spacing:0.1em;font-weight:700;'>
258
+ <span style='width:7px;height:7px;border-radius:50%;background:#0369A1;
259
+ animation:pulse 2s infinite;display:inline-block;'></span>
260
+ Healthcare AI · Högskolan Väst · GDPR-compliant
261
+ </div>
262
+ <h1 style='font-size:34px;font-weight:900;margin:0 0 10px;color:#0F172A;'>VoiceNote AI 🎙️</h1>
263
+ <p style='color:#64748B;margin:0 0 14px;font-size:15px;'>
264
+ OpenAI Whisper · VIPS-dokumentation · WER · SUS · NASA-TLX
265
+ </p>
266
+ <div style='display:inline-flex;align-items:center;gap:8px;background:#F5F3FF;
267
+ border:1.5px solid {technique_color}55;border-radius:100px;padding:5px 16px;
268
+ font-size:12px;color:{technique_color};font-weight:700;'>
269
+ {technique_emoji} RESEARCH MODE: {technique_name}
270
+ </div>
271
+ </div>
272
+ <style>@keyframes pulse{{0%,100%{{opacity:1}}50%{{opacity:0.3}}}}</style>""")
273
+
274
+ with gr.Tabs():
275
+ with gr.TabItem("🎙️ VIPS-dokumentation"):
276
+ with gr.Row():
277
+ with gr.Column(scale=1):
278
+ gr.HTML("""<div style='background:#EFF6FF;border:1.5px solid #BFDBFE;border-radius:12px;padding:14px;margin-bottom:8px;'>
279
+ <p style='color:#1E40AF;font-size:13px;margin:0;'>⓪ Spela in patientsamtalet. Referenstexten används för <strong>WER</strong> — lämna tomt om du inte har en.</p></div>""")
280
+ ljud_input = gr.Audio(sources=["microphone"], type="filepath", label="Spela in patientsamtal", elem_classes=["panel"])
281
+ referens_input = gr.Textbox(label="Referenstext (valfritt, för WER-beräkning)", placeholder="Exakt text som borde ha sagts...", lines=3, elem_classes=["panel"])
282
+
283
+ with gr.Row():
284
+ analysera_btn = gr.Button("① Analysera 🎙️", variant="primary", size="lg")
285
+
286
+ gr.HTML("<div style='border-top:1.5px solid #E2E8F0;margin:18px 0;'></div>")
287
+ gr.HTML("<div style='background:#F0FDF4;border:1.5px solid #A7F3D0;border-radius:12px;padding:14px;margin-bottom:8px;'><p style='color:#047857;font-size:13px;margin:0;'>Eller skriv in text direkt (utan ljudinspelning):</p></div>")
288
+ text_input = gr.Textbox(label="Text-input (direkt VIPS-klassificering)", placeholder="Jag har ont i huvudet och tar Metoprolol...", lines=5, elem_classes=["panel"])
289
+ text_btn = gr.Button("② Klassificera text 📝", size="lg")
290
+
291
+ with gr.Column(scale=1):
292
+ transkription_output = gr.HTML(label="Transkription + WER")
293
+ vips_output = gr.HTML(label="VIPS")
294
+ nedladdning = gr.File(label="💾 Ladda ner VIPS-anteckning (.txt)")
295
+
296
+ gr.HTML("<div style='border-top:2px solid #E2E8F0;margin:24px 0;'></div>")
297
+ gr.HTML("<h3 style='color:#1E40AF;margin-bottom:12px;'>📊 Senaste analyser</h3>")
298
+ historik_output = gr.HTML()
299
+ exportera_btn = gr.Button("📥 Exportera alla analyser (.csv)")
300
+ csv_output = gr.File(label="CSV-export")
301
+
302
+ with gr.TabItem("📋 SUS"):
303
+ gr.HTML("<div style='background:#EFF6FF;border:1.5px solid #BFDBFE;border-radius:12px;padding:16px;margin-bottom:16px;'><p style='color:#1E40AF;margin:0;font-size:14px;'><strong>System Usability Scale (SUS)</strong> — Betygsätt systemet från 1 (instämmer inte) till 5 (instämmer helt).</p></div>")
304
+ sus_inputs = [gr.Radio([1,2,3,4,5], label=q, elem_classes=["panel"]) for q in sus_fragor]
305
+ sus_btn = gr.Button("Beräkna SUS-poäng", variant="primary", size="lg")
306
+ sus_resultat = gr.HTML()
307
+
308
+ with gr.TabItem("🧠 NASA-TLX"):
309
+ gr.HTML("<div style='background:#EFF6FF;border:1.5px solid #BFDBFE;border-radius:12px;padding:16px;margin-bottom:16px;'><p style='color:#1E40AF;margin:0;font-size:14px;'><strong>NASA Task Load Index (TLX)</strong> — Betygsätt varje dimension från 1 (låg) till 20 (hög).</p></div>")
310
+ nasa_inputs = [gr.Slider(1,20,step=1,label=d,elem_classes=["panel"]) for d in nasa_dimensioner]
311
+ nasa_btn = gr.Button("Beräkna NASA-TLX", variant="primary", size="lg")
312
+ nasa_resultat = gr.HTML()
313
+
314
+ # Event handlers
315
+ analysera_btn.click(
316
+ fn=transkribera_ljud,
317
+ inputs=[ljud_input, referens_input],
318
+ outputs=[transkription_output, vips_output, historik_output, nedladdning]
319
+ ).then(
320
+ fn=spara_nedladdning,
321
+ inputs=[nedladdning],
322
+ outputs=[nedladdning]
323
+ )
324
+
325
+ text_btn.click(
326
+ fn=klassificera_text,
327
+ inputs=[text_input, referens_input],
328
+ outputs=[transkription_output, vips_output, historik_output, nedladdning]
329
+ ).then(
330
+ fn=spara_nedladdning,
331
+ inputs=[nedladdning],
332
+ outputs=[nedladdning]
333
+ )
334
+
335
+ exportera_btn.click(
336
+ fn=lambda: exportera_csv(alla_resultat),
337
+ outputs=[csv_output]
338
+ )
339
+
340
+ sus_btn.click(
341
+ fn=berakna_sus,
342
+ inputs=sus_inputs,
343
+ outputs=[sus_resultat]
344
+ )
345
+
346
+ nasa_btn.click(
347
+ fn=berakna_nasa,
348
+ inputs=nasa_inputs,
349
+ outputs=[nasa_resultat]
350
+ )
351
+
352
+ if __name__ == "__main__":
353
+ demo.launch()
config.py ADDED
@@ -0,0 +1,108 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """
2
+ VoiceNote AI - Configuration
3
+ =============================
4
+ All system configuration in one place.
5
+ """
6
+
7
+ import os
8
+ import tempfile
9
+
10
+
11
+ class Config:
12
+ """Configuration class for VoiceNote AI"""
13
+
14
+ # ══════════════════════════════════════════════════════════
15
+ # MODEL SETTINGS
16
+ # ══════════════════════════════════════════════════════════
17
+
18
+ # ASR Model (Automatic Speech Recognition)
19
+ # Using OpenAI's official Whisper model (works with Swedish via language parameter)
20
+ # Options: "openai/whisper-tiny", "openai/whisper-small", "openai/whisper-medium", "openai/whisper-large-v2"
21
+ ASR_MODEL = "openai/whisper-small" # Fast and good for Swedish
22
+
23
+ # LLM Model (Large Language Model for VIPS classification)
24
+ LLM_MODEL = "mistral-large-latest"
25
+
26
+ # ══════════════════════════════════════════════════════════
27
+ # PROMPT TECHNIQUE (for research comparison)
28
+ # ══════════════════════════════════════════════════════════
29
+
30
+ # Options: "few_shot", "chain_of_thought"
31
+ # Change this to compare different prompt techniques
32
+ PROMPT_TECHNIQUE = "few_shot"
33
+
34
+ # ══════════════════════════════════════════════════════════
35
+ # PROCESSING LIMITS
36
+ # ══════════════════════════════════════════════════════════
37
+
38
+ MAX_AUDIO_SECONDS = 30
39
+ CHUNK_LENGTH_S = 30
40
+ STRIDE_LENGTH_S = 5
41
+
42
+ # ══════════════════════════════════════════════════════════
43
+ # LLM SETTINGS
44
+ # ══════════════════════════════════════════════════════════
45
+
46
+ LLM_TEMPERATURE = 0.1
47
+ LLM_MAX_TOKENS = 800 # Increased for Chain-of-Thought
48
+
49
+ # ══════════════════════════════════════════════════════════
50
+ # WER (Word Error Rate) THRESHOLDS
51
+ # ══════════════════════════════════════════════════════════
52
+
53
+ WER_EXCELLENT = 5.0 # < 5% = Excellent
54
+ WER_GOOD = 10.0 # 5-10% = Good
55
+ WER_ACCEPTABLE = 20.0 # 10-20% = Acceptable
56
+
57
+ # ══════════════════════════════════════════════════════════
58
+ # USABILITY THRESHOLDS
59
+ # ══════════════════════════════════════════════════════════
60
+
61
+ # SUS (System Usability Scale) - out of 100
62
+ SUS_EXCELLENT = 85
63
+ SUS_GOOD = 70
64
+ SUS_PASS = 50
65
+
66
+ # NASA-TLX (Task Load Index) - out of 120 (lower is better)
67
+ NASA_LOW = 40
68
+ NASA_MEDIUM = 60
69
+
70
+ # ══════════════════════════════════════════════════════════
71
+ # FILE PATHS
72
+ # ══════════════════════════════════════════════════════════
73
+
74
+ TEMP_DIR = tempfile.gettempdir()
75
+
76
+ # ══════════════════════════════════════════════════════════
77
+ # DISPLAY NAMES (for UI)
78
+ # ══════════════════════════════════════════════════════════
79
+
80
+ @classmethod
81
+ def get_technique_name(cls) -> str:
82
+ """Get full prompt technique name"""
83
+ if cls.PROMPT_TECHNIQUE == "few_shot":
84
+ return "Few-shot Prompting"
85
+ elif cls.PROMPT_TECHNIQUE == "chain_of_thought":
86
+ return "Chain-of-Thought Prompting"
87
+ else:
88
+ return "Unknown Technique"
89
+
90
+ @classmethod
91
+ def get_technique_emoji(cls) -> str:
92
+ """Get emoji for current prompt technique"""
93
+ if cls.PROMPT_TECHNIQUE == "few_shot":
94
+ return "📚"
95
+ elif cls.PROMPT_TECHNIQUE == "chain_of_thought":
96
+ return "🧠"
97
+ else:
98
+ return "❓"
99
+
100
+ @classmethod
101
+ def get_technique_color(cls) -> str:
102
+ """Get color for current prompt technique"""
103
+ if cls.PROMPT_TECHNIQUE == "few_shot":
104
+ return "#7C3AED" # Purple
105
+ elif cls.PROMPT_TECHNIQUE == "chain_of_thought":
106
+ return "#0369A1" # Blue
107
+ else:
108
+ return "#6B7280" # Gray
gdpr_filter.py ADDED
@@ -0,0 +1,159 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """
2
+ VoiceNote AI - GDPR Filter
3
+ ===========================
4
+ Anonymizes personal information from text.
5
+ Dual-layer protection: input + output anonymization.
6
+ """
7
+
8
+ import re
9
+ import logging
10
+
11
+ logger = logging.getLogger(__name__)
12
+
13
+
14
+ class GDPRFilter:
15
+ """GDPR-compliant anonymization with dual-layer protection"""
16
+
17
+ # Replacement tokens
18
+ REPLACEMENTS = {
19
+ 'personnummer': '[PERSONNR]',
20
+ 'telefon': '[TELEFON]',
21
+ 'epost': '[EPOST]',
22
+ 'datum': '[DATUM]',
23
+ 'adress': '[ADRESS]',
24
+ 'namn': '[NAMN]',
25
+ }
26
+
27
+ # Words to exclude from name anonymization (Swedish medical/common terms)
28
+ EXCLUDE_WORDS = {
29
+ # Common Swedish words
30
+ 'Patienten', 'Patient', 'Jag', 'Mig', 'Min', 'Mitt', 'Mina',
31
+ 'Du', 'Dig', 'Din', 'Ditt', 'Dina',
32
+ 'Han', 'Hon', 'Hen', 'Honom', 'Henne',
33
+ 'Vi', 'Oss', 'Vår', 'Vårt', 'Våra',
34
+ 'De', 'Dem', 'Deras',
35
+
36
+ # Medical terms
37
+ 'Läkaren', 'Läkare', 'Sjuksköterskan', 'Sköterska',
38
+ 'Doktor', 'Doktorn', 'Sjuksköterska',
39
+
40
+ # Time/location words that might be capitalized
41
+ 'Sverige', 'Svensk', 'Svenska',
42
+ 'Måndag', 'Tisdag', 'Onsdag', 'Torsdag', 'Fredag', 'Lördag', 'Söndag',
43
+ 'Januari', 'Februari', 'Mars', 'April', 'Maj', 'Juni',
44
+ 'Juli', 'Augusti', 'September', 'Oktober', 'November', 'December',
45
+
46
+ # Common actions/verbs that might be capitalized
47
+ 'Ta', 'Tar', 'Tog', 'Tagit',
48
+ 'Gå', 'Går', 'Gick', 'Gått',
49
+ 'Känner', 'Kände', 'Känt',
50
+ }
51
+
52
+ @classmethod
53
+ def anonymize(cls, text: str) -> str:
54
+ """
55
+ Anonymize personal information in text.
56
+ Returns anonymized text with PII replaced by tokens.
57
+ """
58
+ if not text or not text.strip():
59
+ return text
60
+
61
+ anonymized = text
62
+ stats = {}
63
+
64
+ # 1. Personnummer (Swedish personal ID)
65
+ personnr_pattern = r'\b\d{6,8}[-–]\d{4}\b'
66
+ count = len(re.findall(personnr_pattern, anonymized))
67
+ if count > 0:
68
+ anonymized = re.sub(personnr_pattern, cls.REPLACEMENTS['personnummer'], anonymized)
69
+ stats['personnummer'] = count
70
+
71
+ # 2. Phone numbers
72
+ telefon_patterns = [
73
+ r'\b0\d{1,3}[-\s]?\d{5,8}\b', # Swedish landline/mobile
74
+ r'\b\+46[-\s]?\d{1,3}[-\s]?\d{5,8}\b', # International format
75
+ ]
76
+ phone_count = 0
77
+ for pattern in telefon_patterns:
78
+ matches = re.findall(pattern, anonymized)
79
+ phone_count += len(matches)
80
+ anonymized = re.sub(pattern, cls.REPLACEMENTS['telefon'], anonymized)
81
+ if phone_count > 0:
82
+ stats['telefon'] = phone_count
83
+
84
+ # 3. Email addresses
85
+ epost_pattern = r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b'
86
+ count = len(re.findall(epost_pattern, anonymized))
87
+ if count > 0:
88
+ anonymized = re.sub(epost_pattern, cls.REPLACEMENTS['epost'], anonymized)
89
+ stats['epost'] = count
90
+
91
+ # 4. Dates
92
+ datum_patterns = [
93
+ r'\b\d{4}-\d{2}-\d{2}\b', # YYYY-MM-DD
94
+ r'\b\d{2}/\d{2}/\d{4}\b', # DD/MM/YYYY
95
+ r'\b\d{1,2}\s+(januari|februari|mars|april|maj|juni|juli|augusti|september|oktober|november|december)\s+\d{4}\b',
96
+ ]
97
+ date_count = 0
98
+ for pattern in datum_patterns:
99
+ matches = re.findall(pattern, anonymized, re.IGNORECASE)
100
+ date_count += len(matches)
101
+ anonymized = re.sub(pattern, cls.REPLACEMENTS['datum'], anonymized, flags=re.IGNORECASE)
102
+ if date_count > 0:
103
+ stats['datum'] = date_count
104
+
105
+ # 5. Addresses (Swedish format)
106
+ adress_patterns = [
107
+ r'\b[A-ZÅÄÖ][a-zåäö]+(?:gatan|vägen|stigen|platsen|torget)\s+\d+[A-Za-z]?\b',
108
+ r'\b\d{3}\s?\d{2}\s+[A-ZÅÄÖ][a-zåäö]+\b', # Postal code + city
109
+ ]
110
+ address_count = 0
111
+ for pattern in adress_patterns:
112
+ matches = re.findall(pattern, anonymized)
113
+ address_count += len(matches)
114
+ anonymized = re.sub(pattern, cls.REPLACEMENTS['adress'], anonymized)
115
+ if address_count > 0:
116
+ stats['adress'] = address_count
117
+
118
+ # 6. Names (capitalized words, with exclusions)
119
+ # Only anonymize standalone capitalized words that aren't common words
120
+ words = anonymized.split()
121
+ name_count = 0
122
+ for i, word in enumerate(words):
123
+ # Check if word starts with capital and is >2 chars
124
+ if len(word) > 2 and word[0].isupper() and word[1:].islower():
125
+ # Extract just the word (remove punctuation)
126
+ clean_word = re.sub(r'[^\w]', '', word)
127
+ name = clean_word
128
+
129
+ # Check if it's in exclusion list
130
+ if name not in cls.EXCLUDE_WORDS:
131
+ anonymized = anonymized.replace(name, cls.REPLACEMENTS['namn'], 1)
132
+ name_count += 1
133
+
134
+ if name_count > 0:
135
+ stats['namn'] = name_count
136
+
137
+ if stats:
138
+ logger.debug(f"GDPR filter stats: {stats}")
139
+
140
+ return anonymized
141
+
142
+ @classmethod
143
+ def validate_anonymization(cls, text: str) -> bool:
144
+ """
145
+ Check if text still contains personal information.
146
+ Returns True if text appears anonymized, False otherwise.
147
+ """
148
+ # Check for obvious patterns
149
+ dangerous_patterns = [
150
+ r'\b\d{6,8}[-–]\d{4}\b', # Personnummer
151
+ r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b', # Email
152
+ ]
153
+
154
+ for pattern in dangerous_patterns:
155
+ if re.search(pattern, text):
156
+ logger.warning("Potential PII found after anonymization!")
157
+ return False
158
+
159
+ return True
models.py ADDED
@@ -0,0 +1,134 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """
2
+ VoiceNote AI - Models
3
+ =====================
4
+ ASR (Whisper) and LLM (Mistral) model interfaces.
5
+ """
6
+
7
+ import os
8
+ import torch
9
+ import spaces
10
+ import requests
11
+ import logging
12
+ from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor, pipeline
13
+ from config import Config
14
+
15
+ logger = logging.getLogger(__name__)
16
+
17
+
18
+ # ══════════════════════════════════════════════════════════
19
+ # MISTRAL AI CLIENT (HTTP API)
20
+ # ══════════════════════════════════════════════════════════
21
+
22
+ class MistralClient:
23
+ """Mistral AI client using direct HTTP API (more reliable than SDK)"""
24
+
25
+ def __init__(self):
26
+ self.api_key = os.environ.get("MISTRAL_API_KEY", "")
27
+ if not self.api_key:
28
+ logger.warning("MISTRAL_API_KEY not found in environment")
29
+
30
+ self.api_url = "https://api.mistral.ai/v1/chat/completions"
31
+ logger.info("Mistral client initialized with HTTP API")
32
+
33
+ def chat(self, prompt: str) -> str:
34
+ """Send chat request to Mistral AI via HTTP"""
35
+ try:
36
+ headers = {
37
+ "Content-Type": "application/json",
38
+ "Authorization": f"Bearer {self.api_key}"
39
+ }
40
+
41
+ payload = {
42
+ "model": Config.LLM_MODEL,
43
+ "messages": [{"role": "user", "content": prompt}],
44
+ "temperature": Config.LLM_TEMPERATURE,
45
+ "max_tokens": Config.LLM_MAX_TOKENS
46
+ }
47
+
48
+ response = requests.post(
49
+ self.api_url,
50
+ headers=headers,
51
+ json=payload,
52
+ timeout=30
53
+ )
54
+
55
+ if response.status_code != 200:
56
+ error_msg = f"Mistral API error: {response.status_code}"
57
+ logger.error(f"{error_msg} - {response.text}")
58
+ raise Exception(error_msg)
59
+
60
+ result = response.json()
61
+ return result["choices"][0]["message"]["content"]
62
+
63
+ except Exception as e:
64
+ logger.error(f"Mistral API error: {e}")
65
+ raise Exception(f"LLM fel: {str(e)}")
66
+
67
+
68
+ # ══════════════════════════════════════════════════════════
69
+ # ASR MODEL (WHISPER)
70
+ # ══════════════════════════════════════════════════════════
71
+
72
+ class ASRModel:
73
+ """Automatic Speech Recognition using Whisper"""
74
+
75
+ def __init__(self):
76
+ logger.info(f"Loading ASR model: {Config.ASR_MODEL}")
77
+
78
+ device = "cuda:0" if torch.cuda.is_available() else "cpu"
79
+ torch_dtype = torch.float16 if torch.cuda.is_available() else torch.float32
80
+
81
+ try:
82
+ model = AutoModelForSpeechSeq2Seq.from_pretrained(
83
+ Config.ASR_MODEL,
84
+ torch_dtype=torch_dtype,
85
+ low_cpu_mem_usage=True,
86
+ use_safetensors=True,
87
+ )
88
+ model.to(device)
89
+
90
+ processor = AutoProcessor.from_pretrained(Config.ASR_MODEL)
91
+
92
+ self.pipe = pipeline(
93
+ "automatic-speech-recognition",
94
+ model=model,
95
+ tokenizer=processor.tokenizer,
96
+ feature_extractor=processor.feature_extractor,
97
+ chunk_length_s=Config.CHUNK_LENGTH_S,
98
+ stride_length_s=Config.STRIDE_LENGTH_S,
99
+ torch_dtype=torch_dtype,
100
+ device=device,
101
+ )
102
+
103
+ logger.info(f"ASR model loaded successfully on {device}")
104
+
105
+ except Exception as e:
106
+ logger.error(f"Failed to load ASR model: {e}")
107
+ raise
108
+
109
+ @spaces.GPU(duration=60)
110
+ def transcribe(self, audio_path: str) -> str:
111
+ """
112
+ Transcribe audio file to Swedish text.
113
+ Returns transcribed text.
114
+ """
115
+ try:
116
+ if not audio_path:
117
+ raise ValueError("No audio file provided")
118
+
119
+ logger.info(f"Transcribing audio: {audio_path}")
120
+
121
+ result = self.pipe(
122
+ audio_path,
123
+ generate_kwargs={"language": "sv"}, # Swedish language
124
+ return_timestamps=False,
125
+ )
126
+
127
+ text = result["text"].strip()
128
+ logger.info(f"Transcription successful: {len(text)} characters")
129
+
130
+ return text
131
+
132
+ except Exception as e:
133
+ logger.error(f"Transcription error: {e}")
134
+ raise Exception(f"Taligenkänning misslyckades: {str(e)}")
requirements.txt ADDED
@@ -0,0 +1,30 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # VoiceNote AI - Requirements
2
+ # Optimized for HuggingFace Spaces
3
+
4
+ # Core ML/AI frameworks
5
+ torch
6
+ transformers
7
+ accelerate
8
+ safetensors
9
+
10
+ # HTTP requests (for Mistral API)
11
+ requests
12
+
13
+ # Speech processing
14
+ jiwer
15
+ librosa
16
+ soundfile
17
+
18
+ # Web interface
19
+ gradio
20
+
21
+ # Utilities
22
+ python-dotenv
23
+
24
+ # HuggingFace Spaces (CRITICAL!)
25
+ spaces
26
+
27
+ # Optional: Development dependencies (uncomment if needed locally)
28
+ # pytest
29
+ # black
30
+ # flake8
utils.py ADDED
@@ -0,0 +1,260 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """
2
+ VoiceNote AI - Utilities
3
+ =========================
4
+ WER calculator and HTML formatting functions.
5
+ """
6
+
7
+ import os
8
+ import csv
9
+ import logging
10
+ from typing import Optional, Dict
11
+ from datetime import datetime
12
+ from jiwer import wer as compute_wer
13
+ from config import Config
14
+
15
+ logger = logging.getLogger(__name__)
16
+
17
+
18
+ # ══════════════════════════════════════════════════════════
19
+ # WER CALCULATOR
20
+ # ══════════════════════════════════════════════════════════
21
+
22
+ class WERCalculator:
23
+ """Word Error Rate calculator with validation"""
24
+
25
+ @staticmethod
26
+ def calculate(reference: str, hypothesis: str) -> Optional[float]:
27
+ """
28
+ Calculate WER score.
29
+ Returns percentage (0-100) or None if invalid input.
30
+ """
31
+ if not reference or not reference.strip():
32
+ return None
33
+
34
+ if not hypothesis or not hypothesis.strip():
35
+ logger.warning("Empty hypothesis for WER calculation")
36
+ return 100.0 # All words are errors
37
+
38
+ try:
39
+ score = compute_wer(
40
+ reference.lower().strip(),
41
+ hypothesis.lower().strip()
42
+ )
43
+ percentage = round(score * 100, 1)
44
+ logger.info(f"WER calculated: {percentage}%")
45
+ return percentage
46
+ except Exception as e:
47
+ logger.error(f"WER calculation error: {e}")
48
+ return None
49
+
50
+ @staticmethod
51
+ def get_quality_label(wer: Optional[float]) -> str:
52
+ """Get quality label for WER score"""
53
+ if wer is None:
54
+ return "N/A"
55
+
56
+ if wer < Config.WER_EXCELLENT:
57
+ return "Utmärkt"
58
+ elif wer < Config.WER_GOOD:
59
+ return "Bra"
60
+ elif wer < Config.WER_ACCEPTABLE:
61
+ return "Godkänd"
62
+ else:
63
+ return "Behöver förbättras"
64
+
65
+
66
+ # ══════════════════════════════════════════════════════════
67
+ # HTML FORMATTERS
68
+ # ══════════════════════════════════════════════════════════
69
+
70
+ def formatera_vips_html(vips: dict) -> str:
71
+ """Format VIPS dictionary as colored HTML"""
72
+ colors = {
73
+ 'V': ('#10B981', '#ECFDF5'), # Green
74
+ 'I': ('#F59E0B', '#FFFBEB'), # Amber
75
+ 'P': ('#3B82F6', '#EFF6FF'), # Blue
76
+ 'S': ('#EF4444', '#FEF2F2'), # Red
77
+ }
78
+
79
+ categories = {
80
+ 'V': 'Välbefinnande',
81
+ 'I': 'Integritet',
82
+ 'P': 'Prevention',
83
+ 'S': 'Säkerhet',
84
+ }
85
+
86
+ html_parts = []
87
+ for key in ['V', 'I', 'P', 'S']:
88
+ fg, bg = colors[key]
89
+ cat = categories[key]
90
+ text = vips.get(key, "Ingen relevant information.")
91
+
92
+ html_parts.append(f"""
93
+ <div style='background:{bg};border-left:4px solid {fg};padding:12px 16px;margin-bottom:10px;border-radius:8px;'>
94
+ <div style='color:{fg};font-weight:700;font-size:13px;margin-bottom:4px;'>{key} — {cat}</div>
95
+ <div style='color:#1F2937;font-size:14px;line-height:1.6;'>{text}</div>
96
+ </div>
97
+ """)
98
+
99
+ return "".join(html_parts)
100
+
101
+
102
+ def wer_badge(wer_poang: Optional[float]) -> str:
103
+ """Create WER badge HTML"""
104
+ if wer_poang is None:
105
+ return ""
106
+
107
+ if wer_poang < Config.WER_EXCELLENT:
108
+ color, bg = "#059669", "#ECFDF5"
109
+ label = "Utmärkt ✅"
110
+ elif wer_poang < Config.WER_GOOD:
111
+ color, bg = "#0369A1", "#EFF6FF"
112
+ label = "Bra ✅"
113
+ elif wer_poang < Config.WER_ACCEPTABLE:
114
+ color, bg = "#D97706", "#FFFBEB"
115
+ label = "Godkänd ⚠️"
116
+ else:
117
+ color, bg = "#DC2626", "#FEF2F2"
118
+ label = "Behöver förbättras ❌"
119
+
120
+ return f"""
121
+ <div style='background:{bg};border:2px solid {color}55;border-radius:14px;
122
+ padding:20px;margin-bottom:12px;text-align:center;'>
123
+ <div style='font-size:42px;font-weight:900;color:{color};'>{wer_poang:.1f}%</div>
124
+ <div style='color:{color};font-size:16px;font-weight:700;margin-top:4px;'>WER: {label}</div>
125
+ </div>"""
126
+
127
+
128
+ def formatera_historik_html(historik: list) -> str:
129
+ """Format history list as HTML"""
130
+ if not historik:
131
+ return "<p style='color:#64748B;text-align:center;padding:40px;'>Ingen historik ännu.</p>"
132
+
133
+ items = []
134
+ for post in reversed(historik[-5:]): # Last 5 entries
135
+ tid = post.get('tid', 'N/A')
136
+ wer = post.get('wer_poang')
137
+ wer_text = f"{wer:.1f}%" if wer is not None else "N/A"
138
+
139
+ technique = post.get('prompt_technique', 'unknown')
140
+ technique_emoji = "📚" if technique == "few_shot" else "🧠"
141
+
142
+ items.append(f"""
143
+ <div style='background:#F8FAFC;border:1.5px solid #E2E8F0;border-radius:10px;
144
+ padding:12px;margin-bottom:8px;'>
145
+ <div style='display:flex;justify-content:space-between;align-items:center;'>
146
+ <span style='color:#475569;font-weight:600;'>{technique_emoji} {tid}</span>
147
+ <span style='color:#64748B;font-size:13px;'>WER: {wer_text}</span>
148
+ </div>
149
+ </div>
150
+ """)
151
+
152
+ return "".join(items)
153
+
154
+
155
+ # ══════════════════════════════════════════════════════════
156
+ # FILE EXPORT FUNCTIONS
157
+ # ══════════════════════════════════════════════════════════
158
+
159
+ def spara_nedladdning(innehall: str) -> Optional[str]:
160
+ """Save text file for download"""
161
+ try:
162
+ timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
163
+ path = os.path.join(Config.TEMP_DIR, f"vips_anteckning_{timestamp}.txt")
164
+ with open(path, "w", encoding="utf-8") as f:
165
+ f.write(innehall)
166
+ logger.info(f"Saved text file: {path}")
167
+ return path
168
+ except Exception as e:
169
+ logger.error(f"Error saving text file: {e}")
170
+ return None
171
+
172
+
173
+ def exportera_csv(alla_resultat: list) -> Optional[str]:
174
+ """Export all results as CSV"""
175
+ if not alla_resultat:
176
+ return None
177
+ try:
178
+ path = os.path.join(Config.TEMP_DIR, "voicenote_resultat.csv")
179
+ with open(path, "w", newline="", encoding="utf-8") as f:
180
+ fieldnames = ["tid", "prompt_technique", "asr_tid", "llm_tid", "total_tid", "wer_poang", "transkription", "referens", "vips"]
181
+ writer = csv.DictWriter(f, fieldnames=fieldnames, extrasaction="ignore")
182
+ writer.writeheader()
183
+ writer.writerows(alla_resultat)
184
+ logger.info(f"Exported CSV: {path}")
185
+ return path
186
+ except Exception as e:
187
+ logger.error(f"CSV export error: {e}")
188
+ return None
189
+
190
+
191
+ # ══════════════════════════════════════════════════════════
192
+ # USABILITY CALCULATORS
193
+ # ══════════════════════════════════════════════════════════
194
+
195
+ def berakna_sus(*svar) -> str:
196
+ """Calculate SUS score and return HTML"""
197
+ try:
198
+ pl = []
199
+ for i, val in enumerate(svar):
200
+ try:
201
+ v = int(val)
202
+ except (TypeError, ValueError):
203
+ return "<div style='color:#DC2626;padding:14px;background:#FEF2F2;border-radius:10px;font-weight:600;'>⚠️ Fyll i alla 10 frågor.</div>"
204
+ pl.append(v - 1 if i % 2 == 0 else 5 - v)
205
+
206
+ total = sum(pl) * 2.5
207
+
208
+ # Determine color and label
209
+ if total >= Config.SUS_GOOD:
210
+ f, b = "#059669", "#ECFDF5"
211
+ elif total >= Config.SUS_PASS:
212
+ f, b = "#D97706", "#FFFBEB"
213
+ else:
214
+ f, b = "#DC2626", "#FEF2F2"
215
+
216
+ e = ("Utmärkt ✅" if total >= Config.SUS_EXCELLENT else
217
+ "Bra ✅" if total >= Config.SUS_GOOD else
218
+ "Godkänd ⚠️" if total >= Config.SUS_PASS else
219
+ "Underkänd ❌")
220
+
221
+ return f"""<div style='background:{b};border:2px solid {f}55;border-radius:14px;padding:28px;text-align:center;'>
222
+ <div style='font-size:60px;font-weight:900;color:{f};'>{total:.1f}</div>
223
+ <div style='color:{f};font-size:18px;font-weight:700;margin-top:6px;'>{e}</div>
224
+ </div>"""
225
+ except Exception as ex:
226
+ logger.error(f"SUS calculation error: {ex}")
227
+ return "<div style='color:#DC2626;'>Fel vid beräkning</div>"
228
+
229
+
230
+ def berakna_nasa(*svar) -> str:
231
+ """Calculate NASA-TLX score and return HTML"""
232
+ try:
233
+ values = []
234
+ for val in svar:
235
+ try:
236
+ v = int(val)
237
+ values.append(v)
238
+ except (TypeError, ValueError):
239
+ return "<div style='color:#DC2626;padding:14px;background:#FEF2F2;border-radius:10px;font-weight:600;'>⚠️ Fyll i alla 6 dimensioner.</div>"
240
+
241
+ total = sum(values)
242
+
243
+ if total <= Config.NASA_LOW:
244
+ f, b = "#059669", "#ECFDF5"
245
+ e = "Låg belastning ✅"
246
+ elif total <= Config.NASA_MEDIUM:
247
+ f, b = "#D97706", "#FFFBEB"
248
+ e = "Medel belastning ⚠️"
249
+ else:
250
+ f, b = "#DC2626", "#FEF2F2"
251
+ e = "Hög belastning ❌"
252
+
253
+ return f"""<div style='background:{b};border:2px solid {f}55;border-radius:14px;padding:28px;text-align:center;'>
254
+ <div style='font-size:60px;font-weight:900;color:{f};'>{total}</div>
255
+ <div style='color:{f};font-size:14px;margin-top:4px;'>/ 120 poäng</div>
256
+ <div style='color:{f};font-size:18px;font-weight:700;margin-top:10px;'>{e}</div>
257
+ </div>"""
258
+ except Exception as ex:
259
+ logger.error(f"NASA-TLX calculation error: {ex}")
260
+ return "<div style='color:#DC2626;'>Fel vid beräkning</div>"
vips_classifier.py ADDED
@@ -0,0 +1,181 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """
2
+ VoiceNote AI - VIPS Classifier
3
+ ===============================
4
+ VIPS classification with dual GDPR protection.
5
+ Supports both Few-shot and Chain-of-Thought prompting.
6
+ """
7
+
8
+ import logging
9
+ from typing import Dict
10
+ from config import Config
11
+ from gdpr_filter import GDPRFilter
12
+ from models import MistralClient
13
+
14
+ logger = logging.getLogger(__name__)
15
+
16
+
17
+ class VIPSClassifier:
18
+ """VIPS classification with support for different prompt techniques"""
19
+
20
+ def __init__(self, llm_client: MistralClient):
21
+ self.llm = llm_client
22
+
23
+ @staticmethod
24
+ def build_prompt_few_shot(text: str) -> str:
25
+ """
26
+ Few-shot prompting: Ger 2-3 konkreta exempel före den verkliga uppgiften.
27
+
28
+ Fördelar:
29
+ - Tydliga mönster för modellen att följa
30
+ - Bättre formatering
31
+ - Konsekvent output
32
+
33
+ Forskningsreferens: Brown et al. (2020) - Language Models are Few-Shot Learners
34
+ """
35
+ return f"""Du är ett dokumentationssystem för sjuksköterskor.
36
+ Din uppgift är att klassificera patientinformation enligt VIPS-modellen.
37
+
38
+ OBLIGATORISKA REGLER:
39
+ 1. Skriv ALDRIG namn, personnummer, telefonnummer eller adresser.
40
+ 2. Hitta INTE på information. Dokumentera ENBART vad patienten faktiskt säger.
41
+ 3. Lägg INTE till medicinska råd, diagnoser eller rekommendationer.
42
+ 4. Om en kategori saknar information, skriv exakt: Ingen relevant information.
43
+ 5. Svara i exakt VIPS-format — fyra rader, en per kategori, inget annat.
44
+
45
+ VIPS-KATEGORIER:
46
+ V (Välbefinnande) — Symtom, smärta, mående och känslor
47
+ I (Integritet) — Vanor, önskemål och preferenser
48
+ P (Prevention) — Förebyggande åtgärder som nämnts
49
+ S (Säkerhet) — Risker, mediciner och säkerhetsaspekter
50
+
51
+ EXEMPEL 1:
52
+ Samtal: "Jag har haft huvudvärk i två dagar. Jag tar Metoprolol dagligen."
53
+ Svar:
54
+ V: Patienten rapporterar huvudvärk sedan två dagar.
55
+ I: Ingen relevant information.
56
+ P: Ingen relevant information.
57
+ S: Patienten tar Metoprolol dagligen.
58
+
59
+ EXEMPEL 2:
60
+ Samtal: "Jag sover dåligt på nätterna och känner mig orolig. Jag brukar gå på promenader varje dag."
61
+ Svar:
62
+ V: Patienten rapporterar sömnsvårigheter och känslor av oro.
63
+ I: Patienten promenerar dagligen.
64
+ P: Ingen relevant information.
65
+ S: Ingen relevant information.
66
+
67
+ EXEMPEL 3:
68
+ Samtal: "Jag har ont i bröstet när jag går i trappor. Jag röker 10 cigaretter om dagen."
69
+ Svar:
70
+ V: Patienten rapporterar bröstsmärta vid ansträngning.
71
+ I: Ingen relevant information.
72
+ P: Ingen relevant information.
73
+ S: Patienten röker 10 cigaretter dagligen. Risk för hjärt-kärlsjukdom.
74
+
75
+ NU ÄR DET DIN TUR. Klassificera följande samtal:
76
+ SAMTAL:
77
+ "{text}"
78
+
79
+ Klassificera i exakt VIPS-format (V:, I:, P:, S:):"""
80
+
81
+ @staticmethod
82
+ def build_prompt_chain_of_thought(text: str) -> str:
83
+ """
84
+ Chain-of-Thought prompting: Be modellen att tänka steg-för-steg.
85
+
86
+ Fördelar:
87
+ - Bättre resonemang
88
+ - Minskad hallucination
89
+ - Tydligare logik
90
+
91
+ Forskningsreferens: Wei et al. (2022) - Chain-of-Thought Prompting
92
+ Elicits Reasoning in Large Language Models
93
+ """
94
+ return f"""Du är ett dokumentationssystem för sjuksköterskor.
95
+ Din uppgift är att klassificera patientinformation enligt VIPS-modellen.
96
+
97
+ OBLIGATORISKA REGLER:
98
+ 1. Skriv ALDRIG namn, personnummer, telefonnummer eller adresser.
99
+ 2. Hitta INTE på information. Dokumentera ENBART vad patienten faktiskt säger.
100
+ 3. Lägg INTE till medicinska råd, diagnoser eller rekommendationer.
101
+ 4. Om en kategori saknar information, skriv exakt: Ingen relevant information.
102
+
103
+ VIPS-KATEGORIER:
104
+ V (Välbefinnande) — Symtom, smärta, mående och känslor
105
+ I (Integritet) — Vanor, önskemål och preferenser
106
+ P (Prevention) — Förebyggande åtgärder som nämnts
107
+ S (Säkerhet) — Risker, mediciner och säkerhetsaspekter
108
+
109
+ SAMTAL:
110
+ "{text}"
111
+
112
+ STEG-FÖR-STEG ANALYS:
113
+ Tänk igenom detta systematiskt:
114
+
115
+ 1. Läs igenom samtalet noggrant.
116
+ 2. Identifiera ALL information som nämnts (symtom, känslor, vanor, mediciner, risker).
117
+ 3. Sortera varje informationsdel:
118
+ - Handlar det om hur patienten MÅR? → V (Välbefinnande)
119
+ - Handlar det om patientens VANOR/PREFERENSER? → I (Integritet)
120
+ - Handlar det om FÖREBYGGANDE åtgärder? → P (Prevention)
121
+ - Handlar det om RISKER/MEDICINER/SÄKERHET? → S (Säkerhet)
122
+ 4. För varje kategori: Formulera en kort, professionell mening.
123
+ 5. Om kategorin är tom: Skriv "Ingen relevant information."
124
+
125
+ Genomför analysen steg-för-steg, sedan ge ditt svar i exakt VIPS-format:
126
+ V: [din analys]
127
+ I: [din analys]
128
+ P: [din analys]
129
+ S: [din analys]"""
130
+
131
+ def build_prompt(self, text: str) -> str:
132
+ """
133
+ Build prompt based on configured technique.
134
+ Used for research comparison between Few-shot and Chain-of-Thought.
135
+ """
136
+ if Config.PROMPT_TECHNIQUE == "chain_of_thought":
137
+ logger.info("Using Chain-of-Thought prompting")
138
+ return self.build_prompt_chain_of_thought(text)
139
+ else: # Default to few_shot
140
+ logger.info("Using Few-shot prompting")
141
+ return self.build_prompt_few_shot(text)
142
+
143
+ def classify(self, text: str) -> str:
144
+ """
145
+ Classify text according to VIPS model with dual GDPR protection.
146
+ Returns VIPS-formatted text.
147
+ """
148
+ try:
149
+ # Layer 1: Anonymize input
150
+ anonymized_input = GDPRFilter.anonymize(text)
151
+ logger.info("Input anonymized (Layer 1)")
152
+
153
+ # Build prompt using selected technique
154
+ prompt = self.build_prompt(anonymized_input)
155
+ response = self.llm.chat(prompt)
156
+
157
+ # Layer 2: Anonymize output
158
+ anonymized_output = GDPRFilter.anonymize(response)
159
+ logger.info("Output anonymized (Layer 2)")
160
+
161
+ # Validate anonymization
162
+ if not GDPRFilter.validate_anonymization(anonymized_output):
163
+ logger.error("GDPR validation failed!")
164
+
165
+ return anonymized_output
166
+ except Exception as e:
167
+ logger.error(f"VIPS classification error: {e}")
168
+ raise
169
+
170
+ @staticmethod
171
+ def parse_vips(vips_text: str) -> Dict[str, str]:
172
+ """Parse VIPS text into dictionary"""
173
+ vips = {k: "Ingen relevant information." for k in ["V", "I", "P", "S"]}
174
+
175
+ for line in vips_text.strip().split("\n"):
176
+ line = line.strip()
177
+ for key in vips:
178
+ if line.startswith(f"{key}:"):
179
+ vips[key] = line[2:].strip()
180
+
181
+ return vips