Elasticsearch
Spesifikasi teknis cluster Elasticsearch di Nexa, katalog lengkap 13 indeks data master & analitik, struktur mapping dokumen, serta pola query pencarian real-time.
Elasticsearch Document & Analytics Engine
Elasticsearch berfungsi sebagai mesin pencarian teks cepat (Full-Text Search) dan analitik multidimensi terdistribusi (OLAP) pada platform Nexa Intelligence. Seluruh endpoint visualisasi pada frontend membaca data dari indeks terdenormalisasi di Elasticsearch untuk menjamin latensi sub-100ms pada dataset jutaan artikel berita dan percakapan media sosial.
Arsitektur & Klaster Indeks Elasticsearch
Penyimpanan Elasticsearch di Nexa dibagi secara terstruktur ke dalam 4 klaster fungsional yang mencakup total 13 indeks aktif:
Elasticsearch Cluster (Port 9200)
├── 1. Master Ingestion & Raw Documents
│ ├── articles # Master dokumen berita (SPOK, Entitas, Taksonomi, Sentimen)
│ └── documents # Staging repositori korpus dokumen mentah
│
├── 2. Multichannel Social Media (Phone Farm B3)
│ └── postings # Konten medsos (X, FB, IG, TikTok) & metrik engagement
│
├── 3. Pre-Aggregated Dashboard & Decision Indices
│ ├── national_dashboard # Metrik situasi nasional 24 jam & 3 narasi eksekutif
│ ├── province_dashboard # Matriks risiko kamtibmas 38 provinsi
│ ├── issue_dashboard # Klaster isu strategis per leaf taksonomi (5 risk bands)
│ ├── trending_topics # Top 10 lonjakan isu harian (Velocity Spike ≥ 3.5σ)
│ ├── intelligence_feed # Linimasa intelijen real-time berdampingan dengan peta
│ ├── taxonomy_correlation # Matriks korelasi & ko-okurensi lintas 5 pilar taksonomi
│ ├── decision_intelligence # Rekomendasi mitigasi pimpinan dari LLM
│ └── pillar_intelligence # Sintesis analisa 5 pilar kontekstual 4 peran pemangku kepentingan
│
└── 4. Graph Network & Pipeline Observability
├── knowledge_graph # Graf relasi entitas (nodes, edges, centrality metrics)
└── pipeline_metrics # Observabilitas pipeline, throughput, & cakupan fakta SPOKKatalog Lengkap 13 Indeks Elasticsearch
Tabel di bawah ini merangkum seluruh indeks Elasticsearch di platform Nexa:
| No | Nama Indeks | Worker Penghasil | Format Primary Key (_id) | Siklus Pembaruan | Konsumsi API Backend / Frontend |
|---|---|---|---|---|---|
| 1 | articles | D1 AI Worker | {uuid} / {code_hash} | Real-time Stream | /overview/media-exec, /articles, Pencarian Global |
| 2 | documents | Ingestion Staging | {doc_id} | Real-time Ingestion | Staging internal pipeline & audit dokumen mentah |
| 3 | postings | B3 Phone Farm Swarm | {platform}_{post_id} | Kontinu (15 menit) | /media/channels, /dashboard/aktivis |
| 4 | national_dashboard | nexa-aggregation | national_{YYYY-MM-DD} | Harian (01:00 WIB) | /api/v1/overview/national, /overview/intelligence-analysis |
| 5 | province_dashboard | nexa-aggregation | {province_code}_{YYYY-MM-DD} | Harian (01:15 WIB) | /overview/provinces, Modal Briefing 38 Wilayah |
| 6 | issue_dashboard | nexa-aggregation | {issue_id}_{YYYY-MM-DD} | Harian (01:30 WIB) | /issues/clusters, Dashboard Manajemen Isu |
| 7 | trending_topics | nexa-aggregation | trending_{rank}_{YYYY-MM-DD} | Tiap 3 Jam / Harian | /trending-topics, Kartu KPI Topik Trending |
| 8 | intelligence_feed | nexa-aggregation / D1 | feed_{article_id} | Real-time Stream | /overview/feed, Panel Live Feed Berita Terkini |
| 9 | taxonomy_correlation | nexa-aggregation | corr_{source}_{target}_{date} | Harian (02:00 WIB) | /taxonomy/correlation, Matriks Analisa Korelasi |
| 10 | decision_intelligence | decision-worker | decision_{issue}_{YYYY-MM-DD} | Harian (03:00 WIB) | /decision/recommendations, Panel Rekomendasi Intelijen |
| 11 | pillar_intelligence | decision-worker | pillar_{pillar}_{YYYY-MM-DD} | Harian (02:30 WIB) | /overview/intelligence-analysis, Kartu Analisa 5 Pilar |
| 12 | knowledge_graph | nexa-knowledge-graph | kg_{cluster_id}_{YYYY-MM-DD} | Harian (H-1 Batch) | /knowledge-graph, Visualisasi Graf Interaktif 2D/3D |
| 13 | pipeline_metrics | nexa-aggregation | pipeline_{YYYY-MM-DD} | Harian (H-1 Batch) | /overview/pipeline-metrics, Panel Status Pipeline |
1. Mapping Indeks Master articles
Indeks ini memadukan analisis teks berbahasa Indonesia (Custom Indonesian Analyzer) dengan struktur JSON beraneka dimensi:
{
"settings": {
"number_of_shards": 3,
"number_of_replicas": 1,
"analysis": {
"analyzer": {
"indonesian_analyzer": {
"type": "custom",
"tokenizer": "standard",
"filter": ["lowercase", "indonesian_stop", "indonesian_stemmer"]
}
},
"filter": {
"indonesian_stop": { "type": "stop", "stopwords": "_indonesian_" },
"indonesian_stemmer": { "type": "stemmer", "language": "indonesian" }
}
}
},
"mappings": {
"properties": {
"id": { "type": "keyword" },
"code": { "type": "keyword" },
"url": { "type": "keyword" },
"domain": { "type": "keyword" },
"source_name": { "type": "keyword" },
"title": { "type": "text", "analyzer": "indonesian_analyzer", "fields": { "keyword": { "type": "keyword" } } },
"content": { "type": "text", "analyzer": "indonesian_analyzer" },
"summary": { "type": "text" },
"published_at": { "type": "date" },
"processed_at": { "type": "date" },
"sector_id": { "type": "integer" },
"sector_name": { "type": "keyword" },
"sentiment": {
"properties": {
"label": { "type": "keyword" },
"score": { "type": "float" },
"confidence": { "type": "float" }
}
},
"risk_score": { "type": "float" },
"severity": { "type": "integer" },
"taxonomy": {
"properties": {
"level_1": { "type": "keyword" },
"level_2": { "type": "keyword" },
"level_3": { "type": "keyword" },
"level_4": { "type": "keyword" },
"level_5": { "type": "keyword" },
"pillar": { "type": "keyword" },
"lineages": { "type": "keyword" }
}
},
"entities": {
"type": "nested",
"properties": {
"name": { "type": "keyword" },
"type": { "type": "keyword" },
"role": { "type": "keyword" },
"hierarchy": { "type": "keyword" }
}
},
"geo_location": {
"properties": {
"province": { "type": "keyword" },
"city": { "type": "keyword" },
"coordinates": { "type": "geo_point" }
}
}
}
}
}2. Struktur Indeks Media Sosial & Phone Farm (postings)
Indeks postings menampung konten media sosial multichannel hasil panen armada Phone Farm Swarm (Android APK native):
{
"mappings": {
"properties": {
"id": { "type": "keyword" },
"platform": { "type": "keyword" },
"post_id": { "type": "keyword" },
"author": {
"properties": {
"author_id": { "type": "keyword" },
"username": { "type": "keyword" },
"display_name": { "type": "text" },
"followers_count": { "type": "integer" },
"is_verified": { "type": "boolean" }
}
},
"primary_topic": {
"properties": {
"topic_id": { "type": "integer" },
"topic_name": { "type": "keyword" },
"primary_pillar": { "type": "keyword" }
}
},
"content": { "type": "text", "analyzer": "indonesian_analyzer" },
"likes_count": { "type": "integer" },
"shares_count": { "type": "integer" },
"comments_count": { "type": "integer" },
"sentiment": { "type": "keyword" },
"media_urls": { "type": "keyword" },
"scraped_at": { "type": "date" }
}
}
}3. Struktur Indeks Analitik & Rekomendasi Pimpinan
Indeks hasil sintesis decision-worker yang membagi arahan strategis per pilar ke dalam 4 peran:
{
"mappings": {
"properties": {
"doc_id": { "type": "keyword" },
"target_date": { "type": "date", "format": "yyyy-MM-dd" },
"pillar": { "type": "keyword" },
"label": { "type": "keyword" },
"metrics": {
"properties": {
"article_count": { "type": "integer" },
"negative_count": { "type": "integer" },
"negative_pct": { "type": "float" }
}
},
"insights": {
"properties": {
"pemerintahan": { "properties": { "title": { "type": "keyword" }, "text": { "type": "text" } } },
"pelaku_bisnis": { "properties": { "title": { "type": "keyword" }, "text": { "type": "text" } } },
"lsm_organisasi": { "properties": { "title": { "type": "keyword" }, "text": { "type": "text" } } },
"personal": { "properties": { "title": { "type": "keyword" }, "text": { "type": "text" } } }
}
},
"top_topics": {
"type": "nested",
"properties": { "key": { "type": "keyword" }, "label": { "type": "keyword" }, "count": { "type": "integer" } }
},
"representative_headlines": {
"type": "nested",
"properties": { "article_id": { "type": "keyword" }, "title": { "type": "text" }, "source": { "type": "keyword" } }
}
}
}
}4. Pola Query Agregasi & Point-in-Time (PIT)
A. Point-in-Time (PIT) Scan untuk Konsistensi Korpus
Ketika mengeksekusi pemindaian (deep scanning) korpus berita yang sedang mengalami continuous ingestion, sistem membuka Point-in-Time (PIT) guna membekukan status segmen shard:
# 1. Buka snapshot PIT pada indeks articles
POST /articles/_pit?keep_alive=1m
# Output: { "id": "46ToAwMDYXJ0aWNsZXMS..." }
# 2. Query search menggunakan PIT ID dengan paginasi search_after
POST /_search
{
"size": 50,
"query": { "match_all": {} },
"pit": {
"id": "46ToAwMDYXJ0aWNsZXMS...",
"keep_alive": "1m"
},
"sort": [ { "published_at": "desc" }, { "_shard_doc": "asc" } ]
}B. Komputasi Sebaran Sentimen Lintas 5 Pilar
POST /articles/_search
{
"size": 0,
"query": {
"range": {
"published_at": { "gte": "now-24h", "lte": "now" }
}
},
"aggs": {
"by_pillar": {
"terms": { "field": "taxonomy.pillar", "size": 5 },
"aggs": {
"sentiment_breakdown": {
"terms": { "field": "sentiment.label" },
"aggs": {
"avg_severity": { "avg": { "field": "severity" } }
}
}
}
}
}
}5. Kebijakan Index Lifecycle Management (ILM)
- Hot Phase (0–30 Hari): Indeks beroperasi dengan performa tinggi pada NVMe storage untuk read/write aktif dengan 3 shards dan 1 replika.
- Warm Phase (31–90 Hari): Indeks dialihkan ke mode read-only dengan eksekusi force-merge menjadi 1 segment per shard untuk menghemat memori heap JVM.
- Cold & Archive (>90 Hari): Indeks historis di-snapshot secara otomatis ke Object Storage (MinIO) pada bucket
nexa-backups/elasticsearch/guna menjaga efisiensi kapasitas disk.
PostgreSQL
Spesifikasi komprehensif skema database relasional OLTP PostgreSQL, definisi tabel master, relasi entitas, struktur artikel AI, data sektoral, serta strategi indeks performa di Nexa Intelligence Platform.
Redis
Spesifikasi teknis Redis In-Memory Cache di Nexa, pola prefix key, manajemen session JWT, transport Redis Streams 'scrape-tasks', rate limiting, dan distributed locks.