Nexa Intelligence Docs
Data & Storage

Elasticsearch

Spesifikasi teknis cluster Elasticsearch di Nexa, katalog lengkap 13 indeks data master & analitik, struktur mapping dokumen, serta pola query pencarian real-time.

Elasticsearch Document & Analytics Engine

Elasticsearch berfungsi sebagai mesin pencarian teks cepat (Full-Text Search) dan analitik multidimensi terdistribusi (OLAP) pada platform Nexa Intelligence. Seluruh endpoint visualisasi pada frontend membaca data dari indeks terdenormalisasi di Elasticsearch untuk menjamin latensi sub-100ms pada dataset jutaan artikel berita dan percakapan media sosial.


Arsitektur & Klaster Indeks Elasticsearch

Penyimpanan Elasticsearch di Nexa dibagi secara terstruktur ke dalam 4 klaster fungsional yang mencakup total 13 indeks aktif:

Elasticsearch Cluster (Port 9200)
├── 1. Master Ingestion & Raw Documents
│   ├── articles                      # Master dokumen berita (SPOK, Entitas, Taksonomi, Sentimen)
│   └── documents                     # Staging repositori korpus dokumen mentah
│
├── 2. Multichannel Social Media (Phone Farm B3)
│   └── postings                      # Konten medsos (X, FB, IG, TikTok) & metrik engagement
│
├── 3. Pre-Aggregated Dashboard & Decision Indices
│   ├── national_dashboard            # Metrik situasi nasional 24 jam & 3 narasi eksekutif
│   ├── province_dashboard            # Matriks risiko kamtibmas 38 provinsi
│   ├── issue_dashboard               # Klaster isu strategis per leaf taksonomi (5 risk bands)
│   ├── trending_topics               # Top 10 lonjakan isu harian (Velocity Spike ≥ 3.5σ)
│   ├── intelligence_feed             # Linimasa intelijen real-time berdampingan dengan peta
│   ├── taxonomy_correlation          # Matriks korelasi & ko-okurensi lintas 5 pilar taksonomi
│   ├── decision_intelligence         # Rekomendasi mitigasi pimpinan dari LLM
│   └── pillar_intelligence           # Sintesis analisa 5 pilar kontekstual 4 peran pemangku kepentingan
│
└── 4. Graph Network & Pipeline Observability
    ├── knowledge_graph               # Graf relasi entitas (nodes, edges, centrality metrics)
    └── pipeline_metrics              # Observabilitas pipeline, throughput, & cakupan fakta SPOK

Katalog Lengkap 13 Indeks Elasticsearch

Tabel di bawah ini merangkum seluruh indeks Elasticsearch di platform Nexa:

NoNama IndeksWorker PenghasilFormat Primary Key (_id)Siklus PembaruanKonsumsi API Backend / Frontend
1articlesD1 AI Worker{uuid} / {code_hash}Real-time Stream/overview/media-exec, /articles, Pencarian Global
2documentsIngestion Staging{doc_id}Real-time IngestionStaging internal pipeline & audit dokumen mentah
3postingsB3 Phone Farm Swarm{platform}_{post_id}Kontinu (15 menit)/media/channels, /dashboard/aktivis
4national_dashboardnexa-aggregationnational_{YYYY-MM-DD}Harian (01:00 WIB)/api/v1/overview/national, /overview/intelligence-analysis
5province_dashboardnexa-aggregation{province_code}_{YYYY-MM-DD}Harian (01:15 WIB)/overview/provinces, Modal Briefing 38 Wilayah
6issue_dashboardnexa-aggregation{issue_id}_{YYYY-MM-DD}Harian (01:30 WIB)/issues/clusters, Dashboard Manajemen Isu
7trending_topicsnexa-aggregationtrending_{rank}_{YYYY-MM-DD}Tiap 3 Jam / Harian/trending-topics, Kartu KPI Topik Trending
8intelligence_feednexa-aggregation / D1feed_{article_id}Real-time Stream/overview/feed, Panel Live Feed Berita Terkini
9taxonomy_correlationnexa-aggregationcorr_{source}_{target}_{date}Harian (02:00 WIB)/taxonomy/correlation, Matriks Analisa Korelasi
10decision_intelligencedecision-workerdecision_{issue}_{YYYY-MM-DD}Harian (03:00 WIB)/decision/recommendations, Panel Rekomendasi Intelijen
11pillar_intelligencedecision-workerpillar_{pillar}_{YYYY-MM-DD}Harian (02:30 WIB)/overview/intelligence-analysis, Kartu Analisa 5 Pilar
12knowledge_graphnexa-knowledge-graphkg_{cluster_id}_{YYYY-MM-DD}Harian (H-1 Batch)/knowledge-graph, Visualisasi Graf Interaktif 2D/3D
13pipeline_metricsnexa-aggregationpipeline_{YYYY-MM-DD}Harian (H-1 Batch)/overview/pipeline-metrics, Panel Status Pipeline

1. Mapping Indeks Master articles

Indeks ini memadukan analisis teks berbahasa Indonesia (Custom Indonesian Analyzer) dengan struktur JSON beraneka dimensi:

{
  "settings": {
    "number_of_shards": 3,
    "number_of_replicas": 1,
    "analysis": {
      "analyzer": {
        "indonesian_analyzer": {
          "type": "custom",
          "tokenizer": "standard",
          "filter": ["lowercase", "indonesian_stop", "indonesian_stemmer"]
        }
      },
      "filter": {
        "indonesian_stop": { "type": "stop", "stopwords": "_indonesian_" },
        "indonesian_stemmer": { "type": "stemmer", "language": "indonesian" }
      }
    }
  },
  "mappings": {
    "properties": {
      "id": { "type": "keyword" },
      "code": { "type": "keyword" },
      "url": { "type": "keyword" },
      "domain": { "type": "keyword" },
      "source_name": { "type": "keyword" },
      "title": { "type": "text", "analyzer": "indonesian_analyzer", "fields": { "keyword": { "type": "keyword" } } },
      "content": { "type": "text", "analyzer": "indonesian_analyzer" },
      "summary": { "type": "text" },
      "published_at": { "type": "date" },
      "processed_at": { "type": "date" },
      "sector_id": { "type": "integer" },
      "sector_name": { "type": "keyword" },
      "sentiment": {
        "properties": {
          "label": { "type": "keyword" },
          "score": { "type": "float" },
          "confidence": { "type": "float" }
        }
      },
      "risk_score": { "type": "float" },
      "severity": { "type": "integer" },
      "taxonomy": {
        "properties": {
          "level_1": { "type": "keyword" },
          "level_2": { "type": "keyword" },
          "level_3": { "type": "keyword" },
          "level_4": { "type": "keyword" },
          "level_5": { "type": "keyword" },
          "pillar": { "type": "keyword" },
          "lineages": { "type": "keyword" }
        }
      },
      "entities": {
        "type": "nested",
        "properties": {
          "name": { "type": "keyword" },
          "type": { "type": "keyword" },
          "role": { "type": "keyword" },
          "hierarchy": { "type": "keyword" }
        }
      },
      "geo_location": {
        "properties": {
          "province": { "type": "keyword" },
          "city": { "type": "keyword" },
          "coordinates": { "type": "geo_point" }
        }
      }
    }
  }
}

2. Struktur Indeks Media Sosial & Phone Farm (postings)

Indeks postings menampung konten media sosial multichannel hasil panen armada Phone Farm Swarm (Android APK native):

{
  "mappings": {
    "properties": {
      "id": { "type": "keyword" },
      "platform": { "type": "keyword" },
      "post_id": { "type": "keyword" },
      "author": {
        "properties": {
          "author_id": { "type": "keyword" },
          "username": { "type": "keyword" },
          "display_name": { "type": "text" },
          "followers_count": { "type": "integer" },
          "is_verified": { "type": "boolean" }
        }
      },
      "primary_topic": {
        "properties": {
          "topic_id": { "type": "integer" },
          "topic_name": { "type": "keyword" },
          "primary_pillar": { "type": "keyword" }
        }
      },
      "content": { "type": "text", "analyzer": "indonesian_analyzer" },
      "likes_count": { "type": "integer" },
      "shares_count": { "type": "integer" },
      "comments_count": { "type": "integer" },
      "sentiment": { "type": "keyword" },
      "media_urls": { "type": "keyword" },
      "scraped_at": { "type": "date" }
    }
  }
}

3. Struktur Indeks Analitik & Rekomendasi Pimpinan

Indeks hasil sintesis decision-worker yang membagi arahan strategis per pilar ke dalam 4 peran:

{
  "mappings": {
    "properties": {
      "doc_id": { "type": "keyword" },
      "target_date": { "type": "date", "format": "yyyy-MM-dd" },
      "pillar": { "type": "keyword" },
      "label": { "type": "keyword" },
      "metrics": {
        "properties": {
          "article_count": { "type": "integer" },
          "negative_count": { "type": "integer" },
          "negative_pct": { "type": "float" }
        }
      },
      "insights": {
        "properties": {
          "pemerintahan": { "properties": { "title": { "type": "keyword" }, "text": { "type": "text" } } },
          "pelaku_bisnis": { "properties": { "title": { "type": "keyword" }, "text": { "type": "text" } } },
          "lsm_organisasi": { "properties": { "title": { "type": "keyword" }, "text": { "type": "text" } } },
          "personal": { "properties": { "title": { "type": "keyword" }, "text": { "type": "text" } } }
        }
      },
      "top_topics": {
        "type": "nested",
        "properties": { "key": { "type": "keyword" }, "label": { "type": "keyword" }, "count": { "type": "integer" } }
      },
      "representative_headlines": {
        "type": "nested",
        "properties": { "article_id": { "type": "keyword" }, "title": { "type": "text" }, "source": { "type": "keyword" } }
      }
    }
  }
}

4. Pola Query Agregasi & Point-in-Time (PIT)

A. Point-in-Time (PIT) Scan untuk Konsistensi Korpus

Ketika mengeksekusi pemindaian (deep scanning) korpus berita yang sedang mengalami continuous ingestion, sistem membuka Point-in-Time (PIT) guna membekukan status segmen shard:

# 1. Buka snapshot PIT pada indeks articles
POST /articles/_pit?keep_alive=1m

# Output: { "id": "46ToAwMDYXJ0aWNsZXMS..." }

# 2. Query search menggunakan PIT ID dengan paginasi search_after
POST /_search
{
  "size": 50,
  "query": { "match_all": {} },
  "pit": {
    "id": "46ToAwMDYXJ0aWNsZXMS...",
    "keep_alive": "1m"
  },
  "sort": [ { "published_at": "desc" }, { "_shard_doc": "asc" } ]
}

B. Komputasi Sebaran Sentimen Lintas 5 Pilar

POST /articles/_search
{
  "size": 0,
  "query": {
    "range": {
      "published_at": { "gte": "now-24h", "lte": "now" }
    }
  },
  "aggs": {
    "by_pillar": {
      "terms": { "field": "taxonomy.pillar", "size": 5 },
      "aggs": {
        "sentiment_breakdown": {
          "terms": { "field": "sentiment.label" },
          "aggs": {
            "avg_severity": { "avg": { "field": "severity" } }
          }
        }
      }
    }
  }
}

5. Kebijakan Index Lifecycle Management (ILM)

  • Hot Phase (0–30 Hari): Indeks beroperasi dengan performa tinggi pada NVMe storage untuk read/write aktif dengan 3 shards dan 1 replika.
  • Warm Phase (31–90 Hari): Indeks dialihkan ke mode read-only dengan eksekusi force-merge menjadi 1 segment per shard untuk menghemat memori heap JVM.
  • Cold & Archive (>90 Hari): Indeks historis di-snapshot secara otomatis ke Object Storage (MinIO) pada bucket nexa-backups/elasticsearch/ guna menjaga efisiensi kapasitas disk.

On this page