Gleam Lab · Blog Archive
Blog Page 24
Technical exploration and engineering notes, 655 articles in total.
Big Data 191 - Elasticsearch Cluster Planning & Tuning: Node Roles, Shards, Replicas, Write and Search Checklist
Master / Data / Coordinating node responsibilities and production role isolation strategies, capacity planning calculations.
Big Data 192 - DataX 3.0 Architecture & Practice
Scenario: Offline sync MySQL/HDFS/Hive/OTS/ODPS and other heterogeneous data sources, batch migration and data warehouse ETL.
Spark Super WordCount: Text Cleaning & MySQL Persistence
This is article 75 in the Big Data series, on top of basic WordCount add text preprocessing and database persistence, build a near-production word frequency pipeline.
Spark Serialization & RDD Execution Principle
This is article 76 in the Big Data series, systematically reviewing Spark process communication mechanism, serialization strategy and RDD execution principle.
Big Data 189 - Nginx JSON Logs to ELK: ZK + Kafka + Elasticsearch 7.3.0 + Kibana 7.3.0
Configure Nginx logformat json to output structured accesslog (containing @timestamp, requesttime, status, requesturi, ua and other fields).
Filebeat → Kafka → Logstash → Elasticsearch Practice
Filebeat collects Nginx access.log to Kafka, and Logstash consumes, parses embedded JSON by field conditions, enriches metadata, and writes structured logs to Elasticsear...
Big Data 187 - Logstash Filter Plugin Practice
Filter is responsible for parsing, transforming, filtering events. Multiple Filters execute in configured order.
Big Data 188 - Logstash Output Plugin Practice
Output is the final stage of Logstash pipeline, responsible for outputting processed data to target system.
Big Data 185 - Logstash 7 Getting Started: stdin/file Collection, sincedb, start_position & Error Quick Reference
Logstash 7 getting started tutorial, covering stdin/file collection, sincedb mechanism and start_position effect conditions, with error quick reference table
Big Data 186 - Logstash JDBC vs Syslog Input: Principles, Scenarios & Reusable Configurations
Logstash Input plugin comparison, breakdown technical differences between JDBC Input and Syslog collection pipeline, applicable scenarios and key configs.
Spark Scala WordCount Implementation
Implement distributed WordCount using Spark + Scala and Spark + Java, detailed RDD five-step processing flow, Maven project configuration and spark-submit command.
Spark Scala Practice: Pi Estimation & Mutual Friends
Deep dive into Spark RDD programming through two classic cases: Monte Carlo method distributed Pi estimation, and mutual friends analysis in social networks with two appr...
Big Data 183 - Elasticsearch Concurrency Conflicts & Optimistic Lock
Elasticsearch concurrency conflicts (inventory deduction read-modify-write) breakdown write overwrite cause, and gives engineering solution using ES optimistic.
Big Data 184 - Elasticsearch Doc Values Mechanism Detailed
Disk columnar data structure generated at indexing time, optimized for sorting, aggregation and script values
Big Data 181 - Elasticsearch Segment Merge & Disk Directory Breakdown
Explains why refresh causes small segment increase, how segment merge merges small segments into large ones in background and cleans deleted documents.
Big Data 182 - Elasticsearch Inverted Index Underlying Breakdown
Article details core data structure of Elasticsearch inverted index: Terms Dictionary, Posting List, FST (Finite State Transducer) and SkipList how accelerate.
Big Data 179 - Elasticsearch Inverted Index and Read/Write Process
This article deeply analyzes Elasticsearch's inverted index principle based on Lucene, and document read/write flow.
Big Data 180 - Elasticsearch Near Real-Time Search: Segment, Refresh and Flush
Article details core mechanism of Elasticsearch near real-time search, including Lucene Segment, Memory Buffer, File System Cache, Refresh, Flush and Translog.
Spark Action Operations Overview
This is article 72 in the Big Data series, systematically reviewing Spark RDD Action operators.
Elasticsearch Aggregation Practice: Metrics Aggregations & Bucket Aggregations
Covers complete practice of Metrics Aggregations and Bucket Aggregations, applicable to common Elasticsearch 7.x / 8.x versions in 2025.