SNMP테스트서버 문서 원본 보기
←
SNMP테스트서버
둘러보기로 이동
검색으로 이동
문서 편집 작업을 수행할 권한이 없습니다. 다음 이유를 확인해주세요:
요청한 작업은 다음 중 하나의 권한을 가진 사용자에게 제한됩니다:
관리자
, 문서 편집자.
문서의 원본을 보거나 복사할 수 있습니다.
= Rocky Linux 9 통합 모니터링 서버 구축 가이드 - 초보자용 = : Telegraf / InfluxDB / Grafana / Prometheus / Node Exporter / Windows Exporter / Loki / Alloy / rsyslog 작성 목적: 처음 모니터링 서버를 설치하는 사용자가 "왜 이 프로그램이 필요한지", "어떤 순서로 설치하는지", "어디까지 정상이어야 다음 단계로 넘어가는지"를 이해하면서 구축할 수 있도록 작성한다. 본 문서는 실제 운영 과정에서 구성하고 점검한 내용을 바탕으로 정리했으며, 실제 IP 주소와 인증 정보는 문서용 예시 값으로 변경하였다. '''중요:''' 아래 명령을 한 번에 모두 실행하지 않는다. 각 절의 '''정상 확인''' 항목을 통과한 뒤 다음 단계로 진행한다. __TOC__ == 0. 이 문서에서 만들 시스템 == === 0.1 먼저 전체 구조부터 이해하기 === 모니터링을 처음 구성할 때 가장 혼동하기 쉬운 부분은 "Grafana가 모든 데이터를 직접 수집한다"고 생각하는 것이다. Grafana는 주로 '''보여주는 역할'''을 한다. 실제 데이터 수집과 저장은 다른 프로그램이 담당한다. 이 문서에서는 데이터를 크게 두 종류로 나눈다. # '''Metric''' : CPU 사용률, 메모리 사용률, 포트 트래픽, Ping 손실률처럼 숫자로 표현되는 값 # '''Log''' : Syslog, 서버 이벤트, NAS 접근 기록처럼 문자열로 기록되는 이벤트 전체 데이터 흐름은 다음과 같다. <syntaxhighlight lang="bash" line> [Network Switch] | | SNMPv3 v Telegraf | v InfluxDB | +-------------------+ | [Linux / QNAP] | | | | node_exporter | v | Prometheus | | | +-------------------+ | [Network / Server / NAS] | | | | Syslog | v | rsyslog | | | v | Log File | | | v | Alloy | | | v | Loki | | | +-------------------+ | v Grafana </syntaxhighlight> 즉 다음과 같이 기억하면 된다. {| class="wikitable" ! 프로그램 ! 하는 일 ! 쉽게 표현하면 |- | Telegraf | SNMP 장비의 값을 주기적으로 읽음 | 네트워크 장비용 수집기 |- | InfluxDB | Telegraf가 수집한 시계열 값을 저장 | SNMP 데이터 창고 |- | Prometheus | Exporter가 공개한 Metric을 주기적으로 가져와 저장 | 서버 Metric 수집기 + 데이터베이스 |- | Node Exporter | Linux/QNAP의 CPU, 메모리, 디스크, 네트워크 Metric 제공 | Linux 상태 측정기 |- | Windows Exporter | Windows 성능 카운터 Metric 제공 | Windows 상태 측정기 |- | rsyslog | 장비/서버가 보내는 Syslog를 수신하여 파일로 저장 | 로그 수신기 |- | Alloy | 로그 파일을 읽고 필요한 항목을 분리하여 Loki로 전달 | 로그 전달/가공기 |- | Loki | 로그를 저장하고 검색 가능하게 함 | 로그 데이터베이스 |- | Grafana | 위 데이터들을 Dashboard와 Alert로 표현 | 통합 화면 |} === 0.2 왜 하나의 프로그램으로 모두 처리하지 않는가 === SNMP, 서버 Metric, 로그는 데이터 성격이 서로 다르다. 예를 들어 스위치 포트 트래픽은 누적 Counter를 주기적으로 읽어 증가량을 계산해야 한다. 반면 Syslog는 특정 시점에 발생한 문자열 이벤트이다. 따라서 본 구성에서는 각 도구가 가장 잘하는 역할을 분리한다. * Switch SNMP → Telegraf + InfluxDB * Linux / Windows / QNAP OS Metric → Prometheus * Syslog → rsyslog + Alloy + Loki * 최종 화면 → Grafana 이 구조를 이해하면 장애 발생 시 어느 부분을 확인해야 하는지도 쉽게 구분할 수 있다. 예를 들어 Grafana에서 Linux CPU가 보이지 않는다면 다음 순서로 생각한다. <syntaxhighlight lang="bash" line> Grafana 문제인가? ↓ Prometheus에 데이터가 있는가? ↓ Prometheus가 node_exporter를 수집하고 있는가? ↓ node_exporter가 정상 실행 중인가? </syntaxhighlight> == 1. 예제 환경과 주소 == 실제 운영 IP를 문서에 노출하지 않기 위해 RFC 5737 문서용 주소를 사용한다. {| class="wikitable" ! 대상 ! 문서용 IP ! 역할 |- | Monitoring Server | 192.0.2.10 | Grafana / Prometheus / InfluxDB / Telegraf / Loki / Alloy / rsyslog |- | Linux Server | 192.0.2.20 | node_exporter |- | Windows Server | 192.0.2.30 | windows_exporter |- | QNAP NAS | 198.51.100.10 | node_exporter / SNMP / Syslog |- | Switch-01 | 203.0.113.11 | SNMP / Syslog |- | Switch-02 | 203.0.113.12 | SNMP / Syslog |- | Switch-03 | 203.0.113.13 | SNMP / Syslog |} 실제 설치 시 위 주소를 자신의 환경에 맞게 변경한다. == 2. 설치 전에 알아둘 용어 == === 2.1 Metric === Metric은 시간에 따라 변화하는 숫자 데이터이다. 예: * CPU 32% * Memory 71% * Switch Port RX 120 Mbps * Ping Packet Loss 0% * HDD Temperature 43°C 이런 값은 시간 흐름에 따라 그래프로 보는 것이 중요하다. === 2.2 Counter와 Gauge === SNMP에서 자주 만나는 개념이다. '''Gauge'''는 현재 값을 의미한다. 예: <syntaxhighlight lang="bash" line> CPU Usage = 35 Temperature = 44 operStatus = 1 </syntaxhighlight> '''Counter'''는 계속 증가하는 누적값이다. 예: <syntaxhighlight lang="bash" line> ifHCInOctets = 1234567890123 </syntaxhighlight> 이 값을 그대로 그래프로 그리면 계속 증가하기만 한다. 그래서 실제 트래픽은 현재 값과 이전 값의 차이를 시간으로 나누어 계산한다. <syntaxhighlight lang="bash" line> 현재 Counter - 이전 Counter --------------------------- 경과 시간 </syntaxhighlight> Octet은 8 bit이므로 bps를 구하려면 다시 8을 곱한다. Grafana/Flux에서는 이를 derivative()로 처리한다. === 2.3 Label과 Tag === 장비가 여러 대일 때 어떤 데이터가 어느 장비인지 구분하기 위한 값이다. 예: <syntaxhighlight lang="bash" line> source=203.0.113.11 if_name=port24 index=24 </syntaxhighlight> Prometheus에서는 주로 '''label''', InfluxDB에서는 '''tag'''라고 부른다. === 2.4 Exporter === Prometheus가 직접 Linux 내부 명령을 실행하는 것은 아니다. Exporter가 시스템 정보를 HTTP Metric 형식으로 공개하고 Prometheus가 이를 가져간다. Node Exporter 예: <syntaxhighlight lang="bash" line> http://192.0.2.20:9100/metrics </syntaxhighlight> Windows Exporter: <syntaxhighlight lang="bash" line> http://192.0.2.30:9182/metrics </syntaxhighlight> == 3. Rocky Linux 9 기본 준비 == 이 장에서는 모니터링 프로그램을 설치하기 전에 OS가 정상인지 확인한다. === 3.1 시스템 상태 확인 === <syntaxhighlight lang="bash" line> cat /etc/rocky-release uname -m ip -br address ip route df -hT free -h getenforce ss -lntup </syntaxhighlight> 각 명령의 의미: {| class="wikitable" ! 명령 ! 확인 내용 |- | cat /etc/rocky-release | Rocky Linux 버전 |- | uname -m | CPU Architecture, 일반적인 x86 서버는 x86_64 |- | ip -br address | 서버 IP 주소 |- | ip route | Gateway 및 Routing |- | df -hT | 디스크 용량 |- | free -h | 메모리 |- | getenforce | SELinux 상태 |- | ss -lntup | 현재 사용 중인 TCP/UDP Port |} '''초보자 주의:''' SELinux가 Enforcing이라고 해서 바로 Disabled로 변경하지 않는다. 이후 permission 문제가 발생하면 원인을 확인하고 필요한 정책 또는 Context를 수정한다. === 3.2 기존 설정 백업 === 기존 운영 서버에 구축하는 경우 설정을 변경하기 전에 백업한다. <syntaxhighlight lang="bash" line> umask 077 MON_BACKUP="/root/monitor-backup-$(date +%Y%m%d-%H%M%S)" install -d -m 700 "$MON_BACKUP" cp -a /etc/rsyslog.conf "$MON_BACKUP/" 2>/dev/null || true cp -a /etc/rsyslog.d "$MON_BACKUP/" 2>/dev/null || true cp -a /etc/chrony.conf "$MON_BACKUP/" 2>/dev/null || true cp -a /etc/firewalld "$MON_BACKUP/" 2>/dev/null || true rpm -qa | sort > "$MON_BACKUP/packages-before.txt" </syntaxhighlight> 백업 경로 확인: <syntaxhighlight lang="bash" line> echo "$MON_BACKUP" ls -al "$MON_BACKUP" </syntaxhighlight> === 3.3 기본 관리 도구 설치 === <syntaxhighlight lang="bash" line> dnf install -y \ curl \ ca-certificates \ gnupg2 \ vim-enhanced \ net-snmp-utils \ rsyslog \ logrotate \ chrony \ tcpdump \ iputils \ sysstat \ policycoreutils-python-utils \ dnf-plugins-core \ jq </syntaxhighlight> 주요 패키지 용도: * net-snmp-utils : snmpget, snmpwalk 테스트 * tcpdump : 실제 Packet 수신 확인 * jq : JSON 출력 가독성 개선 * sysstat : iostat 등 서버 I/O 점검 * chrony : 시간 동기화 * policycoreutils-python-utils : SELinux 진단/정책 작업 === 3.4 시간 동기화 === 모니터링에서는 서버와 장비 시간이 맞지 않으면 장애 시각 비교가 어렵다. Timezone 설정: <syntaxhighlight lang="bash" line> timedatectl set-timezone Asia/Seoul </syntaxhighlight> chronyd 시작: <syntaxhighlight lang="bash" line> systemctl enable --now chronyd systemctl restart chronyd </syntaxhighlight> 확인: <syntaxhighlight lang="bash" line> chronyc tracking chronyc sources -v date -Ins </syntaxhighlight> '''정상 확인''' * chronyd가 active * chronyc sources에서 선택된 NTP Source 존재 * 현재 시간이 실제 시간과 크게 다르지 않음 == 4. InfluxDB와 Telegraf를 설치하는 이유 == 스위치의 SNMP 값을 수집하려면 두 프로그램이 필요하다. Telegraf는 장비에 SNMP Query를 보내는 역할을 한다. InfluxDB는 Telegraf가 가져온 값을 시간 순서대로 저장한다. <syntaxhighlight lang="bash" line> Switch --SNMP--> Telegraf --> InfluxDB --> Grafana </syntaxhighlight> == 5. InfluxDB / Telegraf / Grafana 설치 == === 5.1 InfluxData Repository 추가 === 공식 Repository Key를 내려받는다. <syntaxhighlight lang="bash" line> curl -fL https://repos.influxdata.com/influxdata-archive.key \ -o /etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata gpg --show-keys --with-fingerprint \ /etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata rpm --import /etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata </syntaxhighlight> Repository 파일 생성: <syntaxhighlight lang="bash" line> vi /etc/yum.repos.d/influxdata.repo </syntaxhighlight> <syntaxhighlight lang="bash" line> [influxdata] name=InfluxData Repository - Stable baseurl=https://repos.influxdata.com/stable/$basearch/main enabled=1 gpgcheck=1 gpgkey=file:///etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata sslverify=1 </syntaxhighlight> === 5.2 Grafana Repository 추가 === <syntaxhighlight lang="bash" line> vi /etc/yum.repos.d/grafana.repo </syntaxhighlight> <syntaxhighlight lang="bash" line> [grafana] name=Grafana OSS repository baseurl=https://rpm.grafana.com repo_gpgcheck=1 enabled=1 gpgcheck=1 gpgkey=https://rpm.grafana.com/gpg.key sslverify=1 </syntaxhighlight> === 5.3 Repository 확인 후 설치 === <syntaxhighlight lang="bash" line> dnf makecache dnf list --showduplicates \ telegraf \ influxdb2 \ influxdb2-cli \ grafana </syntaxhighlight> 설치: <syntaxhighlight lang="bash" line> dnf install -y \ telegraf \ influxdb2 \ influxdb2-cli \ grafana </syntaxhighlight> 버전 확인: <syntaxhighlight lang="bash" line> rpm -q telegraf influxdb2 influxdb2-cli grafana telegraf --version influxd version influx version </syntaxhighlight> '''왜 버전을 기록하는가?''' 나중에 설정 문법 또는 Plugin 동작이 달라졌을 때 원인을 판단하기 쉽기 때문이다. 실제 구성 과정에서는 Telegraf 1.40.0 환경에서 동작을 확인하였다. == 6. InfluxDB 초기 설정 == === 6.1 InfluxDB를 localhost에만 Bind === 현재 구성에서는 InfluxDB를 외부 장비가 직접 접근할 필요가 없다. Telegraf와 Grafana가 같은 서버에 있으므로 127.0.0.1에만 Bind하면 공격 표면을 줄일 수 있다. <syntaxhighlight lang="bash" line> install -d -m 755 \ /etc/systemd/system/influxdb.service.d vi /etc/systemd/system/influxdb.service.d/10-listen.conf </syntaxhighlight> <syntaxhighlight lang="bash" line> [Service] Environment="INFLUXD_HTTP_BIND_ADDRESS=127.0.0.1:8086" </syntaxhighlight> 적용: <syntaxhighlight lang="bash" line> systemctl daemon-reload systemctl enable --now influxdb systemctl restart influxdb systemctl status influxdb --no-pager </syntaxhighlight> Listen 확인: <syntaxhighlight lang="bash" line> ss -lntp | grep ':8086' </syntaxhighlight> Health Check: <syntaxhighlight lang="bash" line> curl -fsS http://127.0.0.1:8086/health </syntaxhighlight> '''정상 확인''' 127.0.0.1:8086에서 LISTEN하고 health 응답이 정상이어야 한다. === 6.2 InfluxDB 최초 Setup === <syntaxhighlight lang="bash" line> influx setup </syntaxhighlight> 초기 입력 예: {| class="wikitable" ! 항목 ! 예시 ! 설명 |- | Username | monitor-admin | InfluxDB 관리 계정 |- | Organization | network | 관련 데이터의 논리적 그룹 |- | Bucket | snmp_raw | SNMP 데이터 저장 위치 |- | Retention | 720h | 30일 보존 |} 확인: <syntaxhighlight lang="bash" line> influx bucket list --org network </syntaxhighlight> '''Bucket이란?''' 일반 Database의 Database 또는 Table과 완전히 같지는 않지만, 처음에는 "Metric 저장 공간" 정도로 이해하면 된다. === 6.3 Token을 서비스별로 나누는 이유 === Telegraf는 데이터를 쓰기만 하면 되고 Grafana는 읽기만 하면 된다. 따라서 하나의 관리자 Token을 모든 프로그램에 넣지 않는다. 권장: {| class="wikitable" ! Token ! 권한 |- | telegraf-write | snmp_raw Write |- | grafana-read | snmp_raw Read |- | Operator Token | 관리자 전용 |} 실제 Token 값은 Wiki에 기록하지 않는다. == 7. Telegraf 기본 설정 == === 7.1 Telegraf의 역할 === Telegraf는 일정 시간마다 스위치에 SNMP Query를 보내고 결과를 InfluxDB에 저장한다. <syntaxhighlight lang="bash" line> Telegraf | | UDP/161 SNMP GET v Switch Telegraf | | HTTP 8086 v InfluxDB </syntaxhighlight> === 7.2 인증 정보를 설정 파일과 분리 === SNMP Password와 InfluxDB Token을 여러 설정 파일에 직접 적으면 관리가 어렵다. 환경변수 파일로 분리한다. <syntaxhighlight lang="bash" line> vi /etc/telegraf/monitor.env </syntaxhighlight> <syntaxhighlight lang="bash" line> INFLUX_WRITE_TOKEN='<WRITE_TOKEN>' ZYXEL_SNMP_USER='<SNMP_USER>' ZYXEL_SNMP_AUTH='<AUTH_PASSWORD>' ZYXEL_SNMP_PRIV='<PRIV_PASSWORD>' QNAP_SNMP_USER='<QNAP_SNMP_USER>' QNAP_SNMP_AUTH='<QNAP_AUTH_PASSWORD>' QNAP_SNMP_PRIV='<QNAP_PRIV_PASSWORD>' </syntaxhighlight> 권한 제한: <syntaxhighlight lang="bash" line> chown root:root /etc/telegraf/monitor.env chmod 600 /etc/telegraf/monitor.env </syntaxhighlight> systemd가 환경 파일을 읽도록 설정: <syntaxhighlight lang="bash" line> install -d -m 755 \ /etc/systemd/system/telegraf.service.d vi /etc/systemd/system/telegraf.service.d/10-monitor-env.conf </syntaxhighlight> <syntaxhighlight lang="bash" line> [Service] EnvironmentFile=/etc/telegraf/monitor.env </syntaxhighlight> === 7.3 Telegraf Main 설정 === 원본 백업: <syntaxhighlight lang="bash" line> cp -a /etc/telegraf/telegraf.conf \ /etc/telegraf/telegraf.conf.orig </syntaxhighlight> 편집: <syntaxhighlight lang="bash" line> vi /etc/telegraf/telegraf.conf </syntaxhighlight> 기본 예: <syntaxhighlight lang="bash" line> [agent] interval = "30s" round_interval = true flush_interval = "5s" precision = "1ms" metric_batch_size = 1000 metric_buffer_limit = 20000 omit_hostname = false snmp_translator = "gosmi" [[outputs.influxdb_v2]] urls = ["http://127.0.0.1:8086"] token = "${INFLUX_WRITE_TOKEN}" organization = "network" bucket = "snmp_raw" [[inputs.internal]] [[inputs.cpu]] percpu = false totalcpu = true [[inputs.mem]] [[inputs.disk]] mount_points = ["/"] [[inputs.net]] </syntaxhighlight> 각 값의 의미: * interval = 30s : 기본 수집 주기 * flush_interval = 5s : 모아둔 Metric을 DB에 보내는 주기 * outputs.influxdb_v2 : 수집 결과를 어느 InfluxDB에 저장할지 지정 * inputs.internal : Telegraf 자신의 상태 * inputs.cpu/mem/disk/net : 모니터링 서버 자체 상태 == 8. SNMP가 정상인지 먼저 수동으로 확인 == Telegraf 설정 전에 SNMP 자체가 되는지 확인하는 것이 중요하다. Telegraf가 안 된다고 바로 Telegraf 설정만 수정하면 실제 원인이 SNMP ACL인지 Password인지 구분하기 어렵다. === 8.1 SNMPv3 계정 파일 === <syntaxhighlight lang="bash" line> install -d -m 700 /root/.snmp vi /root/.snmp/snmp.conf </syntaxhighlight> 예: <syntaxhighlight lang="bash" line> defVersion 3 defSecurityName <SNMP_USER> defSecurityLevel authPriv defAuthType SHA defAuthPassphrase <SNMP_AUTH_PASSWORD> defPrivType AES defPrivPassphrase <SNMP_PRIV_PASSWORD> </syntaxhighlight> 권한: <syntaxhighlight lang="bash" line> chmod 600 /root/.snmp/snmp.conf </syntaxhighlight> === 8.2 Uptime 조회 === <syntaxhighlight lang="bash" line> snmpget -v3 \ -t 2 \ -r 0 \ -On \ 203.0.113.11 \ .1.3.6.1.2.1.1.3.0 </syntaxhighlight> 응답이 나오면 최소한 다음이 정상이다. * IP Routing * UDP/161 * SNMP User * Authentication * Privacy Password * SNMP View === 8.3 Interface 이름 확인 === <syntaxhighlight lang="bash" line> snmpwalk -v3 \ -t 2 \ -r 0 \ -On \ 203.0.113.11 \ .1.3.6.1.2.1.31.1.1.1.1 </syntaxhighlight> 여기서 마지막 숫자가 ifIndex이다. 예를 들어 결과가 다음과 같다고 가정한다. <syntaxhighlight lang="bash" line> .1.3.6.1.2.1.31.1.1.1.1.24 = STRING: port24 </syntaxhighlight> 이 경우 ifIndex는 24이다. '''중요:''' 물리 Port 24번이 항상 ifIndex 24라는 의미는 아니다. 실제 SNMP 결과를 기준으로 한다. == 9. Telegraf SNMP 설정 == === 9.1 모델별 파일로 나누는 이유 === 제조사가 같아도 모델마다 CPU/Memory OID와 SNMP 암호화 방식이 다를 수 있다. 따라서 하나의 거대한 파일보다는 모델별 파일로 나눈다. <syntaxhighlight lang="bash" line> /etc/telegraf/telegraf.d/ ├── 10-zyxel-gs1900.conf ├── 11-zyxel-gs1920.conf ├── 20-zyxel-es3128.conf ├── 30-icmp_check.conf └── 40-qnap_nas.conf </syntaxhighlight> === 9.2 GS1920 계열 === 실제 확인된 특성: * SNMPv3 SHA + DES * CPU/Memory Private OID 사용 CPU: <syntaxhighlight lang="bash" line> .1.3.6.1.4.1.890.1.15.3.49.1.7.0 </syntaxhighlight> Memory: <syntaxhighlight lang="bash" line> Total .1.3.6.1.4.1.890.1.15.3.50.1.1.1.3.1 Used .1.3.6.1.4.1.890.1.15.3.50.1.1.1.4.1 Percent .1.3.6.1.4.1.890.1.15.3.50.1.1.1.5.1 </syntaxhighlight> === 9.3 GS1900 계열 === CPU: <syntaxhighlight lang="bash" line> .1.3.6.1.4.1.890.1.15.3.2.4.0 </syntaxhighlight> Memory: <syntaxhighlight lang="bash" line> .1.3.6.1.4.1.890.1.15.3.2.5.0 </syntaxhighlight> 실제 구축 과정에서는 SNMP 응답이 느려 다음 값이 안정적이었다. <syntaxhighlight lang="bash" line> timeout = "5s" retries = 1 </syntaxhighlight> '''왜 Timeout을 늘렸는가?''' 수동 snmpget은 되는데 Telegraf에서만 간헐적으로 누락된다면 Telegraf timeout이 장비 응답시간보다 짧을 수 있다. Timeout을 무조건 크게 설정하기보다는 실제 응답을 보고 조정한다. === 9.4 ES-3128GP === 실제 확인된 SNMPv3 방식: * SHA * AES sysObjectID: <syntaxhighlight lang="bash" line> .1.3.6.1.4.1.7800.1.190 </syntaxhighlight> Port Metric은 수집 가능하지만 CPU/Memory Private OID는 정상 확인되지 않아 Dashboard에서 억지로 0으로 표시하지 않는다. '''수집되지 않는 값과 0은 의미가 다르다.''' CPU를 조회할 수 없는 장비에 CPU 0%라고 표시하면 운영자가 정상 상태라고 오해할 수 있다. === 9.5 기본 Port Metric === 가능하면 제조사 Private MIB보다 표준 IF-MIB를 먼저 사용한다. 주요 항목: * ifName * ifAlias * ifSpeed * ifHCInOctets * ifHCOutOctets * ifOperStatus * ifInErrors * ifOutErrors * ifInDiscards * ifOutDiscards * Broadcast * Multicast == 10. ICMP Ping 수집 == SNMP가 정상이어도 장비 자체가 네트워크에서 사라질 수 있다. 따라서 Ping 결과도 별도로 저장한다. <syntaxhighlight lang="bash" line> vi /etc/telegraf/telegraf.d/30-icmp_check.conf </syntaxhighlight> 예: <syntaxhighlight lang="bash" line> [[inputs.ping]] urls = [ "203.0.113.11", "203.0.113.12", "203.0.113.13" ] method = "native" count = 3 deadline = 2.0 interval = 10.0 </syntaxhighlight> === 10.1 CAP_NET_RAW가 필요한 이유 === native Ping은 Raw Socket을 사용한다. root로 telegraf --test를 실행하면 정상인데 systemd 서비스에서는 Ping이 실패할 수 있다. 이 경우 서비스에 CAP_NET_RAW를 부여한다. <syntaxhighlight lang="bash" line> systemctl edit telegraf </syntaxhighlight> <syntaxhighlight lang="bash" line> [Service] CapabilityBoundingSet=CAP_NET_RAW AmbientCapabilities=CAP_NET_RAW </syntaxhighlight> 반영: <syntaxhighlight lang="bash" line> systemctl daemon-reload systemctl restart telegraf </syntaxhighlight> == 11. Telegraf 설정 검사 == 서비스를 재시작하기 전에 설정 오류를 확인한다. <syntaxhighlight lang="bash" line> telegraf \ --config /etc/telegraf/telegraf.conf \ --config-directory /etc/telegraf/telegraf.d \ --test </syntaxhighlight> '''주의:''' --test는 Metric을 화면에 출력하지만 일반적으로 Output 저장까지 검증하는 용도가 아니다. 서비스 계정 조건까지 확인하려면 다음 방식이 더 정확하다. <syntaxhighlight lang="bash" line> systemd-run \ --unit=telegraf-config-check \ --wait \ --pipe \ --collect \ -p User=telegraf \ -p Group=telegraf \ -p EnvironmentFile=/etc/telegraf/monitor.env \ -p CapabilityBoundingSet=CAP_NET_RAW \ -p AmbientCapabilities=CAP_NET_RAW \ /usr/bin/telegraf \ --config /etc/telegraf/telegraf.conf \ --config-directory /etc/telegraf/telegraf.d \ --test </syntaxhighlight> 정상 후 서비스 시작: <syntaxhighlight lang="bash" line> systemctl enable --now telegraf systemctl restart telegraf systemctl status telegraf --no-pager journalctl -u telegraf \ -n 100 \ --no-pager </syntaxhighlight> 로그에서 주의할 문자열: <syntaxhighlight lang="bash" line> timeout unauthorized permission denied buffer drop address already in use </syntaxhighlight> == 12. Grafana 설치 후 처음 해야 할 작업 == Grafana는 데이터를 직접 수집하지 않는다. 먼저 InfluxDB와 Prometheus 같은 데이터소스를 등록해야 한다. === 12.1 Grafana 서비스 === <syntaxhighlight lang="bash" line> systemctl enable --now grafana-server systemctl status grafana-server --no-pager curl -fsS \ http://127.0.0.1:3000/api/health </syntaxhighlight> 웹 접속: <syntaxhighlight lang="bash" line> http://192.0.2.10:3000 </syntaxhighlight> 내부망에서만 사용할 경우에도 방화벽 접근 범위를 관리망으로 제한하는 것이 좋다. === 12.2 InfluxDB Datasource 등록 === Grafana 메뉴: <syntaxhighlight lang="bash" line> Connections → Data sources → Add data source → InfluxDB </syntaxhighlight> 설정: {| class="wikitable" ! 항목 ! 값 |- | Query Language | Flux |- | URL | http://127.0.0.1:8086 |- | Organization | network |- | Token | grafana-read Token |- | Default Bucket | snmp_raw |} Save & Test 성공 여부를 확인한다. == 13. 첫 번째 SNMP Dashboard 만들기 == 처음부터 모든 장비를 한 화면에 넣지 않는다. '''한 대의 스위치 + 한 포트'''로 정상 그래프를 만든 뒤 확대하는 것이 가장 쉽다. === 13.1 원시 Counter 확인 === Explore에서 다음과 같이 최근 데이터를 확인한다. <syntaxhighlight lang="bash" line> from(bucket: "snmp_raw") |> range(start: -10m) |> limit(n: 20) </syntaxhighlight> 여기서 실제 Measurement, source, index, if_name 값을 확인한다. === 13.2 Port RX/TX 계산 === 누적 Octet Counter를 bps로 변환한다. <syntaxhighlight lang="bash" line> from(bucket: "snmp_raw") |> range(start: v.timeRangeStart, stop: v.timeRangeStop) |> filter(fn: (r) => r._measurement == "switch_port" ) |> filter(fn: (r) => r.source == "203.0.113.11" ) |> filter(fn: (r) => r._field == "in_octets" or r._field == "out_octets" ) |> derivative( unit: 1s, nonNegative: true ) |> map(fn: (r) => ({ r with _value: r._value * 8.0 })) </syntaxhighlight> Grafana Unit: <syntaxhighlight lang="bash" line> bits/sec </syntaxhighlight> === 13.3 Error / Discard === Error나 Discard 역시 Counter이므로 증가율을 표시한다. <syntaxhighlight lang="bash" line> from(bucket: "snmp_raw") |> range(start: v.timeRangeStart, stop: v.timeRangeStop) |> filter(fn: (r) => r._measurement == "switch_port" ) |> filter(fn: (r) => r._field == "in_errors" or r._field == "out_errors" or r._field == "in_discards" or r._field == "out_discards" ) |> derivative( unit: 1s, nonNegative: true ) </syntaxhighlight> '''해석''' * Error 증가 → 물리계층, Frame, Interface 문제 가능 * Discard 증가 → Queue, Buffer, QoS, 혼잡 등에 의해 폐기될 가능성 * 값이 0인 상태가 일반적이지만 장비 특성과 Traffic에 따라 해석해야 함 === 13.4 operStatus === operStatus는 Counter가 아니다. 따라서 derivative를 사용하지 않는다. 일반적인 값: <syntaxhighlight lang="bash" line> 1 = up 2 = down </syntaxhighlight> Grafana Value Mapping으로 사람이 읽기 쉬운 문자열로 변환한다. == 14. Prometheus를 추가하는 이유 == SNMP는 네트워크 장비 상태에 좋지만 Linux/Windows 서버 자원 모니터링에는 Exporter + Prometheus가 더 편하다. Prometheus 구조: <syntaxhighlight lang="bash" line> Node Exporter ----\ \ Windows Exporter ---> Prometheus ---> Grafana / QNAP Exporter -----/ </syntaxhighlight> Prometheus는 일정 주기마다 Exporter HTTP Endpoint에 접속하여 Metric을 가져간다. 이를 '''Pull 방식'''이라고 한다. == 15. Prometheus 설치 == === 15.1 사용자와 디렉터리 === <syntaxhighlight lang="bash" line> useradd \ --system \ --no-create-home \ --shell /sbin/nologin \ prometheus install -d \ -o prometheus \ -g prometheus \ /etc/prometheus \ /var/lib/prometheus </syntaxhighlight> Prometheus 공식 바이너리를 준비한 뒤: <syntaxhighlight lang="bash" line> install -m 0755 \ prometheus \ /usr/local/bin/prometheus install -m 0755 \ promtool \ /usr/local/bin/promtool </syntaxhighlight> 버전 확인: <syntaxhighlight lang="bash" line> prometheus --version promtool --version </syntaxhighlight> === 15.2 prometheus.yml === <syntaxhighlight lang="bash" line> vi /etc/prometheus/prometheus.yml </syntaxhighlight> <syntaxhighlight lang="bash" line> global: scrape_interval: 30s scrape_configs: - job_name: prometheus static_configs: - targets: - "127.0.0.1:9090" labels: server_name: monitoring-server - job_name: node static_configs: - targets: - "127.0.0.1:9100" labels: server_name: monitoring-server - targets: - "192.0.2.20:9100" labels: server_name: linux-server - job_name: windows static_configs: - targets: - "192.0.2.30:9182" labels: server_name: windows-server - job_name: qnap static_configs: - targets: - "198.51.100.10:9100" labels: server_name: nas-01 </syntaxhighlight> '''scrape_interval = 30s'''는 Prometheus가 30초마다 각 Target을 조회한다는 의미이다. === 15.3 설정 검사 === YAML은 들여쓰기에 매우 민감하다. 실제 구축 중에도 job을 scrape_configs 바깥에 잘못 넣으면 Prometheus가 시작하지 못했다. 반드시 검사한다. <syntaxhighlight lang="bash" line> promtool check config \ /etc/prometheus/prometheus.yml </syntaxhighlight> === 15.4 systemd 등록 === <syntaxhighlight lang="bash" line> vi /etc/systemd/system/prometheus.service </syntaxhighlight> <syntaxhighlight lang="bash" line> [Unit] Description=Prometheus Wants=network-online.target After=network-online.target [Service] User=prometheus Group=prometheus ExecStart=/usr/local/bin/prometheus \ --config.file=/etc/prometheus/prometheus.yml \ --storage.tsdb.path=/var/lib/prometheus \ --storage.tsdb.retention.time=30d \ --storage.tsdb.retention.size=20GB \ --web.listen-address=127.0.0.1:9090 Restart=always [Install] WantedBy=multi-user.target </syntaxhighlight> 권한: <syntaxhighlight lang="bash" line> chown -R prometheus:prometheus \ /etc/prometheus \ /var/lib/prometheus </syntaxhighlight> 시작: <syntaxhighlight lang="bash" line> systemctl daemon-reload systemctl enable --now prometheus systemctl status prometheus --no-pager </syntaxhighlight> Health: <syntaxhighlight lang="bash" line> curl -s \ http://127.0.0.1:9090/-/healthy </syntaxhighlight> == 16. Linux Node Exporter 설치 == === 16.1 Node Exporter가 하는 일 === Node Exporter는 Linux Kernel과 /proc, /sys 등의 정보를 읽어 Prometheus 형식으로 공개한다. 대표 Metric: * CPU * Memory * Load * Filesystem * Disk I/O * Network Traffic * Network Error/Drop * Uptime * File Descriptor * 일부 Hardware Metric === 16.2 서비스 계정 === <syntaxhighlight lang="bash" line> useradd \ --system \ --no-create-home \ --shell /sbin/nologin \ node_exporter </syntaxhighlight> 바이너리 설치: <syntaxhighlight lang="bash" line> install -m 0755 \ node_exporter \ /usr/local/bin/node_exporter </syntaxhighlight> === 16.3 systemd === <syntaxhighlight lang="bash" line> vi /etc/systemd/system/node_exporter.service </syntaxhighlight> <syntaxhighlight lang="bash" line> [Unit] Description=Node Exporter After=network-online.target Wants=network-online.target [Service] User=node_exporter Group=node_exporter ExecStart=/usr/local/bin/node_exporter Restart=always [Install] WantedBy=multi-user.target </syntaxhighlight> <syntaxhighlight lang="bash" line> systemctl daemon-reload systemctl enable --now node_exporter systemctl status node_exporter --no-pager </syntaxhighlight> Metric 확인: <syntaxhighlight lang="bash" line> curl -s \ http://127.0.0.1:9100/metrics \ | head </syntaxhighlight> 다른 서버에 설치한 경우 Prometheus 서버에서 확인: <syntaxhighlight lang="bash" line> curl -s \ http://192.0.2.20:9100/metrics \ | head </syntaxhighlight> '''정상 확인''' Metric 문자열이 여러 줄 출력되면 Exporter 자체는 정상이다. Prometheus에서 다음 Query도 확인한다. <syntaxhighlight lang="bash" line> up </syntaxhighlight> 값: <syntaxhighlight lang="bash" line> 1 = 정상 Scrape 0 = Scrape 실패 </syntaxhighlight> == 17. Windows Exporter == Windows에서는 windows_exporter를 설치한다. 기본 Port: <syntaxhighlight lang="bash" line> 9182/tcp </syntaxhighlight> Prometheus에서 다음 주소를 조회할 수 있어야 한다. <syntaxhighlight lang="bash" line> http://192.0.2.30:9182/metrics </syntaxhighlight> Windows PowerShell 테스트: <syntaxhighlight lang="bash" line> Invoke-WebRequest \ -UseBasicParsing \ http://127.0.0.1:9182/metrics </syntaxhighlight> 성능 카운터가 비정상인 경우 실제 구축 과정에서 다음 복구 절차를 사용하였다. <syntaxhighlight lang="bash" line> lodctr /R winmgmt /resyncperf </syntaxhighlight> 그 후: <syntaxhighlight lang="bash" line> Restart-Service windows_exporter </syntaxhighlight> == 18. Grafana에 Prometheus 등록 == Grafana 메뉴: <syntaxhighlight lang="bash" line> Connections → Data sources → Prometheus </syntaxhighlight> URL: <syntaxhighlight lang="bash" line> http://127.0.0.1:9090 </syntaxhighlight> Save & Test 후 Explore에서: <syntaxhighlight lang="bash" line> up </syntaxhighlight> Target별 1이 표시되면 정상이다. == 19. Linux Server Dashboard 이해하기 == 처음에는 다음 항목만 만든다. * Uptime * CPU * Memory * Load Average * Filesystem * Network RX/TX 이후 익숙해지면 다음을 추가한다. * Disk I/O * Network Error/Drop * inode * Swap * Process * Application Log === 19.1 CPU Usage === Node Exporter는 CPU 시간을 Mode별 누적으로 제공한다. idle을 제외한 비율로 CPU 사용률을 계산할 수 있다. <syntaxhighlight lang="bash" line> 100 - ( avg by(instance) ( rate( node_cpu_seconds_total{ mode="idle" }[5m] ) ) * 100 ) </syntaxhighlight> === 19.2 Memory Usage === <syntaxhighlight lang="bash" line> 100 * ( 1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes ) </syntaxhighlight> === 19.3 Filesystem Usage === Filesystem은 tmpfs, overlay, snapshot 등 운영자가 원하지 않는 Mount가 함께 보일 수 있다. 따라서 실제 데이터 Mount Point를 확인한 뒤 필터링한다. == 20. QNAP TS-264 모니터링 == QNAP은 한 가지 수집 방식만 사용하면 정보가 부족하다. 따라서 세 경로를 함께 사용한다. <syntaxhighlight lang="bash" line> QNAP | +-- node_exporter --> Prometheus | +-- SNMP ----------> Telegraf --> InfluxDB | +-- Syslog --------> rsyslog --> Alloy --> Loki </syntaxhighlight> 역할: {| class="wikitable" ! 항목 ! 데이터 소스 |- | CPU / Memory / Network | Node Exporter |- | Filesystem | Node Exporter |- | Disk I/O | Node Exporter |- | HDD Model / 온도 / 상태 | SNMP |- | RAID / Volume | SNMP |- | Event Log | Syslog |- | Access Log | Syslog |} === 20.1 QNAP SNMPv3 === 실제 TS-264 환경에서는 다음 조합을 사용하였다. * Authentication: SHA * Privacy: DES QNAP은 해당 구성에서 AES가 동작하지 않았다. 따라서 다른 장비에서 AES가 된다는 이유로 QNAP도 AES라고 가정하지 않는다. === 20.2 주요 Measurement === <syntaxhighlight lang="bash" line> QNAP_TS264 qnap_disk qnap_raid qnap_storage_pool qnap_volume </syntaxhighlight> === 20.3 Volume 값이 11 GiB로 잘못 보이는 문제 === 실제 QTS에서 DataVol1은 약 11.33 TiB인데 SNMP 원시값을 Grafana에서 byte로 바로 처리하면 약 11.3 GiB로 표시되는 문제가 있었다. 실제 장비의 qnap_volume capacity/free 값이 KiB 성격으로 반환되어 1024 배 보정이 필요하였다. <syntaxhighlight lang="bash" line> |> map(fn: (r) => ({ r with capacity_bytes: uint(v: r.capacity_bytes) * uint(v: 1024), free_bytes: uint(v: r.free_bytes) * uint(v: 1024), used_bytes: ( uint(v: r.capacity_bytes) - uint(v: r.free_bytes) ) * uint(v: 1024), used_percent: if float(v: r.capacity_bytes) > 0.0 then ( float(v: r.capacity_bytes) - float(v: r.free_bytes) ) / float(v: r.capacity_bytes) * 100.0 else 0.0 })) </syntaxhighlight> '''왜 Percent에는 1024가 필요 없는가?''' 분자와 분모가 같은 단위이므로 비율 계산에서는 단위가 상쇄된다. === 20.4 실제 Data Volume 찾기 === Node Exporter가 QNAP 내부의 Snapshot Mount까지 모두 보여주므로 처음에는 여러 개의 11.33 TiB Filesystem이 보일 수 있다. 실제 사용자 Data Volume은 다음이었다. <syntaxhighlight lang="bash" line> mountpoint="/share/CACHEDEV1_DATA" device="/dev/mapper/cachedev1" fstype="ext4" </syntaxhighlight> QNAP Snapshot: <syntaxhighlight lang="bash" line> /mnt/snapshot/1/10001 /mnt/snapshot/1/10002 ... </syntaxhighlight> Dashboard에서는 Snapshot을 제외하고 실제 Volume만 표시한다. 사용률: <syntaxhighlight lang="bash" line> 100 * ( 1 - node_filesystem_avail_bytes{ instance="198.51.100.10:9100", mountpoint="/share/CACHEDEV1_DATA" } / node_filesystem_size_bytes{ instance="198.51.100.10:9100", mountpoint="/share/CACHEDEV1_DATA" } ) </syntaxhighlight> 실제 QTS 화면과 약 2.86%로 일치하는 것을 확인하였다. === 20.5 QNAP Network는 bond0만 표시 === QNAP에는 내부 Interface와 Virtual Interface가 여러 개 보일 수 있다. 실제 외부 Traffic을 담당하는 bond0만 필터링한다. RX: <syntaxhighlight lang="bash" line> rate( node_network_receive_bytes_total{ instance="198.51.100.10:9100", device="bond0" }[$__rate_interval] ) * 8 </syntaxhighlight> TX: <syntaxhighlight lang="bash" line> rate( node_network_transmit_bytes_total{ instance="198.51.100.10:9100", device="bond0" }[$__rate_interval] ) * 8 </syntaxhighlight> === 20.6 Network Error / Drop === RX Error: <syntaxhighlight lang="bash" line> rate( node_network_receive_errs_total{ instance="198.51.100.10:9100", device="bond0" }[$__rate_interval] ) </syntaxhighlight> TX Error: <syntaxhighlight lang="bash" line> rate( node_network_transmit_errs_total{ instance="198.51.100.10:9100", device="bond0" }[$__rate_interval] ) </syntaxhighlight> RX Drop: <syntaxhighlight lang="bash" line> rate( node_network_receive_drop_total{ instance="198.51.100.10:9100", device="bond0" }[$__rate_interval] ) </syntaxhighlight> TX Drop: <syntaxhighlight lang="bash" line> rate( node_network_transmit_drop_total{ instance="198.51.100.10:9100", device="bond0" }[$__rate_interval] ) </syntaxhighlight> === 20.7 Disk I/O === QNAP에는 md, dm, loop 등 내부 Device가 많으므로 물리 Disk sd*만 표시한다. Read Bytes/sec: <syntaxhighlight lang="bash" line> rate( node_disk_read_bytes_total{ instance="198.51.100.10:9100", device=~"sd[a-z]+" }[$__rate_interval] ) </syntaxhighlight> Write Bytes/sec: <syntaxhighlight lang="bash" line> rate( node_disk_written_bytes_total{ instance="198.51.100.10:9100", device=~"sd[a-z]+" }[$__rate_interval] ) </syntaxhighlight> Read IOPS: <syntaxhighlight lang="bash" line> rate( node_disk_reads_completed_total{ instance="198.51.100.10:9100", device=~"sd[a-z]+" }[$__rate_interval] ) </syntaxhighlight> Write IOPS: <syntaxhighlight lang="bash" line> rate( node_disk_writes_completed_total{ instance="198.51.100.10:9100", device=~"sd[a-z]+" }[$__rate_interval] ) </syntaxhighlight> == 21. Syslog를 왜 별도로 구성하는가 == Metric만으로는 다음과 같은 내용을 알기 어렵다. * Port가 왜 Down 되었는가 * 사용자가 NAS에 어떤 파일을 접근했는가 * 서비스가 언제 재시작되었는가 * 인증 실패가 발생했는가 이런 이벤트는 Log가 필요하다. 본 구성의 로그 흐름: <syntaxhighlight lang="bash" line> Device | | Syslog v rsyslog | v Log File | v Alloy | v Loki | v Grafana </syntaxhighlight> == 22. rsyslog 구성 == === 22.1 Network Syslog === 네트워크 장비 로그: <syntaxhighlight lang="bash" line> /var/log/network-syslog/events.log </syntaxhighlight> 기본 수신: <syntaxhighlight lang="bash" line> TCP/UDP 514 </syntaxhighlight> === 22.2 Server Syslog === 서버용 로그를 네트워크 장비와 분리하면 Grafana에서 Query하기 쉽다. 예: <syntaxhighlight lang="bash" line> /var/log/server-syslog/events.log </syntaxhighlight> 수신 Port: <syntaxhighlight lang="bash" line> TCP 5514 </syntaxhighlight> === 22.3 QNAP Event와 Access Log 분리 === QNAP QuLog의 두 성격이 다르므로 포트부터 분리한다. {| class="wikitable" ! 종류 ! 포트 ! 파일 ! 의미 |- | Event | TCP 5515 | /var/log/qnap/event.log | 시스템/서비스 이벤트 |- | Access | TCP 5516 | /var/log/qnap/access.log | 사용자/파일 접근 |} Template: <syntaxhighlight lang="bash" line> template(name="QnapSyslogLine" type="string" string="%timegenerated:::date-rfc3339% src=%fromhost-ip% severity=%syslogseverity-text% host=%hostname% %syslogtag%%msg:::sp-if-no-1st-sp%%msg%\n") </syntaxhighlight> 이렇게 저장하면 한 줄이 대략 다음 구조가 된다. <syntaxhighlight lang="bash" line> 2026-09-13T23:40:03+09:00 \ src=198.51.100.10 \ severity=info \ host=nas-01 \ qulogd: ... </syntaxhighlight> Event ruleset: <syntaxhighlight lang="bash" line> ruleset(name="QnapEventLog") { action( type="omfile" file="/var/log/qnap/event.log" template="QnapSyslogLine" fileOwner="root" fileGroup="alloy" fileCreateMode="0640" dirOwner="root" dirGroup="alloy" dirCreateMode="0750" createDirs="on" ) stop } input( type="imtcp" port="5515" ruleset="QnapEventLog" ) </syntaxhighlight> Access ruleset: <syntaxhighlight lang="bash" line> ruleset(name="QnapAccessLog") { action( type="omfile" file="/var/log/qnap/access.log" template="QnapSyslogLine" fileOwner="root" fileGroup="alloy" fileCreateMode="0640" dirOwner="root" dirGroup="alloy" dirCreateMode="0750" createDirs="on" ) stop } input( type="imtcp" port="5516" ruleset="QnapAccessLog" ) </syntaxhighlight> === 22.4 rsyslog 설정 검사 === 재시작 전에 반드시 검사한다. <syntaxhighlight lang="bash" line> rsyslogd -N1 </syntaxhighlight> 오류가 없으면: <syntaxhighlight lang="bash" line> systemctl restart rsyslog ss -lntp \ | grep -E ':5514|:5515|:5516' </syntaxhighlight> 실제 로그 확인: <syntaxhighlight lang="bash" line> tail -f \ /var/log/qnap/access.log </syntaxhighlight> '''문제 분리 방법''' <syntaxhighlight lang="bash" line> tcpdump에는 Packet이 안 보임 → 장비 송신/방화벽/라우팅 확인 tcpdump에는 보이지만 파일이 안 생김 → rsyslog 설정 확인 파일은 생기지만 Grafana에 안 보임 → Alloy/Loki 확인 </syntaxhighlight> == 23. Loki와 Alloy의 역할 == rsyslog가 파일을 만드는 것만으로 Grafana가 그 파일을 검색할 수 있는 것은 아니다. Loki가 로그 저장소 역할을 하고 Alloy가 파일을 읽어 Loki에 넣는다. <syntaxhighlight lang="bash" line> /var/log/... | v Alloy | v Loki | v Grafana </syntaxhighlight> 현재 구성에서는 Loki가 localhost 3100에서 동작한다. <syntaxhighlight lang="bash" line> 127.0.0.1:3100 </syntaxhighlight> Alloy 로컬 관리 Endpoint: <syntaxhighlight lang="bash" line> 127.0.0.1:12345 </syntaxhighlight> == 24. Alloy 설정 이해하기 == === 24.1 Network Syslog === <syntaxhighlight lang="bash" line> loki.source.file "network_syslog" { targets = [ { __path__ = "/var/log/network-syslog/events.log", job = "network-syslog", }, ] forward_to = [ loki.process.network_syslog.receiver ] } </syntaxhighlight> '''source.file'''은 어떤 파일을 읽을지 지정한다. '''job'''은 Loki에서 로그 종류를 구분하기 위한 Label이다. 그 다음 regex로 한 줄을 분해한다. <syntaxhighlight lang="bash" line> loki.process "network_syslog" { stage.regex { expression = `^(?P<received_at>\S+) src=(?P<device_ip>\S+) severity=(?P<severity>\S+) host=(?P<device_host>\S+) (?P<message>.*)$` } stage.timestamp { source = "received_at" format = "RFC3339Nano" action_on_failure = "skip" } stage.labels { values = { device_ip = "" severity = "" } } forward_to = [ loki.write.local.receiver ] } </syntaxhighlight> 여기서 다음 Label이 생긴다. <syntaxhighlight lang="bash" line> job device_ip severity </syntaxhighlight> '''device_host를 Label로 추가하려면''' <syntaxhighlight lang="bash" line> stage.labels { values = { device_ip = "" device_host = "" severity = "" } } </syntaxhighlight> 이는 이후 Alert Mail에서 Hostname까지 표시하고 싶을 때 유용하다. === 24.2 QNAP Access === <syntaxhighlight lang="bash" line> loki.source.file "qnap_access" { targets = [ { __path__ = "/var/log/qnap/access.log", job = "qnap-access", }, ] forward_to = [ loki.process.qnap_access.receiver ] } loki.process "qnap_access" { stage.regex { expression = `^(?P<received_at>\S+) src=(?P<nas_ip>\S+) severity=(?P<severity>\S+) host=(?P<nas_host>\S+) (?P<message>.*)$` } stage.timestamp { source = "received_at" format = "RFC3339Nano" action_on_failure = "skip" } stage.labels { values = { nas_ip = "" nas_host = "" severity = "" log_type = "access" } } forward_to = [ loki.write.local.receiver ] } </syntaxhighlight> === 24.3 Loki Write === <syntaxhighlight lang="bash" line> loki.write "local" { endpoint { url = "http://127.0.0.1:3100/loki/api/v1/push" } } </syntaxhighlight> 즉 Alloy에서 처리한 로그를 로컬 Loki로 전송한다. === 24.4 설정 검사 === <syntaxhighlight lang="bash" line> alloy validate \ /etc/alloy/config.alloy </syntaxhighlight> 정상이라면 아무 오류 없이 종료된다. 반영: <syntaxhighlight lang="bash" line> systemctl restart alloy systemctl status alloy \ --no-pager </syntaxhighlight> Loki에 실제 Job이 만들어졌는지 확인: <syntaxhighlight lang="bash" line> curl -s \ 'http://127.0.0.1:3100/loki/api/v1/label/job/values' \ | jq </syntaxhighlight> 중요한 점은 Alloy 설정에 job을 적었다고 바로 Loki에 Label이 생기는 것이 아니라 '''실제 로그가 최소 한 번 Loki에 저장되어야''' 조회 결과에 나타난다는 것이다. == 25. Alloy Permission 문제 해결 사례 == 실제 구축 과정에서 QNAP 로그 파일은 존재하지만 Alloy가 다음 오류를 출력하였다. <syntaxhighlight lang="bash" line> failed to tail file stat failed permission denied </syntaxhighlight> 먼저 경로 전체 권한 확인: <syntaxhighlight lang="bash" line> namei -l \ /var/log/qnap/access.log </syntaxhighlight> 정상 예: <syntaxhighlight lang="bash" line> drwxr-x--- root alloy qnap -rw-r----- root alloy access.log </syntaxhighlight> Alloy 계정으로 직접 읽기: <syntaxhighlight lang="bash" line> sudo -u alloy \ head /var/log/qnap/access.log </syntaxhighlight> 그래도 실패하면 SELinux 확인: <syntaxhighlight lang="bash" line> getenforce ausearch \ -m AVC \ -ts recent \ | grep -Ei 'alloy|qnap' ls -Zd /var/log/qnap ls -Z /var/log/qnap/access.log </syntaxhighlight> '''문제 해결 원칙''' 권한 문제가 있다고 SELinux를 바로 끄지 않는다. 다음 순서로 본다. # 파일 Owner/Group # 파일 Mode # 상위 Directory execute 권한 # Alloy Service User # SELinux Context / AVC == 26. Grafana에 Loki 추가 == Grafana: <syntaxhighlight lang="bash" line> Connections → Data sources → Add data source → Loki </syntaxhighlight> URL: <syntaxhighlight lang="bash" line> http://127.0.0.1:3100 </syntaxhighlight> Explore에서 Network 로그 확인: <syntaxhighlight lang="bash" line> {job="network-syslog"} </syntaxhighlight> QNAP Access: <syntaxhighlight lang="bash" line> {job="qnap-access"} </syntaxhighlight> QNAP Event: <syntaxhighlight lang="bash" line> {job="qnap-event"} </syntaxhighlight> == 27. Grafana Alert를 처음 구성할 때 알아둘 점 == Grafana Alert는 Dashboard의 그래프와 달리 최종적으로 '''숫자 하나 또는 Alert Instance별 숫자'''를 평가해야 한다. Range Query를 그대로 Alert Condition에 사용하면 다음 오류가 발생할 수 있다. <syntaxhighlight lang="bash" line> looks like time series data, only reduced data can be alerted on </syntaxhighlight> 따라서 다음 중 하나를 사용한다. # Query를 Instant로 구성 # Range Query → Reduce → Threshold == 28. Syslog Critical Alert == Level 3(Error) 이상: <syntaxhighlight lang="bash" line> sum by (device_ip, severity) ( count_over_time( { job="network-syslog", severity=~"emerg|alert|crit|err" }[1m] ) ) </syntaxhighlight> 권장: <syntaxhighlight lang="bash" line> A = Loki Instant Query B = Threshold A IS ABOVE 0 </syntaxhighlight> Summary: <syntaxhighlight lang="bash" line> [Syslog 경고] {{ $labels.device_ip }} {{ $labels.severity }} </syntaxhighlight> Description: <syntaxhighlight lang="bash" line> 장비 {{ $labels.device_ip }} 에서 Syslog Level 3 이상 로그가 발생했습니다. Severity: {{ $labels.severity }} 최근 1분 발생 건수: {{ $values.A.Value }} </syntaxhighlight> === 28.1 Alert Mail에 실제 로그 내용이 없는 이유 === count_over_time()은 문자열 로그를 숫자로 집계한다. 따라서 Query 결과에는 주로 다음만 남는다. * Label * 발생 건수 로그 message 전체를 Loki Label로 만들면 종류가 지나치게 많아져 Cardinality 문제가 발생할 수 있으므로 권장하지 않는다. 메일에는 다음 정도를 넣고 실제 내용은 Grafana Explore에서 확인하는 구조가 안정적이다. * Device IP * Hostname * Severity * 발생 건수 * Grafana Link == 29. ICMP Down Alert == Flux: <syntaxhighlight lang="bash" line> from(bucket: "snmp_raw") |> range(start: -5m) |> filter(fn: (r) => r._measurement == "ping" and r._field == "percent_packet_loss" ) |> group( columns: ["url"] ) |> last() |> keep( columns: [ "_time", "_value", "url" ] ) </syntaxhighlight> Alert 구조: <syntaxhighlight lang="bash" line> A = Flux Query B = Reduce Last Strict C = Threshold B > 99 Pending = 2m </syntaxhighlight> No Data 정책: <syntaxhighlight lang="bash" line> Keep Last State </syntaxhighlight> '''왜 Keep Last State인가?''' Metric이나 Log Source가 순간적으로 No Data가 되었을 때 별도의 DatasourceNoData 메일이 발생하면서 실제 장비 Label이 없는 알림이 발송될 수 있다. No Data 자체를 별도 장애로 관리할 필요가 있다면 별도의 수집 상태 Alert를 구성하는 것이 더 명확하다. == 30. 최종 Dashboard 구성 방법 == 처음 설치한 사용자는 Dashboard를 한 번에 완성하려 하지 말고 다음 단계로 만든다. # 데이터소스 Save & Test # Explore에서 실제 데이터 확인 # Stat 패널 하나 생성 # Time Series 하나 생성 # 장비 한 대 정상 확인 # 변수 추가 # 여러 장비로 확대 # Alert 추가 === 30.1 Network Switch Dashboard === 권장 Row: {| class="wikitable" ! Row ! 표시 내용 ! 목적 |- | Device 상태 | Device Name / Uptime / CPU / Memory / Ping | 장비 자체 상태 |- | Port 상태 | ifName / ifAlias / Speed / operStatus | Link 상태 |- | Traffic | RX / TX bps | 대역폭 사용량 |- | Packet | Unicast / Broadcast / Multicast PPS | Broadcast 폭주 및 Packet 패턴 |- | Error | Error / Discard | 품질 및 혼잡 징후 |- | Syslog | warning / err / crit | 이벤트 원인 확인 |} === 30.2 Linux Dashboard === 권장: * Uptime * CPU * Memory * Load * Disk Capacity * Disk I/O * Network RX/TX * Network Error/Drop * Application/Security Log === 30.3 Windows Dashboard === 권장: * Exporter UP * Uptime * CPU * Memory * Logical Disk * Disk I/O * Network * Windows Event 연동 여부 === 30.4 QNAP Dashboard === 현재 실제 운영 구성을 기준으로 다음과 같이 구성한다. 상단: * Node Exporter UP * Uptime * CPU Usage * Memory Usage * HDD 최고 온도 중단: * CPU / Memory / Load * bond0 RX / TX * Disk 상태 * DataVol1 Volume * DataVol1 Filesystem 사용률 * bond0 Error / Drop * Physical Disk I/O / IOPS 하단: * QNAP Event Log * QNAP Access Log 제거한 항목: * 상단 RAID 상태 * Storage Pool 상태 * Storage Pool 사용률 * RAID 상세 Table * Storage Pool 상세 Table 삭제 이유는 Dashboard를 운영자가 빠르게 읽을 수 있도록 단순화하기 위해서이다. RAID/Storage Pool 상세는 필요 시 별도의 상세 Dashboard에서 확인할 수 있다. == 31. 초보자가 자주 만나는 문제 == === 31.1 Telegraf --test는 되는데 서비스는 안 됨 === root 권한에서는 Ping이 되지만 telegraf 서비스 계정에서는 Raw Socket 권한이 없을 수 있다. 확인: <syntaxhighlight lang="bash" line> journalctl -u telegraf \ -n 100 \ --no-pager </syntaxhighlight> native Ping이면 CAP_NET_RAW 설정을 확인한다. === 31.2 snmpget은 되는데 Telegraf에서 일부 장비만 누락 === 가능한 원인: * timeout이 너무 짧음 * retries가 0 * OID가 모델과 다름 * DES/AES 조합이 장비와 다름 * SNMP View 제한 먼저 동일 OID를 snmpget으로 확인한다. === 31.3 Prometheus가 재시작 반복 === 가장 먼저: <syntaxhighlight lang="bash" line> promtool check config \ /etc/prometheus/prometheus.yml </syntaxhighlight> YAML 들여쓰기를 확인한다. 그 다음: <syntaxhighlight lang="bash" line> journalctl -u prometheus \ -n 100 \ --no-pager </syntaxhighlight> === 31.4 Grafana에서 Filesystem이 너무 많이 보임 === Linux/QNAP에는 tmpfs, loop, snapshot, container mount 등이 존재할 수 있다. 전체를 보여주기보다 운영자가 실제 사용하는 Mount Point만 필터링한다. QNAP 예: <syntaxhighlight lang="bash" line> mountpoint="/share/CACHEDEV1_DATA" </syntaxhighlight> === 31.5 Loki에 Job이 안 보임 === 설정에 job이 있다고 바로 나타나는 것이 아니다. 실제 Log가 Loki에 한 번 이상 저장되어야 한다. 확인 순서: <syntaxhighlight lang="bash" line> tail -n 10 \ /var/log/qnap/access.log journalctl -u alloy \ -n 100 \ --no-pager curl -s \ 'http://127.0.0.1:3100/loki/api/v1/label/job/values' \ | jq </syntaxhighlight> == 32. 구축 완료 후 점검 Checklist == {| class="wikitable" ! 확인 ! 항목 |- | □ | Rocky Linux 시간 동기화 정상 |- | □ | InfluxDB health 정상 |- | □ | Telegraf 서비스 정상 |- | □ | SNMP 장비별 응답 확인 |- | □ | InfluxDB에 SNMP 데이터 저장 |- | □ | Prometheus health 정상 |- | □ | Node Exporter Target UP |- | □ | Windows Exporter Target UP |- | □ | QNAP Node Exporter Target UP |- | □ | rsyslog Network 로그 수신 |- | □ | QNAP Event / Access Log 분리 수신 |- | □ | Alloy Permission 정상 |- | □ | Loki job 확인 |- | □ | Grafana InfluxDB Save & Test |- | □ | Grafana Prometheus Save & Test |- | □ | Grafana Loki Save & Test |- | □ | Switch Traffic 그래프 정상 |- | □ | Error/Discard 그래프 정상 |- | □ | Linux/Windows Dashboard 정상 |- | □ | QNAP DataVol1 용량 QTS와 일치 |- | □ | QNAP Snapshot 제외 |- | □ | QNAP bond0만 Network 표시 |- | □ | Syslog Alert Mail에 장비 IP / Severity 표시 |- | □ | ICMP Down Alert 정상 |} == 33. 전체 장애 점검 순서 == 모니터링이 안 될 때 Grafana부터 무작정 수정하지 않는다. === 33.1 SNMP === <syntaxhighlight lang="bash" line> Switch ↓ snmpget ↓ Telegraf --test ↓ Telegraf Service ↓ InfluxDB ↓ Grafana Explore ↓ Dashboard </syntaxhighlight> === 33.2 Prometheus === <syntaxhighlight lang="bash" line> Exporter /metrics ↓ Prometheus Target ↓ Prometheus Query ↓ Grafana Explore ↓ Dashboard </syntaxhighlight> === 33.3 Syslog === <syntaxhighlight lang="bash" line> Device 송신 ↓ tcpdump ↓ rsyslog ↓ Log File ↓ Alloy ↓ Loki ↓ Grafana Explore ↓ Dashboard / Alert </syntaxhighlight> 이 순서대로 확인하면 문제 지점을 빠르게 좁힐 수 있다. == 34. 운영 원칙 요약 == * 처음부터 모든 장비를 추가하지 않는다. * 한 장비, 한 Port를 먼저 완성한다. * 설정 변경 후 항상 해당 서비스의 검사 명령을 실행한다. * Counter와 현재 상태값을 구분한다. * No Data와 0을 같은 의미로 처리하지 않는다. * SNMP Password와 Token을 Wiki에 저장하지 않는다. * Grafana는 수집기가 아니라 시각화 계층임을 기억한다. * QNAP처럼 제조사별 특이사항은 실제 값과 Vendor 화면을 대조한다. * Alert는 처음부터 너무 많이 만들지 않는다. * 정상 상태의 기준 데이터를 먼저 쌓은 후 임계값을 결정한다. * 장애 시에는 수집 경로를 앞단부터 순서대로 확인한다. == 35. 최종 구성 요약 == <syntaxhighlight lang="bash" line> +----------------------+ | Grafana | +----------+-----------+ | +------------------+------------------+ | | | v v v Prometheus InfluxDB Loki ^ ^ ^ | | | Exporters Telegraf Alloy ^ ^ ^ | | | Linux / Windows / QNAP SNMP Device Syslog Files ^ | rsyslog ^ | Network / Server / NAS </syntaxhighlight> 이 구조가 완성되면 다음 세 종류의 정보를 한 Grafana에서 확인할 수 있다. # 장비와 서버의 현재 상태 # 시간에 따른 성능 변화 # 장애 시점의 이벤트 로그 초보자는 먼저 "데이터가 어디서 생성되어 어디를 거쳐 Grafana까지 도착하는지"를 이해한 뒤 설정값을 수정하는 것이 가장 중요하다.
SNMP테스트서버
문서로 돌아갑니다.
둘러보기 메뉴
개인 도구
로그인
associated-pages
문서
토론
한국어
보기
읽기
원본 보기
역사 보기
더 보기
검색
둘러보기
대문
최근 바뀜
임의 문서로
미디어위키 도움말
특수 문서 목록
도구
여기를 가리키는 문서
가리키는 글의 최근 바뀜
문서 정보