SNMP테스트서버: 두 판 사이의 차이
편집 요약 없음 |
편집 요약 없음 |
||
| (같은 사용자의 중간 판 하나는 보이지 않습니다) | |||
| 1번째 줄: | 1번째 줄: | ||
= Rocky Linux 9 | = Rocky Linux 9 통합 모니터링 서버 구축 가이드 - 초보자용 = | ||
: Telegraf / InfluxDB / Grafana / Prometheus / Node Exporter / Windows Exporter / Loki / Alloy / rsyslog | |||
작성 목적: 처음 모니터링 서버를 설치하는 사용자가 "왜 이 프로그램이 필요한지", "어떤 순서로 설치하는지", "어디까지 정상이어야 다음 단계로 넘어가는지"를 이해하면서 구축할 수 있도록 작성한다. | |||
본 문서는 실제 운영 과정에서 구성하고 점검한 내용을 바탕으로 정리했으며, 실제 IP 주소와 인증 정보는 문서용 예시 값으로 변경하였다. | |||
''' | '''중요:''' 아래 명령을 한 번에 모두 실행하지 않는다. 각 절의 '''정상 확인''' 항목을 통과한 뒤 다음 단계로 진행한다. | ||
__TOC__ | |||
== 0. 이 문서에서 만들 시스템 == | |||
=== 0.1 먼저 전체 구조부터 이해하기 === | |||
모니터링을 처음 구성할 때 가장 혼동하기 쉬운 부분은 "Grafana가 모든 데이터를 직접 수집한다"고 생각하는 것이다. | |||
Grafana는 주로 '''보여주는 역할'''을 한다. 실제 데이터 수집과 저장은 다른 프로그램이 담당한다. | |||
이 문서에서는 데이터를 크게 두 종류로 나눈다. | |||
# '''Metric''' : CPU 사용률, 메모리 사용률, 포트 트래픽, Ping 손실률처럼 숫자로 표현되는 값 | |||
# '''Log''' : Syslog, 서버 이벤트, NAS 접근 기록처럼 문자열로 기록되는 이벤트 | |||
전체 데이터 흐름은 다음과 같다. | |||
<syntaxhighlight lang="bash" line> | |||
[Network Switch] | |||
| | |||
| SNMPv3 | |||
v | |||
Telegraf | |||
| | |||
v | |||
InfluxDB | |||
| | |||
+-------------------+ | |||
| | |||
[Linux / QNAP] | | |||
| | | |||
| node_exporter | | |||
v | | |||
Prometheus | | |||
| | | |||
+-------------------+ | |||
| | |||
[Network / Server / NAS] | | |||
| | | |||
| Syslog | | |||
v | | |||
rsyslog | | |||
| | | |||
v | | |||
Log File | | |||
| | | |||
v | | |||
Alloy | | |||
| | | |||
v | | |||
Loki | | |||
| | | |||
+-------------------+ | |||
| | |||
v | |||
Grafana | |||
</syntaxhighlight> | |||
즉 다음과 같이 기억하면 된다. | |||
{| class="wikitable" | {| class="wikitable" | ||
! 프로그램 | |||
! 하는 일 | |||
! 쉽게 표현하면 | |||
|- | |||
| Telegraf | |||
| SNMP 장비의 값을 주기적으로 읽음 | |||
| 네트워크 장비용 수집기 | |||
|- | |||
| InfluxDB | |||
| Telegraf가 수집한 시계열 값을 저장 | |||
| SNMP 데이터 창고 | |||
|- | |||
| Prometheus | |||
| Exporter가 공개한 Metric을 주기적으로 가져와 저장 | |||
| 서버 Metric 수집기 + 데이터베이스 | |||
|- | |- | ||
| Node Exporter | |||
| Linux/QNAP의 CPU, 메모리, 디스크, 네트워크 Metric 제공 | |||
| Linux 상태 측정기 | |||
|- | |- | ||
| | | Windows Exporter | ||
| | | Windows 성능 카운터 Metric 제공 | ||
| | | Windows 상태 측정기 | ||
|- | |- | ||
| | | rsyslog | ||
| | | 장비/서버가 보내는 Syslog를 수신하여 파일로 저장 | ||
| | | 로그 수신기 | ||
|- | |- | ||
| | | Alloy | ||
| | | 로그 파일을 읽고 필요한 항목을 분리하여 Loki로 전달 | ||
| | | 로그 전달/가공기 | ||
|- | |- | ||
| | | Loki | ||
| | | 로그를 저장하고 검색 가능하게 함 | ||
| | | 로그 데이터베이스 | ||
|- | |- | ||
| | | Grafana | ||
| | | 위 데이터들을 Dashboard와 Alert로 표현 | ||
| | | 통합 화면 | ||
|} | |} | ||
SNMP | === 0.2 왜 하나의 프로그램으로 모두 처리하지 않는가 === | ||
SNMP, 서버 Metric, 로그는 데이터 성격이 서로 다르다. | |||
예를 들어 스위치 포트 트래픽은 누적 Counter를 주기적으로 읽어 증가량을 계산해야 한다. 반면 Syslog는 특정 시점에 발생한 문자열 이벤트이다. | |||
따라서 본 구성에서는 각 도구가 가장 잘하는 역할을 분리한다. | |||
* Switch SNMP → Telegraf + InfluxDB | |||
* Linux / Windows / QNAP OS Metric → Prometheus | |||
* Syslog → rsyslog + Alloy + Loki | |||
* 최종 화면 → Grafana | |||
이 구조를 이해하면 장애 발생 시 어느 부분을 확인해야 하는지도 쉽게 구분할 수 있다. | |||
예를 들어 Grafana에서 Linux CPU가 보이지 않는다면 다음 순서로 생각한다. | |||
<syntaxhighlight lang="bash" line> | |||
Grafana 문제인가? | |||
↓ | |||
Prometheus에 데이터가 있는가? | |||
↓ | |||
Prometheus가 node_exporter를 수집하고 있는가? | |||
↓ | |||
node_exporter가 정상 실행 중인가? | |||
</syntaxhighlight> | |||
== 1. 예제 환경과 주소 == | |||
실제 운영 IP를 문서에 노출하지 않기 위해 RFC 5737 문서용 주소를 사용한다. | |||
{| class="wikitable" | {| class="wikitable" | ||
! 대상 | |||
! 문서용 IP | |||
! 역할 | |||
|- | |- | ||
| Monitoring Server | |||
| 192.0.2.10 | |||
| Grafana / Prometheus / InfluxDB / Telegraf / Loki / Alloy / rsyslog | |||
|- | |- | ||
| | | Linux Server | ||
| | | 192.0.2.20 | ||
| node_exporter | |||
|- | |- | ||
| | | Windows Server | ||
| | | 192.0.2.30 | ||
| windows_exporter | |||
|- | |- | ||
| | | QNAP NAS | ||
| | | 198.51.100.10 | ||
| node_exporter / SNMP / Syslog | |||
|- | |- | ||
| | | Switch-01 | ||
| | | 203.0.113.11 | ||
| SNMP / Syslog | |||
|- | |- | ||
| | | Switch-02 | ||
| SNMP | | 203.0.113.12 | ||
| SNMP / Syslog | |||
|- | |- | ||
| | | Switch-03 | ||
| 203.0.113.13 | |||
| SNMP / Syslog | |||
| | |||
| | |||
|} | |} | ||
실제 설치 시 위 주소를 자신의 환경에 맞게 변경한다. | |||
== 2. 설치 전에 알아둘 용어 == | |||
=== | === 2.1 Metric === | ||
Metric은 시간에 따라 변화하는 숫자 데이터이다. | |||
예: | |||
* CPU 32% | |||
* Memory 71% | |||
* Switch Port RX 120 Mbps | |||
* Ping Packet Loss 0% | |||
* HDD Temperature 43°C | |||
이런 값은 시간 흐름에 따라 그래프로 보는 것이 중요하다. | |||
=== 2.2 Counter와 Gauge === | |||
SNMP에서 자주 만나는 개념이다. | |||
'''Gauge'''는 현재 값을 의미한다. | |||
예: | |||
<syntaxhighlight lang="bash" line> | |||
CPU Usage = 35 | |||
Temperature = 44 | |||
operStatus = 1 | |||
</syntaxhighlight> | |||
'''Counter'''는 계속 증가하는 누적값이다. | |||
예: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
ifHCInOctets = 1234567890123 | |||
</syntaxhighlight> | |||
이 값을 그대로 그래프로 그리면 계속 증가하기만 한다. | |||
그래서 실제 트래픽은 현재 값과 이전 값의 차이를 시간으로 나누어 계산한다. | |||
<syntaxhighlight lang="bash" line> | |||
현재 Counter - 이전 Counter | |||
--------------------------- | |||
경과 시간 | |||
</syntaxhighlight> | |||
Octet은 8 bit이므로 bps를 구하려면 다시 8을 곱한다. | |||
Grafana/Flux에서는 이를 derivative()로 처리한다. | |||
=== 2.3 Label과 Tag === | |||
장비가 여러 대일 때 어떤 데이터가 어느 장비인지 구분하기 위한 값이다. | |||
예: | |||
<syntaxhighlight lang="bash" line> | |||
source=203.0.113.11 | |||
if_name=port24 | |||
index=24 | |||
</syntaxhighlight> | |||
Prometheus에서는 주로 '''label''', InfluxDB에서는 '''tag'''라고 부른다. | |||
=== 2.4 Exporter === | |||
Prometheus가 직접 Linux 내부 명령을 실행하는 것은 아니다. | |||
Exporter가 시스템 정보를 HTTP Metric 형식으로 공개하고 Prometheus가 이를 가져간다. | |||
Node Exporter 예: | |||
<syntaxhighlight lang="bash" line> | |||
http://192.0.2.20:9100/metrics | |||
</syntaxhighlight> | |||
Windows Exporter: | |||
<syntaxhighlight lang="bash" line> | |||
http://192.0.2.30:9182/metrics | |||
</syntaxhighlight> | </syntaxhighlight> | ||
== 3. 기본 | == 3. Rocky Linux 9 기본 준비 == | ||
이 장에서는 모니터링 프로그램을 설치하기 전에 OS가 정상인지 확인한다. | |||
=== 3.1 시스템 상태 확인 === | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
cat /etc/rocky-release | cat /etc/rocky-release | ||
uname -m | uname -m | ||
ip -br address | ip -br address | ||
ip route | ip route | ||
df -hT | df -hT | ||
free -h | free -h | ||
getenforce | getenforce | ||
ss -lntup | ss -lntup | ||
</syntaxhighlight> | |||
각 명령의 의미: | |||
{| class="wikitable" | |||
! 명령 | |||
! 확인 내용 | |||
|- | |||
| cat /etc/rocky-release | |||
| Rocky Linux 버전 | |||
|- | |||
| uname -m | |||
| CPU Architecture, 일반적인 x86 서버는 x86_64 | |||
|- | |||
| ip -br address | |||
| 서버 IP 주소 | |||
|- | |||
| ip route | |||
| Gateway 및 Routing | |||
|- | |||
| df -hT | |||
| 디스크 용량 | |||
|- | |||
| free -h | |||
| 메모리 | |||
|- | |||
| getenforce | |||
| SELinux 상태 | |||
|- | |||
| ss -lntup | |||
| 현재 사용 중인 TCP/UDP Port | |||
|} | |||
'''초보자 주의:''' SELinux가 Enforcing이라고 해서 바로 Disabled로 변경하지 않는다. 이후 permission 문제가 발생하면 원인을 확인하고 필요한 정책 또는 Context를 수정한다. | |||
=== 3.2 기존 설정 백업 === | |||
기존 운영 서버에 구축하는 경우 설정을 변경하기 전에 백업한다. | |||
<syntaxhighlight lang="bash" line> | |||
umask 077 | |||
MON_BACKUP="/root/monitor-backup-$(date +%Y%m%d-%H%M%S)" | |||
install -d -m 700 "$MON_BACKUP" | |||
cp -a /etc/rsyslog.conf "$MON_BACKUP/" 2>/dev/null || true | |||
cp -a /etc/rsyslog.d "$MON_BACKUP/" 2>/dev/null || true | |||
cp -a /etc/chrony.conf "$MON_BACKUP/" 2>/dev/null || true | |||
cp -a /etc/firewalld "$MON_BACKUP/" 2>/dev/null || true | |||
rpm -qa | sort > "$MON_BACKUP/packages-before.txt" | |||
</syntaxhighlight> | </syntaxhighlight> | ||
백업 경로 확인: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
echo "$MON_BACKUP" | |||
ls -al "$MON_BACKUP" | |||
</syntaxhighlight> | </syntaxhighlight> | ||
=== 3.3 기본 관리 도구 설치 === | |||
<syntaxhighlight lang="bash" line> | |||
dnf install -y \ | |||
curl \ | |||
ca-certificates \ | |||
gnupg2 \ | |||
vim-enhanced \ | |||
net-snmp-utils \ | |||
rsyslog \ | |||
logrotate \ | |||
chrony \ | |||
tcpdump \ | |||
iputils \ | |||
sysstat \ | |||
policycoreutils-python-utils \ | |||
dnf-plugins-core \ | |||
jq | |||
</syntaxhighlight> | |||
주요 패키지 용도: | |||
* net-snmp-utils : snmpget, snmpwalk 테스트 | |||
* tcpdump : 실제 Packet 수신 확인 | |||
* jq : JSON 출력 가독성 개선 | |||
* sysstat : iostat 등 서버 I/O 점검 | |||
* chrony : 시간 동기화 | |||
* policycoreutils-python-utils : SELinux 진단/정책 작업 | |||
=== 3.4 시간 동기화 === | |||
모니터링에서는 서버와 장비 시간이 맞지 않으면 장애 시각 비교가 어렵다. | |||
Timezone 설정: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
timedatectl set-timezone Asia/Seoul | timedatectl set-timezone Asia/Seoul | ||
</syntaxhighlight> | |||
chronyd 시작: | |||
<syntaxhighlight lang="bash" line> | |||
systemctl enable --now chronyd | systemctl enable --now chronyd | ||
systemctl restart chronyd | systemctl restart chronyd | ||
</syntaxhighlight> | |||
확인: | |||
<syntaxhighlight lang="bash" line> | |||
chronyc tracking | chronyc tracking | ||
chronyc sources -v | chronyc sources -v | ||
date -Ins | date -Ins | ||
</syntaxhighlight> | </syntaxhighlight> | ||
== 4. | '''정상 확인''' | ||
* chronyd가 active | |||
* chronyc sources에서 선택된 NTP Source 존재 | |||
* 현재 시간이 실제 시간과 크게 다르지 않음 | |||
== 4. InfluxDB와 Telegraf를 설치하는 이유 == | |||
스위치의 SNMP 값을 수집하려면 두 프로그램이 필요하다. | |||
Telegraf는 장비에 SNMP Query를 보내는 역할을 한다. | |||
InfluxDB는 Telegraf가 가져온 값을 시간 순서대로 저장한다. | |||
<syntaxhighlight lang="bash" line> | |||
Switch --SNMP--> Telegraf --> InfluxDB --> Grafana | |||
</syntaxhighlight> | |||
== 5. InfluxDB / Telegraf / Grafana 설치 == | |||
=== 5.1 InfluxData Repository 추가 === | |||
공식 Repository Key를 내려받는다. | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
curl -fL https://repos.influxdata.com/influxdata-archive.key \ | curl -fL https://repos.influxdata.com/influxdata-archive.key \ | ||
-o /etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata | -o /etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata | ||
gpg --show-keys --with-fingerprint /etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata | |||
gpg --show-keys --with-fingerprint \ | |||
/etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata | |||
rpm --import /etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata | |||
</syntaxhighlight> | </syntaxhighlight> | ||
Repository 파일 생성: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
vi /etc/yum.repos.d/influxdata.repo | vi /etc/yum.repos.d/influxdata.repo | ||
</syntaxhighlight> | </syntaxhighlight> | ||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
| 181번째 줄: | 443번째 줄: | ||
sslverify=1 | sslverify=1 | ||
</syntaxhighlight> | </syntaxhighlight> | ||
=== | |||
=== 5.2 Grafana Repository 추가 === | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
vi /etc/yum.repos.d/grafana.repo | vi /etc/yum.repos.d/grafana.repo | ||
</syntaxhighlight> | </syntaxhighlight> | ||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
[grafana] | [grafana] | ||
| 196번째 줄: | 460번째 줄: | ||
sslverify=1 | sslverify=1 | ||
</syntaxhighlight> | </syntaxhighlight> | ||
=== 5.3 Repository 확인 후 설치 === | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
dnf makecache | dnf makecache | ||
dnf list --showduplicates telegraf influxdb2 influxdb2-cli grafana | |||
dnf install telegraf influxdb2 influxdb2-cli grafana | dnf list --showduplicates \ | ||
rpm -q telegraf influxdb2 influxdb2-cli grafana | telegraf \ | ||
influxdb2 \ | |||
influxdb2-cli \ | |||
grafana | |||
</syntaxhighlight> | |||
설치: | |||
<syntaxhighlight lang="bash" line> | |||
dnf install -y \ | |||
telegraf \ | |||
influxdb2 \ | |||
influxdb2-cli \ | |||
grafana | |||
</syntaxhighlight> | |||
버전 확인: | |||
<syntaxhighlight lang="bash" line> | |||
rpm -q telegraf influxdb2 influxdb2-cli grafana | |||
telegraf --version | telegraf --version | ||
influxd version | influxd version | ||
influx version | influx version | ||
</syntaxhighlight> | </syntaxhighlight> | ||
'''왜 버전을 기록하는가?''' | |||
나중에 설정 문법 또는 Plugin 동작이 달라졌을 때 원인을 판단하기 쉽기 때문이다. | |||
=== | 실제 구성 과정에서는 Telegraf 1.40.0 환경에서 동작을 확인하였다. | ||
== 6. InfluxDB 초기 설정 == | |||
=== 6.1 InfluxDB를 localhost에만 Bind === | |||
현재 구성에서는 InfluxDB를 외부 장비가 직접 접근할 필요가 없다. | |||
Telegraf와 Grafana가 같은 서버에 있으므로 127.0.0.1에만 Bind하면 공격 표면을 줄일 수 있다. | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
install -d -m 755 /etc/systemd/system/influxdb.service.d | install -d -m 755 \ | ||
vi /etc/systemd/system/influxdb.service.d/10 | /etc/systemd/system/influxdb.service.d | ||
vi /etc/systemd/system/influxdb.service.d/10-listen.conf | |||
</syntaxhighlight> | </syntaxhighlight> | ||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
[Service] | [Service] | ||
Environment="INFLUXD_HTTP_BIND_ADDRESS=127.0.0.1:8086" | Environment="INFLUXD_HTTP_BIND_ADDRESS=127.0.0.1:8086" | ||
</syntaxhighlight> | </syntaxhighlight> | ||
적용: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
systemctl daemon-reload | systemctl daemon-reload | ||
systemctl enable --now influxdb | systemctl enable --now influxdb | ||
systemctl restart influxdb | systemctl restart influxdb | ||
systemctl status influxdb --no-pager | systemctl status influxdb --no-pager | ||
</syntaxhighlight> | |||
Listen 확인: | |||
<syntaxhighlight lang="bash" line> | |||
ss -lntp | grep ':8086' | ss -lntp | grep ':8086' | ||
</syntaxhighlight> | |||
Health Check: | |||
<syntaxhighlight lang="bash" line> | |||
curl -fsS http://127.0.0.1:8086/health | curl -fsS http://127.0.0.1:8086/health | ||
</syntaxhighlight> | </syntaxhighlight> | ||
=== | '''정상 확인''' | ||
127.0.0.1:8086에서 LISTEN하고 health 응답이 정상이어야 한다. | |||
=== 6.2 InfluxDB 최초 Setup === | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
influx setup | influx setup | ||
</syntaxhighlight> | </syntaxhighlight> | ||
초기 입력 예: | |||
{| class="wikitable" | {| class="wikitable" | ||
! 항목 | ! 항목 | ||
! | ! 예시 | ||
! 설명 | |||
|- | |- | ||
| Username | | Username | ||
| | | monitor-admin | ||
| InfluxDB 관리 계정 | |||
| | |||
|- | |- | ||
| Organization | | Organization | ||
| | | network | ||
| 관련 데이터의 논리적 그룹 | |||
|- | |- | ||
| Bucket | | Bucket | ||
| | | snmp_raw | ||
| SNMP 데이터 저장 위치 | |||
|- | |- | ||
| Retention | | Retention | ||
| | | 720h | ||
| 30일 보존 | |||
|} | |} | ||
확인: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
influx bucket list --org network | influx bucket list --org network | ||
</syntaxhighlight> | </syntaxhighlight> | ||
'''Bucket이란?''' | |||
일반 Database의 Database 또는 Table과 완전히 같지는 않지만, 처음에는 "Metric 저장 공간" 정도로 이해하면 된다. | |||
=== 6.3 Token을 서비스별로 나누는 이유 === | |||
Telegraf는 데이터를 쓰기만 하면 되고 Grafana는 읽기만 하면 된다. | |||
따라서 하나의 관리자 Token을 모든 프로그램에 넣지 않는다. | |||
권장: | |||
{| class="wikitable" | {| class="wikitable" | ||
! Token | |||
! 권한 | |||
|- | |- | ||
| telegraf-write | |||
| snmp_raw Write | |||
| | |||
| | |||
|- | |- | ||
| | | grafana-read | ||
| | | snmp_raw Read | ||
|- | |- | ||
| | | Operator Token | ||
| 관리자 | | 관리자 전용 | ||
|} | |} | ||
실제 Token 값은 Wiki에 기록하지 않는다. | |||
== | == 7. Telegraf 기본 설정 == | ||
=== | === 7.1 Telegraf의 역할 === | ||
Telegraf는 일정 시간마다 스위치에 SNMP Query를 보내고 결과를 InfluxDB에 저장한다. | |||
<syntaxhighlight lang="bash" line> | |||
Telegraf | |||
| | |||
| UDP/161 SNMP GET | |||
v | |||
Switch | |||
Telegraf | |||
| | |||
| HTTP 8086 | |||
v | |||
InfluxDB | |||
</syntaxhighlight> | |||
=== 7.2 인증 정보를 설정 파일과 분리 === | |||
SNMP Password와 InfluxDB Token을 여러 설정 파일에 직접 적으면 관리가 어렵다. | |||
환경변수 파일로 분리한다. | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
vi /etc/telegraf/monitor.env | |||
vi / | |||
</syntaxhighlight> | </syntaxhighlight> | ||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
INFLUX_WRITE_TOKEN='<WRITE_TOKEN>' | |||
ZYXEL_SNMP_USER='<SNMP_USER>' | |||
ZYXEL_SNMP_AUTH='<AUTH_PASSWORD>' | |||
ZYXEL_SNMP_PRIV='<PRIV_PASSWORD>' | |||
QNAP_SNMP_USER='<QNAP_SNMP_USER>' | |||
QNAP_SNMP_AUTH='<QNAP_AUTH_PASSWORD>' | |||
QNAP_SNMP_PRIV='<QNAP_PRIV_PASSWORD>' | |||
</syntaxhighlight> | |||
권한 제한: | |||
<syntaxhighlight lang="bash" line> | |||
chown root:root /etc/telegraf/monitor.env | |||
chmod 600 /etc/telegraf/monitor.env | |||
</syntaxhighlight> | </syntaxhighlight> | ||
systemd가 환경 파일을 읽도록 설정: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
install -d -m 755 \ | |||
/etc/systemd/system/telegraf.service.d | |||
vi /etc/systemd/system/telegraf.service.d/10-monitor-env.conf | |||
</syntaxhighlight> | </syntaxhighlight> | ||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
[Service] | |||
EnvironmentFile=/etc/telegraf/monitor.env | |||
</syntaxhighlight> | </syntaxhighlight> | ||
=== | === 7.3 Telegraf Main 설정 === | ||
원본 백업: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
cp -a /etc/telegraf/telegraf.conf \ | |||
/etc/telegraf/telegraf.conf.orig | |||
</syntaxhighlight> | </syntaxhighlight> | ||
편집: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
vi /etc/telegraf/telegraf.conf | vi /etc/telegraf/telegraf.conf | ||
</syntaxhighlight> | </syntaxhighlight> | ||
기본 예: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
[agent] | [agent] | ||
| 377번째 줄: | 708번째 줄: | ||
[[inputs.internal]] | [[inputs.internal]] | ||
[[inputs.cpu]] | [[inputs.cpu]] | ||
percpu = false | percpu = false | ||
totalcpu = true | totalcpu = true | ||
[[inputs.mem]] | [[inputs.mem]] | ||
[[inputs.disk]] | [[inputs.disk]] | ||
mount_points = ["/"] | mount_points = ["/"] | ||
[[inputs.net]] | [[inputs.net]] | ||
</syntaxhighlight> | </syntaxhighlight> | ||
각 값의 의미: | |||
* interval = 30s : 기본 수집 주기 | |||
* flush_interval = 5s : 모아둔 Metric을 DB에 보내는 주기 | |||
* outputs.influxdb_v2 : 수집 결과를 어느 InfluxDB에 저장할지 지정 | |||
* inputs.internal : Telegraf 자신의 상태 | |||
* inputs.cpu/mem/disk/net : 모니터링 서버 자체 상태 | |||
== 8. SNMP가 정상인지 먼저 수동으로 확인 == | |||
Telegraf 설정 전에 SNMP 자체가 되는지 확인하는 것이 중요하다. | |||
Telegraf가 안 된다고 바로 Telegraf 설정만 수정하면 실제 원인이 SNMP ACL인지 Password인지 구분하기 어렵다. | |||
=== 8.1 SNMPv3 계정 파일 === | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
vi / | install -d -m 700 /root/.snmp | ||
vi /root/.snmp/snmp.conf | |||
</syntaxhighlight> | </syntaxhighlight> | ||
예: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
defVersion 3 | |||
defSecurityName <SNMP_USER> | |||
defSecurityLevel authPriv | |||
defAuthType SHA | |||
defAuthPassphrase <SNMP_AUTH_PASSWORD> | |||
defPrivType AES | |||
defPrivPassphrase <SNMP_PRIV_PASSWORD> | |||
</syntaxhighlight> | |||
권한: | |||
<syntaxhighlight lang="bash" line> | |||
chmod 600 /root/.snmp/snmp.conf | |||
</syntaxhighlight> | |||
=== 8.2 Uptime 조회 === | |||
<syntaxhighlight lang="bash" line> | |||
snmpget -v3 \ | |||
-t 2 \ | |||
-r 0 \ | |||
-On \ | |||
203.0.113.11 \ | |||
.1.3.6.1.2.1.1.3.0 | |||
</syntaxhighlight> | |||
응답이 나오면 최소한 다음이 정상이다. | |||
* IP Routing | |||
* UDP/161 | |||
* SNMP User | |||
* Authentication | |||
* Privacy Password | |||
* SNMP View | |||
=== 8.3 Interface 이름 확인 === | |||
<syntaxhighlight lang="bash" line> | |||
snmpwalk -v3 \ | |||
-t 2 \ | |||
-r 0 \ | |||
-On \ | |||
203.0.113.11 \ | |||
.1.3.6.1.2.1.31.1.1.1.1 | |||
</syntaxhighlight> | </syntaxhighlight> | ||
여기서 마지막 숫자가 ifIndex이다. | |||
예를 들어 결과가 다음과 같다고 가정한다. | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
.1.3.6.1.2.1.31.1.1.1.1.24 = STRING: port24 | |||
</syntaxhighlight> | </syntaxhighlight> | ||
이 경우 ifIndex는 24이다. | |||
'''중요:''' 물리 Port 24번이 항상 ifIndex 24라는 의미는 아니다. 실제 SNMP 결과를 기준으로 한다. | |||
== 9. Telegraf SNMP 설정 == | |||
=== 9.1 모델별 파일로 나누는 이유 === | |||
제조사가 같아도 모델마다 CPU/Memory OID와 SNMP 암호화 방식이 다를 수 있다. | |||
따라서 하나의 거대한 파일보다는 모델별 파일로 나눈다. | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
/etc/telegraf/telegraf.d/ | |||
├── 10-zyxel-gs1900.conf | |||
├── 11-zyxel-gs1920.conf | |||
├── 20-zyxel-es3128.conf | |||
├── 30-icmp_check.conf | |||
└── 40-qnap_nas.conf | |||
</syntaxhighlight> | </syntaxhighlight> | ||
=== | === 9.2 GS1920 계열 === | ||
실제 확인된 특성: | |||
* SNMPv3 SHA + DES | |||
* CPU/Memory Private OID 사용 | |||
CPU: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
.1.3.6.1.4.1.890.1.15.3.49.1.7.0 | |||
</syntaxhighlight> | |||
Memory: | |||
<syntaxhighlight lang="bash" line> | |||
Total .1.3.6.1.4.1.890.1.15.3.50.1.1.1.3.1 | |||
Used .1.3.6.1.4.1.890.1.15.3.50.1.1.1.4.1 | |||
Percent .1.3.6.1.4.1.890.1.15.3.50.1.1.1.5.1 | |||
</syntaxhighlight> | |||
=== 9.3 GS1900 계열 === | |||
CPU: | |||
<syntaxhighlight lang="bash" line> | |||
.1.3.6.1.4.1.890.1.15.3.2.4.0 | |||
</syntaxhighlight> | </syntaxhighlight> | ||
= | Memory: | ||
<syntaxhighlight lang="bash" line> | |||
.1.3.6.1.4.1.890.1.15.3.2.5.0 | |||
</syntaxhighlight> | |||
실제 구축 과정에서는 SNMP 응답이 느려 다음 값이 안정적이었다. | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
timeout = "5s" | |||
retries = 1 | |||
</syntaxhighlight> | |||
'''왜 Timeout을 늘렸는가?''' | |||
수동 snmpget은 되는데 Telegraf에서만 간헐적으로 누락된다면 Telegraf timeout이 장비 응답시간보다 짧을 수 있다. | |||
Timeout을 무조건 크게 설정하기보다는 실제 응답을 보고 조정한다. | |||
=== 9.4 ES-3128GP === | |||
실제 확인된 SNMPv3 방식: | |||
* SHA | |||
* AES | |||
sysObjectID: | |||
<syntaxhighlight lang="bash" line> | |||
.1.3.6.1.4.1.7800.1.190 | |||
</syntaxhighlight> | </syntaxhighlight> | ||
Port Metric은 수집 가능하지만 CPU/Memory Private OID는 정상 확인되지 않아 Dashboard에서 억지로 0으로 표시하지 않는다. | |||
'''수집되지 않는 값과 0은 의미가 다르다.''' | |||
< | CPU를 조회할 수 없는 장비에 CPU 0%라고 표시하면 운영자가 정상 상태라고 오해할 수 있다. | ||
=== 9.5 기본 Port Metric === | |||
가능하면 제조사 Private MIB보다 표준 IF-MIB를 먼저 사용한다. | |||
주요 항목: | |||
* ifName | |||
* ifAlias | |||
* ifSpeed | |||
* ifHCInOctets | |||
* ifHCOutOctets | |||
* ifOperStatus | |||
* ifInErrors | |||
* ifOutErrors | |||
* ifInDiscards | |||
* ifOutDiscards | |||
* Broadcast | |||
* Multicast | |||
== 10. ICMP Ping 수집 == | |||
SNMP가 정상이어도 장비 자체가 네트워크에서 사라질 수 있다. | |||
따라서 Ping 결과도 별도로 저장한다. | |||
<syntaxhighlight lang="bash" line> | |||
vi /etc/telegraf/telegraf.d/30-icmp_check.conf | |||
</syntaxhighlight> | |||
예: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
[[inputs.ping]] | [[inputs.ping]] | ||
urls = [" | urls = [ | ||
"203.0.113.11", | |||
"203.0.113.12", | |||
"203.0.113.13" | |||
] | |||
method = "native" | method = "native" | ||
count = 3 | |||
count = | deadline = 2.0 | ||
deadline = | interval = 10.0 | ||
interval = | |||
</syntaxhighlight> | </syntaxhighlight> | ||
=== 10.1 CAP_NET_RAW가 필요한 이유 === | |||
native Ping은 Raw Socket을 사용한다. | |||
root로 telegraf --test를 실행하면 정상인데 systemd 서비스에서는 Ping이 실패할 수 있다. | |||
이 경우 서비스에 CAP_NET_RAW를 부여한다. | |||
<syntaxhighlight lang="bash" line> | |||
systemctl edit telegraf | |||
</syntaxhighlight> | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
[Service] | |||
CapabilityBoundingSet=CAP_NET_RAW | CapabilityBoundingSet=CAP_NET_RAW | ||
AmbientCapabilities=CAP_NET_RAW | AmbientCapabilities=CAP_NET_RAW | ||
</syntaxhighlight> | </syntaxhighlight> | ||
반영: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
systemctl daemon-reload | systemctl daemon-reload | ||
systemctl restart telegraf | |||
</syntaxhighlight> | </syntaxhighlight> | ||
== 11. Telegraf 설정 검사 == | |||
서비스를 재시작하기 전에 설정 오류를 확인한다. | |||
<syntaxhighlight lang="bash" line> | |||
telegraf \ | |||
--config /etc/telegraf/telegraf.conf \ | |||
--config-directory /etc/telegraf/telegraf.d \ | |||
--test | |||
</syntaxhighlight> | |||
'''주의:''' --test는 Metric을 화면에 출력하지만 일반적으로 Output 저장까지 검증하는 용도가 아니다. | |||
서비스 계정 조건까지 확인하려면 다음 방식이 더 정확하다. | |||
<syntaxhighlight lang="bash" line> | |||
systemd-run \ | |||
--unit=telegraf-config-check \ | |||
--wait \ | |||
--pipe \ | |||
--collect \ | |||
-p User=telegraf \ | |||
-p Group=telegraf \ | |||
-p EnvironmentFile=/etc/telegraf/monitor.env \ | |||
-p CapabilityBoundingSet=CAP_NET_RAW \ | |||
-p AmbientCapabilities=CAP_NET_RAW \ | |||
/usr/bin/telegraf \ | |||
--config /etc/telegraf/telegraf.conf \ | |||
--config-directory /etc/telegraf/telegraf.d \ | |||
--test | |||
</syntaxhighlight> | |||
정상 후 서비스 시작: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
systemctl enable --now telegraf | systemctl enable --now telegraf | ||
systemctl restart telegraf | systemctl restart telegraf | ||
systemctl status telegraf --no-pager | systemctl status telegraf --no-pager | ||
journalctl -u telegraf -n 100 --no-pager | |||
journalctl -u telegraf \ | |||
-n 100 \ | |||
--no-pager | |||
</syntaxhighlight> | |||
로그에서 주의할 문자열: | |||
<syntaxhighlight lang="bash" line> | |||
timeout | |||
unauthorized | |||
permission denied | |||
buffer | |||
drop | |||
address already in use | |||
</syntaxhighlight> | </syntaxhighlight> | ||
== | == 12. Grafana 설치 후 처음 해야 할 작업 == | ||
Grafana는 데이터를 직접 수집하지 않는다. | |||
먼저 InfluxDB와 Prometheus 같은 데이터소스를 등록해야 한다. | |||
=== 12.1 Grafana 서비스 === | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
systemctl enable --now grafana-server | |||
systemctl status grafana-server --no-pager | |||
curl -fsS \ | |||
http://127.0.0.1:3000/api/health | |||
</syntaxhighlight> | </syntaxhighlight> | ||
웹 접속: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
http://192.0.2.10:3000 | |||
</syntaxhighlight> | |||
내부망에서만 사용할 경우에도 방화벽 접근 범위를 관리망으로 제한하는 것이 좋다. | |||
=== 12.2 InfluxDB Datasource 등록 === | |||
Grafana 메뉴: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
Connections | |||
→ Data sources | |||
→ Add data source | |||
→ InfluxDB | |||
</syntaxhighlight> | </syntaxhighlight> | ||
설정: | |||
{| class="wikitable" | {| class="wikitable" | ||
! 항목 | |||
! | |||
! 값 | ! 값 | ||
|- | |- | ||
| Query Language | |||
| Query | |||
| Flux | | Flux | ||
|- | |- | ||
| URL | | URL | ||
| | | http://127.0.0.1:8086 | ||
|- | |- | ||
| Organization | | Organization | ||
| | | network | ||
|- | |- | ||
| Token | | Token | ||
| grafana-read | | grafana-read Token | ||
|- | |- | ||
| Default Bucket | | Default Bucket | ||
| | | snmp_raw | ||
|} | |} | ||
Save & | Save & Test 성공 여부를 확인한다. | ||
=== | == 13. 첫 번째 SNMP Dashboard 만들기 == | ||
처음부터 모든 장비를 한 화면에 넣지 않는다. | |||
'''한 대의 스위치 + 한 포트'''로 정상 그래프를 만든 뒤 확대하는 것이 가장 쉽다. | |||
=== 13.1 원시 Counter 확인 === | |||
Explore에서 다음과 같이 최근 데이터를 확인한다. | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
from(bucket: "snmp_raw") | from(bucket: "snmp_raw") | ||
|> range(start: -10m) | |> range(start: -10m) | ||
|> limit(n: 20) | |||
|> limit(n: | |||
</syntaxhighlight> | </syntaxhighlight> | ||
여기서 실제 Measurement, source, index, if_name 값을 확인한다. | |||
=== 13.2 Port RX/TX 계산 === | |||
누적 Octet Counter를 bps로 변환한다. | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
from(bucket: "snmp_raw") | from(bucket: "snmp_raw") | ||
|> range(start: v.timeRangeStart, stop: v.timeRangeStop) | |> range(start: v.timeRangeStart, stop: v.timeRangeStop) | ||
|> filter(fn: (r) => r._measurement == " | |> filter(fn: (r) => | ||
|> filter(fn: (r) => r.source == " | r._measurement == "switch_port" | ||
|> filter(fn: (r) => r._field == "in_octets" or r._field == "out_octets") | ) | ||
|> derivative(unit: 1s, nonNegative: true) | |> filter(fn: (r) => | ||
|> map(fn: (r) => ({r with _value: r._value * 8.0})) | r.source == "203.0.113.11" | ||
) | |||
|> filter(fn: (r) => | |||
r._field == "in_octets" or | |||
r._field == "out_octets" | |||
) | |||
|> derivative( | |||
unit: 1s, | |||
nonNegative: true | |||
) | |||
|> map(fn: (r) => ({ | |||
r with | |||
_value: r._value * 8.0 | |||
})) | |||
</syntaxhighlight> | </syntaxhighlight> | ||
< | Grafana Unit: | ||
<syntaxhighlight lang="bash" line> | |||
bits/sec | |||
</syntaxhighlight> | |||
=== | === 13.3 Error / Discard === | ||
Error나 Discard 역시 Counter이므로 증가율을 표시한다. | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
|> filter(fn: (r) => r._field == " | from(bucket: "snmp_raw") | ||
|> range(start: v.timeRangeStart, stop: v.timeRangeStop) | |||
|> filter(fn: (r) => | |||
r._measurement == "switch_port" | |||
) | |||
|> filter(fn: (r) => | |||
r._field == "in_errors" or | |||
r._field == "out_errors" or | |||
r._field == "in_discards" or | |||
r._field == "out_discards" | |||
) | |||
|> derivative( | |||
unit: 1s, | |||
nonNegative: true | |||
) | |||
</syntaxhighlight> | |||
'''해석''' | |||
* Error 증가 → 물리계층, Frame, Interface 문제 가능 | |||
* Discard 증가 → Queue, Buffer, QoS, 혼잡 등에 의해 폐기될 가능성 | |||
* 값이 0인 상태가 일반적이지만 장비 특성과 Traffic에 따라 해석해야 함 | |||
=== 13.4 operStatus === | |||
operStatus는 Counter가 아니다. | |||
따라서 derivative를 사용하지 않는다. | |||
일반적인 값: | |||
<syntaxhighlight lang="bash" line> | |||
1 = up | |||
2 = down | |||
</syntaxhighlight> | |||
Grafana Value Mapping으로 사람이 읽기 쉬운 문자열로 변환한다. | |||
== 14. Prometheus를 추가하는 이유 == | |||
SNMP는 네트워크 장비 상태에 좋지만 Linux/Windows 서버 자원 모니터링에는 Exporter + Prometheus가 더 편하다. | |||
Prometheus 구조: | |||
<syntaxhighlight lang="bash" line> | |||
Node Exporter ----\ | |||
\ | |||
Windows Exporter ---> Prometheus ---> Grafana | |||
/ | |||
QNAP Exporter -----/ | |||
</syntaxhighlight> | |||
Prometheus는 일정 주기마다 Exporter HTTP Endpoint에 접속하여 Metric을 가져간다. | |||
이를 '''Pull 방식'''이라고 한다. | |||
== 15. Prometheus 설치 == | |||
=== 15.1 사용자와 디렉터리 === | |||
<syntaxhighlight lang="bash" line> | |||
useradd \ | |||
--system \ | |||
--no-create-home \ | |||
--shell /sbin/nologin \ | |||
prometheus | |||
install -d \ | |||
-o prometheus \ | |||
-g prometheus \ | |||
/etc/prometheus \ | |||
/var/lib/prometheus | |||
</syntaxhighlight> | |||
Prometheus 공식 바이너리를 준비한 뒤: | |||
<syntaxhighlight lang="bash" line> | |||
install -m 0755 \ | |||
prometheus \ | |||
/usr/local/bin/prometheus | |||
install -m 0755 \ | |||
promtool \ | |||
/usr/local/bin/promtool | |||
</syntaxhighlight> | |||
버전 확인: | |||
<syntaxhighlight lang="bash" line> | |||
prometheus --version | |||
promtool --version | |||
</syntaxhighlight> | |||
=== 15.2 prometheus.yml === | |||
<syntaxhighlight lang="bash" line> | |||
vi /etc/prometheus/prometheus.yml | |||
</syntaxhighlight> | |||
<syntaxhighlight lang="bash" line> | |||
global: | |||
scrape_interval: 30s | |||
scrape_configs: | |||
- job_name: prometheus | |||
static_configs: | |||
- targets: | |||
- "127.0.0.1:9090" | |||
labels: | |||
server_name: monitoring-server | |||
- job_name: node | |||
static_configs: | |||
- targets: | |||
- "127.0.0.1:9100" | |||
labels: | |||
server_name: monitoring-server | |||
- targets: | |||
- "192.0.2.20:9100" | |||
labels: | |||
server_name: linux-server | |||
- job_name: windows | |||
static_configs: | |||
- targets: | |||
- "192.0.2.30:9182" | |||
labels: | |||
server_name: windows-server | |||
- job_name: qnap | |||
static_configs: | |||
- targets: | |||
- "198.51.100.10:9100" | |||
labels: | |||
server_name: nas-01 | |||
</syntaxhighlight> | |||
'''scrape_interval = 30s'''는 Prometheus가 30초마다 각 Target을 조회한다는 의미이다. | |||
=== 15.3 설정 검사 === | |||
YAML은 들여쓰기에 매우 민감하다. | |||
실제 구축 중에도 job을 scrape_configs 바깥에 잘못 넣으면 Prometheus가 시작하지 못했다. | |||
반드시 검사한다. | |||
<syntaxhighlight lang="bash" line> | |||
promtool check config \ | |||
/etc/prometheus/prometheus.yml | |||
</syntaxhighlight> | </syntaxhighlight> | ||
=== 15.4 systemd 등록 === | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
vi /etc/systemd/system/prometheus.service | |||
</syntaxhighlight> | </syntaxhighlight> | ||
=== | <syntaxhighlight lang="bash" line> | ||
[Unit] | |||
Description=Prometheus | |||
Wants=network-online.target | |||
After=network-online.target | |||
[Service] | |||
User=prometheus | |||
Group=prometheus | |||
ExecStart=/usr/local/bin/prometheus \ | |||
--config.file=/etc/prometheus/prometheus.yml \ | |||
--storage.tsdb.path=/var/lib/prometheus \ | |||
--storage.tsdb.retention.time=30d \ | |||
--storage.tsdb.retention.size=20GB \ | |||
--web.listen-address=127.0.0.1:9090 | |||
= | Restart=always | ||
= | [Install] | ||
WantedBy=multi-user.target | |||
</syntaxhighlight> | |||
권한: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
chown -R prometheus:prometheus \ | |||
/etc/prometheus \ | |||
/var/lib/prometheus | |||
</syntaxhighlight> | </syntaxhighlight> | ||
시작: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
systemctl daemon-reload | |||
systemctl enable --now prometheus | |||
systemctl status prometheus --no-pager | |||
</syntaxhighlight> | </syntaxhighlight> | ||
Health: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
curl -s \ | |||
http://127.0.0.1:9090/-/healthy | |||
</syntaxhighlight> | |||
== 16. Linux Node Exporter 설치 == | |||
=== 16.1 Node Exporter가 하는 일 === | |||
Node Exporter는 Linux Kernel과 /proc, /sys 등의 정보를 읽어 Prometheus 형식으로 공개한다. | |||
대표 Metric: | |||
* CPU | |||
* Memory | |||
* Load | |||
* Filesystem | |||
* Disk I/O | |||
* Network Traffic | |||
* Network Error/Drop | |||
* Uptime | |||
* File Descriptor | |||
* 일부 Hardware Metric | |||
=== 16.2 서비스 계정 === | |||
<syntaxhighlight lang="bash" line> | |||
useradd \ | |||
--system \ | |||
--no-create-home \ | |||
--shell /sbin/nologin \ | |||
node_exporter | |||
</syntaxhighlight> | </syntaxhighlight> | ||
바이너리 설치: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
install -m 0755 \ | |||
node_exporter \ | |||
/usr/local/bin/node_exporter | |||
</syntaxhighlight> | |||
=== 16.3 systemd === | |||
<syntaxhighlight lang="bash" line> | |||
vi /etc/systemd/system/node_exporter.service | |||
</syntaxhighlight> | </syntaxhighlight> | ||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
[Unit] | |||
Description=Node Exporter | |||
After=network-online.target | |||
Wants=network-online.target | |||
[Service] | |||
User=node_exporter | |||
Group=node_exporter | |||
ExecStart=/usr/local/bin/node_exporter | |||
Restart=always | |||
[Install] | |||
WantedBy=multi-user.target | |||
</syntaxhighlight> | </syntaxhighlight> | ||
= | <syntaxhighlight lang="bash" line> | ||
systemctl daemon-reload | |||
systemctl enable --now node_exporter | |||
systemctl status node_exporter --no-pager | |||
</syntaxhighlight> | |||
Metric 확인: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
/ | curl -s \ | ||
http://127.0.0.1:9100/metrics \ | |||
| head | |||
</syntaxhighlight> | </syntaxhighlight> | ||
다른 서버에 설치한 경우 Prometheus 서버에서 확인: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
curl -s \ | |||
http://192.0.2.20:9100/metrics \ | |||
| head | |||
</syntaxhighlight> | </syntaxhighlight> | ||
'''정상 확인''' | |||
Metric 문자열이 여러 줄 출력되면 Exporter 자체는 정상이다. | |||
Prometheus에서 다음 Query도 확인한다. | |||
<syntaxhighlight lang="bash" line> | |||
up | |||
</syntaxhighlight> | |||
값: | |||
<syntaxhighlight lang="bash" line> | |||
1 = 정상 Scrape | |||
0 = Scrape 실패 | |||
</syntaxhighlight> | |||
== 17. Windows Exporter == | |||
Windows에서는 windows_exporter를 설치한다. | |||
기본 Port: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
9182/tcp | |||
</syntaxhighlight> | </syntaxhighlight> | ||
Prometheus에서 다음 주소를 조회할 수 있어야 한다. | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
http://192.0.2.30:9182/metrics | |||
</syntaxhighlight> | </syntaxhighlight> | ||
Windows PowerShell 테스트: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
Invoke-WebRequest \ | |||
-UseBasicParsing \ | |||
http://127.0.0.1:9182/metrics | |||
</syntaxhighlight> | </syntaxhighlight> | ||
성능 카운터가 비정상인 경우 실제 구축 과정에서 다음 복구 절차를 사용하였다. | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
lodctr /R | |||
winmgmt /resyncperf | |||
</syntaxhighlight> | </syntaxhighlight> | ||
그 후: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
Restart-Service windows_exporter | |||
</syntaxhighlight> | </syntaxhighlight> | ||
== 18. Grafana에 Prometheus 등록 == | |||
Grafana 메뉴: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
Connections | |||
→ Data sources | |||
→ Prometheus | |||
</syntaxhighlight> | </syntaxhighlight> | ||
URL: | |||
<syntaxhighlight lang="bash" line> | |||
http://127.0.0.1:9090 | |||
</syntaxhighlight> | |||
Save & Test 후 Explore에서: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
up | |||
</syntaxhighlight> | |||
Target별 1이 표시되면 정상이다. | |||
== 19. Linux Server Dashboard 이해하기 == | |||
처음에는 다음 항목만 만든다. | |||
* Uptime | |||
* CPU | |||
* Memory | |||
* Load Average | |||
* Filesystem | |||
* Network RX/TX | |||
이후 익숙해지면 다음을 추가한다. | |||
* Disk I/O | |||
* Network Error/Drop | |||
* inode | |||
* Swap | |||
* Process | |||
* Application Log | |||
=== 19.1 CPU Usage === | |||
Node Exporter는 CPU 시간을 Mode별 누적으로 제공한다. | |||
idle을 제외한 비율로 CPU 사용률을 계산할 수 있다. | |||
<syntaxhighlight lang="bash" line> | |||
100 - ( | |||
avg by(instance) ( | |||
rate( | |||
node_cpu_seconds_total{ | |||
mode="idle" | |||
}[5m] | |||
) | |||
) * 100 | |||
) | |||
</syntaxhighlight> | |||
=== 19.2 Memory Usage === | |||
<syntaxhighlight lang="bash" line> | |||
100 * ( | |||
1 - | |||
node_memory_MemAvailable_bytes | |||
/ | |||
node_memory_MemTotal_bytes | |||
) | |||
</syntaxhighlight> | |||
=== 19.3 Filesystem Usage === | |||
Filesystem은 tmpfs, overlay, snapshot 등 운영자가 원하지 않는 Mount가 함께 보일 수 있다. | |||
따라서 실제 데이터 Mount Point를 확인한 뒤 필터링한다. | |||
== 20. QNAP TS-264 모니터링 == | |||
QNAP은 한 가지 수집 방식만 사용하면 정보가 부족하다. | |||
따라서 세 경로를 함께 사용한다. | |||
<syntaxhighlight lang="bash" line> | |||
QNAP | |||
| | |||
+-- node_exporter --> Prometheus | |||
| | |||
+-- SNMP ----------> Telegraf --> InfluxDB | |||
| | |||
+-- Syslog --------> rsyslog --> Alloy --> Loki | |||
</syntaxhighlight> | </syntaxhighlight> | ||
역할: | |||
{| class="wikitable" | {| class="wikitable" | ||
! 항목 | ! 항목 | ||
! | ! 데이터 소스 | ||
|- | |- | ||
| | | CPU / Memory / Network | ||
| | | Node Exporter | ||
|- | |- | ||
| | | Filesystem | ||
| | | Node Exporter | ||
|- | |- | ||
| | | Disk I/O | ||
| | | Node Exporter | ||
|- | |- | ||
| | | HDD Model / 온도 / 상태 | ||
| | | SNMP | ||
|- | |- | ||
| RAID / Volume | |||
| SNMP | | SNMP | ||
|- | |- | ||
| | | Event Log | ||
| | | Syslog | ||
|- | |- | ||
| | | Access Log | ||
| | | Syslog | ||
|} | |} | ||
=== 20.1 QNAP SNMPv3 === | |||
실제 TS-264 환경에서는 다음 조합을 사용하였다. | |||
* Authentication: SHA | |||
* Privacy: DES | |||
QNAP은 해당 구성에서 AES가 동작하지 않았다. | |||
따라서 다른 장비에서 AES가 된다는 이유로 QNAP도 AES라고 가정하지 않는다. | |||
=== 20.2 주요 Measurement === | |||
<syntaxhighlight lang="bash" line> | |||
QNAP_TS264 | |||
qnap_disk | |||
qnap_raid | |||
qnap_storage_pool | |||
qnap_volume | |||
</syntaxhighlight> | |||
=== 20.3 Volume 값이 11 GiB로 잘못 보이는 문제 === | |||
실제 QTS에서 DataVol1은 약 11.33 TiB인데 SNMP 원시값을 Grafana에서 byte로 바로 처리하면 약 11.3 GiB로 표시되는 문제가 있었다. | |||
실제 장비의 qnap_volume capacity/free 값이 KiB 성격으로 반환되어 1024 배 보정이 필요하였다. | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
|> map(fn: (r) => ({ | |||
r with | |||
capacity_bytes: | |||
uint(v: r.capacity_bytes) | |||
* uint(v: 1024), | |||
free_bytes: | |||
uint(v: r.free_bytes) | |||
* uint(v: 1024), | |||
used_bytes: | |||
( | |||
uint(v: r.capacity_bytes) | |||
- uint(v: r.free_bytes) | |||
) | |||
* uint(v: 1024), | |||
used_percent: | |||
if float(v: r.capacity_bytes) > 0.0 then | |||
( | |||
float(v: r.capacity_bytes) | |||
- float(v: r.free_bytes) | |||
) | |||
/ | |||
float(v: r.capacity_bytes) | |||
* 100.0 | |||
else | |||
0.0 | |||
})) | |||
</syntaxhighlight> | </syntaxhighlight> | ||
'''왜 Percent에는 1024가 필요 없는가?''' | |||
분자와 분모가 같은 단위이므로 비율 계산에서는 단위가 상쇄된다. | |||
=== 20.4 실제 Data Volume 찾기 === | |||
Node Exporter가 QNAP 내부의 Snapshot Mount까지 모두 보여주므로 처음에는 여러 개의 11.33 TiB Filesystem이 보일 수 있다. | |||
실제 사용자 Data Volume은 다음이었다. | |||
=== | <syntaxhighlight lang="bash" line> | ||
mountpoint="/share/CACHEDEV1_DATA" | |||
device="/dev/mapper/cachedev1" | |||
fstype="ext4" | |||
</syntaxhighlight> | |||
QNAP Snapshot: | |||
<syntaxhighlight lang="bash" line> | |||
/mnt/snapshot/1/10001 | |||
/mnt/snapshot/1/10002 | |||
... | |||
</syntaxhighlight> | |||
Dashboard에서는 Snapshot을 제외하고 실제 Volume만 표시한다. | |||
사용률: | |||
<syntaxhighlight lang="bash" line> | |||
100 * ( | |||
1 - | |||
node_filesystem_avail_bytes{ | |||
instance="198.51.100.10:9100", | |||
mountpoint="/share/CACHEDEV1_DATA" | |||
} | |||
/ | |||
node_filesystem_size_bytes{ | |||
instance="198.51.100.10:9100", | |||
mountpoint="/share/CACHEDEV1_DATA" | |||
} | |||
) | |||
</syntaxhighlight> | |||
실제 QTS 화면과 약 2.86%로 일치하는 것을 확인하였다. | |||
=== 20.5 QNAP Network는 bond0만 표시 === | |||
QNAP에는 내부 Interface와 Virtual Interface가 여러 개 보일 수 있다. | |||
실제 외부 Traffic을 담당하는 bond0만 필터링한다. | |||
RX: | |||
<syntaxhighlight lang="bash" line> | |||
rate( | |||
node_network_receive_bytes_total{ | |||
instance="198.51.100.10:9100", | |||
device="bond0" | |||
}[$__rate_interval] | |||
) * 8 | |||
</syntaxhighlight> | |||
TX: | |||
<syntaxhighlight lang="bash" line> | |||
rate( | |||
node_network_transmit_bytes_total{ | |||
instance="198.51.100.10:9100", | |||
device="bond0" | |||
}[$__rate_interval] | |||
) * 8 | |||
</syntaxhighlight> | |||
=== 20.6 Network Error / Drop === | |||
RX Error: | |||
<syntaxhighlight lang="bash" line> | |||
rate( | |||
node_network_receive_errs_total{ | |||
instance="198.51.100.10:9100", | |||
device="bond0" | |||
}[$__rate_interval] | |||
) | |||
</syntaxhighlight> | |||
TX Error: | |||
<syntaxhighlight lang="bash" line> | |||
rate( | |||
node_network_transmit_errs_total{ | |||
instance="198.51.100.10:9100", | |||
device="bond0" | |||
}[$__rate_interval] | |||
) | |||
</syntaxhighlight> | |||
RX Drop: | |||
<syntaxhighlight lang="bash" line> | |||
rate( | |||
node_network_receive_drop_total{ | |||
instance="198.51.100.10:9100", | |||
device="bond0" | |||
}[$__rate_interval] | |||
) | |||
</syntaxhighlight> | |||
TX Drop: | |||
<syntaxhighlight lang="bash" line> | |||
rate( | |||
node_network_transmit_drop_total{ | |||
instance="198.51.100.10:9100", | |||
device="bond0" | |||
}[$__rate_interval] | |||
) | |||
</syntaxhighlight> | |||
=== 20.7 Disk I/O === | |||
QNAP에는 md, dm, loop 등 내부 Device가 많으므로 물리 Disk sd*만 표시한다. | |||
Read Bytes/sec: | |||
<syntaxhighlight lang="bash" line> | |||
rate( | |||
node_disk_read_bytes_total{ | |||
instance="198.51.100.10:9100", | |||
device=~"sd[a-z]+" | |||
}[$__rate_interval] | |||
) | |||
</syntaxhighlight> | |||
Write Bytes/sec: | |||
<syntaxhighlight lang="bash" line> | |||
rate( | |||
node_disk_written_bytes_total{ | |||
instance="198.51.100.10:9100", | |||
device=~"sd[a-z]+" | |||
}[$__rate_interval] | |||
) | |||
</syntaxhighlight> | |||
Read IOPS: | |||
<syntaxhighlight lang="bash" line> | |||
rate( | |||
node_disk_reads_completed_total{ | |||
instance="198.51.100.10:9100", | |||
device=~"sd[a-z]+" | |||
}[$__rate_interval] | |||
) | |||
</syntaxhighlight> | |||
Write IOPS: | |||
<syntaxhighlight lang="bash" line> | |||
rate( | |||
node_disk_writes_completed_total{ | |||
instance="198.51.100.10:9100", | |||
device=~"sd[a-z]+" | |||
}[$__rate_interval] | |||
) | |||
</syntaxhighlight> | |||
== 21. Syslog를 왜 별도로 구성하는가 == | |||
Metric만으로는 다음과 같은 내용을 알기 어렵다. | |||
* Port가 왜 Down 되었는가 | |||
* 사용자가 NAS에 어떤 파일을 접근했는가 | |||
* 서비스가 언제 재시작되었는가 | |||
* 인증 실패가 발생했는가 | |||
이런 이벤트는 Log가 필요하다. | |||
본 구성의 로그 흐름: | |||
<syntaxhighlight lang="bash" line> | |||
Device | |||
| | |||
| Syslog | |||
v | |||
rsyslog | |||
| | |||
v | |||
Log File | |||
| | |||
v | |||
Alloy | |||
| | |||
v | |||
Loki | |||
| | |||
v | |||
Grafana | |||
</syntaxhighlight> | |||
== 22. rsyslog 구성 == | |||
=== 22.1 Network Syslog === | |||
네트워크 장비 로그: | |||
<syntaxhighlight lang="bash" line> | |||
/var/log/network-syslog/events.log | |||
</syntaxhighlight> | |||
기본 수신: | |||
<syntaxhighlight lang="bash" line> | |||
TCP/UDP 514 | |||
</syntaxhighlight> | |||
=== 22.2 Server Syslog === | |||
서버용 로그를 네트워크 장비와 분리하면 Grafana에서 Query하기 쉽다. | |||
예: | |||
<syntaxhighlight lang="bash" line> | |||
/var/log/server-syslog/events.log | |||
</syntaxhighlight> | |||
수신 Port: | |||
<syntaxhighlight lang="bash" line> | |||
TCP 5514 | |||
</syntaxhighlight> | |||
=== 22.3 QNAP Event와 Access Log 분리 === | |||
QNAP QuLog의 두 성격이 다르므로 포트부터 분리한다. | |||
{| class="wikitable" | {| class="wikitable" | ||
! 종류 | |||
! 포트 | |||
! 파일 | |||
! 의미 | |||
|- | |- | ||
| Event | |||
| TCP 5515 | |||
| /var/log/qnap/event.log | |||
| 시스템/서비스 이벤트 | |||
|- | |- | ||
| | | Access | ||
| | | TCP 5516 | ||
| | | /var/log/qnap/access.log | ||
| 사용자/파일 접근 | |||
| | |||
|} | |} | ||
Template: | |||
<syntaxhighlight lang="bash" line> | |||
template(name="QnapSyslogLine" type="string" | |||
string="%timegenerated:::date-rfc3339% src=%fromhost-ip% severity=%syslogseverity-text% host=%hostname% %syslogtag%%msg:::sp-if-no-1st-sp%%msg%\n") | |||
</syntaxhighlight> | |||
이렇게 저장하면 한 줄이 대략 다음 구조가 된다. | |||
<syntaxhighlight lang="bash" line> | |||
2026-09-13T23:40:03+09:00 \ | |||
src=198.51.100.10 \ | |||
severity=info \ | |||
host=nas-01 \ | |||
qulogd: ... | |||
</syntaxhighlight> | |||
Event ruleset: | |||
<syntaxhighlight lang="bash" line> | |||
ruleset(name="QnapEventLog") { | |||
action( | |||
type="omfile" | |||
file="/var/log/qnap/event.log" | |||
template="QnapSyslogLine" | |||
fileOwner="root" | |||
fileGroup="alloy" | |||
fileCreateMode="0640" | |||
dirOwner="root" | |||
dirGroup="alloy" | |||
dirCreateMode="0750" | |||
createDirs="on" | |||
) | |||
stop | |||
} | |||
input( | |||
type="imtcp" | |||
port="5515" | |||
ruleset="QnapEventLog" | |||
) | |||
</syntaxhighlight> | |||
Access ruleset: | |||
<syntaxhighlight lang="bash" line> | |||
ruleset(name="QnapAccessLog") { | |||
action( | |||
type="omfile" | |||
file="/var/log/qnap/access.log" | |||
template="QnapSyslogLine" | |||
fileOwner="root" | |||
fileGroup="alloy" | |||
fileCreateMode="0640" | |||
dirOwner="root" | |||
dirGroup="alloy" | |||
dirCreateMode="0750" | |||
createDirs="on" | |||
) | |||
stop | |||
} | |||
input( | |||
type="imtcp" | |||
port="5516" | |||
ruleset="QnapAccessLog" | |||
) | |||
</syntaxhighlight> | |||
=== 22.4 rsyslog 설정 검사 === | |||
재시작 전에 반드시 검사한다. | |||
<syntaxhighlight lang="bash" line> | |||
rsyslogd -N1 | |||
</syntaxhighlight> | |||
오류가 없으면: | |||
<syntaxhighlight lang="bash" line> | |||
systemctl restart rsyslog | |||
ss -lntp \ | |||
| grep -E ':5514|:5515|:5516' | |||
</syntaxhighlight> | |||
실제 로그 확인: | |||
<syntaxhighlight lang="bash" line> | |||
tail -f \ | |||
/var/log/qnap/access.log | |||
</syntaxhighlight> | |||
'''문제 분리 방법''' | |||
<syntaxhighlight lang="bash" line> | |||
tcpdump에는 Packet이 안 보임 | |||
→ 장비 송신/방화벽/라우팅 확인 | |||
tcpdump에는 보이지만 파일이 안 생김 | |||
→ rsyslog 설정 확인 | |||
파일은 생기지만 Grafana에 안 보임 | |||
→ Alloy/Loki 확인 | |||
</syntaxhighlight> | |||
== 23. Loki와 Alloy의 역할 == | |||
rsyslog가 파일을 만드는 것만으로 Grafana가 그 파일을 검색할 수 있는 것은 아니다. | |||
Loki가 로그 저장소 역할을 하고 Alloy가 파일을 읽어 Loki에 넣는다. | |||
<syntaxhighlight lang="bash" line> | |||
/var/log/... | |||
| | |||
v | |||
Alloy | |||
| | |||
v | |||
Loki | |||
| | |||
v | |||
Grafana | |||
</syntaxhighlight> | |||
현재 구성에서는 Loki가 localhost 3100에서 동작한다. | |||
<syntaxhighlight lang="bash" line> | |||
127.0.0.1:3100 | |||
</syntaxhighlight> | |||
Alloy 로컬 관리 Endpoint: | |||
<syntaxhighlight lang="bash" line> | |||
127.0.0.1:12345 | |||
</syntaxhighlight> | |||
== 24. Alloy 설정 이해하기 == | |||
=== 24.1 Network Syslog === | |||
<syntaxhighlight lang="bash" line> | |||
loki.source.file "network_syslog" { | |||
targets = [ | |||
{ | |||
__path__ = "/var/log/network-syslog/events.log", | |||
job = "network-syslog", | |||
}, | |||
] | |||
forward_to = [ | |||
loki.process.network_syslog.receiver | |||
] | |||
} | |||
</syntaxhighlight> | |||
'''source.file'''은 어떤 파일을 읽을지 지정한다. | |||
'''job'''은 Loki에서 로그 종류를 구분하기 위한 Label이다. | |||
그 다음 regex로 한 줄을 분해한다. | |||
<syntaxhighlight lang="bash" line> | |||
loki.process "network_syslog" { | |||
stage.regex { | |||
expression = `^(?P<received_at>\S+) src=(?P<device_ip>\S+) severity=(?P<severity>\S+) host=(?P<device_host>\S+) (?P<message>.*)$` | |||
} | |||
stage.timestamp { | |||
source = "received_at" | |||
format = "RFC3339Nano" | |||
action_on_failure = "skip" | |||
} | |||
stage.labels { | |||
values = { | |||
device_ip = "" | |||
severity = "" | |||
} | |||
} | |||
forward_to = [ | |||
loki.write.local.receiver | |||
] | |||
} | |||
</syntaxhighlight> | |||
여기서 다음 Label이 생긴다. | |||
<syntaxhighlight lang="bash" line> | |||
job | |||
device_ip | |||
severity | |||
</syntaxhighlight> | |||
'''device_host를 Label로 추가하려면''' | |||
<syntaxhighlight lang="bash" line> | |||
stage.labels { | |||
values = { | |||
device_ip = "" | |||
device_host = "" | |||
severity = "" | |||
} | |||
} | |||
</syntaxhighlight> | |||
이는 이후 Alert Mail에서 Hostname까지 표시하고 싶을 때 유용하다. | |||
=== 24.2 QNAP Access === | |||
<syntaxhighlight lang="bash" line> | |||
loki.source.file "qnap_access" { | |||
targets = [ | |||
{ | |||
__path__ = "/var/log/qnap/access.log", | |||
job = "qnap-access", | |||
}, | |||
] | |||
forward_to = [ | |||
loki.process.qnap_access.receiver | |||
] | |||
} | |||
loki.process "qnap_access" { | |||
=== | stage.regex { | ||
expression = `^(?P<received_at>\S+) src=(?P<nas_ip>\S+) severity=(?P<severity>\S+) host=(?P<nas_host>\S+) (?P<message>.*)$` | |||
} | |||
stage.timestamp { | |||
source = "received_at" | |||
format = "RFC3339Nano" | |||
action_on_failure = "skip" | |||
} | |||
stage.labels { | |||
values = { | |||
nas_ip = "" | |||
nas_host = "" | |||
severity = "" | |||
log_type = "access" | |||
} | |||
} | |||
forward_to = [ | |||
loki.write.local.receiver | |||
] | |||
} | |||
</syntaxhighlight> | |||
=== 24.3 Loki Write === | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
loki.write "local" { | |||
endpoint { | |||
url = "http://127.0.0.1:3100/loki/api/v1/push" | |||
} | |||
} | |||
</syntaxhighlight> | </syntaxhighlight> | ||
즉 Alloy에서 처리한 로그를 로컬 Loki로 전송한다. | |||
=== 24.4 설정 검사 === | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
alloy validate \ | |||
/etc/alloy/config.alloy | |||
</syntaxhighlight> | </syntaxhighlight> | ||
정상이라면 아무 오류 없이 종료된다. | |||
반영: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
systemctl restart alloy | |||
systemctl status alloy \ | |||
--no-pager | |||
</syntaxhighlight> | </syntaxhighlight> | ||
< | |||
Loki에 실제 Job이 만들어졌는지 확인: | |||
<syntaxhighlight lang="bash" line> | |||
curl -s \ | |||
'http://127.0.0.1:3100/loki/api/v1/label/job/values' \ | |||
| jq | |||
</syntaxhighlight> | |||
중요한 점은 Alloy 설정에 job을 적었다고 바로 Loki에 Label이 생기는 것이 아니라 '''실제 로그가 최소 한 번 Loki에 저장되어야''' 조회 결과에 나타난다는 것이다. | |||
== 25. Alloy Permission 문제 해결 사례 == | |||
실제 구축 과정에서 QNAP 로그 파일은 존재하지만 Alloy가 다음 오류를 출력하였다. | |||
<syntaxhighlight lang="bash" line> | |||
failed to tail file | |||
stat failed | |||
permission denied | |||
</syntaxhighlight> | |||
먼저 경로 전체 권한 확인: | |||
<syntaxhighlight lang="bash" line> | |||
namei -l \ | |||
/var/log/qnap/access.log | |||
</syntaxhighlight> | |||
정상 예: | |||
<syntaxhighlight lang="bash" line> | |||
drwxr-x--- root alloy qnap | |||
-rw-r----- root alloy access.log | |||
</syntaxhighlight> | |||
Alloy 계정으로 직접 읽기: | |||
<syntaxhighlight lang="bash" line> | |||
sudo -u alloy \ | |||
head /var/log/qnap/access.log | |||
</syntaxhighlight> | |||
그래도 실패하면 SELinux 확인: | |||
<syntaxhighlight lang="bash" line> | |||
getenforce | |||
ausearch \ | |||
-m AVC \ | |||
-ts recent \ | |||
| grep -Ei 'alloy|qnap' | |||
ls -Zd /var/log/qnap | |||
ls -Z /var/log/qnap/access.log | |||
</syntaxhighlight> | |||
'''문제 해결 원칙''' | |||
권한 문제가 있다고 SELinux를 바로 끄지 않는다. | |||
다음 순서로 본다. | |||
# 파일 Owner/Group | |||
# 파일 Mode | |||
# 상위 Directory execute 권한 | |||
# Alloy Service User | |||
# SELinux Context / AVC | |||
== 26. Grafana에 Loki 추가 == | |||
Grafana: | |||
<syntaxhighlight lang="bash" line> | |||
Connections | |||
→ Data sources | |||
→ Add data source | |||
→ Loki | |||
</syntaxhighlight> | |||
URL: | |||
<syntaxhighlight lang="bash" line> | |||
http://127.0.0.1:3100 | |||
</syntaxhighlight> | |||
Explore에서 Network 로그 확인: | |||
<syntaxhighlight lang="bash" line> | |||
{job="network-syslog"} | |||
</syntaxhighlight> | |||
QNAP Access: | |||
<syntaxhighlight lang="bash" line> | |||
{job="qnap-access"} | |||
</syntaxhighlight> | |||
QNAP Event: | |||
<syntaxhighlight lang="bash" line> | |||
{job="qnap-event"} | |||
</syntaxhighlight> | |||
== 27. Grafana Alert를 처음 구성할 때 알아둘 점 == | |||
Grafana Alert는 Dashboard의 그래프와 달리 최종적으로 '''숫자 하나 또는 Alert Instance별 숫자'''를 평가해야 한다. | |||
Range Query를 그대로 Alert Condition에 사용하면 다음 오류가 발생할 수 있다. | |||
<syntaxhighlight lang="bash" line> | |||
looks like time series data, | |||
only reduced data can be alerted on | |||
</syntaxhighlight> | |||
따라서 다음 중 하나를 사용한다. | |||
# Query를 Instant로 구성 | |||
# Range Query → Reduce → Threshold | |||
== 28. Syslog Critical Alert == | |||
Level 3(Error) 이상: | |||
<syntaxhighlight lang="bash" line> | |||
sum by (device_ip, severity) ( | |||
count_over_time( | |||
{ | |||
job="network-syslog", | |||
severity=~"emerg|alert|crit|err" | |||
}[1m] | |||
) | |||
) | |||
</syntaxhighlight> | |||
권장: | |||
<syntaxhighlight lang="bash" line> | |||
A = Loki Instant Query | |||
B = Threshold | |||
A IS ABOVE 0 | |||
</syntaxhighlight> | |||
Summary: | |||
<syntaxhighlight lang="bash" line> | |||
[Syslog 경고] | |||
{{ $labels.device_ip }} | |||
{{ $labels.severity }} | |||
</syntaxhighlight> | |||
Description: | |||
<syntaxhighlight lang="bash" line> | |||
장비 {{ $labels.device_ip }} 에서 | |||
Syslog Level 3 이상 로그가 발생했습니다. | |||
Severity: | |||
{{ $labels.severity }} | |||
최근 1분 발생 건수: | |||
{{ $values.A.Value }} | |||
</syntaxhighlight> | |||
=== 28.1 Alert Mail에 실제 로그 내용이 없는 이유 === | |||
count_over_time()은 문자열 로그를 숫자로 집계한다. | |||
따라서 Query 결과에는 주로 다음만 남는다. | |||
* Label | |||
* 발생 건수 | |||
로그 message 전체를 Loki Label로 만들면 종류가 지나치게 많아져 Cardinality 문제가 발생할 수 있으므로 권장하지 않는다. | |||
메일에는 다음 정도를 넣고 실제 내용은 Grafana Explore에서 확인하는 구조가 안정적이다. | |||
* Device IP | |||
* Hostname | |||
* Severity | |||
* 발생 건수 | |||
* Grafana Link | |||
== 29. ICMP Down Alert == | |||
Flux: | |||
<syntaxhighlight lang="bash" line> | |||
from(bucket: "snmp_raw") | |||
|> range(start: -5m) | |||
|> filter(fn: (r) => | |||
r._measurement == "ping" and | |||
r._field == "percent_packet_loss" | |||
) | |||
|> group( | |||
columns: ["url"] | |||
) | |||
|> last() | |||
|> keep( | |||
columns: [ | |||
"_time", | |||
"_value", | |||
"url" | |||
] | |||
) | |||
</syntaxhighlight> | |||
Alert 구조: | |||
<syntaxhighlight lang="bash" line> | |||
A = Flux Query | |||
B = Reduce | |||
Last | |||
Strict | |||
C = Threshold | |||
B > 99 | |||
Pending = 2m | |||
</syntaxhighlight> | |||
No Data 정책: | |||
<syntaxhighlight lang="bash" line> | |||
Keep Last State | |||
</syntaxhighlight> | |||
'''왜 Keep Last State인가?''' | |||
Metric이나 Log Source가 순간적으로 No Data가 되었을 때 별도의 DatasourceNoData 메일이 발생하면서 실제 장비 Label이 없는 알림이 발송될 수 있다. | |||
No Data 자체를 별도 장애로 관리할 필요가 있다면 별도의 수집 상태 Alert를 구성하는 것이 더 명확하다. | |||
== 30. 최종 Dashboard 구성 방법 == | |||
처음 설치한 사용자는 Dashboard를 한 번에 완성하려 하지 말고 다음 단계로 만든다. | |||
# 데이터소스 Save & Test | |||
# Explore에서 실제 데이터 확인 | |||
# Stat 패널 하나 생성 | |||
# Time Series 하나 생성 | |||
# 장비 한 대 정상 확인 | |||
# 변수 추가 | |||
# 여러 장비로 확대 | |||
# Alert 추가 | |||
=== 30.1 Network Switch Dashboard === | |||
권장 Row: | |||
{| class="wikitable" | {| class="wikitable" | ||
! Row | |||
! 표시 내용 | |||
! 목적 | |||
|- | |- | ||
| Device 상태 | |||
| Device Name / Uptime / CPU / Memory / Ping | |||
| | | 장비 자체 상태 | ||
| | |||
| | |||
|- | |- | ||
| | | Port 상태 | ||
| | | ifName / ifAlias / Speed / operStatus | ||
| Link 상태 | |||
|- | |- | ||
| | | Traffic | ||
| | | RX / TX bps | ||
| 대역폭 사용량 | |||
|- | |- | ||
| | | Packet | ||
| | | Unicast / Broadcast / Multicast PPS | ||
| Broadcast 폭주 및 Packet 패턴 | |||
|- | |- | ||
| | | Error | ||
| | | Error / Discard | ||
| 품질 및 혼잡 징후 | |||
|- | |- | ||
| | | Syslog | ||
| | | warning / err / crit | ||
| 이벤트 원인 확인 | |||
|} | |} | ||
=== 30.2 Linux Dashboard === | |||
권장: | |||
* Uptime | |||
* CPU | |||
* Memory | |||
* Load | |||
* Disk Capacity | |||
* Disk I/O | |||
* Network RX/TX | |||
* Network Error/Drop | |||
* Application/Security Log | |||
=== 30.3 Windows Dashboard === | |||
권장: | |||
* Exporter UP | |||
* Uptime | |||
* CPU | |||
* Memory | |||
* Logical Disk | |||
* Disk I/O | |||
* Network | |||
* Windows Event 연동 여부 | |||
=== 30.4 QNAP Dashboard === | |||
현재 실제 운영 구성을 기준으로 다음과 같이 구성한다. | |||
상단: | |||
* Node Exporter UP | |||
* Uptime | |||
* CPU Usage | |||
* Memory Usage | |||
* HDD 최고 온도 | |||
중단: | |||
* CPU / Memory / Load | |||
* bond0 RX / TX | |||
* Disk 상태 | |||
* DataVol1 Volume | |||
* DataVol1 Filesystem 사용률 | |||
* bond0 Error / Drop | |||
* Physical Disk I/O / IOPS | |||
하단: | |||
* QNAP Event Log | |||
* QNAP Access Log | |||
제거한 항목: | |||
* 상단 RAID 상태 | |||
* Storage Pool 상태 | |||
* Storage Pool 사용률 | |||
* RAID 상세 Table | |||
* Storage Pool 상세 Table | |||
삭제 이유는 Dashboard를 운영자가 빠르게 읽을 수 있도록 단순화하기 위해서이다. | |||
RAID/Storage Pool 상세는 필요 시 별도의 상세 Dashboard에서 확인할 수 있다. | |||
== 31. 초보자가 자주 만나는 문제 == | |||
=== 31.1 Telegraf --test는 되는데 서비스는 안 됨 === | |||
root 권한에서는 Ping이 되지만 telegraf 서비스 계정에서는 Raw Socket 권한이 없을 수 있다. | |||
확인: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
journalctl -u telegraf \ | |||
-n 100 \ | |||
--no-pager | |||
</syntaxhighlight> | </syntaxhighlight> | ||
native Ping이면 CAP_NET_RAW 설정을 확인한다. | |||
=== 31.2 snmpget은 되는데 Telegraf에서 일부 장비만 누락 === | |||
가능한 원인: | |||
* timeout이 너무 짧음 | |||
* retries가 0 | |||
* OID가 모델과 다름 | |||
* DES/AES 조합이 장비와 다름 | |||
* SNMP View 제한 | |||
먼저 동일 OID를 snmpget으로 확인한다. | |||
=== | === 31.3 Prometheus가 재시작 반복 === | ||
가장 먼저: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
promtool check config \ | |||
/etc/prometheus/prometheus.yml | |||
</syntaxhighlight> | </syntaxhighlight> | ||
YAML 들여쓰기를 확인한다. | |||
그 다음: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
journalctl -u prometheus \ | |||
-n 100 \ | |||
--no-pager | |||
</syntaxhighlight> | </syntaxhighlight> | ||
=== | === 31.4 Grafana에서 Filesystem이 너무 많이 보임 === | ||
Linux/QNAP에는 tmpfs, loop, snapshot, container mount 등이 존재할 수 있다. | |||
전체를 보여주기보다 운영자가 실제 사용하는 Mount Point만 필터링한다. | |||
QNAP 예: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
mountpoint="/share/CACHEDEV1_DATA" | |||
</syntaxhighlight> | </syntaxhighlight> | ||
=== | === 31.5 Loki에 Job이 안 보임 === | ||
설정에 job이 있다고 바로 나타나는 것이 아니다. | |||
실제 Log가 Loki에 한 번 이상 저장되어야 한다. | |||
== | 확인 순서: | ||
<syntaxhighlight lang="bash" line> | |||
tail -n 10 \ | |||
/var/log/qnap/access.log | |||
journalctl -u alloy \ | |||
-n 100 \ | |||
--no-pager | |||
curl -s \ | |||
'http://127.0.0.1:3100/loki/api/v1/label/job/values' \ | |||
| jq | |||
</syntaxhighlight> | |||
== 32. 구축 완료 후 점검 Checklist == | |||
{| class="wikitable" | {| class="wikitable" | ||
! 확인 | |||
! 항목 | |||
|- | |- | ||
| □ | |||
| Rocky Linux 시간 동기화 정상 | |||
|- | |- | ||
| | | □ | ||
| | | InfluxDB health 정상 | ||
|- | |- | ||
| | | □ | ||
| | | Telegraf 서비스 정상 | ||
|- | |- | ||
| | | □ | ||
| | | SNMP 장비별 응답 확인 | ||
|- | |- | ||
| | | □ | ||
| | | InfluxDB에 SNMP 데이터 저장 | ||
|- | |- | ||
| | | □ | ||
| | | Prometheus health 정상 | ||
|- | |- | ||
| | | □ | ||
| | | Node Exporter Target UP | ||
|- | |- | ||
| | | □ | ||
| | | Windows Exporter Target UP | ||
|- | |- | ||
| | | □ | ||
| | | QNAP Node Exporter Target UP | ||
|- | |- | ||
| | | □ | ||
| | | rsyslog Network 로그 수신 | ||
| | |- | ||
| □ | |||
| QNAP Event / Access Log 분리 수신 | |||
|- | |||
| □ | |||
| Alloy Permission 정상 | |||
|- | |||
| □ | |||
| Loki job 확인 | |||
|- | |||
| □ | |||
| Grafana InfluxDB Save & Test | |||
|- | |- | ||
| □ | |||
| Grafana Prometheus Save & Test | |||
|- | |- | ||
| | | □ | ||
| | | Grafana Loki Save & Test | ||
|- | |- | ||
| | | □ | ||
| | | Switch Traffic 그래프 정상 | ||
|- | |- | ||
| | | □ | ||
| | | Error/Discard 그래프 정상 | ||
|- | |- | ||
| | | □ | ||
| | | Linux/Windows Dashboard 정상 | ||
|- | |- | ||
| | | □ | ||
| | | QNAP DataVol1 용량 QTS와 일치 | ||
|- | |- | ||
| | | □ | ||
| | | QNAP Snapshot 제외 | ||
|- | |- | ||
| | | □ | ||
| | | QNAP bond0만 Network 표시 | ||
|- | |- | ||
| | | □ | ||
| | | Syslog Alert Mail에 장비 IP / Severity 표시 | ||
|- | |- | ||
| | | □ | ||
| | | ICMP Down Alert 정상 | ||
|} | |} | ||
== 33. 전체 장애 점검 순서 == | |||
모니터링이 안 될 때 Grafana부터 무작정 수정하지 않는다. | |||
=== 33.1 SNMP === | |||
<syntaxhighlight lang="bash" line> | |||
Switch | |||
↓ | |||
snmpget | |||
↓ | |||
Telegraf --test | |||
↓ | |||
Telegraf Service | |||
↓ | |||
InfluxDB | |||
↓ | |||
Grafana Explore | |||
↓ | |||
Dashboard | |||
</syntaxhighlight> | |||
=== 33.2 Prometheus === | |||
<syntaxhighlight lang="bash" line> | |||
Exporter /metrics | |||
↓ | |||
Prometheus Target | |||
↓ | |||
Prometheus Query | |||
↓ | |||
Grafana Explore | |||
↓ | |||
Dashboard | |||
</syntaxhighlight> | |||
=== 33.3 Syslog === | |||
<syntaxhighlight lang="bash" line> | |||
Device 송신 | |||
↓ | |||
tcpdump | |||
↓ | |||
rsyslog | |||
↓ | |||
Log File | |||
↓ | |||
Alloy | |||
↓ | |||
Loki | |||
↓ | |||
Grafana Explore | |||
↓ | |||
Dashboard / Alert | |||
</syntaxhighlight> | |||
이 순서대로 확인하면 문제 지점을 빠르게 좁힐 수 있다. | |||
== 34. 운영 원칙 요약 == | |||
* 처음부터 모든 장비를 추가하지 않는다. | |||
* 한 장비, 한 Port를 먼저 완성한다. | |||
* 설정 변경 후 항상 해당 서비스의 검사 명령을 실행한다. | |||
* Counter와 현재 상태값을 구분한다. | |||
* No Data와 0을 같은 의미로 처리하지 않는다. | |||
* SNMP Password와 Token을 Wiki에 저장하지 않는다. | |||
* Grafana는 수집기가 아니라 시각화 계층임을 기억한다. | |||
* QNAP처럼 제조사별 특이사항은 실제 값과 Vendor 화면을 대조한다. | |||
* Alert는 처음부터 너무 많이 만들지 않는다. | |||
* 정상 상태의 기준 데이터를 먼저 쌓은 후 임계값을 결정한다. | |||
* 장애 시에는 수집 경로를 앞단부터 순서대로 확인한다. | |||
== 35. 최종 구성 요약 == | |||
<syntaxhighlight lang="bash" line> | |||
+----------------------+ | |||
| Grafana | | |||
+----------+-----------+ | |||
| | |||
+------------------+------------------+ | |||
| | | | |||
v v v | |||
Prometheus InfluxDB Loki | |||
^ ^ ^ | |||
| | | | |||
Exporters Telegraf Alloy | |||
^ ^ ^ | |||
| | | | |||
Linux / Windows / QNAP SNMP Device Syslog Files | |||
^ | |||
| | |||
rsyslog | |||
^ | |||
| | |||
Network / Server / NAS | |||
</syntaxhighlight> | |||
이 구조가 완성되면 다음 세 종류의 정보를 한 Grafana에서 확인할 수 있다. | |||
# 장비와 서버의 현재 상태 | |||
# 시간에 따른 성능 변화 | |||
# 장애 시점의 이벤트 로그 | |||
초보자는 먼저 "데이터가 어디서 생성되어 어디를 거쳐 Grafana까지 도착하는지"를 이해한 뒤 설정값을 수정하는 것이 가장 중요하다. | |||
2026년 9월 23일 (수) 22:28 기준 최신판
Rocky Linux 9 통합 모니터링 서버 구축 가이드 - 초보자용
- Telegraf / InfluxDB / Grafana / Prometheus / Node Exporter / Windows Exporter / Loki / Alloy / rsyslog
작성 목적: 처음 모니터링 서버를 설치하는 사용자가 "왜 이 프로그램이 필요한지", "어떤 순서로 설치하는지", "어디까지 정상이어야 다음 단계로 넘어가는지"를 이해하면서 구축할 수 있도록 작성한다.
본 문서는 실제 운영 과정에서 구성하고 점검한 내용을 바탕으로 정리했으며, 실제 IP 주소와 인증 정보는 문서용 예시 값으로 변경하였다.
중요: 아래 명령을 한 번에 모두 실행하지 않는다. 각 절의 정상 확인 항목을 통과한 뒤 다음 단계로 진행한다.
0. 이 문서에서 만들 시스템
0.1 먼저 전체 구조부터 이해하기
모니터링을 처음 구성할 때 가장 혼동하기 쉬운 부분은 "Grafana가 모든 데이터를 직접 수집한다"고 생각하는 것이다.
Grafana는 주로 보여주는 역할을 한다. 실제 데이터 수집과 저장은 다른 프로그램이 담당한다.
이 문서에서는 데이터를 크게 두 종류로 나눈다.
- Metric : CPU 사용률, 메모리 사용률, 포트 트래픽, Ping 손실률처럼 숫자로 표현되는 값
- Log : Syslog, 서버 이벤트, NAS 접근 기록처럼 문자열로 기록되는 이벤트
전체 데이터 흐름은 다음과 같다.
[Network Switch]
|
| SNMPv3
v
Telegraf
|
v
InfluxDB
|
+-------------------+
|
[Linux / QNAP] |
| |
| node_exporter |
v |
Prometheus |
| |
+-------------------+
|
[Network / Server / NAS] |
| |
| Syslog |
v |
rsyslog |
| |
v |
Log File |
| |
v |
Alloy |
| |
v |
Loki |
| |
+-------------------+
|
v
Grafana
즉 다음과 같이 기억하면 된다.
| 프로그램 | 하는 일 | 쉽게 표현하면 |
|---|---|---|
| Telegraf | SNMP 장비의 값을 주기적으로 읽음 | 네트워크 장비용 수집기 |
| InfluxDB | Telegraf가 수집한 시계열 값을 저장 | SNMP 데이터 창고 |
| Prometheus | Exporter가 공개한 Metric을 주기적으로 가져와 저장 | 서버 Metric 수집기 + 데이터베이스 |
| Node Exporter | Linux/QNAP의 CPU, 메모리, 디스크, 네트워크 Metric 제공 | Linux 상태 측정기 |
| Windows Exporter | Windows 성능 카운터 Metric 제공 | Windows 상태 측정기 |
| rsyslog | 장비/서버가 보내는 Syslog를 수신하여 파일로 저장 | 로그 수신기 |
| Alloy | 로그 파일을 읽고 필요한 항목을 분리하여 Loki로 전달 | 로그 전달/가공기 |
| Loki | 로그를 저장하고 검색 가능하게 함 | 로그 데이터베이스 |
| Grafana | 위 데이터들을 Dashboard와 Alert로 표현 | 통합 화면 |
0.2 왜 하나의 프로그램으로 모두 처리하지 않는가
SNMP, 서버 Metric, 로그는 데이터 성격이 서로 다르다.
예를 들어 스위치 포트 트래픽은 누적 Counter를 주기적으로 읽어 증가량을 계산해야 한다. 반면 Syslog는 특정 시점에 발생한 문자열 이벤트이다.
따라서 본 구성에서는 각 도구가 가장 잘하는 역할을 분리한다.
- Switch SNMP → Telegraf + InfluxDB
- Linux / Windows / QNAP OS Metric → Prometheus
- Syslog → rsyslog + Alloy + Loki
- 최종 화면 → Grafana
이 구조를 이해하면 장애 발생 시 어느 부분을 확인해야 하는지도 쉽게 구분할 수 있다.
예를 들어 Grafana에서 Linux CPU가 보이지 않는다면 다음 순서로 생각한다.
Grafana 문제인가?
↓
Prometheus에 데이터가 있는가?
↓
Prometheus가 node_exporter를 수집하고 있는가?
↓
node_exporter가 정상 실행 중인가?
1. 예제 환경과 주소
실제 운영 IP를 문서에 노출하지 않기 위해 RFC 5737 문서용 주소를 사용한다.
| 대상 | 문서용 IP | 역할 |
|---|---|---|
| Monitoring Server | 192.0.2.10 | Grafana / Prometheus / InfluxDB / Telegraf / Loki / Alloy / rsyslog |
| Linux Server | 192.0.2.20 | node_exporter |
| Windows Server | 192.0.2.30 | windows_exporter |
| QNAP NAS | 198.51.100.10 | node_exporter / SNMP / Syslog |
| Switch-01 | 203.0.113.11 | SNMP / Syslog |
| Switch-02 | 203.0.113.12 | SNMP / Syslog |
| Switch-03 | 203.0.113.13 | SNMP / Syslog |
실제 설치 시 위 주소를 자신의 환경에 맞게 변경한다.
2. 설치 전에 알아둘 용어
2.1 Metric
Metric은 시간에 따라 변화하는 숫자 데이터이다.
예:
- CPU 32%
- Memory 71%
- Switch Port RX 120 Mbps
- Ping Packet Loss 0%
- HDD Temperature 43°C
이런 값은 시간 흐름에 따라 그래프로 보는 것이 중요하다.
2.2 Counter와 Gauge
SNMP에서 자주 만나는 개념이다.
Gauge는 현재 값을 의미한다.
예:
CPU Usage = 35
Temperature = 44
operStatus = 1
Counter는 계속 증가하는 누적값이다.
예:
ifHCInOctets = 1234567890123
이 값을 그대로 그래프로 그리면 계속 증가하기만 한다.
그래서 실제 트래픽은 현재 값과 이전 값의 차이를 시간으로 나누어 계산한다.
현재 Counter - 이전 Counter
---------------------------
경과 시간
Octet은 8 bit이므로 bps를 구하려면 다시 8을 곱한다.
Grafana/Flux에서는 이를 derivative()로 처리한다.
2.3 Label과 Tag
장비가 여러 대일 때 어떤 데이터가 어느 장비인지 구분하기 위한 값이다.
예:
source=203.0.113.11
if_name=port24
index=24
Prometheus에서는 주로 label, InfluxDB에서는 tag라고 부른다.
2.4 Exporter
Prometheus가 직접 Linux 내부 명령을 실행하는 것은 아니다.
Exporter가 시스템 정보를 HTTP Metric 형식으로 공개하고 Prometheus가 이를 가져간다.
Node Exporter 예:
http://192.0.2.20:9100/metrics
Windows Exporter:
http://192.0.2.30:9182/metrics
3. Rocky Linux 9 기본 준비
이 장에서는 모니터링 프로그램을 설치하기 전에 OS가 정상인지 확인한다.
3.1 시스템 상태 확인
cat /etc/rocky-release
uname -m
ip -br address
ip route
df -hT
free -h
getenforce
ss -lntup
각 명령의 의미:
| 명령 | 확인 내용 |
|---|---|
| cat /etc/rocky-release | Rocky Linux 버전 |
| uname -m | CPU Architecture, 일반적인 x86 서버는 x86_64 |
| ip -br address | 서버 IP 주소 |
| ip route | Gateway 및 Routing |
| df -hT | 디스크 용량 |
| free -h | 메모리 |
| getenforce | SELinux 상태 |
| ss -lntup | 현재 사용 중인 TCP/UDP Port |
초보자 주의: SELinux가 Enforcing이라고 해서 바로 Disabled로 변경하지 않는다. 이후 permission 문제가 발생하면 원인을 확인하고 필요한 정책 또는 Context를 수정한다.
3.2 기존 설정 백업
기존 운영 서버에 구축하는 경우 설정을 변경하기 전에 백업한다.
umask 077
MON_BACKUP="/root/monitor-backup-$(date +%Y%m%d-%H%M%S)"
install -d -m 700 "$MON_BACKUP"
cp -a /etc/rsyslog.conf "$MON_BACKUP/" 2>/dev/null || true
cp -a /etc/rsyslog.d "$MON_BACKUP/" 2>/dev/null || true
cp -a /etc/chrony.conf "$MON_BACKUP/" 2>/dev/null || true
cp -a /etc/firewalld "$MON_BACKUP/" 2>/dev/null || true
rpm -qa | sort > "$MON_BACKUP/packages-before.txt"
백업 경로 확인:
echo "$MON_BACKUP"
ls -al "$MON_BACKUP"
3.3 기본 관리 도구 설치
dnf install -y \
curl \
ca-certificates \
gnupg2 \
vim-enhanced \
net-snmp-utils \
rsyslog \
logrotate \
chrony \
tcpdump \
iputils \
sysstat \
policycoreutils-python-utils \
dnf-plugins-core \
jq
주요 패키지 용도:
- net-snmp-utils : snmpget, snmpwalk 테스트
- tcpdump : 실제 Packet 수신 확인
- jq : JSON 출력 가독성 개선
- sysstat : iostat 등 서버 I/O 점검
- chrony : 시간 동기화
- policycoreutils-python-utils : SELinux 진단/정책 작업
3.4 시간 동기화
모니터링에서는 서버와 장비 시간이 맞지 않으면 장애 시각 비교가 어렵다.
Timezone 설정:
timedatectl set-timezone Asia/Seoul
chronyd 시작:
systemctl enable --now chronyd
systemctl restart chronyd
확인:
chronyc tracking
chronyc sources -v
date -Ins
정상 확인
- chronyd가 active
- chronyc sources에서 선택된 NTP Source 존재
- 현재 시간이 실제 시간과 크게 다르지 않음
4. InfluxDB와 Telegraf를 설치하는 이유
스위치의 SNMP 값을 수집하려면 두 프로그램이 필요하다.
Telegraf는 장비에 SNMP Query를 보내는 역할을 한다.
InfluxDB는 Telegraf가 가져온 값을 시간 순서대로 저장한다.
Switch --SNMP--> Telegraf --> InfluxDB --> Grafana
5. InfluxDB / Telegraf / Grafana 설치
5.1 InfluxData Repository 추가
공식 Repository Key를 내려받는다.
curl -fL https://repos.influxdata.com/influxdata-archive.key \
-o /etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata
gpg --show-keys --with-fingerprint \
/etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata
rpm --import /etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata
Repository 파일 생성:
vi /etc/yum.repos.d/influxdata.repo
[influxdata]
name=InfluxData Repository - Stable
baseurl=https://repos.influxdata.com/stable/$basearch/main
enabled=1
gpgcheck=1
gpgkey=file:///etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata
sslverify=1
5.2 Grafana Repository 추가
vi /etc/yum.repos.d/grafana.repo
[grafana]
name=Grafana OSS repository
baseurl=https://rpm.grafana.com
repo_gpgcheck=1
enabled=1
gpgcheck=1
gpgkey=https://rpm.grafana.com/gpg.key
sslverify=1
5.3 Repository 확인 후 설치
dnf makecache
dnf list --showduplicates \
telegraf \
influxdb2 \
influxdb2-cli \
grafana
설치:
dnf install -y \
telegraf \
influxdb2 \
influxdb2-cli \
grafana
버전 확인:
rpm -q telegraf influxdb2 influxdb2-cli grafana
telegraf --version
influxd version
influx version
왜 버전을 기록하는가?
나중에 설정 문법 또는 Plugin 동작이 달라졌을 때 원인을 판단하기 쉽기 때문이다.
실제 구성 과정에서는 Telegraf 1.40.0 환경에서 동작을 확인하였다.
6. InfluxDB 초기 설정
6.1 InfluxDB를 localhost에만 Bind
현재 구성에서는 InfluxDB를 외부 장비가 직접 접근할 필요가 없다.
Telegraf와 Grafana가 같은 서버에 있으므로 127.0.0.1에만 Bind하면 공격 표면을 줄일 수 있다.
install -d -m 755 \
/etc/systemd/system/influxdb.service.d
vi /etc/systemd/system/influxdb.service.d/10-listen.conf
[Service]
Environment="INFLUXD_HTTP_BIND_ADDRESS=127.0.0.1:8086"
적용:
systemctl daemon-reload
systemctl enable --now influxdb
systemctl restart influxdb
systemctl status influxdb --no-pager
Listen 확인:
ss -lntp | grep ':8086'
Health Check:
curl -fsS http://127.0.0.1:8086/health
정상 확인
127.0.0.1:8086에서 LISTEN하고 health 응답이 정상이어야 한다.
6.2 InfluxDB 최초 Setup
influx setup
초기 입력 예:
| 항목 | 예시 | 설명 |
|---|---|---|
| Username | monitor-admin | InfluxDB 관리 계정 |
| Organization | network | 관련 데이터의 논리적 그룹 |
| Bucket | snmp_raw | SNMP 데이터 저장 위치 |
| Retention | 720h | 30일 보존 |
확인:
influx bucket list --org network
Bucket이란?
일반 Database의 Database 또는 Table과 완전히 같지는 않지만, 처음에는 "Metric 저장 공간" 정도로 이해하면 된다.
6.3 Token을 서비스별로 나누는 이유
Telegraf는 데이터를 쓰기만 하면 되고 Grafana는 읽기만 하면 된다.
따라서 하나의 관리자 Token을 모든 프로그램에 넣지 않는다.
권장:
| Token | 권한 |
|---|---|
| telegraf-write | snmp_raw Write |
| grafana-read | snmp_raw Read |
| Operator Token | 관리자 전용 |
실제 Token 값은 Wiki에 기록하지 않는다.
7. Telegraf 기본 설정
7.1 Telegraf의 역할
Telegraf는 일정 시간마다 스위치에 SNMP Query를 보내고 결과를 InfluxDB에 저장한다.
Telegraf
|
| UDP/161 SNMP GET
v
Switch
Telegraf
|
| HTTP 8086
v
InfluxDB
7.2 인증 정보를 설정 파일과 분리
SNMP Password와 InfluxDB Token을 여러 설정 파일에 직접 적으면 관리가 어렵다.
환경변수 파일로 분리한다.
vi /etc/telegraf/monitor.env
INFLUX_WRITE_TOKEN='<WRITE_TOKEN>'
ZYXEL_SNMP_USER='<SNMP_USER>'
ZYXEL_SNMP_AUTH='<AUTH_PASSWORD>'
ZYXEL_SNMP_PRIV='<PRIV_PASSWORD>'
QNAP_SNMP_USER='<QNAP_SNMP_USER>'
QNAP_SNMP_AUTH='<QNAP_AUTH_PASSWORD>'
QNAP_SNMP_PRIV='<QNAP_PRIV_PASSWORD>'
권한 제한:
chown root:root /etc/telegraf/monitor.env
chmod 600 /etc/telegraf/monitor.env
systemd가 환경 파일을 읽도록 설정:
install -d -m 755 \
/etc/systemd/system/telegraf.service.d
vi /etc/systemd/system/telegraf.service.d/10-monitor-env.conf
[Service]
EnvironmentFile=/etc/telegraf/monitor.env
7.3 Telegraf Main 설정
원본 백업:
cp -a /etc/telegraf/telegraf.conf \
/etc/telegraf/telegraf.conf.orig
편집:
vi /etc/telegraf/telegraf.conf
기본 예:
[agent]
interval = "30s"
round_interval = true
flush_interval = "5s"
precision = "1ms"
metric_batch_size = 1000
metric_buffer_limit = 20000
omit_hostname = false
snmp_translator = "gosmi"
[[outputs.influxdb_v2]]
urls = ["http://127.0.0.1:8086"]
token = "${INFLUX_WRITE_TOKEN}"
organization = "network"
bucket = "snmp_raw"
[[inputs.internal]]
[[inputs.cpu]]
percpu = false
totalcpu = true
[[inputs.mem]]
[[inputs.disk]]
mount_points = ["/"]
[[inputs.net]]
각 값의 의미:
- interval = 30s : 기본 수집 주기
- flush_interval = 5s : 모아둔 Metric을 DB에 보내는 주기
- outputs.influxdb_v2 : 수집 결과를 어느 InfluxDB에 저장할지 지정
- inputs.internal : Telegraf 자신의 상태
- inputs.cpu/mem/disk/net : 모니터링 서버 자체 상태
8. SNMP가 정상인지 먼저 수동으로 확인
Telegraf 설정 전에 SNMP 자체가 되는지 확인하는 것이 중요하다.
Telegraf가 안 된다고 바로 Telegraf 설정만 수정하면 실제 원인이 SNMP ACL인지 Password인지 구분하기 어렵다.
8.1 SNMPv3 계정 파일
install -d -m 700 /root/.snmp
vi /root/.snmp/snmp.conf
예:
defVersion 3
defSecurityName <SNMP_USER>
defSecurityLevel authPriv
defAuthType SHA
defAuthPassphrase <SNMP_AUTH_PASSWORD>
defPrivType AES
defPrivPassphrase <SNMP_PRIV_PASSWORD>
권한:
chmod 600 /root/.snmp/snmp.conf
8.2 Uptime 조회
snmpget -v3 \
-t 2 \
-r 0 \
-On \
203.0.113.11 \
.1.3.6.1.2.1.1.3.0
응답이 나오면 최소한 다음이 정상이다.
- IP Routing
- UDP/161
- SNMP User
- Authentication
- Privacy Password
- SNMP View
8.3 Interface 이름 확인
snmpwalk -v3 \
-t 2 \
-r 0 \
-On \
203.0.113.11 \
.1.3.6.1.2.1.31.1.1.1.1
여기서 마지막 숫자가 ifIndex이다.
예를 들어 결과가 다음과 같다고 가정한다.
.1.3.6.1.2.1.31.1.1.1.1.24 = STRING: port24
이 경우 ifIndex는 24이다.
중요: 물리 Port 24번이 항상 ifIndex 24라는 의미는 아니다. 실제 SNMP 결과를 기준으로 한다.
9. Telegraf SNMP 설정
9.1 모델별 파일로 나누는 이유
제조사가 같아도 모델마다 CPU/Memory OID와 SNMP 암호화 방식이 다를 수 있다.
따라서 하나의 거대한 파일보다는 모델별 파일로 나눈다.
/etc/telegraf/telegraf.d/
├── 10-zyxel-gs1900.conf
├── 11-zyxel-gs1920.conf
├── 20-zyxel-es3128.conf
├── 30-icmp_check.conf
└── 40-qnap_nas.conf
9.2 GS1920 계열
실제 확인된 특성:
- SNMPv3 SHA + DES
- CPU/Memory Private OID 사용
CPU:
.1.3.6.1.4.1.890.1.15.3.49.1.7.0
Memory:
Total .1.3.6.1.4.1.890.1.15.3.50.1.1.1.3.1
Used .1.3.6.1.4.1.890.1.15.3.50.1.1.1.4.1
Percent .1.3.6.1.4.1.890.1.15.3.50.1.1.1.5.1
9.3 GS1900 계열
CPU:
.1.3.6.1.4.1.890.1.15.3.2.4.0
Memory:
.1.3.6.1.4.1.890.1.15.3.2.5.0
실제 구축 과정에서는 SNMP 응답이 느려 다음 값이 안정적이었다.
timeout = "5s"
retries = 1
왜 Timeout을 늘렸는가?
수동 snmpget은 되는데 Telegraf에서만 간헐적으로 누락된다면 Telegraf timeout이 장비 응답시간보다 짧을 수 있다.
Timeout을 무조건 크게 설정하기보다는 실제 응답을 보고 조정한다.
9.4 ES-3128GP
실제 확인된 SNMPv3 방식:
- SHA
- AES
sysObjectID:
.1.3.6.1.4.1.7800.1.190
Port Metric은 수집 가능하지만 CPU/Memory Private OID는 정상 확인되지 않아 Dashboard에서 억지로 0으로 표시하지 않는다.
수집되지 않는 값과 0은 의미가 다르다.
CPU를 조회할 수 없는 장비에 CPU 0%라고 표시하면 운영자가 정상 상태라고 오해할 수 있다.
9.5 기본 Port Metric
가능하면 제조사 Private MIB보다 표준 IF-MIB를 먼저 사용한다.
주요 항목:
- ifName
- ifAlias
- ifSpeed
- ifHCInOctets
- ifHCOutOctets
- ifOperStatus
- ifInErrors
- ifOutErrors
- ifInDiscards
- ifOutDiscards
- Broadcast
- Multicast
10. ICMP Ping 수집
SNMP가 정상이어도 장비 자체가 네트워크에서 사라질 수 있다.
따라서 Ping 결과도 별도로 저장한다.
vi /etc/telegraf/telegraf.d/30-icmp_check.conf
예:
[[inputs.ping]]
urls = [
"203.0.113.11",
"203.0.113.12",
"203.0.113.13"
]
method = "native"
count = 3
deadline = 2.0
interval = 10.0
10.1 CAP_NET_RAW가 필요한 이유
native Ping은 Raw Socket을 사용한다.
root로 telegraf --test를 실행하면 정상인데 systemd 서비스에서는 Ping이 실패할 수 있다.
이 경우 서비스에 CAP_NET_RAW를 부여한다.
systemctl edit telegraf
[Service]
CapabilityBoundingSet=CAP_NET_RAW
AmbientCapabilities=CAP_NET_RAW
반영:
systemctl daemon-reload
systemctl restart telegraf
11. Telegraf 설정 검사
서비스를 재시작하기 전에 설정 오류를 확인한다.
telegraf \
--config /etc/telegraf/telegraf.conf \
--config-directory /etc/telegraf/telegraf.d \
--test
주의: --test는 Metric을 화면에 출력하지만 일반적으로 Output 저장까지 검증하는 용도가 아니다.
서비스 계정 조건까지 확인하려면 다음 방식이 더 정확하다.
systemd-run \
--unit=telegraf-config-check \
--wait \
--pipe \
--collect \
-p User=telegraf \
-p Group=telegraf \
-p EnvironmentFile=/etc/telegraf/monitor.env \
-p CapabilityBoundingSet=CAP_NET_RAW \
-p AmbientCapabilities=CAP_NET_RAW \
/usr/bin/telegraf \
--config /etc/telegraf/telegraf.conf \
--config-directory /etc/telegraf/telegraf.d \
--test
정상 후 서비스 시작:
systemctl enable --now telegraf
systemctl restart telegraf
systemctl status telegraf --no-pager
journalctl -u telegraf \
-n 100 \
--no-pager
로그에서 주의할 문자열:
timeout
unauthorized
permission denied
buffer
drop
address already in use
12. Grafana 설치 후 처음 해야 할 작업
Grafana는 데이터를 직접 수집하지 않는다.
먼저 InfluxDB와 Prometheus 같은 데이터소스를 등록해야 한다.
12.1 Grafana 서비스
systemctl enable --now grafana-server
systemctl status grafana-server --no-pager
curl -fsS \
http://127.0.0.1:3000/api/health
웹 접속:
http://192.0.2.10:3000
내부망에서만 사용할 경우에도 방화벽 접근 범위를 관리망으로 제한하는 것이 좋다.
12.2 InfluxDB Datasource 등록
Grafana 메뉴:
Connections
→ Data sources
→ Add data source
→ InfluxDB
설정:
| 항목 | 값 |
|---|---|
| Query Language | Flux |
| URL | http://127.0.0.1:8086 |
| Organization | network |
| Token | grafana-read Token |
| Default Bucket | snmp_raw |
Save & Test 성공 여부를 확인한다.
13. 첫 번째 SNMP Dashboard 만들기
처음부터 모든 장비를 한 화면에 넣지 않는다.
한 대의 스위치 + 한 포트로 정상 그래프를 만든 뒤 확대하는 것이 가장 쉽다.
13.1 원시 Counter 확인
Explore에서 다음과 같이 최근 데이터를 확인한다.
from(bucket: "snmp_raw")
|> range(start: -10m)
|> limit(n: 20)
여기서 실제 Measurement, source, index, if_name 값을 확인한다.
13.2 Port RX/TX 계산
누적 Octet Counter를 bps로 변환한다.
from(bucket: "snmp_raw")
|> range(start: v.timeRangeStart, stop: v.timeRangeStop)
|> filter(fn: (r) =>
r._measurement == "switch_port"
)
|> filter(fn: (r) =>
r.source == "203.0.113.11"
)
|> filter(fn: (r) =>
r._field == "in_octets" or
r._field == "out_octets"
)
|> derivative(
unit: 1s,
nonNegative: true
)
|> map(fn: (r) => ({
r with
_value: r._value * 8.0
}))
Grafana Unit:
bits/sec
13.3 Error / Discard
Error나 Discard 역시 Counter이므로 증가율을 표시한다.
from(bucket: "snmp_raw")
|> range(start: v.timeRangeStart, stop: v.timeRangeStop)
|> filter(fn: (r) =>
r._measurement == "switch_port"
)
|> filter(fn: (r) =>
r._field == "in_errors" or
r._field == "out_errors" or
r._field == "in_discards" or
r._field == "out_discards"
)
|> derivative(
unit: 1s,
nonNegative: true
)
해석
- Error 증가 → 물리계층, Frame, Interface 문제 가능
- Discard 증가 → Queue, Buffer, QoS, 혼잡 등에 의해 폐기될 가능성
- 값이 0인 상태가 일반적이지만 장비 특성과 Traffic에 따라 해석해야 함
13.4 operStatus
operStatus는 Counter가 아니다.
따라서 derivative를 사용하지 않는다.
일반적인 값:
1 = up
2 = down
Grafana Value Mapping으로 사람이 읽기 쉬운 문자열로 변환한다.
14. Prometheus를 추가하는 이유
SNMP는 네트워크 장비 상태에 좋지만 Linux/Windows 서버 자원 모니터링에는 Exporter + Prometheus가 더 편하다.
Prometheus 구조:
Node Exporter ----\
\
Windows Exporter ---> Prometheus ---> Grafana
/
QNAP Exporter -----/
Prometheus는 일정 주기마다 Exporter HTTP Endpoint에 접속하여 Metric을 가져간다.
이를 Pull 방식이라고 한다.
15. Prometheus 설치
15.1 사용자와 디렉터리
useradd \
--system \
--no-create-home \
--shell /sbin/nologin \
prometheus
install -d \
-o prometheus \
-g prometheus \
/etc/prometheus \
/var/lib/prometheus
Prometheus 공식 바이너리를 준비한 뒤:
install -m 0755 \
prometheus \
/usr/local/bin/prometheus
install -m 0755 \
promtool \
/usr/local/bin/promtool
버전 확인:
prometheus --version
promtool --version
15.2 prometheus.yml
vi /etc/prometheus/prometheus.yml
global:
scrape_interval: 30s
scrape_configs:
- job_name: prometheus
static_configs:
- targets:
- "127.0.0.1:9090"
labels:
server_name: monitoring-server
- job_name: node
static_configs:
- targets:
- "127.0.0.1:9100"
labels:
server_name: monitoring-server
- targets:
- "192.0.2.20:9100"
labels:
server_name: linux-server
- job_name: windows
static_configs:
- targets:
- "192.0.2.30:9182"
labels:
server_name: windows-server
- job_name: qnap
static_configs:
- targets:
- "198.51.100.10:9100"
labels:
server_name: nas-01
scrape_interval = 30s는 Prometheus가 30초마다 각 Target을 조회한다는 의미이다.
15.3 설정 검사
YAML은 들여쓰기에 매우 민감하다.
실제 구축 중에도 job을 scrape_configs 바깥에 잘못 넣으면 Prometheus가 시작하지 못했다.
반드시 검사한다.
promtool check config \
/etc/prometheus/prometheus.yml
15.4 systemd 등록
vi /etc/systemd/system/prometheus.service
[Unit]
Description=Prometheus
Wants=network-online.target
After=network-online.target
[Service]
User=prometheus
Group=prometheus
ExecStart=/usr/local/bin/prometheus \
--config.file=/etc/prometheus/prometheus.yml \
--storage.tsdb.path=/var/lib/prometheus \
--storage.tsdb.retention.time=30d \
--storage.tsdb.retention.size=20GB \
--web.listen-address=127.0.0.1:9090
Restart=always
[Install]
WantedBy=multi-user.target
권한:
chown -R prometheus:prometheus \
/etc/prometheus \
/var/lib/prometheus
시작:
systemctl daemon-reload
systemctl enable --now prometheus
systemctl status prometheus --no-pager
Health:
curl -s \
http://127.0.0.1:9090/-/healthy
16. Linux Node Exporter 설치
16.1 Node Exporter가 하는 일
Node Exporter는 Linux Kernel과 /proc, /sys 등의 정보를 읽어 Prometheus 형식으로 공개한다.
대표 Metric:
- CPU
- Memory
- Load
- Filesystem
- Disk I/O
- Network Traffic
- Network Error/Drop
- Uptime
- File Descriptor
- 일부 Hardware Metric
16.2 서비스 계정
useradd \
--system \
--no-create-home \
--shell /sbin/nologin \
node_exporter
바이너리 설치:
install -m 0755 \
node_exporter \
/usr/local/bin/node_exporter
16.3 systemd
vi /etc/systemd/system/node_exporter.service
[Unit]
Description=Node Exporter
After=network-online.target
Wants=network-online.target
[Service]
User=node_exporter
Group=node_exporter
ExecStart=/usr/local/bin/node_exporter
Restart=always
[Install]
WantedBy=multi-user.target
systemctl daemon-reload
systemctl enable --now node_exporter
systemctl status node_exporter --no-pager
Metric 확인:
curl -s \
http://127.0.0.1:9100/metrics \
| head
다른 서버에 설치한 경우 Prometheus 서버에서 확인:
curl -s \
http://192.0.2.20:9100/metrics \
| head
정상 확인
Metric 문자열이 여러 줄 출력되면 Exporter 자체는 정상이다.
Prometheus에서 다음 Query도 확인한다.
up
값:
1 = 정상 Scrape
0 = Scrape 실패
17. Windows Exporter
Windows에서는 windows_exporter를 설치한다.
기본 Port:
9182/tcp
Prometheus에서 다음 주소를 조회할 수 있어야 한다.
http://192.0.2.30:9182/metrics
Windows PowerShell 테스트:
Invoke-WebRequest \
-UseBasicParsing \
http://127.0.0.1:9182/metrics
성능 카운터가 비정상인 경우 실제 구축 과정에서 다음 복구 절차를 사용하였다.
lodctr /R
winmgmt /resyncperf
그 후:
Restart-Service windows_exporter
18. Grafana에 Prometheus 등록
Grafana 메뉴:
Connections
→ Data sources
→ Prometheus
URL:
http://127.0.0.1:9090
Save & Test 후 Explore에서:
up
Target별 1이 표시되면 정상이다.
19. Linux Server Dashboard 이해하기
처음에는 다음 항목만 만든다.
- Uptime
- CPU
- Memory
- Load Average
- Filesystem
- Network RX/TX
이후 익숙해지면 다음을 추가한다.
- Disk I/O
- Network Error/Drop
- inode
- Swap
- Process
- Application Log
19.1 CPU Usage
Node Exporter는 CPU 시간을 Mode별 누적으로 제공한다.
idle을 제외한 비율로 CPU 사용률을 계산할 수 있다.
100 - (
avg by(instance) (
rate(
node_cpu_seconds_total{
mode="idle"
}[5m]
)
) * 100
)
19.2 Memory Usage
100 * (
1 -
node_memory_MemAvailable_bytes
/
node_memory_MemTotal_bytes
)19.3 Filesystem Usage
Filesystem은 tmpfs, overlay, snapshot 등 운영자가 원하지 않는 Mount가 함께 보일 수 있다.
따라서 실제 데이터 Mount Point를 확인한 뒤 필터링한다.
20. QNAP TS-264 모니터링
QNAP은 한 가지 수집 방식만 사용하면 정보가 부족하다.
따라서 세 경로를 함께 사용한다.
QNAP
|
+-- node_exporter --> Prometheus
|
+-- SNMP ----------> Telegraf --> InfluxDB
|
+-- Syslog --------> rsyslog --> Alloy --> Loki역할:
| 항목 | 데이터 소스 |
|---|---|
| CPU / Memory / Network | Node Exporter |
| Filesystem | Node Exporter |
| Disk I/O | Node Exporter |
| HDD Model / 온도 / 상태 | SNMP |
| RAID / Volume | SNMP |
| Event Log | Syslog |
| Access Log | Syslog |
20.1 QNAP SNMPv3
실제 TS-264 환경에서는 다음 조합을 사용하였다.
- Authentication: SHA
- Privacy: DES
QNAP은 해당 구성에서 AES가 동작하지 않았다.
따라서 다른 장비에서 AES가 된다는 이유로 QNAP도 AES라고 가정하지 않는다.
20.2 주요 Measurement
QNAP_TS264
qnap_disk
qnap_raid
qnap_storage_pool
qnap_volume20.3 Volume 값이 11 GiB로 잘못 보이는 문제
실제 QTS에서 DataVol1은 약 11.33 TiB인데 SNMP 원시값을 Grafana에서 byte로 바로 처리하면 약 11.3 GiB로 표시되는 문제가 있었다.
실제 장비의 qnap_volume capacity/free 값이 KiB 성격으로 반환되어 1024 배 보정이 필요하였다.
|> map(fn: (r) => ({
r with
capacity_bytes:
uint(v: r.capacity_bytes)
* uint(v: 1024),
free_bytes:
uint(v: r.free_bytes)
* uint(v: 1024),
used_bytes:
(
uint(v: r.capacity_bytes)
- uint(v: r.free_bytes)
)
* uint(v: 1024),
used_percent:
if float(v: r.capacity_bytes) > 0.0 then
(
float(v: r.capacity_bytes)
- float(v: r.free_bytes)
)
/
float(v: r.capacity_bytes)
* 100.0
else
0.0
}))왜 Percent에는 1024가 필요 없는가?
분자와 분모가 같은 단위이므로 비율 계산에서는 단위가 상쇄된다.
20.4 실제 Data Volume 찾기
Node Exporter가 QNAP 내부의 Snapshot Mount까지 모두 보여주므로 처음에는 여러 개의 11.33 TiB Filesystem이 보일 수 있다.
실제 사용자 Data Volume은 다음이었다.
mountpoint="/share/CACHEDEV1_DATA"
device="/dev/mapper/cachedev1"
fstype="ext4"QNAP Snapshot:
/mnt/snapshot/1/10001
/mnt/snapshot/1/10002
...Dashboard에서는 Snapshot을 제외하고 실제 Volume만 표시한다.
사용률:
100 * (
1 -
node_filesystem_avail_bytes{
instance="198.51.100.10:9100",
mountpoint="/share/CACHEDEV1_DATA"
}
/
node_filesystem_size_bytes{
instance="198.51.100.10:9100",
mountpoint="/share/CACHEDEV1_DATA"
}
)실제 QTS 화면과 약 2.86%로 일치하는 것을 확인하였다.
20.5 QNAP Network는 bond0만 표시
QNAP에는 내부 Interface와 Virtual Interface가 여러 개 보일 수 있다.
실제 외부 Traffic을 담당하는 bond0만 필터링한다.
RX:
rate(
node_network_receive_bytes_total{
instance="198.51.100.10:9100",
device="bond0"
}[$__rate_interval]
) * 8TX:
rate(
node_network_transmit_bytes_total{
instance="198.51.100.10:9100",
device="bond0"
}[$__rate_interval]
) * 820.6 Network Error / Drop
RX Error:
rate(
node_network_receive_errs_total{
instance="198.51.100.10:9100",
device="bond0"
}[$__rate_interval]
)TX Error:
rate(
node_network_transmit_errs_total{
instance="198.51.100.10:9100",
device="bond0"
}[$__rate_interval]
)RX Drop:
rate(
node_network_receive_drop_total{
instance="198.51.100.10:9100",
device="bond0"
}[$__rate_interval]
)TX Drop:
rate(
node_network_transmit_drop_total{
instance="198.51.100.10:9100",
device="bond0"
}[$__rate_interval]
)20.7 Disk I/O
QNAP에는 md, dm, loop 등 내부 Device가 많으므로 물리 Disk sd*만 표시한다.
Read Bytes/sec:
rate(
node_disk_read_bytes_total{
instance="198.51.100.10:9100",
device=~"sd[a-z]+"
}[$__rate_interval]
)Write Bytes/sec:
rate(
node_disk_written_bytes_total{
instance="198.51.100.10:9100",
device=~"sd[a-z]+"
}[$__rate_interval]
)Read IOPS:
rate(
node_disk_reads_completed_total{
instance="198.51.100.10:9100",
device=~"sd[a-z]+"
}[$__rate_interval]
)Write IOPS:
rate(
node_disk_writes_completed_total{
instance="198.51.100.10:9100",
device=~"sd[a-z]+"
}[$__rate_interval]
)21. Syslog를 왜 별도로 구성하는가
Metric만으로는 다음과 같은 내용을 알기 어렵다.
- Port가 왜 Down 되었는가
- 사용자가 NAS에 어떤 파일을 접근했는가
- 서비스가 언제 재시작되었는가
- 인증 실패가 발생했는가
이런 이벤트는 Log가 필요하다.
본 구성의 로그 흐름:
Device
|
| Syslog
v
rsyslog
|
v
Log File
|
v
Alloy
|
v
Loki
|
v
Grafana22. rsyslog 구성
22.1 Network Syslog
네트워크 장비 로그:
/var/log/network-syslog/events.log기본 수신:
TCP/UDP 51422.2 Server Syslog
서버용 로그를 네트워크 장비와 분리하면 Grafana에서 Query하기 쉽다.
예:
/var/log/server-syslog/events.log수신 Port:
TCP 551422.3 QNAP Event와 Access Log 분리
QNAP QuLog의 두 성격이 다르므로 포트부터 분리한다.
| 종류 | 포트 | 파일 | 의미 |
|---|---|---|---|
| Event | TCP 5515 | /var/log/qnap/event.log | 시스템/서비스 이벤트 |
| Access | TCP 5516 | /var/log/qnap/access.log | 사용자/파일 접근 |
Template:
template(name="QnapSyslogLine" type="string"
string="%timegenerated:::date-rfc3339% src=%fromhost-ip% severity=%syslogseverity-text% host=%hostname% %syslogtag%%msg:::sp-if-no-1st-sp%%msg%\n")이렇게 저장하면 한 줄이 대략 다음 구조가 된다.
2026-09-13T23:40:03+09:00 \
src=198.51.100.10 \
severity=info \
host=nas-01 \
qulogd: ...Event ruleset:
ruleset(name="QnapEventLog") {
action(
type="omfile"
file="/var/log/qnap/event.log"
template="QnapSyslogLine"
fileOwner="root"
fileGroup="alloy"
fileCreateMode="0640"
dirOwner="root"
dirGroup="alloy"
dirCreateMode="0750"
createDirs="on"
)
stop
}
input(
type="imtcp"
port="5515"
ruleset="QnapEventLog"
)Access ruleset:
ruleset(name="QnapAccessLog") {
action(
type="omfile"
file="/var/log/qnap/access.log"
template="QnapSyslogLine"
fileOwner="root"
fileGroup="alloy"
fileCreateMode="0640"
dirOwner="root"
dirGroup="alloy"
dirCreateMode="0750"
createDirs="on"
)
stop
}
input(
type="imtcp"
port="5516"
ruleset="QnapAccessLog"
)22.4 rsyslog 설정 검사
재시작 전에 반드시 검사한다.
rsyslogd -N1오류가 없으면:
systemctl restart rsyslog
ss -lntp \
| grep -E ':5514|:5515|:5516'실제 로그 확인:
tail -f \
/var/log/qnap/access.log문제 분리 방법
tcpdump에는 Packet이 안 보임
→ 장비 송신/방화벽/라우팅 확인
tcpdump에는 보이지만 파일이 안 생김
→ rsyslog 설정 확인
파일은 생기지만 Grafana에 안 보임
→ Alloy/Loki 확인23. Loki와 Alloy의 역할
rsyslog가 파일을 만드는 것만으로 Grafana가 그 파일을 검색할 수 있는 것은 아니다.
Loki가 로그 저장소 역할을 하고 Alloy가 파일을 읽어 Loki에 넣는다.
/var/log/...
|
v
Alloy
|
v
Loki
|
v
Grafana현재 구성에서는 Loki가 localhost 3100에서 동작한다.
127.0.0.1:3100Alloy 로컬 관리 Endpoint:
127.0.0.1:1234524. Alloy 설정 이해하기
24.1 Network Syslog
loki.source.file "network_syslog" {
targets = [
{
__path__ = "/var/log/network-syslog/events.log",
job = "network-syslog",
},
]
forward_to = [
loki.process.network_syslog.receiver
]
}source.file은 어떤 파일을 읽을지 지정한다.
job은 Loki에서 로그 종류를 구분하기 위한 Label이다.
그 다음 regex로 한 줄을 분해한다.
loki.process "network_syslog" {
stage.regex {
expression = `^(?P<received_at>\S+) src=(?P<device_ip>\S+) severity=(?P<severity>\S+) host=(?P<device_host>\S+) (?P<message>.*)$`
}
stage.timestamp {
source = "received_at"
format = "RFC3339Nano"
action_on_failure = "skip"
}
stage.labels {
values = {
device_ip = ""
severity = ""
}
}
forward_to = [
loki.write.local.receiver
]
}여기서 다음 Label이 생긴다.
job
device_ip
severitydevice_host를 Label로 추가하려면
stage.labels {
values = {
device_ip = ""
device_host = ""
severity = ""
}
}이는 이후 Alert Mail에서 Hostname까지 표시하고 싶을 때 유용하다.
24.2 QNAP Access
loki.source.file "qnap_access" {
targets = [
{
__path__ = "/var/log/qnap/access.log",
job = "qnap-access",
},
]
forward_to = [
loki.process.qnap_access.receiver
]
}
loki.process "qnap_access" {
stage.regex {
expression = `^(?P<received_at>\S+) src=(?P<nas_ip>\S+) severity=(?P<severity>\S+) host=(?P<nas_host>\S+) (?P<message>.*)$`
}
stage.timestamp {
source = "received_at"
format = "RFC3339Nano"
action_on_failure = "skip"
}
stage.labels {
values = {
nas_ip = ""
nas_host = ""
severity = ""
log_type = "access"
}
}
forward_to = [
loki.write.local.receiver
]
}24.3 Loki Write
loki.write "local" {
endpoint {
url = "http://127.0.0.1:3100/loki/api/v1/push"
}
}즉 Alloy에서 처리한 로그를 로컬 Loki로 전송한다.
24.4 설정 검사
alloy validate \
/etc/alloy/config.alloy정상이라면 아무 오류 없이 종료된다.
반영:
systemctl restart alloy
systemctl status alloy \
--no-pagerLoki에 실제 Job이 만들어졌는지 확인:
curl -s \
'http://127.0.0.1:3100/loki/api/v1/label/job/values' \
| jq중요한 점은 Alloy 설정에 job을 적었다고 바로 Loki에 Label이 생기는 것이 아니라 실제 로그가 최소 한 번 Loki에 저장되어야 조회 결과에 나타난다는 것이다.
25. Alloy Permission 문제 해결 사례
실제 구축 과정에서 QNAP 로그 파일은 존재하지만 Alloy가 다음 오류를 출력하였다.
failed to tail file
stat failed
permission denied먼저 경로 전체 권한 확인:
namei -l \
/var/log/qnap/access.log정상 예:
drwxr-x--- root alloy qnap
-rw-r----- root alloy access.logAlloy 계정으로 직접 읽기:
sudo -u alloy \
head /var/log/qnap/access.log그래도 실패하면 SELinux 확인:
getenforce
ausearch \
-m AVC \
-ts recent \
| grep -Ei 'alloy|qnap'
ls -Zd /var/log/qnap
ls -Z /var/log/qnap/access.log문제 해결 원칙
권한 문제가 있다고 SELinux를 바로 끄지 않는다.
다음 순서로 본다.
- 파일 Owner/Group
- 파일 Mode
- 상위 Directory execute 권한
- Alloy Service User
- SELinux Context / AVC
26. Grafana에 Loki 추가
Grafana:
Connections
→ Data sources
→ Add data source
→ LokiURL:
http://127.0.0.1:3100Explore에서 Network 로그 확인:
{job="network-syslog"}QNAP Access:
{job="qnap-access"}QNAP Event:
{job="qnap-event"}27. Grafana Alert를 처음 구성할 때 알아둘 점
Grafana Alert는 Dashboard의 그래프와 달리 최종적으로 숫자 하나 또는 Alert Instance별 숫자를 평가해야 한다.
Range Query를 그대로 Alert Condition에 사용하면 다음 오류가 발생할 수 있다.
looks like time series data,
only reduced data can be alerted on따라서 다음 중 하나를 사용한다.
- Query를 Instant로 구성
- Range Query → Reduce → Threshold
28. Syslog Critical Alert
Level 3(Error) 이상:
sum by (device_ip, severity) (
count_over_time(
{
job="network-syslog",
severity=~"emerg|alert|crit|err"
}[1m]
)
)권장:
A = Loki Instant Query
B = Threshold
A IS ABOVE 0Summary:
[Syslog 경고]
{{ $labels.device_ip }}
{{ $labels.severity }}Description:
장비 {{ $labels.device_ip }} 에서
Syslog Level 3 이상 로그가 발생했습니다.
Severity:
{{ $labels.severity }}
최근 1분 발생 건수:
{{ $values.A.Value }}28.1 Alert Mail에 실제 로그 내용이 없는 이유
count_over_time()은 문자열 로그를 숫자로 집계한다.
따라서 Query 결과에는 주로 다음만 남는다.
- Label
- 발생 건수
로그 message 전체를 Loki Label로 만들면 종류가 지나치게 많아져 Cardinality 문제가 발생할 수 있으므로 권장하지 않는다.
메일에는 다음 정도를 넣고 실제 내용은 Grafana Explore에서 확인하는 구조가 안정적이다.
- Device IP
- Hostname
- Severity
- 발생 건수
- Grafana Link
29. ICMP Down Alert
Flux:
from(bucket: "snmp_raw")
|> range(start: -5m)
|> filter(fn: (r) =>
r._measurement == "ping" and
r._field == "percent_packet_loss"
)
|> group(
columns: ["url"]
)
|> last()
|> keep(
columns: [
"_time",
"_value",
"url"
]
)Alert 구조:
A = Flux Query
B = Reduce
Last
Strict
C = Threshold
B > 99
Pending = 2mNo Data 정책:
Keep Last State왜 Keep Last State인가?
Metric이나 Log Source가 순간적으로 No Data가 되었을 때 별도의 DatasourceNoData 메일이 발생하면서 실제 장비 Label이 없는 알림이 발송될 수 있다.
No Data 자체를 별도 장애로 관리할 필요가 있다면 별도의 수집 상태 Alert를 구성하는 것이 더 명확하다.
30. 최종 Dashboard 구성 방법
처음 설치한 사용자는 Dashboard를 한 번에 완성하려 하지 말고 다음 단계로 만든다.
- 데이터소스 Save & Test
- Explore에서 실제 데이터 확인
- Stat 패널 하나 생성
- Time Series 하나 생성
- 장비 한 대 정상 확인
- 변수 추가
- 여러 장비로 확대
- Alert 추가
30.1 Network Switch Dashboard
권장 Row:
| Row | 표시 내용 | 목적 |
|---|---|---|
| Device 상태 | Device Name / Uptime / CPU / Memory / Ping | 장비 자체 상태 |
| Port 상태 | ifName / ifAlias / Speed / operStatus | Link 상태 |
| Traffic | RX / TX bps | 대역폭 사용량 |
| Packet | Unicast / Broadcast / Multicast PPS | Broadcast 폭주 및 Packet 패턴 |
| Error | Error / Discard | 품질 및 혼잡 징후 |
| Syslog | warning / err / crit | 이벤트 원인 확인 |
30.2 Linux Dashboard
권장:
- Uptime
- CPU
- Memory
- Load
- Disk Capacity
- Disk I/O
- Network RX/TX
- Network Error/Drop
- Application/Security Log
30.3 Windows Dashboard
권장:
- Exporter UP
- Uptime
- CPU
- Memory
- Logical Disk
- Disk I/O
- Network
- Windows Event 연동 여부
30.4 QNAP Dashboard
현재 실제 운영 구성을 기준으로 다음과 같이 구성한다.
상단:
- Node Exporter UP
- Uptime
- CPU Usage
- Memory Usage
- HDD 최고 온도
중단:
- CPU / Memory / Load
- bond0 RX / TX
- Disk 상태
- DataVol1 Volume
- DataVol1 Filesystem 사용률
- bond0 Error / Drop
- Physical Disk I/O / IOPS
하단:
- QNAP Event Log
- QNAP Access Log
제거한 항목:
- 상단 RAID 상태
- Storage Pool 상태
- Storage Pool 사용률
- RAID 상세 Table
- Storage Pool 상세 Table
삭제 이유는 Dashboard를 운영자가 빠르게 읽을 수 있도록 단순화하기 위해서이다.
RAID/Storage Pool 상세는 필요 시 별도의 상세 Dashboard에서 확인할 수 있다.
31. 초보자가 자주 만나는 문제
31.1 Telegraf --test는 되는데 서비스는 안 됨
root 권한에서는 Ping이 되지만 telegraf 서비스 계정에서는 Raw Socket 권한이 없을 수 있다.
확인:
journalctl -u telegraf \
-n 100 \
--no-pagernative Ping이면 CAP_NET_RAW 설정을 확인한다.
31.2 snmpget은 되는데 Telegraf에서 일부 장비만 누락
가능한 원인:
- timeout이 너무 짧음
- retries가 0
- OID가 모델과 다름
- DES/AES 조합이 장비와 다름
- SNMP View 제한
먼저 동일 OID를 snmpget으로 확인한다.
31.3 Prometheus가 재시작 반복
가장 먼저:
promtool check config \
/etc/prometheus/prometheus.ymlYAML 들여쓰기를 확인한다.
그 다음:
journalctl -u prometheus \
-n 100 \
--no-pager31.4 Grafana에서 Filesystem이 너무 많이 보임
Linux/QNAP에는 tmpfs, loop, snapshot, container mount 등이 존재할 수 있다.
전체를 보여주기보다 운영자가 실제 사용하는 Mount Point만 필터링한다.
QNAP 예:
mountpoint="/share/CACHEDEV1_DATA"31.5 Loki에 Job이 안 보임
설정에 job이 있다고 바로 나타나는 것이 아니다.
실제 Log가 Loki에 한 번 이상 저장되어야 한다.
확인 순서:
tail -n 10 \
/var/log/qnap/access.log
journalctl -u alloy \
-n 100 \
--no-pager
curl -s \
'http://127.0.0.1:3100/loki/api/v1/label/job/values' \
| jq32. 구축 완료 후 점검 Checklist
| 확인 | 항목 |
|---|---|
| □ | Rocky Linux 시간 동기화 정상 |
| □ | InfluxDB health 정상 |
| □ | Telegraf 서비스 정상 |
| □ | SNMP 장비별 응답 확인 |
| □ | InfluxDB에 SNMP 데이터 저장 |
| □ | Prometheus health 정상 |
| □ | Node Exporter Target UP |
| □ | Windows Exporter Target UP |
| □ | QNAP Node Exporter Target UP |
| □ | rsyslog Network 로그 수신 |
| □ | QNAP Event / Access Log 분리 수신 |
| □ | Alloy Permission 정상 |
| □ | Loki job 확인 |
| □ | Grafana InfluxDB Save & Test |
| □ | Grafana Prometheus Save & Test |
| □ | Grafana Loki Save & Test |
| □ | Switch Traffic 그래프 정상 |
| □ | Error/Discard 그래프 정상 |
| □ | Linux/Windows Dashboard 정상 |
| □ | QNAP DataVol1 용량 QTS와 일치 |
| □ | QNAP Snapshot 제외 |
| □ | QNAP bond0만 Network 표시 |
| □ | Syslog Alert Mail에 장비 IP / Severity 표시 |
| □ | ICMP Down Alert 정상 |
33. 전체 장애 점검 순서
모니터링이 안 될 때 Grafana부터 무작정 수정하지 않는다.
33.1 SNMP
Switch
↓
snmpget
↓
Telegraf --test
↓
Telegraf Service
↓
InfluxDB
↓
Grafana Explore
↓
Dashboard33.2 Prometheus
Exporter /metrics
↓
Prometheus Target
↓
Prometheus Query
↓
Grafana Explore
↓
Dashboard33.3 Syslog
Device 송신
↓
tcpdump
↓
rsyslog
↓
Log File
↓
Alloy
↓
Loki
↓
Grafana Explore
↓
Dashboard / Alert이 순서대로 확인하면 문제 지점을 빠르게 좁힐 수 있다.
34. 운영 원칙 요약
- 처음부터 모든 장비를 추가하지 않는다.
- 한 장비, 한 Port를 먼저 완성한다.
- 설정 변경 후 항상 해당 서비스의 검사 명령을 실행한다.
- Counter와 현재 상태값을 구분한다.
- No Data와 0을 같은 의미로 처리하지 않는다.
- SNMP Password와 Token을 Wiki에 저장하지 않는다.
- Grafana는 수집기가 아니라 시각화 계층임을 기억한다.
- QNAP처럼 제조사별 특이사항은 실제 값과 Vendor 화면을 대조한다.
- Alert는 처음부터 너무 많이 만들지 않는다.
- 정상 상태의 기준 데이터를 먼저 쌓은 후 임계값을 결정한다.
- 장애 시에는 수집 경로를 앞단부터 순서대로 확인한다.
35. 최종 구성 요약
+----------------------+
| Grafana |
+----------+-----------+
|
+------------------+------------------+
| | |
v v v
Prometheus InfluxDB Loki
^ ^ ^
| | |
Exporters Telegraf Alloy
^ ^ ^
| | |
Linux / Windows / QNAP SNMP Device Syslog Files
^
|
rsyslog
^
|
Network / Server / NAS이 구조가 완성되면 다음 세 종류의 정보를 한 Grafana에서 확인할 수 있다.
- 장비와 서버의 현재 상태
- 시간에 따른 성능 변화
- 장애 시점의 이벤트 로그
초보자는 먼저 "데이터가 어디서 생성되어 어디를 거쳐 Grafana까지 도착하는지"를 이해한 뒤 설정값을 수정하는 것이 가장 중요하다.