SNMP테스트서버: 두 판 사이의 차이
편집 요약 없음 |
편집 요약 없음 |
||
| 1번째 줄: | 1번째 줄: | ||
= Rocky Linux 9 | = Rocky Linux 9 통합 모니터링 서버 구축 가이드 = | ||
: Telegraf / InfluxDB / Grafana / Prometheus / Node Exporter / Windows Exporter / Loki / Alloy / rsyslog | |||
작성 기준: 실제 구성 및 장애 처리 기록 기반 | |||
문서 형식: MediaWiki | |||
IP 주소: 문서용 예시 주소로 치환 | |||
운영환경 적용 전 실제 장비 주소, 계정, 토큰, OID를 확인한다. | |||
__TOC__ | |||
== 0. 목적 및 최종 구성 == | |||
본 문서는 Rocky Linux 9 기반 모니터링 서버를 처음 구축하는 단계부터 네트워크 장비 SNMP, 서버 Metric, Syslog, QNAP NAS 모니터링 및 Grafana Dashboard 구성까지 정리한다. | |||
최종 구조는 다음과 같다. | |||
<syntaxhighlight lang="bash" line> | |||
Network Switch | |||
| | |||
+-- SNMPv3 UDP/161 | |||
| | | |||
| +--> Telegraf | |||
| | | |||
| +--> InfluxDB OSS 2.x | |||
| | |||
+-- Syslog TCP/UDP | |||
| | |||
+--> rsyslog | |||
| | |||
+--> Log File | |||
| | |||
+--> Grafana Alloy | |||
| | |||
+--> Loki | |||
Linux / QNAP | |||
| | |||
+-- node_exporter :9100 | |||
| | |||
+--> Prometheus | |||
Windows | |||
| | |||
+-- windows_exporter :9182 | |||
| | |||
+--> Prometheus | |||
Prometheus + InfluxDB + Loki | |||
| | |||
+--> Grafana | |||
</syntaxhighlight> | |||
| | |||
=== 0.1 문서용 IP 주소 === | |||
본 문서에서는 실제 운영 IP를 노출하지 않고 RFC 문서용 주소를 사용한다. | |||
{| class="wikitable" | {| class="wikitable" | ||
! 대상 | |||
! 예시 주소 | |||
! 용도 | |||
|- | |- | ||
| Monitoring Server | |||
| 192.0.2.10 | |||
| Grafana / Prometheus / InfluxDB / Telegraf / Loki / Alloy / rsyslog | |||
|- | |- | ||
| | | Linux Server | ||
| | | 192.0.2.20 | ||
| node_exporter | |||
|- | |- | ||
| | | Windows Server | ||
| | | 192.0.2.30 | ||
| windows_exporter | |||
|- | |- | ||
| | | QNAP NAS | ||
| | | 198.51.100.10 | ||
| node_exporter / SNMP / Syslog | |||
|- | |- | ||
| | | Switch-01 | ||
| | | 203.0.113.11 | ||
| SNMP / Syslog | |||
|- | |- | ||
| | | Switch-02 | ||
| SNMP | | 203.0.113.12 | ||
| SNMP / Syslog | |||
|- | |- | ||
| | | Switch-03 | ||
| 203.0.113.13 | |||
| SNMP / Syslog | |||
| | |||
| | |||
|} | |} | ||
== 1. Rocky Linux 기본 준비 == | |||
=== 1.1 시스템 상태 확인 === | |||
<syntaxhighlight lang="bash" line> | |||
cat /etc/rocky-release | |||
uname -m | |||
ip -br address | |||
ip route | |||
df -hT | |||
free -h | |||
getenforce | |||
ss -lntup | |||
</syntaxhighlight> | |||
SELinux와 firewalld는 초기부터 비활성화하지 않는다. | |||
=== 1.2 백업 디렉터리 생성 === | |||
기존 서버에 추가 설치하는 경우 설정 파일 백업을 먼저 수행한다. | |||
<syntaxhighlight lang="bash" line> | |||
umask 077 | |||
MON_BACKUP="/root/monitor-backup-$(date +%Y%m%d-%H%M%S)" | |||
install -d -m 700 "$MON_BACKUP" | |||
cp -a /etc/rsyslog.conf "$MON_BACKUP/" 2>/dev/null || true | |||
cp -a /etc/rsyslog.d "$MON_BACKUP/" 2>/dev/null || true | |||
cp -a /etc/chrony.conf "$MON_BACKUP/" 2>/dev/null || true | |||
cp -a /etc/firewalld "$MON_BACKUP/" 2>/dev/null || true | |||
rpm -qa | sort > "$MON_BACKUP/packages-before.txt" | rpm -qa | sort > "$MON_BACKUP/packages-before.txt" | ||
</syntaxhighlight> | </syntaxhighlight> | ||
== 3 | === 1.3 기본 패키지 설치 === | ||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
dnf install -y \ | |||
curl \ | |||
ca-certificates \ | |||
gnupg2 \ | |||
vim-enhanced \ | |||
net-snmp-utils \ | |||
rsyslog \ | |||
logrotate \ | |||
chrony \ | |||
tcpdump \ | |||
iputils \ | |||
sysstat \ | |||
policycoreutils-python-utils \ | |||
dnf-plugins-core \ | |||
jq | |||
</syntaxhighlight> | |||
=== 1.4 시간 동기화 === | |||
모니터링에서는 장비와 서버 간 시간이 맞아야 장애 시각을 비교할 수 있다. | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
timedatectl set-timezone Asia/Seoul | timedatectl set-timezone Asia/Seoul | ||
systemctl enable --now chronyd | systemctl enable --now chronyd | ||
systemctl restart chronyd | systemctl restart chronyd | ||
chronyc tracking | chronyc tracking | ||
chronyc sources -v | chronyc sources -v | ||
date -Ins | date -Ins | ||
</syntaxhighlight> | </syntaxhighlight> | ||
== | == 2. InfluxDB / Telegraf / Grafana 설치 == | ||
=== 2.1 InfluxData Repository === | |||
=== | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
curl -fL https://repos.influxdata.com/influxdata-archive.key \ | curl -fL https://repos.influxdata.com/influxdata-archive.key \ | ||
-o /etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata | -o /etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata | ||
gpg --show-keys --with-fingerprint /etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata | |||
gpg --show-keys --with-fingerprint \ | |||
/etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata | |||
rpm --import /etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata | |||
</syntaxhighlight> | </syntaxhighlight> | ||
Repository: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
vi /etc/yum.repos.d/influxdata.repo | vi /etc/yum.repos.d/influxdata.repo | ||
</syntaxhighlight> | </syntaxhighlight> | ||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
| 181번째 줄: | 189번째 줄: | ||
sslverify=1 | sslverify=1 | ||
</syntaxhighlight> | </syntaxhighlight> | ||
=== | |||
=== 2.2 Grafana Repository === | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
vi /etc/yum.repos.d/grafana.repo | vi /etc/yum.repos.d/grafana.repo | ||
</syntaxhighlight> | </syntaxhighlight> | ||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
[grafana] | [grafana] | ||
| 196번째 줄: | 206번째 줄: | ||
sslverify=1 | sslverify=1 | ||
</syntaxhighlight> | </syntaxhighlight> | ||
=== 2.3 패키지 설치 === | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
dnf makecache | dnf makecache | ||
dnf list --showduplicates telegraf influxdb2 influxdb2-cli grafana | |||
dnf install telegraf influxdb2 influxdb2-cli grafana | dnf list --showduplicates \ | ||
telegraf \ | |||
influxdb2 \ | |||
influxdb2-cli \ | |||
grafana | |||
dnf install -y \ | |||
telegraf \ | |||
influxdb2 \ | |||
influxdb2-cli \ | |||
grafana | |||
</syntaxhighlight> | |||
버전 확인: | |||
<syntaxhighlight lang="bash" line> | |||
rpm -q telegraf influxdb2 influxdb2-cli grafana rsyslog | rpm -q telegraf influxdb2 influxdb2-cli grafana rsyslog | ||
telegraf --version | telegraf --version | ||
influxd version | influxd version | ||
influx version | influx version | ||
</syntaxhighlight> | </syntaxhighlight> | ||
실제 구축 과정에서는 Telegraf 1.40.0 환경에서 동작을 확인하였다. | |||
== | == 3. InfluxDB 초기 구성 == | ||
=== | === 3.1 Localhost Bind === | ||
InfluxDB API는 Grafana 및 Telegraf가 같은 서버에 있으므로 localhost만 수신하도록 구성한다. | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
install -d -m 755 /etc/systemd/system/influxdb.service.d | install -d -m 755 /etc/systemd/system/influxdb.service.d | ||
vi /etc/systemd/system/influxdb.service.d/10 | |||
vi /etc/systemd/system/influxdb.service.d/10-listen.conf | |||
</syntaxhighlight> | </syntaxhighlight> | ||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
[Service] | [Service] | ||
Environment="INFLUXD_HTTP_BIND_ADDRESS=127.0.0.1:8086" | Environment="INFLUXD_HTTP_BIND_ADDRESS=127.0.0.1:8086" | ||
</syntaxhighlight> | </syntaxhighlight> | ||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
systemctl daemon-reload | systemctl daemon-reload | ||
systemctl enable --now influxdb | systemctl enable --now influxdb | ||
systemctl restart influxdb | systemctl restart influxdb | ||
systemctl status influxdb --no-pager | systemctl status influxdb --no-pager | ||
ss -lntp | grep ':8086' | ss -lntp | grep ':8086' | ||
curl -fsS http://127.0.0.1:8086/health | curl -fsS http://127.0.0.1:8086/health | ||
</syntaxhighlight> | </syntaxhighlight> | ||
=== | === 3.2 Initial Setup === | ||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
influx setup | influx setup | ||
</syntaxhighlight> | </syntaxhighlight> | ||
예시: | |||
{| class="wikitable" | {| class="wikitable" | ||
! 설정 | |||
! | ! 값 | ||
! | |||
|- | |- | ||
| Organization | | Organization | ||
| | | network | ||
|- | |- | ||
| Bucket | | Bucket | ||
| | | snmp_raw | ||
|- | |- | ||
| Retention | | Retention | ||
| | | 720h | ||
|} | |} | ||
Bucket 확인: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
influx bucket list --org network | influx bucket list --org network | ||
</syntaxhighlight> | </syntaxhighlight> | ||
=== 3.3 Token 분리 === | |||
서비스별 Token 권한을 분리한다. | |||
* telegraf-write : snmp_raw Write | |||
* grafana-read : snmp_raw Read | |||
* Operator Token : 관리자용 | |||
운영 문서에 실제 Token 값을 기록하지 않는다. | |||
== | == 4. Telegraf 기본 구성 == | ||
=== | === 4.1 환경 변수 파일 === | ||
<syntaxhighlight lang="bash" line> | |||
vi /etc/telegraf/monitor.env | |||
</syntaxhighlight> | |||
<syntaxhighlight lang="bash" line> | |||
INFLUX_WRITE_TOKEN='<WRITE_TOKEN>' | |||
=== | ZYXEL_SNMP_USER='<SNMP_USER>' | ||
ZYXEL_SNMP_AUTH='<AUTH_PASSWORD>' | |||
ZYXEL_SNMP_PRIV='<PRIV_PASSWORD>' | |||
QNAP_SNMP_USER='<QNAP_SNMP_USER>' | |||
QNAP_SNMP_AUTH='<QNAP_AUTH_PASSWORD>' | |||
QNAP_SNMP_PRIV='<QNAP_PRIV_PASSWORD>' | |||
< | |||
</syntaxhighlight> | </syntaxhighlight> | ||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
chown root:root /etc/telegraf/monitor.env | |||
chmod 600 /etc/telegraf/monitor.env | |||
install -d -m 755 /etc/systemd/system/telegraf.service.d | |||
vi /etc/systemd/system/telegraf.service.d/10-monitor-env.conf | |||
</syntaxhighlight> | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
[Service] | |||
EnvironmentFile=/etc/telegraf/monitor.env | |||
</syntaxhighlight> | </syntaxhighlight> | ||
== | === 4.2 Main Configuration === | ||
= | <syntaxhighlight lang="bash" line> | ||
cp -a /etc/telegraf/telegraf.conf \ | |||
/etc/telegraf/telegraf.conf.orig | |||
vi /etc/telegraf/telegraf.conf | vi /etc/telegraf/telegraf.conf | ||
</syntaxhighlight> | </syntaxhighlight> | ||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
[agent] | [agent] | ||
| 377번째 줄: | 365번째 줄: | ||
[[inputs.internal]] | [[inputs.internal]] | ||
[[inputs.cpu]] | [[inputs.cpu]] | ||
percpu = false | percpu = false | ||
totalcpu = true | totalcpu = true | ||
[[inputs.mem]] | [[inputs.mem]] | ||
[[inputs.disk]] | [[inputs.disk]] | ||
mount_points = ["/"] | mount_points = ["/"] | ||
[[inputs.net]] | [[inputs.net]] | ||
</syntaxhighlight> | </syntaxhighlight> | ||
== 5. SNMPv3 사전 확인 == | |||
SNMP는 v3 authPriv 사용을 기본으로 한다. | |||
Net-SNMP 테스트 파일: | |||
<syntaxhighlight lang="bash" line> | |||
install -d -m 700 /root/.snmp | |||
vi /root/.snmp/snmp.conf | |||
</syntaxhighlight> | |||
<syntaxhighlight lang="bash" line> | |||
defVersion 3 | |||
defSecurityName <SNMP_USER> | |||
defSecurityLevel authPriv | |||
defAuthType SHA | |||
defAuthPassphrase <SNMP_AUTH_PASSWORD> | |||
defPrivType AES | |||
defPrivPassphrase <SNMP_PRIV_PASSWORD> | |||
</syntaxhighlight> | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
chmod 600 /root/.snmp/snmp.conf | |||
snmpget -v3 -t 2 -r 0 -On \ | |||
203.0.113.11 \ | |||
.1.3.6.1.2.1.1.3.0 | |||
snmpwalk -v3 -t 2 -r 0 -On \ | |||
203.0.113.11 \ | |||
.1.3.6.1.2.1.31.1.1.1.1 | |||
</syntaxhighlight> | </syntaxhighlight> | ||
ifName OID 마지막 값이 ifIndex이다. | |||
== 6. Telegraf SNMP 구성 == | |||
설정 파일은 장비 모델별로 분리한다. | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
/etc/telegraf/telegraf.d/ | |||
├── 10-zyxel-gs1900.conf | |||
├── 11-zyxel-gs1920.conf | |||
├── 20-zyxel-es3128.conf | |||
├── 30-icmp_check.conf | |||
├── 40-qnap_nas.conf | |||
└── sflow.conf | |||
</syntaxhighlight> | </syntaxhighlight> | ||
=== 6.1 GS1920 계열 === | |||
예시 대상: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
203.0.113.21 | |||
203.0.113.22 | |||
203.0.113.23 | |||
</syntaxhighlight> | </syntaxhighlight> | ||
SNMPv3: | |||
* SHA | |||
* DES | |||
CPU OID: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
.1.3.6.1.4.1.890.1.15.3.49.1.7.0 | |||
</syntaxhighlight> | </syntaxhighlight> | ||
Memory: | |||
< | <syntaxhighlight lang="bash" line> | ||
Total .1.3.6.1.4.1.890.1.15.3.50.1.1.1.3.1 | |||
Used .1.3.6.1.4.1.890.1.15.3.50.1.1.1.4.1 | |||
Percent .1.3.6.1.4.1.890.1.15.3.50.1.1.1.5.1 | |||
</syntaxhighlight> | |||
=== 6.2 GS1900 계열 === | |||
CPU: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
.1.3.6.1.4.1.890.1.15.3.2.4.0 | |||
</syntaxhighlight> | |||
Memory: | |||
<syntaxhighlight lang="bash" line> | |||
.1.3.6.1.4.1.890.1.15.3.2.5.0 | |||
</syntaxhighlight> | |||
실제 구축에서는 2초 timeout / retries 0에서 응답 누락이 발생하여 다음과 같이 완화하였다. | |||
<syntaxhighlight lang="bash" line> | |||
timeout = "5s" | |||
retries = 1 | |||
</syntaxhighlight> | </syntaxhighlight> | ||
=== | === 6.3 ES-3128GP === | ||
SNMPv3: | |||
* SHA | |||
* AES | |||
sysObjectID: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
.1.3.6.1.4.1.7800.1.190 | |||
</syntaxhighlight> | |||
CPU / Memory private OID는 장비에서 정상 확인되지 않아 Dashboard에서는 N/A 처리한다. | |||
=== 6.4 IF-MIB 주요 항목 === | |||
수집 대상: | |||
* ifName | |||
* ifAlias | |||
* ifSpeed | |||
* ifHCInOctets | |||
* ifHCOutOctets | |||
* ifOperStatus | |||
* ifInErrors | |||
* ifOutErrors | |||
* ifInDiscards | |||
* ifOutDiscards | |||
* Broadcast | |||
* Multicast | |||
누적 Counter는 Grafana에서 그대로 표시하지 않고 derivative/rate 계산 후 표시한다. | |||
== 7. ICMP 수집 == | |||
Telegraf ping input을 사용한다. | |||
<syntaxhighlight lang="bash" line> | |||
vi /etc/telegraf/telegraf.d/30-icmp_check.conf | |||
</syntaxhighlight> | </syntaxhighlight> | ||
예시 대상: | |||
= | <syntaxhighlight lang="bash" line> | ||
203.0.113.11 | |||
203.0.113.12 | |||
203.0.113.13 | |||
</syntaxhighlight> | |||
권장 설정: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
[[inputs.ping]] | [[inputs.ping]] | ||
urls = [" | urls = [ | ||
"203.0.113.11", | |||
"203.0.113.12", | |||
"203.0.113.13" | |||
] | |||
method = "native" | method = "native" | ||
count = 3 | |||
count = | deadline = 2.0 | ||
deadline = | interval = 10.0 | ||
interval = | |||
</syntaxhighlight> | </syntaxhighlight> | ||
< | native ping은 CAP_NET_RAW가 필요하다. | ||
<syntaxhighlight lang="bash" line> | |||
systemctl edit telegraf | |||
</syntaxhighlight> | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
[Service] | |||
CapabilityBoundingSet=CAP_NET_RAW | CapabilityBoundingSet=CAP_NET_RAW | ||
AmbientCapabilities=CAP_NET_RAW | AmbientCapabilities=CAP_NET_RAW | ||
</syntaxhighlight> | </syntaxhighlight> | ||
=== | <syntaxhighlight lang="bash" line> | ||
systemctl daemon-reload | |||
systemctl restart telegraf | |||
</syntaxhighlight> | |||
== 8. Telegraf Test 및 서비스 시작 == | |||
설정 검사: | |||
<syntaxhighlight lang="bash" line> | |||
telegraf \ | |||
--config /etc/telegraf/telegraf.conf \ | |||
--config-directory /etc/telegraf/telegraf.d \ | |||
--test | |||
</syntaxhighlight> | |||
실제 서비스 계정으로 확인: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
systemctl daemon-reload | systemctl daemon-reload | ||
systemd-run --unit= | |||
systemd-run \ | |||
--unit=telegraf-config-check \ | |||
--wait \ | |||
--pipe \ | |||
--collect \ | |||
-p User=telegraf \ | |||
-p Group=telegraf \ | |||
-p EnvironmentFile=/etc/telegraf/monitor.env \ | |||
-p CapabilityBoundingSet=CAP_NET_RAW \ | |||
-p AmbientCapabilities=CAP_NET_RAW \ | |||
/usr/bin/telegraf \ | |||
--config /etc/telegraf/telegraf.conf \ | |||
--config-directory /etc/telegraf/telegraf.d \ | |||
--test | |||
</syntaxhighlight> | </syntaxhighlight> | ||
서비스 시작: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
systemctl enable --now telegraf | systemctl enable --now telegraf | ||
systemctl restart telegraf | systemctl restart telegraf | ||
systemctl status telegraf --no-pager | systemctl status telegraf --no-pager | ||
journalctl -u telegraf -n 100 --no-pager | journalctl -u telegraf -n 100 --no-pager | ||
</syntaxhighlight> | </syntaxhighlight> | ||
== | == 9. Grafana 설치 및 InfluxDB 연결 == | ||
=== | === 9.1 Grafana 서비스 === | ||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
vi /etc/grafana/grafana.ini | vi /etc/grafana/grafana.ini | ||
</syntaxhighlight> | </syntaxhighlight> | ||
내부망 운영 환경에 맞게 listen 주소를 설정한다. | |||
예시: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
[server] | [server] | ||
http_addr = | http_addr = 0.0.0.0 | ||
http_port = 3000 | http_port = 3000 | ||
| 629번째 줄: | 626번째 줄: | ||
enabled = false | enabled = false | ||
</syntaxhighlight> | </syntaxhighlight> | ||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
systemctl enable --now grafana-server | systemctl enable --now grafana-server | ||
systemctl restart grafana-server | systemctl restart grafana-server | ||
systemctl status grafana-server --no-pager | systemctl status grafana-server --no-pager | ||
curl -fsS http://127.0.0.1:3000/api/health | curl -fsS http://127.0.0.1:3000/api/health | ||
</syntaxhighlight> | </syntaxhighlight> | ||
=== | === 9.2 InfluxDB Datasource === | ||
Connections → Data sources → Add data source → InfluxDB | Grafana: | ||
<syntaxhighlight lang="bash" line> | |||
Connections | |||
→ Data sources | |||
→ Add data source | |||
→ InfluxDB | |||
</syntaxhighlight> | |||
{| class="wikitable" | {| class="wikitable" | ||
! 설정 | ! 설정 | ||
! 값 | ! 값 | ||
|- | |- | ||
| Query language | | Query language | ||
| 654번째 줄: | 655번째 줄: | ||
|- | |- | ||
| URL | | URL | ||
| | | http://127.0.0.1:8086 | ||
|- | |- | ||
| Organization | | Organization | ||
| | | network | ||
|- | |- | ||
| Token | | Token | ||
| grafana-read | | grafana-read | ||
|- | |- | ||
| Default Bucket | | Default Bucket | ||
| | | snmp_raw | ||
|- | |- | ||
| Min time interval | | Min time interval | ||
| | | 1s | ||
|} | |} | ||
== 10. SNMP Dashboard 기본 구성 == | |||
=== | === 10.1 Traffic === | ||
Flux: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
from(bucket: "snmp_raw") | from(bucket: "snmp_raw") | ||
|> range(start: | |> range(start: v.timeRangeStart, stop: v.timeRangeStop) | ||
|> filter(fn: (r) => r._measurement == " | |> filter(fn: (r) => r._measurement == "switch_port") | ||
|> | |> filter(fn: (r) => r.source == "203.0.113.11") | ||
|> filter(fn: (r) => | |||
r._field == "in_octets" or | |||
r._field == "out_octets" | |||
) | |||
|> derivative(unit: 1s, nonNegative: true) | |||
|> map(fn: (r) => ({ | |||
r with _value: r._value * 8.0 | |||
})) | |||
</syntaxhighlight> | </syntaxhighlight> | ||
= | Grafana Unit: | ||
<syntaxhighlight lang="bash" line> | |||
bits/sec | |||
</syntaxhighlight> | |||
=== 10.2 Error / Discard === | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
|> filter(fn: (r) => | |||
r._field == "in_errors" or | |||
|> filter(fn: (r) => r. | r._field == "out_errors" or | ||
r._field == "in_discards" or | |||
r._field == "out_discards" | |||
) | |||
|> derivative(unit: 1s, nonNegative: true) | |> derivative(unit: 1s, nonNegative: true) | ||
</syntaxhighlight> | </syntaxhighlight> | ||
=== 10.3 Broadcast / Multicast === | |||
Broadcast/Multicast는 별도 패널로 표시하여 폭주 및 비정상 증가를 확인한다. | |||
== 11. Prometheus 설치 == | |||
실제 구축에서는 Prometheus 3.13.3을 사용하였다. | |||
=== 11.1 사용자 및 디렉터리 === | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
useradd \ | |||
--system \ | |||
--no-create-home \ | |||
--shell /sbin/nologin \ | |||
prometheus | |||
install -d -o prometheus -g prometheus \ | |||
/etc/prometheus \ | |||
/var/lib/prometheus | |||
</syntaxhighlight> | </syntaxhighlight> | ||
Prometheus 바이너리는 공식 릴리스에서 다운로드하고 checksum 확인 후 설치한다. | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
install -m 0755 prometheus /usr/local/bin/prometheus | |||
install -m 0755 promtool /usr/local/bin/promtool | |||
</syntaxhighlight> | </syntaxhighlight> | ||
=== | === 11.2 Prometheus 설정 === | ||
<syntaxhighlight lang="bash" line> | |||
vi /etc/prometheus/prometheus.yml | |||
</syntaxhighlight> | |||
<syntaxhighlight lang="bash" line> | |||
global: | |||
scrape_interval: 30s | |||
scrape_configs: | |||
- job_name: prometheus | |||
static_configs: | |||
- targets: | |||
- "127.0.0.1:9090" | |||
labels: | |||
server_name: monitoring-server | |||
- job_name: node | |||
static_configs: | |||
- targets: | |||
- "127.0.0.1:9100" | |||
labels: | |||
server_name: monitoring-server | |||
- targets: | |||
- "192.0.2.20:9100" | |||
labels: | |||
server_name: linux-server | |||
- job_name: windows | |||
static_configs: | |||
- targets: | |||
- "192.0.2.30:9182" | |||
labels: | |||
server_name: windows-server | |||
- job_name: qnap | |||
static_configs: | |||
- targets: | |||
- "198.51.100.10:9100" | |||
labels: | |||
server_name: nas-01 | |||
</syntaxhighlight> | </syntaxhighlight> | ||
=== 11.3 systemd === | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
vi /etc/systemd/system/prometheus.service | |||
vi /etc/ | |||
</syntaxhighlight> | </syntaxhighlight> | ||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
[Unit] | |||
Description=Prometheus | |||
Wants=network-online.target | |||
After=network-online.target | |||
[Service] | |||
User=prometheus | |||
Group=prometheus | |||
ExecStart=/usr/local/bin/prometheus \ | |||
--config.file=/etc/prometheus/prometheus.yml \ | |||
--storage.tsdb.path=/var/lib/prometheus \ | |||
--storage.tsdb.retention.time=30d \ | |||
--storage.tsdb.retention.size=20GB \ | |||
--web.listen-address=127.0.0.1:9090 | |||
Restart=always | |||
[Install] | |||
WantedBy=multi-user.target | |||
</syntaxhighlight> | </syntaxhighlight> | ||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
chown -R prometheus:prometheus \ | |||
systemctl enable --now | /etc/prometheus \ | ||
systemctl | /var/lib/prometheus | ||
promtool check config /etc/prometheus/prometheus.yml | |||
systemctl daemon-reload | |||
systemctl enable --now prometheus | |||
systemctl status prometheus --no-pager | |||
</syntaxhighlight> | </syntaxhighlight> | ||
== 12. Node Exporter 설치 == | |||
실제 구축에서는 Node Exporter 1.12.1 환경을 사용하였다. | |||
=== 12.1 Linux Node Exporter === | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
useradd \ | |||
--system \ | |||
--no-create-home \ | |||
--shell /sbin/nologin \ | |||
node_exporter | |||
install -m 0755 node_exporter \ | |||
/usr/local/bin/node_exporter | |||
</syntaxhighlight> | </syntaxhighlight> | ||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
/ | vi /etc/systemd/system/node_exporter.service | ||
</syntaxhighlight> | </syntaxhighlight> | ||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
[Unit] | |||
Description=Node Exporter | |||
After=network-online.target | |||
Wants=network-online.target | |||
[Service] | |||
User=node_exporter | |||
Group=node_exporter | |||
ExecStart=/usr/local/bin/node_exporter | |||
Restart=always | |||
[Install] | |||
WantedBy=multi-user.target | |||
</syntaxhighlight> | </syntaxhighlight> | ||
= | <syntaxhighlight lang="bash" line> | ||
systemctl daemon-reload | |||
systemctl enable --now node_exporter | |||
curl -s http://127.0.0.1:9100/metrics | head | |||
</syntaxhighlight> | |||
외부 서버의 node_exporter는 Prometheus 서버 주소만 9100/tcp에 접근하도록 방화벽을 제한한다. | |||
== 13. Windows Exporter == | |||
Windows Server에서는 windows_exporter를 사용한다. | |||
기본 포트: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
9182/tcp | |||
</syntaxhighlight> | </syntaxhighlight> | ||
Prometheus Target: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
192.0.2.30:9182 | |||
</syntaxhighlight> | </syntaxhighlight> | ||
성능 카운터 이상 시 다음 명령으로 복구한 사례가 있다. | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
lodctr /R | |||
winmgmt /resyncperf | |||
</syntaxhighlight> | </syntaxhighlight> | ||
PowerShell: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
Restart-Service windows_exporter | |||
</syntaxhighlight> | </syntaxhighlight> | ||
확인: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
Invoke-WebRequest http://127.0.0.1:9182/metrics | |||
</syntaxhighlight> | </syntaxhighlight> | ||
== 14. Grafana Prometheus Datasource == | |||
Grafana: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
Connections | |||
→ Data sources | |||
→ Prometheus | |||
</syntaxhighlight> | </syntaxhighlight> | ||
URL: | |||
<syntaxhighlight lang="bash" line> | |||
http://127.0.0.1:9090 | |||
</syntaxhighlight> | |||
Save & Test 후 다음 PromQL로 확인한다. | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
up | |||
</syntaxhighlight> | </syntaxhighlight> | ||
=== | == 15. Server Dashboard 구성 == | ||
Linux Dashboard: | |||
* Uptime | |||
* CPU Usage | |||
* Memory Usage | |||
* Load Average | |||
* Filesystem | |||
* Disk Read / Write | |||
* Network RX / TX | |||
* Network Error / Drop | |||
* Service 상태 | |||
Windows Dashboard: | |||
* CPU | |||
* Memory | |||
* Disk | |||
* Network | |||
* Uptime | |||
* Exporter 상태 | |||
Disk 용량은 실제 GB/GiB 값으로 표시하며 불필요한 Bar Gauge는 제거한다. | |||
== 16. QNAP TS-264 모니터링 == | |||
QNAP은 세 종류의 데이터를 함께 사용한다. | |||
{| class="wikitable" | {| class="wikitable" | ||
! 데이터 | |||
! 수집 방법 | |||
|- | |- | ||
| CPU / Memory / Network / Filesystem / Disk I/O | |||
| node_exporter → Prometheus | |||
|- | |- | ||
| | | HDD / RAID / Storage / Volume | ||
| | | SNMP → Telegraf → InfluxDB | ||
|- | |- | ||
| | | Event / Access Log | ||
| Syslog → rsyslog → Alloy → Loki | |||
| | |||
|} | |} | ||
=== 16.1 QNAP SNMP === | |||
QNAP SNMPv3: | |||
* SHA | |||
* DES | |||
Measurement: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
QNAP_TS264 | |||
qnap_disk | |||
qnap_raid | |||
qnap_storage_pool | |||
qnap_volume | |||
</syntaxhighlight> | </syntaxhighlight> | ||
Disk: | |||
* disk_id | |||
* manufacturer | |||
* model | |||
* disk_type | |||
* disk_status | |||
* temperature | |||
* capacity_bytes | |||
RAID: | |||
* RAID ID | |||
* RAID Name | |||
* RAID Status | |||
* RAID Level | |||
* Capacity | |||
== | === 16.2 Volume 단위 보정 === | ||
QNAP qnap_volume의 capacity/free 값은 실제 장비에서 KiB 형태로 반환되는 것으로 확인되었다. | |||
원시값을 Grafana에서 byte로 바로 해석하면 약 11.3 GiB로 잘못 표시된다. | |||
Flux: | |||
<syntaxhighlight lang="bash" line> | |||
|> map(fn: (r) => ({ | |||
r with | |||
capacity_bytes: | |||
uint(v: r.capacity_bytes) * uint(v: 1024), | |||
free_bytes: | |||
uint(v: r.free_bytes) * uint(v: 1024), | |||
used_bytes: | |||
( | |||
uint(v: r.capacity_bytes) | |||
- uint(v: r.free_bytes) | |||
) * uint(v: 1024), | |||
used_percent: | |||
if float(v: r.capacity_bytes) > 0.0 then | |||
( | |||
float(v: r.capacity_bytes) | |||
- float(v: r.free_bytes) | |||
) | |||
/ float(v: r.capacity_bytes) | |||
* 100.0 | |||
else 0.0 | |||
})) | |||
</syntaxhighlight> | |||
=== 16.3 실제 Data Volume === | |||
Node Exporter 기준 실제 사용자 Data Volume: | |||
<syntaxhighlight lang="bash" line> | |||
mountpoint="/share/CACHEDEV1_DATA" | |||
device="/dev/mapper/cachedev1" | |||
fstype="ext4" | |||
</syntaxhighlight> | |||
Snapshot: | |||
<syntaxhighlight lang="bash" line> | |||
/mnt/snapshot/... | |||
</syntaxhighlight> | |||
위 Snapshot 경로는 Dashboard Filesystem 패널에서 제외한다. | |||
DataVol1 사용률: | |||
<syntaxhighlight lang="bash" line> | |||
100 * ( | |||
1 - | |||
node_filesystem_avail_bytes{ | |||
instance="198.51.100.10:9100", | |||
mountpoint="/share/CACHEDEV1_DATA" | |||
} | |||
/ | |||
node_filesystem_size_bytes{ | |||
instance="198.51.100.10:9100", | |||
mountpoint="/share/CACHEDEV1_DATA" | |||
} | |||
) | |||
</syntaxhighlight> | |||
=== 16.4 Network는 bond0만 표시 === | |||
RX: | |||
<syntaxhighlight lang="bash" line> | |||
rate( | |||
node_network_receive_bytes_total{ | |||
instance="198.51.100.10:9100", | |||
device="bond0" | |||
}[$__rate_interval] | |||
) * 8 | |||
</syntaxhighlight> | |||
TX: | |||
<syntaxhighlight lang="bash" line> | |||
rate( | |||
node_network_transmit_bytes_total{ | |||
instance="198.51.100.10:9100", | |||
device="bond0" | |||
}[$__rate_interval] | |||
) * 8 | |||
</syntaxhighlight> | |||
=== 16.5 Network Error / Drop === | |||
<syntaxhighlight lang="bash" line> | |||
rate(node_network_receive_errs_total{ | |||
instance="198.51.100.10:9100", | |||
device="bond0" | |||
}[$__rate_interval]) | |||
</syntaxhighlight> | |||
<syntaxhighlight lang="bash" line> | |||
rate(node_network_transmit_errs_total{ | |||
instance="198.51.100.10:9100", | |||
device="bond0" | |||
}[$__rate_interval]) | |||
</syntaxhighlight> | |||
<syntaxhighlight lang="bash" line> | |||
rate(node_network_receive_drop_total{ | |||
instance="198.51.100.10:9100", | |||
device="bond0" | |||
}[$__rate_interval]) | |||
</syntaxhighlight> | |||
<syntaxhighlight lang="bash" line> | |||
rate(node_network_transmit_drop_total{ | |||
instance="198.51.100.10:9100", | |||
device="bond0" | |||
}[$__rate_interval]) | |||
</syntaxhighlight> | |||
=== 16.6 Disk I/O === | |||
Read: | |||
<syntaxhighlight lang="bash" line> | |||
rate(node_disk_read_bytes_total{ | |||
instance="198.51.100.10:9100", | |||
device=~"sd[a-z]+" | |||
}[$__rate_interval]) | |||
</syntaxhighlight> | |||
Write: | |||
<syntaxhighlight lang="bash" line> | |||
rate(node_disk_written_bytes_total{ | |||
instance="198.51.100.10:9100", | |||
device=~"sd[a-z]+" | |||
}[$__rate_interval]) | |||
</syntaxhighlight> | |||
Read IOPS: | |||
<syntaxhighlight lang="bash" line> | |||
rate(node_disk_reads_completed_total{ | |||
instance="198.51.100.10:9100", | |||
device=~"sd[a-z]+" | |||
}[$__rate_interval]) | |||
</syntaxhighlight> | |||
Write IOPS: | |||
<syntaxhighlight lang="bash" line> | |||
rate(node_disk_writes_completed_total{ | |||
instance="198.51.100.10:9100", | |||
device=~"sd[a-z]+" | |||
}[$__rate_interval]) | |||
</syntaxhighlight> | |||
== 17. rsyslog 구성 == | |||
=== 17.1 Network Syslog === | |||
예시 로그 파일: | |||
<syntaxhighlight lang="bash" line> | |||
/var/log/network-syslog/events.log | |||
</syntaxhighlight> | |||
네트워크 장비: | |||
<syntaxhighlight lang="bash" line> | |||
UDP/TCP 514 | |||
</syntaxhighlight> | |||
=== 17.2 Server Syslog === | |||
<syntaxhighlight lang="bash" line> | |||
/var/log/server-syslog/events.log | |||
</syntaxhighlight> | |||
예시 포트: | |||
<syntaxhighlight lang="bash" line> | |||
5514/tcp | |||
</syntaxhighlight> | |||
=== 17.3 QNAP Event / Access 분리 === | |||
QNAP 로그는 Event와 Access를 별도 포트로 분리한다. | |||
{| class="wikitable" | {| class="wikitable" | ||
! 로그 | |||
! Port | |||
! File | |||
|- | |- | ||
| Event | |||
| TCP 5515 | |||
| /var/log/qnap/event.log | |||
|- | |- | ||
| | | Access | ||
| | | TCP 5516 | ||
| | | /var/log/qnap/access.log | ||
|} | |} | ||
Template: | |||
<syntaxhighlight lang="bash" line> | |||
template(name="QnapSyslogLine" type="string" | |||
string="%timegenerated:::date-rfc3339% src=%fromhost-ip% severity=%syslogseverity-text% host=%hostname% %syslogtag%%msg:::sp-if-no-1st-sp%%msg%\n") | |||
</syntaxhighlight> | |||
Event: | |||
<syntaxhighlight lang="bash" line> | |||
ruleset(name="QnapEventLog") { | |||
action( | |||
type="omfile" | |||
file="/var/log/qnap/event.log" | |||
template="QnapSyslogLine" | |||
fileOwner="root" | |||
fileGroup="alloy" | |||
fileCreateMode="0640" | |||
dirOwner="root" | |||
dirGroup="alloy" | |||
dirCreateMode="0750" | |||
createDirs="on" | |||
) | |||
stop | |||
} | |||
input( | |||
type="imtcp" | |||
port="5515" | |||
ruleset="QnapEventLog" | |||
) | |||
</syntaxhighlight> | |||
Access: | |||
<syntaxhighlight lang="bash" line> | |||
ruleset(name="QnapAccessLog") { | |||
action( | |||
type="omfile" | |||
file="/var/log/qnap/access.log" | |||
template="QnapSyslogLine" | |||
fileOwner="root" | |||
fileGroup="alloy" | |||
fileCreateMode="0640" | |||
dirOwner="root" | |||
dirGroup="alloy" | |||
dirCreateMode="0750" | |||
createDirs="on" | |||
) | |||
stop | |||
} | |||
input( | |||
type="imtcp" | |||
port="5516" | |||
ruleset="QnapAccessLog" | |||
) | |||
</syntaxhighlight> | |||
검증: | |||
<syntaxhighlight lang="bash" line> | |||
rsyslogd -N1 | |||
systemctl restart rsyslog | |||
ss -lntp | grep -E ':5514|:5515|:5516' | |||
</syntaxhighlight> | |||
== 18. Loki / Alloy 구성 == | |||
Loki Endpoint: | |||
<syntaxhighlight lang="bash" line> | |||
http://127.0.0.1:3100 | |||
</syntaxhighlight> | |||
Alloy UI: | |||
<syntaxhighlight lang="bash" line> | |||
127.0.0.1:12345 | |||
</syntaxhighlight> | |||
=== 18.1 Alloy 기본 설정 === | |||
<syntaxhighlight lang="bash" line> | |||
logging { | |||
level = "info" | |||
} | |||
</syntaxhighlight> | |||
=== 18.2 Network Syslog === | |||
<syntaxhighlight lang="bash" line> | |||
loki.source.file "network_syslog" { | |||
targets = [ | |||
{ | |||
__path__ = "/var/log/network-syslog/events.log", | |||
job = "network-syslog", | |||
}, | |||
] | |||
forward_to = [loki.process.network_syslog.receiver] | |||
} | |||
loki.process "network_syslog" { | |||
stage.regex { | |||
expression = `^(?P<received_at>\S+) src=(?P<device_ip>\S+) severity=(?P<severity>\S+) host=(?P<device_host>\S+) (?P<message>.*)$` | |||
} | |||
stage.timestamp { | |||
source = "received_at" | |||
format = "RFC3339Nano" | |||
action_on_failure = "skip" | |||
} | |||
stage.labels { | |||
values = { | |||
device_ip = "", | |||
severity = "", | |||
} | |||
} | |||
forward_to = [loki.write.local.receiver] | |||
} | |||
</syntaxhighlight> | |||
=== 18.3 Server Syslog === | |||
<syntaxhighlight lang="bash" line> | |||
loki.source.file "server_syslog" { | |||
targets = [ | |||
{ | |||
__path__ = "/var/log/server-syslog/events.log", | |||
job = "server-syslog", | |||
}, | |||
] | |||
forward_to = [loki.process.server_syslog.receiver] | |||
} | |||
loki.process "server_syslog" { | |||
stage.regex { | |||
expression = `^(?P<received_at>\S+) src=(?P<server_ip>\S+) severity=(?P<severity>\S+) host=(?P<server_host>\S+) (?P<source>[^:\s\[]+)(?:\[\d+\])?:?\s+(?P<message>.*)$` | |||
} | |||
stage.timestamp { | |||
source = "received_at" | |||
format = "RFC3339Nano" | |||
action_on_failure = "skip" | |||
} | |||
stage.labels { | |||
values = { | |||
server_ip = "", | |||
server_host = "", | |||
severity = "", | |||
source = "", | |||
} | |||
} | |||
forward_to = [loki.write.local.receiver] | |||
} | |||
</syntaxhighlight> | |||
=== 18.4 QNAP Event === | |||
<syntaxhighlight lang="bash" line> | |||
loki.source.file "qnap_event" { | |||
targets = [ | |||
{ | |||
__path__ = "/var/log/qnap/event.log", | |||
job = "qnap-event", | |||
}, | |||
] | |||
forward_to = [loki.process.qnap_event.receiver] | |||
} | |||
loki.process "qnap_event" { | |||
stage.regex { | |||
expression = `^(?P<received_at>\S+) src=(?P<nas_ip>\S+) severity=(?P<severity>\S+) host=(?P<nas_host>\S+) (?P<message>.*)$` | |||
} | |||
stage.timestamp { | |||
source = "received_at" | |||
format = "RFC3339Nano" | |||
action_on_failure = "skip" | |||
} | |||
stage.labels { | |||
values = { | |||
nas_ip = "", | |||
nas_host = "", | |||
severity = "", | |||
log_type = "event", | |||
} | |||
} | |||
forward_to = [loki.write.local.receiver] | |||
} | |||
</syntaxhighlight> | |||
=== 18.5 QNAP Access === | |||
<syntaxhighlight lang="bash" line> | |||
loki.source.file "qnap_access" { | |||
targets = [ | |||
{ | |||
__path__ = "/var/log/qnap/access.log", | |||
job = "qnap-access", | |||
}, | |||
] | |||
forward_to = [loki.process.qnap_access.receiver] | |||
} | |||
loki.process "qnap_access" { | |||
stage.regex { | |||
expression = `^(?P<received_at>\S+) src=(?P<nas_ip>\S+) severity=(?P<severity>\S+) host=(?P<nas_host>\S+) (?P<message>.*)$` | |||
} | |||
stage.timestamp { | |||
source = "received_at" | |||
format = "RFC3339Nano" | |||
action_on_failure = "skip" | |||
} | |||
stage.labels { | |||
values = { | |||
nas_ip = "", | |||
nas_host = "", | |||
severity = "", | |||
log_type = "access", | |||
} | |||
} | |||
forward_to = [loki.write.local.receiver] | |||
} | |||
</syntaxhighlight> | |||
=== 18.6 Loki Write === | |||
<syntaxhighlight lang="bash" line> | |||
loki.write "local" { | |||
endpoint { | |||
url = "http://127.0.0.1:3100/loki/api/v1/push" | |||
} | |||
} | |||
</syntaxhighlight> | |||
검증: | |||
<syntaxhighlight lang="bash" line> | |||
alloy validate /etc/alloy/config.alloy | |||
systemctl restart alloy | |||
systemctl status alloy --no-pager | |||
</syntaxhighlight> | |||
Loki Job 확인: | |||
<syntaxhighlight lang="bash" line> | |||
curl -s \ | |||
'http://127.0.0.1:3100/loki/api/v1/label/job/values' \ | |||
| jq | |||
</syntaxhighlight> | |||
정상 예: | |||
<syntaxhighlight lang="bash" line> | |||
network-syslog | |||
server-syslog | |||
qnap-event | |||
qnap-access | |||
</syntaxhighlight> | |||
=== 18.7 Alloy Permission 문제 === | |||
다음과 같은 오류가 발생할 수 있다. | |||
<syntaxhighlight lang="bash" line> | |||
failed to tail file | |||
stat failed | |||
permission denied | |||
</syntaxhighlight> | |||
확인: | |||
<syntaxhighlight lang="bash" line> | |||
namei -l /var/log/qnap/access.log | |||
systemctl show alloy \ | |||
-p User \ | |||
-p Group | |||
sudo -u alloy \ | |||
head /var/log/qnap/access.log | |||
</syntaxhighlight> | |||
권장 권한: | |||
<syntaxhighlight lang="bash" line> | |||
drwxr-x--- root alloy /var/log/qnap | |||
-rw-r----- root alloy /var/log/qnap/access.log | |||
-rw-r----- root alloy /var/log/qnap/event.log | |||
</syntaxhighlight> | |||
SELinux 확인: | |||
<syntaxhighlight lang="bash" line> | |||
getenforce | |||
ausearch -m AVC -ts recent \ | |||
| grep -Ei 'alloy|qnap' | |||
ls -Zd /var/log/qnap | |||
ls -Z /var/log/qnap/access.log | |||
</syntaxhighlight> | |||
== 19. Loki Query == | |||
Network: | |||
<syntaxhighlight lang="bash" line> | |||
{job="network-syslog"} | |||
</syntaxhighlight> | |||
Server: | |||
<syntaxhighlight lang="bash" line> | |||
{job="server-syslog"} | |||
</syntaxhighlight> | |||
QNAP Event: | |||
<syntaxhighlight lang="bash" line> | |||
{job="qnap-event"} | |||
</syntaxhighlight> | |||
QNAP Access: | |||
<syntaxhighlight lang="bash" line> | |||
{job="qnap-access"} | |||
</syntaxhighlight> | |||
Critical Network Syslog: | |||
<syntaxhighlight lang="bash" line> | |||
{job="network-syslog",severity=~"emerg|alert|crit|err"} | |||
</syntaxhighlight> | |||
== 20. Grafana Syslog Alert == | |||
Syslog Level 3 이상: | |||
<syntaxhighlight lang="bash" line> | |||
sum by (device_ip, severity) ( | |||
count_over_time( | |||
{ | |||
job="network-syslog", | |||
severity=~"emerg|alert|crit|err" | |||
}[1m] | |||
) | |||
) | |||
</syntaxhighlight> | |||
권장 Alert 구조: | |||
<syntaxhighlight lang="bash" line> | |||
A = Loki Instant Query | |||
B = Threshold | |||
A IS ABOVE 0 | |||
</syntaxhighlight> | |||
Summary: | |||
<syntaxhighlight lang="bash" line> | |||
[Syslog 경고] {{ $labels.device_ip }} - {{ $labels.severity }} | |||
</syntaxhighlight> | |||
Description: | |||
<syntaxhighlight lang="bash" line> | |||
장비 {{ $labels.device_ip }} 에서 Syslog Level 3(Error) 이상의 로그가 발생했습니다. | |||
최근 1분 발생 건수: {{ $values.A.Value }} | |||
Severity: {{ $labels.severity }} | |||
</syntaxhighlight> | |||
Range Query를 그대로 Alert 조건으로 사용할 경우 다음 오류가 발생할 수 있다. | |||
<syntaxhighlight lang="bash" line> | |||
invalid format of evaluation results for the alert definition A: | |||
looks like time series data, only reduced data can be alerted on. | |||
</syntaxhighlight> | |||
Range Query를 유지할 경우 Reduce Expression을 추가해야 한다. | |||
== 21. ICMP Alert == | |||
Flux: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
from(bucket: "snmp_raw") | |||
|> range(start: -5m) | |||
|> filter(fn: (r) => | |||
r._measurement == "ping" and | |||
r._field == "percent_packet_loss" | |||
) | |||
|> group(columns: ["url"]) | |||
|> last() | |||
|> keep(columns: ["_time", "_value", "url"]) | |||
</syntaxhighlight> | </syntaxhighlight> | ||
Alert: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
A = Flux Query | |||
B = Reduce / Last / Strict | |||
C = Threshold > 99 | |||
Pending = 2m | |||
</syntaxhighlight> | </syntaxhighlight> | ||
No Data: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
Keep Last State | |||
</syntaxhighlight> | </syntaxhighlight> | ||
== 22. 최종 Dashboard 구성안 == | |||
=== 22.1 Network Switch Dashboard === | |||
{| class="wikitable" | {| class="wikitable" | ||
! Row | |||
! Panels | |||
|- | |- | ||
| 상태 | |||
| Device / Uptime / CPU / Memory / Ping | |||
|- | |- | ||
| | | Port 상태 | ||
| | | ifName / Alias / Speed / OperStatus | ||
|- | |- | ||
| | | Traffic | ||
| | | RX bps / TX bps | ||
|- | |- | ||
| | | Packet | ||
| | | Unicast / Broadcast / Multicast PPS | ||
|- | |- | ||
| | | Error | ||
| | | RX/TX Error / Discard | ||
|- | |- | ||
| | | Syslog | ||
| | | Warning / Error / Critical Event | ||
|} | |} | ||
=== 22.2 Linux Server Dashboard === | |||
* Uptime | |||
* CPU | |||
* Memory | |||
* Load | |||
* Disk Capacity | |||
* Disk I/O | |||
* Network RX/TX | |||
* Network Error/Drop | |||
* Service Log | |||
* Security Event | |||
=== 22.3 Windows Dashboard === | |||
* Uptime | |||
* CPU | |||
* Memory | |||
* Logical Disk | |||
* Disk I/O | |||
* Network | |||
* Windows Exporter 상태 | |||
=== 22.4 QNAP Dashboard === | |||
최종 구성: | |||
<syntaxhighlight lang="bash" line> | |||
Node Exporter | |||
Uptime | |||
CPU | |||
Memory | |||
HDD 최고 온도 | |||
CPU / Memory / Load | |||
Network RX / TX | |||
- bond0 only | |||
Disk 상태 | |||
Volume | |||
- DataVol1 | |||
- SNMP capacity/free x1024 보정 | |||
Filesystem 사용률 | |||
- /share/CACHEDEV1_DATA only | |||
- Snapshot 제외 | |||
Network Errors / Drops | |||
- bond0 RX Error | |||
- bond0 TX Error | |||
- bond0 RX Drop | |||
- bond0 TX Drop | |||
Disk I/O | |||
- Physical sd* only | |||
- Read B/s | |||
- Write B/s | |||
- Read IOPS | |||
- Write IOPS | |||
QNAP Event Log | |||
QNAP Access Log | |||
</syntaxhighlight> | |||
QNAP Dashboard에서 제거한 항목: | |||
* 상단 RAID 상태 | |||
* Storage Pool 상태 | |||
* Storage Pool 사용률 | |||
* RAID 상세 Table | |||
* Storage Pool 상세 Table | |||
RAID/Storage Pool 정보는 필요 시 별도 상세 Dashboard에서 조회한다. | |||
== 23. 운영 점검 == | |||
전체 서비스: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
systemctl is-active \ | |||
influxdb \ | |||
telegraf \ | |||
grafana-server \ | |||
prometheus \ | |||
loki \ | |||
alloy \ | |||
rsyslog \ | |||
chronyd | |||
</syntaxhighlight> | |||
Listening Port: | |||
<syntaxhighlight lang="bash" line> | |||
ss -lntup | |||
</syntaxhighlight> | </syntaxhighlight> | ||
Prometheus: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
curl -s http://127.0.0.1:9090/-/healthy | |||
</syntaxhighlight> | </syntaxhighlight> | ||
Loki: | |||
<syntaxhighlight lang="bash" line> | |||
curl -s http://127.0.0.1:3100/ready | |||
</syntaxhighlight> | |||
Telegraf: | |||
= | <syntaxhighlight lang="bash" line> | ||
journalctl -u telegraf \ | |||
--since '-10 min' \ | |||
--no-pager | |||
</syntaxhighlight> | |||
Alloy: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
journalctl -u alloy \ | |||
--since '-10 min' \ | |||
--no-pager | |||
</syntaxhighlight> | |||
Filesystem: | |||
<syntaxhighlight lang="bash" line> | |||
df -hT | |||
</syntaxhighlight> | </syntaxhighlight> | ||
System I/O: | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
vmstat 1 5 | |||
iostat -xz 1 5 | |||
</syntaxhighlight> | </syntaxhighlight> | ||
== | == 24. 장애 점검 순서 == | ||
=== Prometheus Target Down === | |||
<syntaxhighlight lang="bash" line> | <syntaxhighlight lang="bash" line> | ||
curl http://TARGET_IP:9100/metrics | |||
journalctl -u prometheus \ | |||
-n 100 \ | |||
--no-pager | |||
</syntaxhighlight> | |||
=== SNMP Timeout === | |||
<syntaxhighlight lang="bash" line> | |||
snmpget -v3 \ | |||
-t 5 \ | |||
-r 1 \ | |||
-On \ | |||
SWITCH_IP \ | |||
.1.3.6.1.2.1.1.3.0 | |||
</syntaxhighlight> | </syntaxhighlight> | ||
확인 항목: | |||
* SNMP User | |||
* SHA/DES/AES 조합 | |||
* ACL | |||
* Source IP | |||
* Timeout | |||
* SNMP View | |||
* 장비 CPU | |||
=== Syslog 미수신 === | |||
= | <syntaxhighlight lang="bash" line> | ||
ss -lntup \ | |||
| grep -E ':514|:5514|:5515|:5516' | |||
tcpdump -ni any \ | |||
host DEVICE_IP | |||
tail -f /var/log/qnap/access.log | |||
</syntaxhighlight> | |||
=== | === Loki 미표시 === | ||
<syntaxhighlight lang="bash" line> | |||
alloy validate \ | |||
/etc/alloy/config.alloy | |||
journalctl -u alloy \ | |||
-n 100 \ | |||
--no-pager | |||
curl -s \ | |||
'http://127.0.0.1:3100/loki/api/v1/label/job/values' \ | |||
| jq | |||
</syntaxhighlight> | |||
== 25. 운영 원칙 == | |||
* SNMP Metric은 Telegraf → InfluxDB로 저장한다. | |||
* Host Metric은 Prometheus로 저장한다. | |||
* Syslog는 rsyslog → Alloy → Loki로 저장한다. | |||
* Grafana는 세 데이터소스를 통합한다. | |||
* SNMP Counter는 누적값을 그대로 표시하지 않고 rate/derivative 계산한다. | |||
* No Data를 0 또는 정상으로 강제 표시하지 않는다. | |||
* QNAP Snapshot filesystem은 Data Volume 사용률에서 제외한다. | |||
* QNAP Volume SNMP capacity/free 값은 장비 특성상 ×1024 보정한다. | |||
* QNAP Network는 실제 활성 bond0만 표시한다. | |||
* Loki message 전체를 label로 만들지 않는다. | |||
* 장비 Syslog Alert는 device_ip / severity와 발생 건수를 메일에 포함한다. | |||
* Alert의 Range Query는 Reduce 또는 Instant Query 구조로 구성한다. | |||
* 모든 Token 및 SNMP 비밀번호는 문서에 실제 값을 기록하지 않는다. | |||
2026년 9월 23일 (수) 20:22 판
Rocky Linux 9 통합 모니터링 서버 구축 가이드
- Telegraf / InfluxDB / Grafana / Prometheus / Node Exporter / Windows Exporter / Loki / Alloy / rsyslog
작성 기준: 실제 구성 및 장애 처리 기록 기반 문서 형식: MediaWiki IP 주소: 문서용 예시 주소로 치환 운영환경 적용 전 실제 장비 주소, 계정, 토큰, OID를 확인한다.
0. 목적 및 최종 구성
본 문서는 Rocky Linux 9 기반 모니터링 서버를 처음 구축하는 단계부터 네트워크 장비 SNMP, 서버 Metric, Syslog, QNAP NAS 모니터링 및 Grafana Dashboard 구성까지 정리한다.
최종 구조는 다음과 같다.
Network Switch
|
+-- SNMPv3 UDP/161
| |
| +--> Telegraf
| |
| +--> InfluxDB OSS 2.x
|
+-- Syslog TCP/UDP
|
+--> rsyslog
|
+--> Log File
|
+--> Grafana Alloy
|
+--> Loki
Linux / QNAP
|
+-- node_exporter :9100
|
+--> Prometheus
Windows
|
+-- windows_exporter :9182
|
+--> Prometheus
Prometheus + InfluxDB + Loki
|
+--> Grafana
0.1 문서용 IP 주소
본 문서에서는 실제 운영 IP를 노출하지 않고 RFC 문서용 주소를 사용한다.
| 대상 | 예시 주소 | 용도 |
|---|---|---|
| Monitoring Server | 192.0.2.10 | Grafana / Prometheus / InfluxDB / Telegraf / Loki / Alloy / rsyslog |
| Linux Server | 192.0.2.20 | node_exporter |
| Windows Server | 192.0.2.30 | windows_exporter |
| QNAP NAS | 198.51.100.10 | node_exporter / SNMP / Syslog |
| Switch-01 | 203.0.113.11 | SNMP / Syslog |
| Switch-02 | 203.0.113.12 | SNMP / Syslog |
| Switch-03 | 203.0.113.13 | SNMP / Syslog |
1. Rocky Linux 기본 준비
1.1 시스템 상태 확인
cat /etc/rocky-release
uname -m
ip -br address
ip route
df -hT
free -h
getenforce
ss -lntup
SELinux와 firewalld는 초기부터 비활성화하지 않는다.
1.2 백업 디렉터리 생성
기존 서버에 추가 설치하는 경우 설정 파일 백업을 먼저 수행한다.
umask 077
MON_BACKUP="/root/monitor-backup-$(date +%Y%m%d-%H%M%S)"
install -d -m 700 "$MON_BACKUP"
cp -a /etc/rsyslog.conf "$MON_BACKUP/" 2>/dev/null || true
cp -a /etc/rsyslog.d "$MON_BACKUP/" 2>/dev/null || true
cp -a /etc/chrony.conf "$MON_BACKUP/" 2>/dev/null || true
cp -a /etc/firewalld "$MON_BACKUP/" 2>/dev/null || true
rpm -qa | sort > "$MON_BACKUP/packages-before.txt"
1.3 기본 패키지 설치
dnf install -y \
curl \
ca-certificates \
gnupg2 \
vim-enhanced \
net-snmp-utils \
rsyslog \
logrotate \
chrony \
tcpdump \
iputils \
sysstat \
policycoreutils-python-utils \
dnf-plugins-core \
jq
1.4 시간 동기화
모니터링에서는 장비와 서버 간 시간이 맞아야 장애 시각을 비교할 수 있다.
timedatectl set-timezone Asia/Seoul
systemctl enable --now chronyd
systemctl restart chronyd
chronyc tracking
chronyc sources -v
date -Ins
2. InfluxDB / Telegraf / Grafana 설치
2.1 InfluxData Repository
curl -fL https://repos.influxdata.com/influxdata-archive.key \
-o /etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata
gpg --show-keys --with-fingerprint \
/etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata
rpm --import /etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata
Repository:
vi /etc/yum.repos.d/influxdata.repo
[influxdata]
name=InfluxData Repository - Stable
baseurl=https://repos.influxdata.com/stable/$basearch/main
enabled=1
gpgcheck=1
gpgkey=file:///etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata
sslverify=1
2.2 Grafana Repository
vi /etc/yum.repos.d/grafana.repo
[grafana]
name=Grafana OSS repository
baseurl=https://rpm.grafana.com
repo_gpgcheck=1
enabled=1
gpgcheck=1
gpgkey=https://rpm.grafana.com/gpg.key
sslverify=1
2.3 패키지 설치
dnf makecache
dnf list --showduplicates \
telegraf \
influxdb2 \
influxdb2-cli \
grafana
dnf install -y \
telegraf \
influxdb2 \
influxdb2-cli \
grafana
버전 확인:
rpm -q telegraf influxdb2 influxdb2-cli grafana rsyslog
telegraf --version
influxd version
influx version
실제 구축 과정에서는 Telegraf 1.40.0 환경에서 동작을 확인하였다.
3. InfluxDB 초기 구성
3.1 Localhost Bind
InfluxDB API는 Grafana 및 Telegraf가 같은 서버에 있으므로 localhost만 수신하도록 구성한다.
install -d -m 755 /etc/systemd/system/influxdb.service.d
vi /etc/systemd/system/influxdb.service.d/10-listen.conf
[Service]
Environment="INFLUXD_HTTP_BIND_ADDRESS=127.0.0.1:8086"
systemctl daemon-reload
systemctl enable --now influxdb
systemctl restart influxdb
systemctl status influxdb --no-pager
ss -lntp | grep ':8086'
curl -fsS http://127.0.0.1:8086/health
3.2 Initial Setup
influx setup
예시:
| 설정 | 값 |
|---|---|
| Organization | network |
| Bucket | snmp_raw |
| Retention | 720h |
Bucket 확인:
influx bucket list --org network
3.3 Token 분리
서비스별 Token 권한을 분리한다.
- telegraf-write : snmp_raw Write
- grafana-read : snmp_raw Read
- Operator Token : 관리자용
운영 문서에 실제 Token 값을 기록하지 않는다.
4. Telegraf 기본 구성
4.1 환경 변수 파일
vi /etc/telegraf/monitor.env
INFLUX_WRITE_TOKEN='<WRITE_TOKEN>'
ZYXEL_SNMP_USER='<SNMP_USER>'
ZYXEL_SNMP_AUTH='<AUTH_PASSWORD>'
ZYXEL_SNMP_PRIV='<PRIV_PASSWORD>'
QNAP_SNMP_USER='<QNAP_SNMP_USER>'
QNAP_SNMP_AUTH='<QNAP_AUTH_PASSWORD>'
QNAP_SNMP_PRIV='<QNAP_PRIV_PASSWORD>'
chown root:root /etc/telegraf/monitor.env
chmod 600 /etc/telegraf/monitor.env
install -d -m 755 /etc/systemd/system/telegraf.service.d
vi /etc/systemd/system/telegraf.service.d/10-monitor-env.conf
[Service]
EnvironmentFile=/etc/telegraf/monitor.env
4.2 Main Configuration
cp -a /etc/telegraf/telegraf.conf \
/etc/telegraf/telegraf.conf.orig
vi /etc/telegraf/telegraf.conf
[agent]
interval = "30s"
round_interval = true
flush_interval = "5s"
precision = "1ms"
metric_batch_size = 1000
metric_buffer_limit = 20000
omit_hostname = false
snmp_translator = "gosmi"
[[outputs.influxdb_v2]]
urls = ["http://127.0.0.1:8086"]
token = "${INFLUX_WRITE_TOKEN}"
organization = "network"
bucket = "snmp_raw"
[[inputs.internal]]
[[inputs.cpu]]
percpu = false
totalcpu = true
[[inputs.mem]]
[[inputs.disk]]
mount_points = ["/"]
[[inputs.net]]
5. SNMPv3 사전 확인
SNMP는 v3 authPriv 사용을 기본으로 한다.
Net-SNMP 테스트 파일:
install -d -m 700 /root/.snmp
vi /root/.snmp/snmp.conf
defVersion 3
defSecurityName <SNMP_USER>
defSecurityLevel authPriv
defAuthType SHA
defAuthPassphrase <SNMP_AUTH_PASSWORD>
defPrivType AES
defPrivPassphrase <SNMP_PRIV_PASSWORD>
chmod 600 /root/.snmp/snmp.conf
snmpget -v3 -t 2 -r 0 -On \
203.0.113.11 \
.1.3.6.1.2.1.1.3.0
snmpwalk -v3 -t 2 -r 0 -On \
203.0.113.11 \
.1.3.6.1.2.1.31.1.1.1.1
ifName OID 마지막 값이 ifIndex이다.
6. Telegraf SNMP 구성
설정 파일은 장비 모델별로 분리한다.
/etc/telegraf/telegraf.d/
├── 10-zyxel-gs1900.conf
├── 11-zyxel-gs1920.conf
├── 20-zyxel-es3128.conf
├── 30-icmp_check.conf
├── 40-qnap_nas.conf
└── sflow.conf
6.1 GS1920 계열
예시 대상:
203.0.113.21
203.0.113.22
203.0.113.23
SNMPv3:
- SHA
- DES
CPU OID:
.1.3.6.1.4.1.890.1.15.3.49.1.7.0
Memory:
Total .1.3.6.1.4.1.890.1.15.3.50.1.1.1.3.1
Used .1.3.6.1.4.1.890.1.15.3.50.1.1.1.4.1
Percent .1.3.6.1.4.1.890.1.15.3.50.1.1.1.5.1
6.2 GS1900 계열
CPU:
.1.3.6.1.4.1.890.1.15.3.2.4.0
Memory:
.1.3.6.1.4.1.890.1.15.3.2.5.0
실제 구축에서는 2초 timeout / retries 0에서 응답 누락이 발생하여 다음과 같이 완화하였다.
timeout = "5s"
retries = 1
6.3 ES-3128GP
SNMPv3:
- SHA
- AES
sysObjectID:
.1.3.6.1.4.1.7800.1.190
CPU / Memory private OID는 장비에서 정상 확인되지 않아 Dashboard에서는 N/A 처리한다.
6.4 IF-MIB 주요 항목
수집 대상:
- ifName
- ifAlias
- ifSpeed
- ifHCInOctets
- ifHCOutOctets
- ifOperStatus
- ifInErrors
- ifOutErrors
- ifInDiscards
- ifOutDiscards
- Broadcast
- Multicast
누적 Counter는 Grafana에서 그대로 표시하지 않고 derivative/rate 계산 후 표시한다.
7. ICMP 수집
Telegraf ping input을 사용한다.
vi /etc/telegraf/telegraf.d/30-icmp_check.conf
예시 대상:
203.0.113.11
203.0.113.12
203.0.113.13
권장 설정:
[[inputs.ping]]
urls = [
"203.0.113.11",
"203.0.113.12",
"203.0.113.13"
]
method = "native"
count = 3
deadline = 2.0
interval = 10.0
native ping은 CAP_NET_RAW가 필요하다.
systemctl edit telegraf
[Service]
CapabilityBoundingSet=CAP_NET_RAW
AmbientCapabilities=CAP_NET_RAW
systemctl daemon-reload
systemctl restart telegraf
8. Telegraf Test 및 서비스 시작
설정 검사:
telegraf \
--config /etc/telegraf/telegraf.conf \
--config-directory /etc/telegraf/telegraf.d \
--test
실제 서비스 계정으로 확인:
systemctl daemon-reload
systemd-run \
--unit=telegraf-config-check \
--wait \
--pipe \
--collect \
-p User=telegraf \
-p Group=telegraf \
-p EnvironmentFile=/etc/telegraf/monitor.env \
-p CapabilityBoundingSet=CAP_NET_RAW \
-p AmbientCapabilities=CAP_NET_RAW \
/usr/bin/telegraf \
--config /etc/telegraf/telegraf.conf \
--config-directory /etc/telegraf/telegraf.d \
--test
서비스 시작:
systemctl enable --now telegraf
systemctl restart telegraf
systemctl status telegraf --no-pager
journalctl -u telegraf -n 100 --no-pager
9. Grafana 설치 및 InfluxDB 연결
9.1 Grafana 서비스
vi /etc/grafana/grafana.ini
내부망 운영 환경에 맞게 listen 주소를 설정한다.
예시:
[server]
http_addr = 0.0.0.0
http_port = 3000
[users]
allow_sign_up = false
[auth.anonymous]
enabled = false
systemctl enable --now grafana-server
systemctl restart grafana-server
systemctl status grafana-server --no-pager
curl -fsS http://127.0.0.1:3000/api/health
9.2 InfluxDB Datasource
Grafana:
Connections
→ Data sources
→ Add data source
→ InfluxDB
| 설정 | 값 |
|---|---|
| Query language | Flux |
| URL | http://127.0.0.1:8086 |
| Organization | network |
| Token | grafana-read |
| Default Bucket | snmp_raw |
| Min time interval | 1s |
10. SNMP Dashboard 기본 구성
10.1 Traffic
Flux:
from(bucket: "snmp_raw")
|> range(start: v.timeRangeStart, stop: v.timeRangeStop)
|> filter(fn: (r) => r._measurement == "switch_port")
|> filter(fn: (r) => r.source == "203.0.113.11")
|> filter(fn: (r) =>
r._field == "in_octets" or
r._field == "out_octets"
)
|> derivative(unit: 1s, nonNegative: true)
|> map(fn: (r) => ({
r with _value: r._value * 8.0
}))
Grafana Unit:
bits/sec
10.2 Error / Discard
|> filter(fn: (r) =>
r._field == "in_errors" or
r._field == "out_errors" or
r._field == "in_discards" or
r._field == "out_discards"
)
|> derivative(unit: 1s, nonNegative: true)
10.3 Broadcast / Multicast
Broadcast/Multicast는 별도 패널로 표시하여 폭주 및 비정상 증가를 확인한다.
11. Prometheus 설치
실제 구축에서는 Prometheus 3.13.3을 사용하였다.
11.1 사용자 및 디렉터리
useradd \
--system \
--no-create-home \
--shell /sbin/nologin \
prometheus
install -d -o prometheus -g prometheus \
/etc/prometheus \
/var/lib/prometheus
Prometheus 바이너리는 공식 릴리스에서 다운로드하고 checksum 확인 후 설치한다.
install -m 0755 prometheus /usr/local/bin/prometheus
install -m 0755 promtool /usr/local/bin/promtool
11.2 Prometheus 설정
vi /etc/prometheus/prometheus.yml
global:
scrape_interval: 30s
scrape_configs:
- job_name: prometheus
static_configs:
- targets:
- "127.0.0.1:9090"
labels:
server_name: monitoring-server
- job_name: node
static_configs:
- targets:
- "127.0.0.1:9100"
labels:
server_name: monitoring-server
- targets:
- "192.0.2.20:9100"
labels:
server_name: linux-server
- job_name: windows
static_configs:
- targets:
- "192.0.2.30:9182"
labels:
server_name: windows-server
- job_name: qnap
static_configs:
- targets:
- "198.51.100.10:9100"
labels:
server_name: nas-01
11.3 systemd
vi /etc/systemd/system/prometheus.service
[Unit]
Description=Prometheus
Wants=network-online.target
After=network-online.target
[Service]
User=prometheus
Group=prometheus
ExecStart=/usr/local/bin/prometheus \
--config.file=/etc/prometheus/prometheus.yml \
--storage.tsdb.path=/var/lib/prometheus \
--storage.tsdb.retention.time=30d \
--storage.tsdb.retention.size=20GB \
--web.listen-address=127.0.0.1:9090
Restart=always
[Install]
WantedBy=multi-user.target
chown -R prometheus:prometheus \
/etc/prometheus \
/var/lib/prometheus
promtool check config /etc/prometheus/prometheus.yml
systemctl daemon-reload
systemctl enable --now prometheus
systemctl status prometheus --no-pager
12. Node Exporter 설치
실제 구축에서는 Node Exporter 1.12.1 환경을 사용하였다.
12.1 Linux Node Exporter
useradd \
--system \
--no-create-home \
--shell /sbin/nologin \
node_exporter
install -m 0755 node_exporter \
/usr/local/bin/node_exporter
vi /etc/systemd/system/node_exporter.service
[Unit]
Description=Node Exporter
After=network-online.target
Wants=network-online.target
[Service]
User=node_exporter
Group=node_exporter
ExecStart=/usr/local/bin/node_exporter
Restart=always
[Install]
WantedBy=multi-user.target
systemctl daemon-reload
systemctl enable --now node_exporter
curl -s http://127.0.0.1:9100/metrics | head
외부 서버의 node_exporter는 Prometheus 서버 주소만 9100/tcp에 접근하도록 방화벽을 제한한다.
13. Windows Exporter
Windows Server에서는 windows_exporter를 사용한다.
기본 포트:
9182/tcp
Prometheus Target:
192.0.2.30:9182
성능 카운터 이상 시 다음 명령으로 복구한 사례가 있다.
lodctr /R
winmgmt /resyncperf
PowerShell:
Restart-Service windows_exporter
확인:
Invoke-WebRequest http://127.0.0.1:9182/metrics
14. Grafana Prometheus Datasource
Grafana:
Connections
→ Data sources
→ Prometheus
URL:
http://127.0.0.1:9090
Save & Test 후 다음 PromQL로 확인한다.
up
15. Server Dashboard 구성
Linux Dashboard:
- Uptime
- CPU Usage
- Memory Usage
- Load Average
- Filesystem
- Disk Read / Write
- Network RX / TX
- Network Error / Drop
- Service 상태
Windows Dashboard:
- CPU
- Memory
- Disk
- Network
- Uptime
- Exporter 상태
Disk 용량은 실제 GB/GiB 값으로 표시하며 불필요한 Bar Gauge는 제거한다.
16. QNAP TS-264 모니터링
QNAP은 세 종류의 데이터를 함께 사용한다.
| 데이터 | 수집 방법 |
|---|---|
| CPU / Memory / Network / Filesystem / Disk I/O | node_exporter → Prometheus |
| HDD / RAID / Storage / Volume | SNMP → Telegraf → InfluxDB |
| Event / Access Log | Syslog → rsyslog → Alloy → Loki |
16.1 QNAP SNMP
QNAP SNMPv3:
- SHA
- DES
Measurement:
QNAP_TS264
qnap_disk
qnap_raid
qnap_storage_pool
qnap_volume
Disk:
- disk_id
- manufacturer
- model
- disk_type
- disk_status
- temperature
- capacity_bytes
RAID:
- RAID ID
- RAID Name
- RAID Status
- RAID Level
- Capacity
16.2 Volume 단위 보정
QNAP qnap_volume의 capacity/free 값은 실제 장비에서 KiB 형태로 반환되는 것으로 확인되었다.
원시값을 Grafana에서 byte로 바로 해석하면 약 11.3 GiB로 잘못 표시된다.
Flux:
|> map(fn: (r) => ({
r with
capacity_bytes:
uint(v: r.capacity_bytes) * uint(v: 1024),
free_bytes:
uint(v: r.free_bytes) * uint(v: 1024),
used_bytes:
(
uint(v: r.capacity_bytes)
- uint(v: r.free_bytes)
) * uint(v: 1024),
used_percent:
if float(v: r.capacity_bytes) > 0.0 then
(
float(v: r.capacity_bytes)
- float(v: r.free_bytes)
)
/ float(v: r.capacity_bytes)
* 100.0
else 0.0
}))
16.3 실제 Data Volume
Node Exporter 기준 실제 사용자 Data Volume:
mountpoint="/share/CACHEDEV1_DATA"
device="/dev/mapper/cachedev1"
fstype="ext4"
Snapshot:
/mnt/snapshot/...
위 Snapshot 경로는 Dashboard Filesystem 패널에서 제외한다.
DataVol1 사용률:
100 * (
1 -
node_filesystem_avail_bytes{
instance="198.51.100.10:9100",
mountpoint="/share/CACHEDEV1_DATA"
}
/
node_filesystem_size_bytes{
instance="198.51.100.10:9100",
mountpoint="/share/CACHEDEV1_DATA"
}
)
16.4 Network는 bond0만 표시
RX:
rate(
node_network_receive_bytes_total{
instance="198.51.100.10:9100",
device="bond0"
}[$__rate_interval]
) * 8
TX:
rate(
node_network_transmit_bytes_total{
instance="198.51.100.10:9100",
device="bond0"
}[$__rate_interval]
) * 8
16.5 Network Error / Drop
rate(node_network_receive_errs_total{
instance="198.51.100.10:9100",
device="bond0"
}[$__rate_interval])
rate(node_network_transmit_errs_total{
instance="198.51.100.10:9100",
device="bond0"
}[$__rate_interval])
rate(node_network_receive_drop_total{
instance="198.51.100.10:9100",
device="bond0"
}[$__rate_interval])
rate(node_network_transmit_drop_total{
instance="198.51.100.10:9100",
device="bond0"
}[$__rate_interval])
16.6 Disk I/O
Read:
rate(node_disk_read_bytes_total{
instance="198.51.100.10:9100",
device=~"sd[a-z]+"
}[$__rate_interval])
Write:
rate(node_disk_written_bytes_total{
instance="198.51.100.10:9100",
device=~"sd[a-z]+"
}[$__rate_interval])
Read IOPS:
rate(node_disk_reads_completed_total{
instance="198.51.100.10:9100",
device=~"sd[a-z]+"
}[$__rate_interval])
Write IOPS:
rate(node_disk_writes_completed_total{
instance="198.51.100.10:9100",
device=~"sd[a-z]+"
}[$__rate_interval])
17. rsyslog 구성
17.1 Network Syslog
예시 로그 파일:
/var/log/network-syslog/events.log
네트워크 장비:
UDP/TCP 514
17.2 Server Syslog
/var/log/server-syslog/events.log
예시 포트:
5514/tcp
17.3 QNAP Event / Access 분리
QNAP 로그는 Event와 Access를 별도 포트로 분리한다.
| 로그 | Port | File |
|---|---|---|
| Event | TCP 5515 | /var/log/qnap/event.log |
| Access | TCP 5516 | /var/log/qnap/access.log |
Template:
template(name="QnapSyslogLine" type="string"
string="%timegenerated:::date-rfc3339% src=%fromhost-ip% severity=%syslogseverity-text% host=%hostname% %syslogtag%%msg:::sp-if-no-1st-sp%%msg%\n")
Event:
ruleset(name="QnapEventLog") {
action(
type="omfile"
file="/var/log/qnap/event.log"
template="QnapSyslogLine"
fileOwner="root"
fileGroup="alloy"
fileCreateMode="0640"
dirOwner="root"
dirGroup="alloy"
dirCreateMode="0750"
createDirs="on"
)
stop
}
input(
type="imtcp"
port="5515"
ruleset="QnapEventLog"
)
Access:
ruleset(name="QnapAccessLog") {
action(
type="omfile"
file="/var/log/qnap/access.log"
template="QnapSyslogLine"
fileOwner="root"
fileGroup="alloy"
fileCreateMode="0640"
dirOwner="root"
dirGroup="alloy"
dirCreateMode="0750"
createDirs="on"
)
stop
}
input(
type="imtcp"
port="5516"
ruleset="QnapAccessLog"
)
검증:
rsyslogd -N1
systemctl restart rsyslog
ss -lntp | grep -E ':5514|:5515|:5516'
18. Loki / Alloy 구성
Loki Endpoint:
http://127.0.0.1:3100
Alloy UI:
127.0.0.1:12345
18.1 Alloy 기본 설정
logging {
level = "info"
}
18.2 Network Syslog
loki.source.file "network_syslog" {
targets = [
{
__path__ = "/var/log/network-syslog/events.log",
job = "network-syslog",
},
]
forward_to = [loki.process.network_syslog.receiver]
}
loki.process "network_syslog" {
stage.regex {
expression = `^(?P<received_at>\S+) src=(?P<device_ip>\S+) severity=(?P<severity>\S+) host=(?P<device_host>\S+) (?P<message>.*)$`
}
stage.timestamp {
source = "received_at"
format = "RFC3339Nano"
action_on_failure = "skip"
}
stage.labels {
values = {
device_ip = "",
severity = "",
}
}
forward_to = [loki.write.local.receiver]
}
18.3 Server Syslog
loki.source.file "server_syslog" {
targets = [
{
__path__ = "/var/log/server-syslog/events.log",
job = "server-syslog",
},
]
forward_to = [loki.process.server_syslog.receiver]
}
loki.process "server_syslog" {
stage.regex {
expression = `^(?P<received_at>\S+) src=(?P<server_ip>\S+) severity=(?P<severity>\S+) host=(?P<server_host>\S+) (?P<source>[^:\s\[]+)(?:\[\d+\])?:?\s+(?P<message>.*)$`
}
stage.timestamp {
source = "received_at"
format = "RFC3339Nano"
action_on_failure = "skip"
}
stage.labels {
values = {
server_ip = "",
server_host = "",
severity = "",
source = "",
}
}
forward_to = [loki.write.local.receiver]
}
18.4 QNAP Event
loki.source.file "qnap_event" {
targets = [
{
__path__ = "/var/log/qnap/event.log",
job = "qnap-event",
},
]
forward_to = [loki.process.qnap_event.receiver]
}
loki.process "qnap_event" {
stage.regex {
expression = `^(?P<received_at>\S+) src=(?P<nas_ip>\S+) severity=(?P<severity>\S+) host=(?P<nas_host>\S+) (?P<message>.*)$`
}
stage.timestamp {
source = "received_at"
format = "RFC3339Nano"
action_on_failure = "skip"
}
stage.labels {
values = {
nas_ip = "",
nas_host = "",
severity = "",
log_type = "event",
}
}
forward_to = [loki.write.local.receiver]
}
18.5 QNAP Access
loki.source.file "qnap_access" {
targets = [
{
__path__ = "/var/log/qnap/access.log",
job = "qnap-access",
},
]
forward_to = [loki.process.qnap_access.receiver]
}
loki.process "qnap_access" {
stage.regex {
expression = `^(?P<received_at>\S+) src=(?P<nas_ip>\S+) severity=(?P<severity>\S+) host=(?P<nas_host>\S+) (?P<message>.*)$`
}
stage.timestamp {
source = "received_at"
format = "RFC3339Nano"
action_on_failure = "skip"
}
stage.labels {
values = {
nas_ip = "",
nas_host = "",
severity = "",
log_type = "access",
}
}
forward_to = [loki.write.local.receiver]
}
18.6 Loki Write
loki.write "local" {
endpoint {
url = "http://127.0.0.1:3100/loki/api/v1/push"
}
}
검증:
alloy validate /etc/alloy/config.alloy
systemctl restart alloy
systemctl status alloy --no-pagerLoki Job 확인:
curl -s \
'http://127.0.0.1:3100/loki/api/v1/label/job/values' \
| jq정상 예:
network-syslog
server-syslog
qnap-event
qnap-access18.7 Alloy Permission 문제
다음과 같은 오류가 발생할 수 있다.
failed to tail file
stat failed
permission denied확인:
namei -l /var/log/qnap/access.log
systemctl show alloy \
-p User \
-p Group
sudo -u alloy \
head /var/log/qnap/access.log권장 권한:
drwxr-x--- root alloy /var/log/qnap
-rw-r----- root alloy /var/log/qnap/access.log
-rw-r----- root alloy /var/log/qnap/event.logSELinux 확인:
getenforce
ausearch -m AVC -ts recent \
| grep -Ei 'alloy|qnap'
ls -Zd /var/log/qnap
ls -Z /var/log/qnap/access.log19. Loki Query
Network:
{job="network-syslog"}Server:
{job="server-syslog"}QNAP Event:
{job="qnap-event"}QNAP Access:
{job="qnap-access"}Critical Network Syslog:
{job="network-syslog",severity=~"emerg|alert|crit|err"}20. Grafana Syslog Alert
Syslog Level 3 이상:
sum by (device_ip, severity) (
count_over_time(
{
job="network-syslog",
severity=~"emerg|alert|crit|err"
}[1m]
)
)권장 Alert 구조:
A = Loki Instant Query
B = Threshold
A IS ABOVE 0Summary:
[Syslog 경고] {{ $labels.device_ip }} - {{ $labels.severity }}Description:
장비 {{ $labels.device_ip }} 에서 Syslog Level 3(Error) 이상의 로그가 발생했습니다.
최근 1분 발생 건수: {{ $values.A.Value }}
Severity: {{ $labels.severity }}Range Query를 그대로 Alert 조건으로 사용할 경우 다음 오류가 발생할 수 있다.
invalid format of evaluation results for the alert definition A:
looks like time series data, only reduced data can be alerted on.Range Query를 유지할 경우 Reduce Expression을 추가해야 한다.
21. ICMP Alert
Flux:
from(bucket: "snmp_raw")
|> range(start: -5m)
|> filter(fn: (r) =>
r._measurement == "ping" and
r._field == "percent_packet_loss"
)
|> group(columns: ["url"])
|> last()
|> keep(columns: ["_time", "_value", "url"])Alert:
A = Flux Query
B = Reduce / Last / Strict
C = Threshold > 99
Pending = 2mNo Data:
Keep Last State22. 최종 Dashboard 구성안
22.1 Network Switch Dashboard
| Row | Panels |
|---|---|
| 상태 | Device / Uptime / CPU / Memory / Ping |
| Port 상태 | ifName / Alias / Speed / OperStatus |
| Traffic | RX bps / TX bps |
| Packet | Unicast / Broadcast / Multicast PPS |
| Error | RX/TX Error / Discard |
| Syslog | Warning / Error / Critical Event |
22.2 Linux Server Dashboard
- Uptime
- CPU
- Memory
- Load
- Disk Capacity
- Disk I/O
- Network RX/TX
- Network Error/Drop
- Service Log
- Security Event
22.3 Windows Dashboard
- Uptime
- CPU
- Memory
- Logical Disk
- Disk I/O
- Network
- Windows Exporter 상태
22.4 QNAP Dashboard
최종 구성:
Node Exporter
Uptime
CPU
Memory
HDD 최고 온도
CPU / Memory / Load
Network RX / TX
- bond0 only
Disk 상태
Volume
- DataVol1
- SNMP capacity/free x1024 보정
Filesystem 사용률
- /share/CACHEDEV1_DATA only
- Snapshot 제외
Network Errors / Drops
- bond0 RX Error
- bond0 TX Error
- bond0 RX Drop
- bond0 TX Drop
Disk I/O
- Physical sd* only
- Read B/s
- Write B/s
- Read IOPS
- Write IOPS
QNAP Event Log
QNAP Access LogQNAP Dashboard에서 제거한 항목:
- 상단 RAID 상태
- Storage Pool 상태
- Storage Pool 사용률
- RAID 상세 Table
- Storage Pool 상세 Table
RAID/Storage Pool 정보는 필요 시 별도 상세 Dashboard에서 조회한다.
23. 운영 점검
전체 서비스:
systemctl is-active \
influxdb \
telegraf \
grafana-server \
prometheus \
loki \
alloy \
rsyslog \
chronydListening Port:
ss -lntupPrometheus:
curl -s http://127.0.0.1:9090/-/healthyLoki:
curl -s http://127.0.0.1:3100/readyTelegraf:
journalctl -u telegraf \
--since '-10 min' \
--no-pagerAlloy:
journalctl -u alloy \
--since '-10 min' \
--no-pagerFilesystem:
df -hTSystem I/O:
vmstat 1 5
iostat -xz 1 524. 장애 점검 순서
Prometheus Target Down
curl http://TARGET_IP:9100/metrics
journalctl -u prometheus \
-n 100 \
--no-pagerSNMP Timeout
snmpget -v3 \
-t 5 \
-r 1 \
-On \
SWITCH_IP \
.1.3.6.1.2.1.1.3.0확인 항목:
- SNMP User
- SHA/DES/AES 조합
- ACL
- Source IP
- Timeout
- SNMP View
- 장비 CPU
Syslog 미수신
ss -lntup \
| grep -E ':514|:5514|:5515|:5516'
tcpdump -ni any \
host DEVICE_IP
tail -f /var/log/qnap/access.logLoki 미표시
alloy validate \
/etc/alloy/config.alloy
journalctl -u alloy \
-n 100 \
--no-pager
curl -s \
'http://127.0.0.1:3100/loki/api/v1/label/job/values' \
| jq25. 운영 원칙
- SNMP Metric은 Telegraf → InfluxDB로 저장한다.
- Host Metric은 Prometheus로 저장한다.
- Syslog는 rsyslog → Alloy → Loki로 저장한다.
- Grafana는 세 데이터소스를 통합한다.
- SNMP Counter는 누적값을 그대로 표시하지 않고 rate/derivative 계산한다.
- No Data를 0 또는 정상으로 강제 표시하지 않는다.
- QNAP Snapshot filesystem은 Data Volume 사용률에서 제외한다.
- QNAP Volume SNMP capacity/free 값은 장비 특성상 ×1024 보정한다.
- QNAP Network는 실제 활성 bond0만 표시한다.
- Loki message 전체를 label로 만들지 않는다.
- 장비 Syslog Alert는 device_ip / severity와 발생 건수를 메일에 포함한다.
- Alert의 Range Query는 Reduce 또는 Instant Query 구조로 구성한다.
- 모든 Token 및 SNMP 비밀번호는 문서에 실제 값을 기록하지 않는다.