SNMP테스트서버
Rocky Linux 9 통합 모니터링 서버 구축 가이드
- Telegraf / InfluxDB / Grafana / Prometheus / Node Exporter / Windows Exporter / Loki / Alloy / rsyslog
작성 기준: 실제 구성 및 장애 처리 기록 기반 문서 형식: MediaWiki IP 주소: 문서용 예시 주소로 치환 운영환경 적용 전 실제 장비 주소, 계정, 토큰, OID를 확인한다.
0. 목적 및 최종 구성
본 문서는 Rocky Linux 9 기반 모니터링 서버를 처음 구축하는 단계부터 네트워크 장비 SNMP, 서버 Metric, Syslog, QNAP NAS 모니터링 및 Grafana Dashboard 구성까지 정리한다.
최종 구조는 다음과 같다.
Network Switch
|
+-- SNMPv3 UDP/161
| |
| +--> Telegraf
| |
| +--> InfluxDB OSS 2.x
|
+-- Syslog TCP/UDP
|
+--> rsyslog
|
+--> Log File
|
+--> Grafana Alloy
|
+--> Loki
Linux / QNAP
|
+-- node_exporter :9100
|
+--> Prometheus
Windows
|
+-- windows_exporter :9182
|
+--> Prometheus
Prometheus + InfluxDB + Loki
|
+--> Grafana
0.1 문서용 IP 주소
본 문서에서는 실제 운영 IP를 노출하지 않고 RFC 문서용 주소를 사용한다.
| 대상 | 예시 주소 | 용도 |
|---|---|---|
| Monitoring Server | 192.0.2.10 | Grafana / Prometheus / InfluxDB / Telegraf / Loki / Alloy / rsyslog |
| Linux Server | 192.0.2.20 | node_exporter |
| Windows Server | 192.0.2.30 | windows_exporter |
| QNAP NAS | 198.51.100.10 | node_exporter / SNMP / Syslog |
| Switch-01 | 203.0.113.11 | SNMP / Syslog |
| Switch-02 | 203.0.113.12 | SNMP / Syslog |
| Switch-03 | 203.0.113.13 | SNMP / Syslog |
1. Rocky Linux 기본 준비
1.1 시스템 상태 확인
cat /etc/rocky-release
uname -m
ip -br address
ip route
df -hT
free -h
getenforce
ss -lntup
SELinux와 firewalld는 초기부터 비활성화하지 않는다.
1.2 백업 디렉터리 생성
기존 서버에 추가 설치하는 경우 설정 파일 백업을 먼저 수행한다.
umask 077
MON_BACKUP="/root/monitor-backup-$(date +%Y%m%d-%H%M%S)"
install -d -m 700 "$MON_BACKUP"
cp -a /etc/rsyslog.conf "$MON_BACKUP/" 2>/dev/null || true
cp -a /etc/rsyslog.d "$MON_BACKUP/" 2>/dev/null || true
cp -a /etc/chrony.conf "$MON_BACKUP/" 2>/dev/null || true
cp -a /etc/firewalld "$MON_BACKUP/" 2>/dev/null || true
rpm -qa | sort > "$MON_BACKUP/packages-before.txt"
1.3 기본 패키지 설치
dnf install -y \
curl \
ca-certificates \
gnupg2 \
vim-enhanced \
net-snmp-utils \
rsyslog \
logrotate \
chrony \
tcpdump \
iputils \
sysstat \
policycoreutils-python-utils \
dnf-plugins-core \
jq
1.4 시간 동기화
모니터링에서는 장비와 서버 간 시간이 맞아야 장애 시각을 비교할 수 있다.
timedatectl set-timezone Asia/Seoul
systemctl enable --now chronyd
systemctl restart chronyd
chronyc tracking
chronyc sources -v
date -Ins
2. InfluxDB / Telegraf / Grafana 설치
2.1 InfluxData Repository
curl -fL https://repos.influxdata.com/influxdata-archive.key \
-o /etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata
gpg --show-keys --with-fingerprint \
/etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata
rpm --import /etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata
Repository:
vi /etc/yum.repos.d/influxdata.repo
[influxdata]
name=InfluxData Repository - Stable
baseurl=https://repos.influxdata.com/stable/$basearch/main
enabled=1
gpgcheck=1
gpgkey=file:///etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata
sslverify=1
2.2 Grafana Repository
vi /etc/yum.repos.d/grafana.repo
[grafana]
name=Grafana OSS repository
baseurl=https://rpm.grafana.com
repo_gpgcheck=1
enabled=1
gpgcheck=1
gpgkey=https://rpm.grafana.com/gpg.key
sslverify=1
2.3 패키지 설치
dnf makecache
dnf list --showduplicates \
telegraf \
influxdb2 \
influxdb2-cli \
grafana
dnf install -y \
telegraf \
influxdb2 \
influxdb2-cli \
grafana
버전 확인:
rpm -q telegraf influxdb2 influxdb2-cli grafana rsyslog
telegraf --version
influxd version
influx version
실제 구축 과정에서는 Telegraf 1.40.0 환경에서 동작을 확인하였다.
3. InfluxDB 초기 구성
3.1 Localhost Bind
InfluxDB API는 Grafana 및 Telegraf가 같은 서버에 있으므로 localhost만 수신하도록 구성한다.
install -d -m 755 /etc/systemd/system/influxdb.service.d
vi /etc/systemd/system/influxdb.service.d/10-listen.conf
[Service]
Environment="INFLUXD_HTTP_BIND_ADDRESS=127.0.0.1:8086"
systemctl daemon-reload
systemctl enable --now influxdb
systemctl restart influxdb
systemctl status influxdb --no-pager
ss -lntp | grep ':8086'
curl -fsS http://127.0.0.1:8086/health
3.2 Initial Setup
influx setup
예시:
| 설정 | 값 |
|---|---|
| Organization | network |
| Bucket | snmp_raw |
| Retention | 720h |
Bucket 확인:
influx bucket list --org network
3.3 Token 분리
서비스별 Token 권한을 분리한다.
- telegraf-write : snmp_raw Write
- grafana-read : snmp_raw Read
- Operator Token : 관리자용
운영 문서에 실제 Token 값을 기록하지 않는다.
4. Telegraf 기본 구성
4.1 환경 변수 파일
vi /etc/telegraf/monitor.env
INFLUX_WRITE_TOKEN='<WRITE_TOKEN>'
ZYXEL_SNMP_USER='<SNMP_USER>'
ZYXEL_SNMP_AUTH='<AUTH_PASSWORD>'
ZYXEL_SNMP_PRIV='<PRIV_PASSWORD>'
QNAP_SNMP_USER='<QNAP_SNMP_USER>'
QNAP_SNMP_AUTH='<QNAP_AUTH_PASSWORD>'
QNAP_SNMP_PRIV='<QNAP_PRIV_PASSWORD>'
chown root:root /etc/telegraf/monitor.env
chmod 600 /etc/telegraf/monitor.env
install -d -m 755 /etc/systemd/system/telegraf.service.d
vi /etc/systemd/system/telegraf.service.d/10-monitor-env.conf
[Service]
EnvironmentFile=/etc/telegraf/monitor.env
4.2 Main Configuration
cp -a /etc/telegraf/telegraf.conf \
/etc/telegraf/telegraf.conf.orig
vi /etc/telegraf/telegraf.conf
[agent]
interval = "30s"
round_interval = true
flush_interval = "5s"
precision = "1ms"
metric_batch_size = 1000
metric_buffer_limit = 20000
omit_hostname = false
snmp_translator = "gosmi"
[[outputs.influxdb_v2]]
urls = ["http://127.0.0.1:8086"]
token = "${INFLUX_WRITE_TOKEN}"
organization = "network"
bucket = "snmp_raw"
[[inputs.internal]]
[[inputs.cpu]]
percpu = false
totalcpu = true
[[inputs.mem]]
[[inputs.disk]]
mount_points = ["/"]
[[inputs.net]]
5. SNMPv3 사전 확인
SNMP는 v3 authPriv 사용을 기본으로 한다.
Net-SNMP 테스트 파일:
install -d -m 700 /root/.snmp
vi /root/.snmp/snmp.conf
defVersion 3
defSecurityName <SNMP_USER>
defSecurityLevel authPriv
defAuthType SHA
defAuthPassphrase <SNMP_AUTH_PASSWORD>
defPrivType AES
defPrivPassphrase <SNMP_PRIV_PASSWORD>
chmod 600 /root/.snmp/snmp.conf
snmpget -v3 -t 2 -r 0 -On \
203.0.113.11 \
.1.3.6.1.2.1.1.3.0
snmpwalk -v3 -t 2 -r 0 -On \
203.0.113.11 \
.1.3.6.1.2.1.31.1.1.1.1
ifName OID 마지막 값이 ifIndex이다.
6. Telegraf SNMP 구성
설정 파일은 장비 모델별로 분리한다.
/etc/telegraf/telegraf.d/
├── 10-zyxel-gs1900.conf
├── 11-zyxel-gs1920.conf
├── 20-zyxel-es3128.conf
├── 30-icmp_check.conf
├── 40-qnap_nas.conf
└── sflow.conf
6.1 GS1920 계열
예시 대상:
203.0.113.21
203.0.113.22
203.0.113.23
SNMPv3:
- SHA
- DES
CPU OID:
.1.3.6.1.4.1.890.1.15.3.49.1.7.0
Memory:
Total .1.3.6.1.4.1.890.1.15.3.50.1.1.1.3.1
Used .1.3.6.1.4.1.890.1.15.3.50.1.1.1.4.1
Percent .1.3.6.1.4.1.890.1.15.3.50.1.1.1.5.1
6.2 GS1900 계열
CPU:
.1.3.6.1.4.1.890.1.15.3.2.4.0
Memory:
.1.3.6.1.4.1.890.1.15.3.2.5.0
실제 구축에서는 2초 timeout / retries 0에서 응답 누락이 발생하여 다음과 같이 완화하였다.
timeout = "5s"
retries = 1
6.3 ES-3128GP
SNMPv3:
- SHA
- AES
sysObjectID:
.1.3.6.1.4.1.7800.1.190
CPU / Memory private OID는 장비에서 정상 확인되지 않아 Dashboard에서는 N/A 처리한다.
6.4 IF-MIB 주요 항목
수집 대상:
- ifName
- ifAlias
- ifSpeed
- ifHCInOctets
- ifHCOutOctets
- ifOperStatus
- ifInErrors
- ifOutErrors
- ifInDiscards
- ifOutDiscards
- Broadcast
- Multicast
누적 Counter는 Grafana에서 그대로 표시하지 않고 derivative/rate 계산 후 표시한다.
7. ICMP 수집
Telegraf ping input을 사용한다.
vi /etc/telegraf/telegraf.d/30-icmp_check.conf
예시 대상:
203.0.113.11
203.0.113.12
203.0.113.13
권장 설정:
[[inputs.ping]]
urls = [
"203.0.113.11",
"203.0.113.12",
"203.0.113.13"
]
method = "native"
count = 3
deadline = 2.0
interval = 10.0
native ping은 CAP_NET_RAW가 필요하다.
systemctl edit telegraf
[Service]
CapabilityBoundingSet=CAP_NET_RAW
AmbientCapabilities=CAP_NET_RAW
systemctl daemon-reload
systemctl restart telegraf
8. Telegraf Test 및 서비스 시작
설정 검사:
telegraf \
--config /etc/telegraf/telegraf.conf \
--config-directory /etc/telegraf/telegraf.d \
--test
실제 서비스 계정으로 확인:
systemctl daemon-reload
systemd-run \
--unit=telegraf-config-check \
--wait \
--pipe \
--collect \
-p User=telegraf \
-p Group=telegraf \
-p EnvironmentFile=/etc/telegraf/monitor.env \
-p CapabilityBoundingSet=CAP_NET_RAW \
-p AmbientCapabilities=CAP_NET_RAW \
/usr/bin/telegraf \
--config /etc/telegraf/telegraf.conf \
--config-directory /etc/telegraf/telegraf.d \
--test
서비스 시작:
systemctl enable --now telegraf
systemctl restart telegraf
systemctl status telegraf --no-pager
journalctl -u telegraf -n 100 --no-pager
9. Grafana 설치 및 InfluxDB 연결
9.1 Grafana 서비스
vi /etc/grafana/grafana.ini
내부망 운영 환경에 맞게 listen 주소를 설정한다.
예시:
[server]
http_addr = 0.0.0.0
http_port = 3000
[users]
allow_sign_up = false
[auth.anonymous]
enabled = false
systemctl enable --now grafana-server
systemctl restart grafana-server
systemctl status grafana-server --no-pager
curl -fsS http://127.0.0.1:3000/api/health
9.2 InfluxDB Datasource
Grafana:
Connections
→ Data sources
→ Add data source
→ InfluxDB
| 설정 | 값 |
|---|---|
| Query language | Flux |
| URL | http://127.0.0.1:8086 |
| Organization | network |
| Token | grafana-read |
| Default Bucket | snmp_raw |
| Min time interval | 1s |
10. SNMP Dashboard 기본 구성
10.1 Traffic
Flux:
from(bucket: "snmp_raw")
|> range(start: v.timeRangeStart, stop: v.timeRangeStop)
|> filter(fn: (r) => r._measurement == "switch_port")
|> filter(fn: (r) => r.source == "203.0.113.11")
|> filter(fn: (r) =>
r._field == "in_octets" or
r._field == "out_octets"
)
|> derivative(unit: 1s, nonNegative: true)
|> map(fn: (r) => ({
r with _value: r._value * 8.0
}))
Grafana Unit:
bits/sec
10.2 Error / Discard
|> filter(fn: (r) =>
r._field == "in_errors" or
r._field == "out_errors" or
r._field == "in_discards" or
r._field == "out_discards"
)
|> derivative(unit: 1s, nonNegative: true)
10.3 Broadcast / Multicast
Broadcast/Multicast는 별도 패널로 표시하여 폭주 및 비정상 증가를 확인한다.
11. Prometheus 설치
실제 구축에서는 Prometheus 3.13.3을 사용하였다.
11.1 사용자 및 디렉터리
useradd \
--system \
--no-create-home \
--shell /sbin/nologin \
prometheus
install -d -o prometheus -g prometheus \
/etc/prometheus \
/var/lib/prometheus
Prometheus 바이너리는 공식 릴리스에서 다운로드하고 checksum 확인 후 설치한다.
install -m 0755 prometheus /usr/local/bin/prometheus
install -m 0755 promtool /usr/local/bin/promtool
11.2 Prometheus 설정
vi /etc/prometheus/prometheus.yml
global:
scrape_interval: 30s
scrape_configs:
- job_name: prometheus
static_configs:
- targets:
- "127.0.0.1:9090"
labels:
server_name: monitoring-server
- job_name: node
static_configs:
- targets:
- "127.0.0.1:9100"
labels:
server_name: monitoring-server
- targets:
- "192.0.2.20:9100"
labels:
server_name: linux-server
- job_name: windows
static_configs:
- targets:
- "192.0.2.30:9182"
labels:
server_name: windows-server
- job_name: qnap
static_configs:
- targets:
- "198.51.100.10:9100"
labels:
server_name: nas-01
11.3 systemd
vi /etc/systemd/system/prometheus.service
[Unit]
Description=Prometheus
Wants=network-online.target
After=network-online.target
[Service]
User=prometheus
Group=prometheus
ExecStart=/usr/local/bin/prometheus \
--config.file=/etc/prometheus/prometheus.yml \
--storage.tsdb.path=/var/lib/prometheus \
--storage.tsdb.retention.time=30d \
--storage.tsdb.retention.size=20GB \
--web.listen-address=127.0.0.1:9090
Restart=always
[Install]
WantedBy=multi-user.target
chown -R prometheus:prometheus \
/etc/prometheus \
/var/lib/prometheus
promtool check config /etc/prometheus/prometheus.yml
systemctl daemon-reload
systemctl enable --now prometheus
systemctl status prometheus --no-pager
12. Node Exporter 설치
실제 구축에서는 Node Exporter 1.12.1 환경을 사용하였다.
12.1 Linux Node Exporter
useradd \
--system \
--no-create-home \
--shell /sbin/nologin \
node_exporter
install -m 0755 node_exporter \
/usr/local/bin/node_exporter
vi /etc/systemd/system/node_exporter.service
[Unit]
Description=Node Exporter
After=network-online.target
Wants=network-online.target
[Service]
User=node_exporter
Group=node_exporter
ExecStart=/usr/local/bin/node_exporter
Restart=always
[Install]
WantedBy=multi-user.target
systemctl daemon-reload
systemctl enable --now node_exporter
curl -s http://127.0.0.1:9100/metrics | head
외부 서버의 node_exporter는 Prometheus 서버 주소만 9100/tcp에 접근하도록 방화벽을 제한한다.
13. Windows Exporter
Windows Server에서는 windows_exporter를 사용한다.
기본 포트:
9182/tcp
Prometheus Target:
192.0.2.30:9182
성능 카운터 이상 시 다음 명령으로 복구한 사례가 있다.
lodctr /R
winmgmt /resyncperf
PowerShell:
Restart-Service windows_exporter
확인:
Invoke-WebRequest http://127.0.0.1:9182/metrics
14. Grafana Prometheus Datasource
Grafana:
Connections
→ Data sources
→ Prometheus
URL:
http://127.0.0.1:9090
Save & Test 후 다음 PromQL로 확인한다.
up
15. Server Dashboard 구성
Linux Dashboard:
- Uptime
- CPU Usage
- Memory Usage
- Load Average
- Filesystem
- Disk Read / Write
- Network RX / TX
- Network Error / Drop
- Service 상태
Windows Dashboard:
- CPU
- Memory
- Disk
- Network
- Uptime
- Exporter 상태
Disk 용량은 실제 GB/GiB 값으로 표시하며 불필요한 Bar Gauge는 제거한다.
16. QNAP TS-264 모니터링
QNAP은 세 종류의 데이터를 함께 사용한다.
| 데이터 | 수집 방법 |
|---|---|
| CPU / Memory / Network / Filesystem / Disk I/O | node_exporter → Prometheus |
| HDD / RAID / Storage / Volume | SNMP → Telegraf → InfluxDB |
| Event / Access Log | Syslog → rsyslog → Alloy → Loki |
16.1 QNAP SNMP
QNAP SNMPv3:
- SHA
- DES
Measurement:
QNAP_TS264
qnap_disk
qnap_raid
qnap_storage_pool
qnap_volume
Disk:
- disk_id
- manufacturer
- model
- disk_type
- disk_status
- temperature
- capacity_bytes
RAID:
- RAID ID
- RAID Name
- RAID Status
- RAID Level
- Capacity
16.2 Volume 단위 보정
QNAP qnap_volume의 capacity/free 값은 실제 장비에서 KiB 형태로 반환되는 것으로 확인되었다.
원시값을 Grafana에서 byte로 바로 해석하면 약 11.3 GiB로 잘못 표시된다.
Flux:
|> map(fn: (r) => ({
r with
capacity_bytes:
uint(v: r.capacity_bytes) * uint(v: 1024),
free_bytes:
uint(v: r.free_bytes) * uint(v: 1024),
used_bytes:
(
uint(v: r.capacity_bytes)
- uint(v: r.free_bytes)
) * uint(v: 1024),
used_percent:
if float(v: r.capacity_bytes) > 0.0 then
(
float(v: r.capacity_bytes)
- float(v: r.free_bytes)
)
/ float(v: r.capacity_bytes)
* 100.0
else 0.0
}))
16.3 실제 Data Volume
Node Exporter 기준 실제 사용자 Data Volume:
mountpoint="/share/CACHEDEV1_DATA"
device="/dev/mapper/cachedev1"
fstype="ext4"
Snapshot:
/mnt/snapshot/...
위 Snapshot 경로는 Dashboard Filesystem 패널에서 제외한다.
DataVol1 사용률:
100 * (
1 -
node_filesystem_avail_bytes{
instance="198.51.100.10:9100",
mountpoint="/share/CACHEDEV1_DATA"
}
/
node_filesystem_size_bytes{
instance="198.51.100.10:9100",
mountpoint="/share/CACHEDEV1_DATA"
}
)
16.4 Network는 bond0만 표시
RX:
rate(
node_network_receive_bytes_total{
instance="198.51.100.10:9100",
device="bond0"
}[$__rate_interval]
) * 8
TX:
rate(
node_network_transmit_bytes_total{
instance="198.51.100.10:9100",
device="bond0"
}[$__rate_interval]
) * 8
16.5 Network Error / Drop
rate(node_network_receive_errs_total{
instance="198.51.100.10:9100",
device="bond0"
}[$__rate_interval])
rate(node_network_transmit_errs_total{
instance="198.51.100.10:9100",
device="bond0"
}[$__rate_interval])
rate(node_network_receive_drop_total{
instance="198.51.100.10:9100",
device="bond0"
}[$__rate_interval])
rate(node_network_transmit_drop_total{
instance="198.51.100.10:9100",
device="bond0"
}[$__rate_interval])
16.6 Disk I/O
Read:
rate(node_disk_read_bytes_total{
instance="198.51.100.10:9100",
device=~"sd[a-z]+"
}[$__rate_interval])
Write:
rate(node_disk_written_bytes_total{
instance="198.51.100.10:9100",
device=~"sd[a-z]+"
}[$__rate_interval])
Read IOPS:
rate(node_disk_reads_completed_total{
instance="198.51.100.10:9100",
device=~"sd[a-z]+"
}[$__rate_interval])
Write IOPS:
rate(node_disk_writes_completed_total{
instance="198.51.100.10:9100",
device=~"sd[a-z]+"
}[$__rate_interval])
17. rsyslog 구성
17.1 Network Syslog
예시 로그 파일:
/var/log/network-syslog/events.log
네트워크 장비:
UDP/TCP 514
17.2 Server Syslog
/var/log/server-syslog/events.log
예시 포트:
5514/tcp
17.3 QNAP Event / Access 분리
QNAP 로그는 Event와 Access를 별도 포트로 분리한다.
| 로그 | Port | File |
|---|---|---|
| Event | TCP 5515 | /var/log/qnap/event.log |
| Access | TCP 5516 | /var/log/qnap/access.log |
Template:
template(name="QnapSyslogLine" type="string"
string="%timegenerated:::date-rfc3339% src=%fromhost-ip% severity=%syslogseverity-text% host=%hostname% %syslogtag%%msg:::sp-if-no-1st-sp%%msg%\n")
Event:
ruleset(name="QnapEventLog") {
action(
type="omfile"
file="/var/log/qnap/event.log"
template="QnapSyslogLine"
fileOwner="root"
fileGroup="alloy"
fileCreateMode="0640"
dirOwner="root"
dirGroup="alloy"
dirCreateMode="0750"
createDirs="on"
)
stop
}
input(
type="imtcp"
port="5515"
ruleset="QnapEventLog"
)
Access:
ruleset(name="QnapAccessLog") {
action(
type="omfile"
file="/var/log/qnap/access.log"
template="QnapSyslogLine"
fileOwner="root"
fileGroup="alloy"
fileCreateMode="0640"
dirOwner="root"
dirGroup="alloy"
dirCreateMode="0750"
createDirs="on"
)
stop
}
input(
type="imtcp"
port="5516"
ruleset="QnapAccessLog"
)
검증:
rsyslogd -N1
systemctl restart rsyslog
ss -lntp | grep -E ':5514|:5515|:5516'
18. Loki / Alloy 구성
Loki Endpoint:
http://127.0.0.1:3100
Alloy UI:
127.0.0.1:12345
18.1 Alloy 기본 설정
logging {
level = "info"
}
18.2 Network Syslog
loki.source.file "network_syslog" {
targets = [
{
__path__ = "/var/log/network-syslog/events.log",
job = "network-syslog",
},
]
forward_to = [loki.process.network_syslog.receiver]
}
loki.process "network_syslog" {
stage.regex {
expression = `^(?P<received_at>\S+) src=(?P<device_ip>\S+) severity=(?P<severity>\S+) host=(?P<device_host>\S+) (?P<message>.*)$`
}
stage.timestamp {
source = "received_at"
format = "RFC3339Nano"
action_on_failure = "skip"
}
stage.labels {
values = {
device_ip = "",
severity = "",
}
}
forward_to = [loki.write.local.receiver]
}
18.3 Server Syslog
loki.source.file "server_syslog" {
targets = [
{
__path__ = "/var/log/server-syslog/events.log",
job = "server-syslog",
},
]
forward_to = [loki.process.server_syslog.receiver]
}
loki.process "server_syslog" {
stage.regex {
expression = `^(?P<received_at>\S+) src=(?P<server_ip>\S+) severity=(?P<severity>\S+) host=(?P<server_host>\S+) (?P<source>[^:\s\[]+)(?:\[\d+\])?:?\s+(?P<message>.*)$`
}
stage.timestamp {
source = "received_at"
format = "RFC3339Nano"
action_on_failure = "skip"
}
stage.labels {
values = {
server_ip = "",
server_host = "",
severity = "",
source = "",
}
}
forward_to = [loki.write.local.receiver]
}
18.4 QNAP Event
loki.source.file "qnap_event" {
targets = [
{
__path__ = "/var/log/qnap/event.log",
job = "qnap-event",
},
]
forward_to = [loki.process.qnap_event.receiver]
}
loki.process "qnap_event" {
stage.regex {
expression = `^(?P<received_at>\S+) src=(?P<nas_ip>\S+) severity=(?P<severity>\S+) host=(?P<nas_host>\S+) (?P<message>.*)$`
}
stage.timestamp {
source = "received_at"
format = "RFC3339Nano"
action_on_failure = "skip"
}
stage.labels {
values = {
nas_ip = "",
nas_host = "",
severity = "",
log_type = "event",
}
}
forward_to = [loki.write.local.receiver]
}
18.5 QNAP Access
loki.source.file "qnap_access" {
targets = [
{
__path__ = "/var/log/qnap/access.log",
job = "qnap-access",
},
]
forward_to = [loki.process.qnap_access.receiver]
}
loki.process "qnap_access" {
stage.regex {
expression = `^(?P<received_at>\S+) src=(?P<nas_ip>\S+) severity=(?P<severity>\S+) host=(?P<nas_host>\S+) (?P<message>.*)$`
}
stage.timestamp {
source = "received_at"
format = "RFC3339Nano"
action_on_failure = "skip"
}
stage.labels {
values = {
nas_ip = "",
nas_host = "",
severity = "",
log_type = "access",
}
}
forward_to = [loki.write.local.receiver]
}
18.6 Loki Write
loki.write "local" {
endpoint {
url = "http://127.0.0.1:3100/loki/api/v1/push"
}
}
검증:
alloy validate /etc/alloy/config.alloy
systemctl restart alloy
systemctl status alloy --no-pagerLoki Job 확인:
curl -s \
'http://127.0.0.1:3100/loki/api/v1/label/job/values' \
| jq정상 예:
network-syslog
server-syslog
qnap-event
qnap-access18.7 Alloy Permission 문제
다음과 같은 오류가 발생할 수 있다.
failed to tail file
stat failed
permission denied확인:
namei -l /var/log/qnap/access.log
systemctl show alloy \
-p User \
-p Group
sudo -u alloy \
head /var/log/qnap/access.log권장 권한:
drwxr-x--- root alloy /var/log/qnap
-rw-r----- root alloy /var/log/qnap/access.log
-rw-r----- root alloy /var/log/qnap/event.logSELinux 확인:
getenforce
ausearch -m AVC -ts recent \
| grep -Ei 'alloy|qnap'
ls -Zd /var/log/qnap
ls -Z /var/log/qnap/access.log19. Loki Query
Network:
{job="network-syslog"}Server:
{job="server-syslog"}QNAP Event:
{job="qnap-event"}QNAP Access:
{job="qnap-access"}Critical Network Syslog:
{job="network-syslog",severity=~"emerg|alert|crit|err"}20. Grafana Syslog Alert
Syslog Level 3 이상:
sum by (device_ip, severity) (
count_over_time(
{
job="network-syslog",
severity=~"emerg|alert|crit|err"
}[1m]
)
)권장 Alert 구조:
A = Loki Instant Query
B = Threshold
A IS ABOVE 0Summary:
[Syslog 경고] {{ $labels.device_ip }} - {{ $labels.severity }}Description:
장비 {{ $labels.device_ip }} 에서 Syslog Level 3(Error) 이상의 로그가 발생했습니다.
최근 1분 발생 건수: {{ $values.A.Value }}
Severity: {{ $labels.severity }}Range Query를 그대로 Alert 조건으로 사용할 경우 다음 오류가 발생할 수 있다.
invalid format of evaluation results for the alert definition A:
looks like time series data, only reduced data can be alerted on.Range Query를 유지할 경우 Reduce Expression을 추가해야 한다.
21. ICMP Alert
Flux:
from(bucket: "snmp_raw")
|> range(start: -5m)
|> filter(fn: (r) =>
r._measurement == "ping" and
r._field == "percent_packet_loss"
)
|> group(columns: ["url"])
|> last()
|> keep(columns: ["_time", "_value", "url"])Alert:
A = Flux Query
B = Reduce / Last / Strict
C = Threshold > 99
Pending = 2mNo Data:
Keep Last State22. 최종 Dashboard 구성안
22.1 Network Switch Dashboard
| Row | Panels |
|---|---|
| 상태 | Device / Uptime / CPU / Memory / Ping |
| Port 상태 | ifName / Alias / Speed / OperStatus |
| Traffic | RX bps / TX bps |
| Packet | Unicast / Broadcast / Multicast PPS |
| Error | RX/TX Error / Discard |
| Syslog | Warning / Error / Critical Event |
22.2 Linux Server Dashboard
- Uptime
- CPU
- Memory
- Load
- Disk Capacity
- Disk I/O
- Network RX/TX
- Network Error/Drop
- Service Log
- Security Event
22.3 Windows Dashboard
- Uptime
- CPU
- Memory
- Logical Disk
- Disk I/O
- Network
- Windows Exporter 상태
22.4 QNAP Dashboard
최종 구성:
Node Exporter
Uptime
CPU
Memory
HDD 최고 온도
CPU / Memory / Load
Network RX / TX
- bond0 only
Disk 상태
Volume
- DataVol1
- SNMP capacity/free x1024 보정
Filesystem 사용률
- /share/CACHEDEV1_DATA only
- Snapshot 제외
Network Errors / Drops
- bond0 RX Error
- bond0 TX Error
- bond0 RX Drop
- bond0 TX Drop
Disk I/O
- Physical sd* only
- Read B/s
- Write B/s
- Read IOPS
- Write IOPS
QNAP Event Log
QNAP Access LogQNAP Dashboard에서 제거한 항목:
- 상단 RAID 상태
- Storage Pool 상태
- Storage Pool 사용률
- RAID 상세 Table
- Storage Pool 상세 Table
RAID/Storage Pool 정보는 필요 시 별도 상세 Dashboard에서 조회한다.
23. 운영 점검
전체 서비스:
systemctl is-active \
influxdb \
telegraf \
grafana-server \
prometheus \
loki \
alloy \
rsyslog \
chronydListening Port:
ss -lntupPrometheus:
curl -s http://127.0.0.1:9090/-/healthyLoki:
curl -s http://127.0.0.1:3100/readyTelegraf:
journalctl -u telegraf \
--since '-10 min' \
--no-pagerAlloy:
journalctl -u alloy \
--since '-10 min' \
--no-pagerFilesystem:
df -hTSystem I/O:
vmstat 1 5
iostat -xz 1 524. 장애 점검 순서
Prometheus Target Down
curl http://TARGET_IP:9100/metrics
journalctl -u prometheus \
-n 100 \
--no-pagerSNMP Timeout
snmpget -v3 \
-t 5 \
-r 1 \
-On \
SWITCH_IP \
.1.3.6.1.2.1.1.3.0확인 항목:
- SNMP User
- SHA/DES/AES 조합
- ACL
- Source IP
- Timeout
- SNMP View
- 장비 CPU
Syslog 미수신
ss -lntup \
| grep -E ':514|:5514|:5515|:5516'
tcpdump -ni any \
host DEVICE_IP
tail -f /var/log/qnap/access.logLoki 미표시
alloy validate \
/etc/alloy/config.alloy
journalctl -u alloy \
-n 100 \
--no-pager
curl -s \
'http://127.0.0.1:3100/loki/api/v1/label/job/values' \
| jq25. 운영 원칙
- SNMP Metric은 Telegraf → InfluxDB로 저장한다.
- Host Metric은 Prometheus로 저장한다.
- Syslog는 rsyslog → Alloy → Loki로 저장한다.
- Grafana는 세 데이터소스를 통합한다.
- SNMP Counter는 누적값을 그대로 표시하지 않고 rate/derivative 계산한다.
- No Data를 0 또는 정상으로 강제 표시하지 않는다.
- QNAP Snapshot filesystem은 Data Volume 사용률에서 제외한다.
- QNAP Volume SNMP capacity/free 값은 장비 특성상 ×1024 보정한다.
- QNAP Network는 실제 활성 bond0만 표시한다.
- Loki message 전체를 label로 만들지 않는다.
- 장비 Syslog Alert는 device_ip / severity와 발생 건수를 메일에 포함한다.
- Alert의 Range Query는 Reduce 또는 Instant Query 구조로 구성한다.
- 모든 Token 및 SNMP 비밀번호는 문서에 실제 값을 기록하지 않는다.