SNMP테스트서버: 두 판 사이의 차이

잉여 위키
둘러보기로 이동 검색으로 이동
편집 요약 없음
편집 요약 없음
1번째 줄: 1번째 줄:
= Rocky Linux 9 · Zyxel 스위치 IPT 장애 모니터링 서버 구성 가이드 =
= Rocky Linux 9 통합 모니터링 서버 구축 가이드 =
: Telegraf / InfluxDB / Grafana / Prometheus / Node Exporter / Windows Exporter / Loki / Alloy / rsyslog


작성일: 2026-09-10 · 구성안 v1.0
작성 기준: 실제 구성 및 장애 처리 기록 기반
문서 형식: MediaWiki
IP 주소: 문서용 예시 주소로 치환
운영환경 적용 전 실제 장비 주소, 계정, 토큰, OID를 확인한다.


== 0. 목적·범위·검증 수준 ==
__TOC__


'''목적:''' Zyxel L3/L2 → PoE 전화기 → PC 환경에서 전화 연결 실패 시 네트워크 혼잡·패킷 폐기·링크/PoE 이벤트가 동반되는지 기록한다. 콜서버의 TCP/UDP 5060 tcpdump는 별도로 수행하며 이 문서는 해당 설정을 변경하지 않는다.
== 0. 목적 및 최종 구성 ==


* [확정] KVM VM / Rocky Linux 9, Zyxel L3/L2, 전화기 뒤 PC 연결, 콜서버 SIP 캡처 병행 계획.
본 문서는 Rocky Linux 9 기반 모니터링 서버를 처음 구축하는 단계부터 네트워크 장비 SNMP, 서버 Metric, Syslog, QNAP NAS 모니터링 및 Grafana Dashboard 구성까지 정리한다.
* [미확인] 스위치 모델·펌웨어·대수·주소·Voice VLAN·콜서버 경로, SNMPv3 알고리즘, sFlow 지원, CPU/PoE 전용 OID.
* [제안] 먼저 '''Telegraf + InfluxDB OSS 2.x + Grafana OSS + rsyslog'''를 RPM/systemd로 직접 설치한다. 전체 포트 30초, 의심 포트/업링크 5초에서 시작하고 필요한 포트만 1초로 단축한다.
* '''sFlow는 조건부 단계:''' Akvorado가 후보지만 모델의 exporter 지원과 배포 릴리스 설정이 확정되지 않았으므로 §12는 준비·배포 검토·수신 검증 절차다. 아직 운영 기동을 보장하는 완성 설정이 아니다. SNMP/Syslog 기본 구성은 독립적으로 진행 가능하다.
* 실제 Rocky VM·Zyxel 장비에서 실행 시험한 문서는 아니다. 구성 예시를 정적 점검했으며 각 단계의 현장 검증을 통과한 후 진행한다.


__TOC__
최종 구조는 다음과 같다.
 
<syntaxhighlight lang="bash" line>
Network Switch
    |
    +-- SNMPv3 UDP/161
    |      |
    |      +--> Telegraf
    |              |
    |              +--> InfluxDB OSS 2.x
    |
    +-- Syslog TCP/UDP
          |
          +--> rsyslog
                    |
                    +--> Log File
                            |
                            +--> Grafana Alloy
                                      |
                                      +--> Loki
 
Linux / QNAP
    |
    +-- node_exporter :9100
            |
            +--> Prometheus


== 1. 구조와 준비값 ==
Windows
    |
    +-- windows_exporter :9182
            |
            +--> Prometheus


{| class="wikitable"
Prometheus + InfluxDB + Loki
|-
            |
! 서비스
            +--> Grafana
! 설치 방식
</syntaxhighlight>
! 연결/저장
|-
| Telegraf 1.x
| RPM / telegraf.service
| 스위치 UDP/161 조회 → 로컬 InfluxDB
|-
| InfluxDB OSS 2.x
| RPM / influxdb.service
| 127.0.0.1:8086, 30일 raw bucket
|-
| Grafana OSS
| RPM / grafana-server.service
| 127.0.0.1:3000, SSH 터널 접속
|-
| rsyslog 8.x
| Rocky RPM / rsyslog.service
| UDP/514 → /var/log/zyxel/&lt;송신IP&gt;/events.log
|-
| Akvorado 배포 스택
| 별도 검토 후 공식 Docker Compose
| sFlow UDP/6343 → ClickHouse 등
|}


SNMP Trap UDP/162는 이번 기본 구성에 포함하지 않는다. Rocky를 조회하는 SNMP agent인 <code>snmpd</code>도 불필요하다. 이 서버는 SNMP 조회자다.
=== 0.1 문서용 IP 주소 ===


=== 준비표 — 모든 <code>&lt;...&gt;</code>를 실제 값으로 교체 ===
본 문서에서는 실제 운영 IP를 노출하지 않고 RFC 문서용 주소를 사용한다.


{| class="wikitable"
{| class="wikitable"
! 대상
! 예시 주소
! 용도
|-
|-
! 자리표시자
| Monitoring Server
! 의미
| 192.0.2.10
| Grafana / Prometheus / InfluxDB / Telegraf / Loki / Alloy / rsyslog
|-
|-
| <code>&lt;COLLECTOR_IP&gt;</code>
| Linux Server
| Rocky VM 고정 IPv4
| 192.0.2.20
| node_exporter
|-
|-
| <code>&lt;FW_ZONE&gt;</code>
| Windows Server
| 관리 NIC에 실제 적용된 firewalld zone
| 192.0.2.30
| windows_exporter
|-
|-
| <code>&lt;SWITCH_SOURCE_CIDR&gt;</code>
| QNAP NAS
| 실제 Syslog 송신 IP /32 또는 승인된 스위치 관리 대역
| 198.51.100.10
| node_exporter / SNMP / Syslog
|-
|-
| <code>&lt;SWITCH_IP&gt;</code> / <code>&lt;SWITCH_NAME&gt;</code>
| Switch-01
| 수집 대상 관리 IP / 식별 이름
| 203.0.113.11
| SNMP / Syslog
|-
|-
| <code>&lt;IFINDEX&gt;</code> / <code>&lt;PORT_LABEL&gt;</code>
| Switch-02
| SNMP 조회로 확인한 인덱스 / 포트명
| 203.0.113.12
| SNMP / Syslog
|-
|-
| <code>&lt;CALLSERVER_IP&gt;</code>
| Switch-03
| 콜서버 IP
| 203.0.113.13
|-
| SNMP / Syslog
| <code>&lt;SNMP_USER&gt;</code> 등
| 읽기 전용 SNMPv3 계정·인증값
|-
| <code>&lt;NTP_SERVER&gt;</code>
| 콜서버/스위치와 맞춘 승인된 시간 서버
|}
|}


소규모 시작 가정: 기본 구성 4 vCPU / RAM 8GB / SSD 100~200GB. Akvorado 동시 운영은 8 vCPU / RAM 16~32GB / SSD 300~500GB부터 부하 시험한다. 이는 보장 사양이 아니며 포트 수·샘플 PPS·저장량으로 조정한다. KVM CPU 과할당, 메모리 ballooning 및 호스트 디스크 포화를 피한다.
== 1. Rocky Linux 기본 준비 ==
 
=== 1.1 시스템 상태 확인 ===


VM은 VirtIO NIC를 KVM 관리 브리지에 연결하고, 스위치까지 라우팅 가능한 고정 IP를 사용한다. SNMP/sFlow에는 미러 포트나 promiscuous 설정이 필요 없다. 같은 VLAN이 필수도 아니다. NAT는 SNMP 접근제어와 sFlow 송신자 식별을 복잡하게 하므로 가능한 관리망 브리지/라우팅 연결을 사용한다.
<syntaxhighlight lang="bash" line>
cat /etc/rocky-release
uname -m
ip -br address
ip route
df -hT
free -h
getenforce
ss -lntup
</syntaxhighlight>


=== 관측 범위 한계 ===
SELinux와 firewalld는 초기부터 비활성화하지 않는다.


* 전화기 연결 물리 포트의 카운터는 전화기와 PC 합산이다. IF-MIB만으로 Voice/Data를 분리할 수 없다.
=== 1.2 백업 디렉터리 생성 ===
* 같은 L2 내부에서 끝나는 트래픽은 L3만 관측해서는 보이지 않을 수 있다.
* 수집 경로가 폭주 구간과 겹치면 장애 때 데이터도 빠진다. OOB가 없으면 코어에 가까운 안정적 관리 경로를 확보한다.
* 1초 SNMP는 1초 구간 평균이며 장비 내부 카운터 갱신이 1초보다 느릴 수 있다. Microburst 부재를 증명하지 못한다.


== 2. 변경 영향·백업 ==
기존 서버에 추가 설치하는 경우 설정 파일 백업을 먼저 수행한다.


전제는 '''신규 전용 VM'''이다. 기존 운영 서버에 적용한다면 같은 이름의 설정을 덮어쓰지 말고 병합한다. 기존 Telegraf 입력/출력, rsyslog 입력 포트, Grafana, InfluxDB 데이터가 있으면 본 초기화 절차를 중단한다.
<syntaxhighlight lang="bash" line>
umask 077


위험 및 롤백 원칙:
MON_BACKUP="/root/monitor-backup-$(date +%Y%m%d-%H%M%S)"


* 패키지 설치 자체로 전화 경로를 변경하지 않는다. 다만 고주기 SNMP/sFlow는 스위치 관리 CPU와 저장량을 증가시킬 수 있다.
install -d -m 700 "$MON_BACKUP"
* rsyslog 재시작 동안 로그 수신에 짧은 공백이 발생할 수 있다.
* 보존기간 만료/로그 회전은 오래된 데이터를 자동 제거한다. 장애 증거는 만료 전에 별도 보관한다.
* 방화벽은 기존 zone/SSH를 유지하고 특정 송신자 규칙만 추가한다. 실패하면 추가한 규칙만 제거한다.
* <code>firewall-cmd --complete-reload</code>는 설정 백업 복원이 아니다. 사용하지 않는다.
* VM 스냅샷은 단기 변경 복구용이며 별도 백업을 대체하지 않는다. 복원하면 스냅샷 이후 수집 데이터가 사라질 수 있다.


root 셸에서 실행한다. 아래 백업 변수는 이 작업용이며 동일 셸에서 유지한다.
cp -a /etc/rsyslog.conf "$MON_BACKUP/" 2>/dev/null || true
cp -a /etc/rsyslog.d "$MON_BACKUP/" 2>/dev/null || true
cp -a /etc/chrony.conf "$MON_BACKUP/" 2>/dev/null || true
cp -a /etc/firewalld "$MON_BACKUP/" 2>/dev/null || true


<syntaxhighlight lang="bash" line>
umask 077
MON_BACKUP="/root/ipt-monitor-backup-$(date +%Y%m%d-%H%M%S)"
install -d -m 700 "$MON_BACKUP"
cp -a /etc/rsyslog.conf "$MON_BACKUP/"
cp -a /etc/rsyslog.d "$MON_BACKUP/"
cp -a /etc/chrony.conf "$MON_BACKUP/"
cp -a /etc/firewalld "$MON_BACKUP/"
firewall-cmd --list-all-zones > "$MON_BACKUP/firewall-runtime.txt"
firewall-cmd --permanent --list-all-zones > "$MON_BACKUP/firewall-permanent.txt"
rpm -qa | sort > "$MON_BACKUP/packages-before.txt"
rpm -qa | sort > "$MON_BACKUP/packages-before.txt"
</syntaxhighlight>
</syntaxhighlight>
firewalld가 미설치/미실행이면 해당 명령 실패를 무시한 채 진행하지 말고, 현재 호스트 방화벽 구성과 SSH 허용부터 확인한다. KVM 콘솔을 확보한다. 스위치 설정도 GUI의 configuration backup으로 별도 내려받는다.


== 3. 기본 패키지·시간 동기화 ==
=== 1.3 기본 패키지 설치 ===


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
cat /etc/rocky-release
dnf install -y \
uname -m
    curl \
ip -br address
    ca-certificates \
ip route
    gnupg2 \
df -hT
    vim-enhanced \
free -h
    net-snmp-utils \
getenforce
    rsyslog \
ss -lntup
    logrotate \
    chrony \
    tcpdump \
    iputils \
    sysstat \
    policycoreutils-python-utils \
    dnf-plugins-core \
    jq
</syntaxhighlight>


dnf install -y curl ca-certificates gnupg2 vim-enhanced \
=== 1.4 시간 동기화 ===
    net-snmp-utils rsyslog logrotate chrony tcpdump iputils \
    sysstat policycoreutils-python-utils dnf-plugins-core
</syntaxhighlight>
네트워크·SELinux는 일괄 비활성화하지 않는다. 신규 OS 업데이트/재부팅은 수집 시작 전 점검 시간에 수행한다.


<code>/etc/chrony.conf</code>의 기존 시간 서버 정책을 확인하고, 승인된 서버로 맞춘다. 다음 한 줄은 설정 파일 예시이며 셸 명령이 아니다.
모니터링에서는 장비와 서버 간 시간이 맞아야 장애 시각을 비교할 수 있다.


<syntaxhighlight lang="bash" line>
server <NTP_SERVER> iburst
</syntaxhighlight>
<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
timedatectl set-timezone Asia/Seoul
timedatectl set-timezone Asia/Seoul
systemctl enable --now chronyd
systemctl enable --now chronyd
systemctl restart chronyd
systemctl restart chronyd
chronyc tracking
chronyc tracking
chronyc sources -v
chronyc sources -v
date -Ins
date -Ins
</syntaxhighlight>
</syntaxhighlight>
정상 기준: 선택된 소스 <code>^*</code>, 동기화 정상, 콜서버 및 스위치와 시간 차이가 분석 해상도보다 충분히 작음. 1초 분석은 가능하면 100ms 미만 오차를 목표로 한다. 큰 시간 보정은 수집 시작 전에 한다.


== 4. 공식 저장소와 패키지 설치 ==
== 2. InfluxDB / Telegraf / Grafana 설치 ==


이 가이드는 '''InfluxDB OSS 2.x / Flux'''용이다. InfluxDB 3의 SQL 설정과 혼용하지 않는다. 최신 버전이라는 의미가 아니라 본 가이드의 호환 대상이다. 저장소에서 설치 가능한 정확한 RPM 버전을 먼저 확인·기록한다.
=== 2.1 InfluxData Repository ===
 
공식 설치 근거: [https://docs.influxdata.com/influxdb/v2/install/?t=Linux InfluxDB 2 RPM 설치], [https://docs.influxdata.com/telegraf/v1/install/ Telegraf RPM 설치], [https://grafana.com/docs/grafana/latest/setup-grafana/installation/redhat-rhel-fedora/ Grafana OSS RPM 설치]. 지속 갱신 문서, 확인일 2026-09-10.
 
=== 4.1 InfluxData 저장소 ===


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
curl -fL https://repos.influxdata.com/influxdata-archive.key \
curl -fL https://repos.influxdata.com/influxdata-archive.key \
     -o /etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata
     -o /etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata
gpg --show-keys --with-fingerprint /etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata
 
gpg --show-keys --with-fingerprint \
    /etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata
 
rpm --import /etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata
</syntaxhighlight>
</syntaxhighlight>
확인한 공식 문서의 기본키 fingerprint: <code>24C975CBA61A024EE1B631787C3D57159FC2F927</code>. 일치하지 않으면 키 교체 공지를 공식 문서에서 확인하기 전에는 import하지 않는다. <code>gpgcheck=0</code> 또는 <code>--nogpgcheck</code>로 우회하지 않는다.
 
Repository:


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
rpm --import /etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata
vi /etc/yum.repos.d/influxdata.repo
vi /etc/yum.repos.d/influxdata.repo
</syntaxhighlight>
</syntaxhighlight>
파일 내용 (<code>$basearch</code>는 DNF가 해석하므로 그대로 둔다):


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
181번째 줄: 189번째 줄:
sslverify=1
sslverify=1
</syntaxhighlight>
</syntaxhighlight>
=== 4.2 Grafana 저장소 ===
 
=== 2.2 Grafana Repository ===


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
vi /etc/yum.repos.d/grafana.repo
vi /etc/yum.repos.d/grafana.repo
</syntaxhighlight>
</syntaxhighlight>
<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
[grafana]
[grafana]
196번째 줄: 206번째 줄:
sslverify=1
sslverify=1
</syntaxhighlight>
</syntaxhighlight>
Grafana 키 import 요청 시 공식 키/문서와 대조한다. 패키지는 <code>grafana</code>이며 <code>grafana-enterprise</code>를 설치하지 않는다.
 
=== 2.3 패키지 설치 ===


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
dnf makecache
dnf makecache
dnf list --showduplicates telegraf influxdb2 influxdb2-cli grafana
 
dnf install telegraf influxdb2 influxdb2-cli grafana
dnf list --showduplicates \
    telegraf \
    influxdb2 \
    influxdb2-cli \
    grafana
 
dnf install -y \
    telegraf \
    influxdb2 \
    influxdb2-cli \
    grafana
</syntaxhighlight>
 
버전 확인:
 
<syntaxhighlight lang="bash" line>
rpm -q telegraf influxdb2 influxdb2-cli grafana rsyslog
rpm -q telegraf influxdb2 influxdb2-cli grafana rsyslog
telegraf --version
telegraf --version
influxd version
influxd version
influx version
influx version
</syntaxhighlight>
</syntaxhighlight>
<code>influxdb2-cli</code>가 저장소에 없다면 임의의 다른 CLI를 설치하지 말고 공식 InfluxDB 2 CLI 설치 방법으로 해당 아키텍처의 서명/체크섬 검증된 패키지를 설치한다. 서버/CLI의 버전 번호는 서로 같지 않을 수 있다.


이후 원격 운영 전 위 버전 출력을 작업 기록에 보관한다. Telegraf/Grafana는 검증된 버전을 변경 관리 대상으로 관리하고 장애 조사 중 자동 업그레이드는 하지 않는다.
실제 구축 과정에서는 Telegraf 1.40.0 환경에서 동작을 확인하였다.


== 5. InfluxDB 초기 설정 및 권한 분리 ==
== 3. InfluxDB 초기 구성 ==


=== 5.1 로컬에서만 API 수신 ===
=== 3.1 Localhost Bind ===
 
InfluxDB API는 Grafana 및 Telegraf가 같은 서버에 있으므로 localhost만 수신하도록 구성한다.


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
install -d -m 755 /etc/systemd/system/influxdb.service.d
install -d -m 755 /etc/systemd/system/influxdb.service.d
vi /etc/systemd/system/influxdb.service.d/10-ipt-listen.conf
 
vi /etc/systemd/system/influxdb.service.d/10-listen.conf
</syntaxhighlight>
</syntaxhighlight>
<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
[Service]
[Service]
Environment="INFLUXD_HTTP_BIND_ADDRESS=127.0.0.1:8086"
Environment="INFLUXD_HTTP_BIND_ADDRESS=127.0.0.1:8086"
</syntaxhighlight>
</syntaxhighlight>
<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
systemctl daemon-reload
systemctl daemon-reload
systemctl enable --now influxdb
systemctl enable --now influxdb
systemctl restart influxdb
systemctl restart influxdb
systemctl status influxdb --no-pager
systemctl status influxdb --no-pager
ss -lntp | grep ':8086'
ss -lntp | grep ':8086'
curl -fsS http://127.0.0.1:8086/health
curl -fsS http://127.0.0.1:8086/health
</syntaxhighlight>
</syntaxhighlight>
정상: <code>127.0.0.1:8086</code>에만 수신, health 정상. 전체 주소로 열리면 서비스/설정의 기존 바인딩 override를 확인한다. VM 방화벽에 8086을 열지 않는다.


=== 5.2 초기화 ===
=== 3.2 Initial Setup ===


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
influx setup
influx setup
</syntaxhighlight>
</syntaxhighlight>
대화형 입력값:
 
예시:


{| class="wikitable"
{| class="wikitable"
|-
! 설정
! 항목
! 값
! 입력
|-
| Username
| 전용 관리자 계정
|-
| Password
| 고유한 강한 암호
|-
|-
| Organization
| Organization
| <code>network</code>
| network
|-
|-
| Bucket
| Bucket
| <code>snmp_raw</code>
| snmp_raw
|-
|-
| Retention
| Retention
| <code>720h</code> (30일)
| 720h
|}
|}


이미 setup 완료로 표시되면 재초기화하지 말고 기존 조직·bucket을 확인한다. <code>influx bucket list</code>로 <code>snmp_raw</code> retention이 무제한이 아닌 720h인지 검증한다. 별도 테스트 환경이 아닌 기존 DB를 삭제해서 다시 시작하지 않는다.
Bucket 확인:


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
influx bucket list --org network
influx bucket list --org network
</syntaxhighlight>
</syntaxhighlight>
=== 5.3 관리자 접속과 서비스 토큰 ===


관리자 PC에서 SSH 터널을 유지한다. <code>127.0.0.1</code>은 PC 자신의 주소이며 SSH가 VM으로 전달한다.
=== 3.3 Token 분리 ===


<syntaxhighlight lang="bash" line>
서비스별 Token 권한을 분리한다.
ssh -N -L 8086:127.0.0.1:8086 -L 3000:127.0.0.1:3000 <SSH_USER>@<COLLECTOR_IP>
</syntaxhighlight>
브라우저에서 <code>http://127.0.0.1:8086</code> 접속 → API Tokens → Custom Token:


{| class="wikitable"
* telegraf-write : snmp_raw Write
|-
* grafana-read : snmp_raw Read
! 토큰 이름
* Operator Token : 관리자용
! 범위
|-
| <code>telegraf-write</code>
| <code>snmp_raw</code> Write만
|-
| <code>grafana-read</code>
| <code>snmp_raw</code> Read만
|-
| 초기 Operator 토큰
| 관리자 초기화·백업용, 서비스에 배포 금지
|}


발급된 토큰은 승인된 비밀 관리 위치에 보관한다. 화면/문서/채팅에 붙여넣지 않는다. 초기 CLI credential도 비밀이며 root만 읽게 유지한다. Grafana 연결에 필요한 읽기 권한과 Telegraf 쓰기 권한을 분리한다.
운영 문서에 실제 Token 값을 기록하지 않는다.


== 6. SNMP 및 MIB 사전 확인 ==
== 4. Telegraf 기본 구성 ==


=== 6.1 Zyxel 측 선행 작업 ===
=== 4.1 환경 변수 파일 ===


모델별 UI/CLI가 다르므로 임의 명령을 넣지 않는다. 관리 화면에서 다음 항목을 구성한다.
<syntaxhighlight lang="bash" line>
vi /etc/telegraf/monitor.env
</syntaxhighlight>


* SNMP 활성화, SNMPv3 읽기 전용 사용자.
<syntaxhighlight lang="bash" line>
* 인증/암호화 알고리즘은 장비와 Rocky가 함께 지원하는 조합. 아래 예시는 SHA/AES이며 다른 알고리즘이면 양쪽을 같이 수정.
INFLUX_WRITE_TOKEN='<WRITE_TOKEN>'
* 조회 허용 원본은 Rocky 관리 IP. NAT가 있다면 장비에 보이는 실제 원본 IP 확인.
* IF-MIB/system 조회 권한. 빈 테이블을 기능 미지원으로 단정하기 전 SNMP view 확인.
* 포트 description에는 전화기 내선/위치나 업링크 상대를 운영 규칙에 따라 기록하되 개인정보 최소화.


=== 6.2 CLI credential을 명령행에 노출하지 않기 ===
ZYXEL_SNMP_USER='<SNMP_USER>'
ZYXEL_SNMP_AUTH='<AUTH_PASSWORD>'
ZYXEL_SNMP_PRIV='<PRIV_PASSWORD>'


아래는 root 조회용 Net-SNMP 사용자 설정이다. 기존 파일이 있으면 백업 후 병합한다. Telegraf는 이 파일을 사용하지 않는다.
QNAP_SNMP_USER='<QNAP_SNMP_USER>'
 
QNAP_SNMP_AUTH='<QNAP_AUTH_PASSWORD>'
<syntaxhighlight lang="bash" line>
QNAP_SNMP_PRIV='<QNAP_PRIV_PASSWORD>'
install -d -m 700 /root/.snmp
vi /root/.snmp/snmp.conf
</syntaxhighlight>
<syntaxhighlight lang="bash" line>
defVersion 3
defSecurityName <SNMP_USER>
defSecurityLevel authPriv
defAuthType SHA
defAuthPassphrase <SNMP_AUTH_PASSWORD>
defPrivType AES
defPrivPassphrase <SNMP_PRIV_PASSWORD>
</syntaxhighlight>
</syntaxhighlight>
<syntaxhighlight lang="bash" line>
chmod 600 /root/.snmp/snmp.conf
snmpget -v3 -t 2 -r 0 -On <SWITCH_IP> .1.3.6.1.2.1.1.3.0
snmpwalk -v3 -t 2 -r 0 -On <SWITCH_IP> .1.3.6.1.2.1.31.1.1.1.1
snmpwalk -v3 -t 2 -r 0 -On <SWITCH_IP> .1.3.6.1.2.1.31.1.1.1.18
</syntaxhighlight>
<code>ifName</code> OID의 마지막 숫자가 ifIndex다. 실제 포트 번호와 같다고 가정하지 않는다. 스위치 IP·포트명·ifIndex·전화기 IP·내선을 기록하고 재부팅/스택 변경 후 재확인한다.


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
snmpget -v3 -t 2 -r 0 -On <SWITCH_IP> \
chown root:root /etc/telegraf/monitor.env
    .1.3.6.1.2.1.31.1.1.1.6.<IFINDEX> \
chmod 600 /etc/telegraf/monitor.env
    .1.3.6.1.2.1.31.1.1.1.10.<IFINDEX> \
    .1.3.6.1.2.1.31.1.1.1.15.<IFINDEX> \
    .1.3.6.1.2.1.2.2.1.8.<IFINDEX>
</syntaxhighlight>
정상: Counter64 / Counter64 / Mbps 단위 속도 / operStatus 값. Timeout이면 경로·ACL·SNMP credential, No Such Instance면 인덱스, No Such Object면 view/지원 확인. 지원하지 않는 필드는 수집 설정에서 제외하고 미수집으로 기록한다. 0으로 대체하지 않는다.


=== 6.3 MIB 파일 정책 ===
install -d -m 755 /etc/systemd/system/telegraf.service.d


기본 수집은 숫자 OID + 명시적 field name을 사용한다. IF-MIB 파일 로딩 실패와 원격 SNMP 조회 실패는 별개다. <code>net-snmp-config</code>는 기본 utils 패키지에 없을 수 있어 필수 명령으로 사용하지 않는다.
vi /etc/systemd/system/telegraf.service.d/10-monitor-env.conf
</syntaxhighlight>


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
rpm -ql net-snmp-libs | grep '/mibs/'
[Service]
snmptranslate -On IF-MIB::ifHCInOctets
EnvironmentFile=/etc/telegraf/monitor.env
</syntaxhighlight>
</syntaxhighlight>
제조사 MIB는 Zyxel 공식 지원 페이지에서 '''정확한 모델/펌웨어용 묶음'''을 받는다. CPU·PoE·큐 드롭 OID는 확인 전 추가하지 않는다. SNMPv3 쓰기 권한 또는 PoE 전원 제어는 수집기에 부여하지 않는다.


== 7. Telegraf 구성 ==
=== 4.2 Main Configuration ===


=== 7.1 메인 파일과 환경 파일 ===
<syntaxhighlight lang="bash" line>
 
cp -a /etc/telegraf/telegraf.conf \
기존 설정 교체 시 수집 공백 발생 가능. 신규 전용 VM에서 기본 예제 설정을 백업한 후 아래 내용으로 교체한다. <code>/etc/telegraf/telegraf.d/</code>에 기존 활성 conf가 있으면 먼저 내용을 확인한다.
    /etc/telegraf/telegraf.conf.orig


<syntaxhighlight lang="bash" line>
cp -a /etc/telegraf/telegraf.conf "$MON_BACKUP/telegraf.conf.package"
ls -l /etc/telegraf/telegraf.d/
vi /etc/telegraf/telegraf.conf
vi /etc/telegraf/telegraf.conf
</syntaxhighlight>
</syntaxhighlight>
<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
[agent]
[agent]
377번째 줄: 365번째 줄:


[[inputs.internal]]
[[inputs.internal]]
[[inputs.cpu]]
[[inputs.cpu]]
   percpu = false
   percpu = false
   totalcpu = true
   totalcpu = true
[[inputs.mem]]
[[inputs.mem]]
[[inputs.disk]]
[[inputs.disk]]
   mount_points = ["/"]
   mount_points = ["/"]
[[inputs.net]]
[[inputs.net]]
</syntaxhighlight>
</syntaxhighlight>
별도 데이터 디스크를 쓰면 실제 마운트 경로를 <code>mount_points</code>에 추가한다. <code>inputs.cpu</code>는 Rocky VM의 CPU이며 스위치 CPU가 아니다.
 
== 5. SNMPv3 사전 확인 ==
 
SNMP는 v3 authPriv 사용을 기본으로 한다.
 
Net-SNMP 테스트 파일:
 
<syntaxhighlight lang="bash" line>
install -d -m 700 /root/.snmp
 
vi /root/.snmp/snmp.conf
</syntaxhighlight>
 
<syntaxhighlight lang="bash" line>
defVersion 3
defSecurityName <SNMP_USER>
defSecurityLevel authPriv
defAuthType SHA
defAuthPassphrase <SNMP_AUTH_PASSWORD>
defPrivType AES
defPrivPassphrase <SNMP_PRIV_PASSWORD>
</syntaxhighlight>


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
vi /etc/telegraf/ipt-monitor.env
chmod 600 /root/.snmp/snmp.conf
 
snmpget -v3 -t 2 -r 0 -On \
    203.0.113.11 \
    .1.3.6.1.2.1.1.3.0
 
snmpwalk -v3 -t 2 -r 0 -On \
    203.0.113.11 \
    .1.3.6.1.2.1.31.1.1.1.1
</syntaxhighlight>
</syntaxhighlight>
ifName OID 마지막 값이 ifIndex이다.
== 6. Telegraf SNMP 구성 ==
설정 파일은 장비 모델별로 분리한다.
<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
INFLUX_WRITE_TOKEN='<TELEGRAF_WRITE_TOKEN>'
/etc/telegraf/telegraf.d/
ZYXEL_SNMP_USER='<SNMP_USER>'
├── 10-zyxel-gs1900.conf
ZYXEL_SNMP_AUTH='<SNMP_AUTH_PASSWORD>'
├── 11-zyxel-gs1920.conf
ZYXEL_SNMP_PRIV='<SNMP_PRIV_PASSWORD>'
├── 20-zyxel-es3128.conf
├── 30-icmp_check.conf
├── 40-qnap_nas.conf
└── sflow.conf
</syntaxhighlight>
</syntaxhighlight>
위 파일은 systemd EnvironmentFile 형식이다. 셸에서 source로 읽지 않는다. 비밀값에 큰따옴표/역슬래시가 있으면 TOML 확장 시 구문에 영향을 줄 수 있으므로 공식 secret store를 사용하거나 이스케이프를 검증한다.
 
=== 6.1 GS1920 계열 ===
 
예시 대상:


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
chown root:root /etc/telegraf/ipt-monitor.env
203.0.113.21
chmod 600 /etc/telegraf/ipt-monitor.env
203.0.113.22
install -d -m 755 /etc/systemd/system/telegraf.service.d
203.0.113.23
vi /etc/systemd/system/telegraf.service.d/10-ipt-monitor.conf
</syntaxhighlight>
</syntaxhighlight>
SNMPv3:
* SHA
* DES
CPU OID:
<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
[Service]
.1.3.6.1.4.1.890.1.15.3.49.1.7.0
EnvironmentFile=/etc/telegraf/ipt-monitor.env
</syntaxhighlight>
</syntaxhighlight>
systemd가 환경 파일을 읽어 서비스에 전달한다. 일반 사용자의 읽기는 막지만 root나 동일 서비스 권한으로부터 완전히 숨기는 방식은 아니다.


=== 7.2 전체 포트 30초 수집 ===
Memory:


<code>/etc/telegraf/telegraf.d/10-zyxel-ports.conf</code>를 작성한다. 한 대의 예시이며 같은 credential을 쓰는 장비는 <code>agents</code>에 추가한다. 인증값이 다르면 입력 블록과 환경변수 이름을 분리한다.
<syntaxhighlight lang="bash" line>
Total  .1.3.6.1.4.1.890.1.15.3.50.1.1.1.3.1
Used    .1.3.6.1.4.1.890.1.15.3.50.1.1.1.4.1
Percent .1.3.6.1.4.1.890.1.15.3.50.1.1.1.5.1
</syntaxhighlight>


'''table 본체에는 oid를 지정하지 않는다.''' 필요한 열만 아래 field로 선택해 전체 테이블의 불필요한 열 조회를 피한다.
=== 6.2 GS1900 계열 ===
 
CPU:


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
[[inputs.snmp]]
.1.3.6.1.4.1.890.1.15.3.2.4.0
  agents = ["udp://<SWITCH_IP>:161"]
</syntaxhighlight>
  version = 3
  sec_name = "${ZYXEL_SNMP_USER}"
  sec_level = "authPriv"
  auth_protocol = "SHA"
  auth_password = "${ZYXEL_SNMP_AUTH}"
  priv_protocol = "AES"
  priv_password = "${ZYXEL_SNMP_PRIV}"
  timeout = "2s"
  retries = 0
  interval = "30s"
  agent_host_tag = "source"
  name = "switch_system"


  [[inputs.snmp.field]]
Memory:
    name = "uptime"
    oid = ".1.3.6.1.2.1.1.3.0"


  [[inputs.snmp.table]]
<syntaxhighlight lang="bash" line>
    name = "switch_port"
.1.3.6.1.4.1.890.1.15.3.2.5.0
    index_as_tag = true
</syntaxhighlight>
    inherit_tags = ["source"]


    [[inputs.snmp.table.field]]
실제 구축에서는 2초 timeout / retries 0에서 응답 누락이 발생하여 다음과 같이 완화하였다.
      name = "if_name"
 
      oid = ".1.3.6.1.2.1.31.1.1.1.1"
<syntaxhighlight lang="bash" line>
      is_tag = true
timeout = "5s"
    [[inputs.snmp.table.field]]
retries = 1
      name = "if_alias"
      oid = ".1.3.6.1.2.1.31.1.1.1.18"
    [[inputs.snmp.table.field]]
      name = "in_octets"
      oid = ".1.3.6.1.2.1.31.1.1.1.6"
    [[inputs.snmp.table.field]]
      name = "out_octets"
      oid = ".1.3.6.1.2.1.31.1.1.1.10"
    [[inputs.snmp.table.field]]
      name = "in_ucast"
      oid = ".1.3.6.1.2.1.31.1.1.1.7"
    [[inputs.snmp.table.field]]
      name = "in_mcast"
      oid = ".1.3.6.1.2.1.31.1.1.1.8"
    [[inputs.snmp.table.field]]
      name = "in_bcast"
      oid = ".1.3.6.1.2.1.31.1.1.1.9"
    [[inputs.snmp.table.field]]
      name = "out_ucast"
      oid = ".1.3.6.1.2.1.31.1.1.1.11"
    [[inputs.snmp.table.field]]
      name = "out_mcast"
      oid = ".1.3.6.1.2.1.31.1.1.1.12"
    [[inputs.snmp.table.field]]
      name = "out_bcast"
      oid = ".1.3.6.1.2.1.31.1.1.1.13"
    [[inputs.snmp.table.field]]
      name = "speed_mbps"
      oid = ".1.3.6.1.2.1.31.1.1.1.15"
    [[inputs.snmp.table.field]]
      name = "discontinuity"
      oid = ".1.3.6.1.2.1.31.1.1.1.19"
    [[inputs.snmp.table.field]]
      name = "oper_status"
      oid = ".1.3.6.1.2.1.2.2.1.8"
    [[inputs.snmp.table.field]]
      name = "in_discards"
      oid = ".1.3.6.1.2.1.2.2.1.13"
    [[inputs.snmp.table.field]]
      name = "out_discards"
      oid = ".1.3.6.1.2.1.2.2.1.19"
    [[inputs.snmp.table.field]]
      name = "in_errors"
      oid = ".1.3.6.1.2.1.2.2.1.14"
    [[inputs.snmp.table.field]]
      name = "out_errors"
      oid = ".1.3.6.1.2.1.2.2.1.20"
</syntaxhighlight>
</syntaxhighlight>
<code>if_alias</code>는 설명 변경에 따른 series 증가를 줄이기 위해 field로 저장한다. 실제 출력에 <code>source</code>, <code>index</code>, <code>if_name</code> 태그가 있는지 확인한다.


=== 7.3 중요 포트 5초 수집 ===
=== 6.3 ES-3128GP ===


<code>/etc/telegraf/telegraf.d/20-zyxel-fast.conf</code>를 작성한다. 필요한 포트 수만큼 블록을 복제하고 IP·ifIndex·태그를 모두 바꾼다. 같은 포트의 중복 설정은 금지하며 일반 수집과 measurement를 분리한다.
SNMPv3:
 
* SHA
* AES
 
sysObjectID:


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
[[inputs.snmp]]
.1.3.6.1.4.1.7800.1.190
  agents = ["udp://<SWITCH_IP>:161"]
</syntaxhighlight>
  version = 3
 
  sec_name = "${ZYXEL_SNMP_USER}"
CPU / Memory private OID는 장비에서 정상 확인되지 않아 Dashboard에서는 N/A 처리한다.
  sec_level = "authPriv"
 
  auth_protocol = "SHA"
=== 6.4 IF-MIB 주요 항목 ===
  auth_password = "${ZYXEL_SNMP_AUTH}"
 
  priv_protocol = "AES"
수집 대상:
  priv_password = "${ZYXEL_SNMP_PRIV}"
 
  timeout = "800ms"
* ifName
  retries = 0
* ifAlias
  interval = "5s"
* ifSpeed
  agent_host_tag = "source"
* ifHCInOctets
  name = "switch_port_fast"
* ifHCOutOctets
* ifOperStatus
* ifInErrors
* ifOutErrors
* ifInDiscards
* ifOutDiscards
* Broadcast
* Multicast
 
누적 Counter는 Grafana에서 그대로 표시하지 않고 derivative/rate 계산 후 표시한다.
 
== 7. ICMP 수집 ==


  [inputs.snmp.tags]
Telegraf ping input을 사용한다.
    if_name = "<PORT_LABEL>"
    index = "<IFINDEX>"


  [[inputs.snmp.field]]
<syntaxhighlight lang="bash" line>
    name = "in_octets"
vi /etc/telegraf/telegraf.d/30-icmp_check.conf
    oid = ".1.3.6.1.2.1.31.1.1.1.6.<IFINDEX>"
  [[inputs.snmp.field]]
    name = "out_octets"
    oid = ".1.3.6.1.2.1.31.1.1.1.10.<IFINDEX>"
  [[inputs.snmp.field]]
    name = "in_ucast"
    oid = ".1.3.6.1.2.1.31.1.1.1.7.<IFINDEX>"
  [[inputs.snmp.field]]
    name = "in_mcast"
    oid = ".1.3.6.1.2.1.31.1.1.1.8.<IFINDEX>"
  [[inputs.snmp.field]]
    name = "in_bcast"
    oid = ".1.3.6.1.2.1.31.1.1.1.9.<IFINDEX>"
  [[inputs.snmp.field]]
    name = "out_ucast"
    oid = ".1.3.6.1.2.1.31.1.1.1.11.<IFINDEX>"
  [[inputs.snmp.field]]
    name = "out_mcast"
    oid = ".1.3.6.1.2.1.31.1.1.1.12.<IFINDEX>"
  [[inputs.snmp.field]]
    name = "out_bcast"
    oid = ".1.3.6.1.2.1.31.1.1.1.13.<IFINDEX>"
  [[inputs.snmp.field]]
    name = "in_discards"
    oid = ".1.3.6.1.2.1.2.2.1.13.<IFINDEX>"
  [[inputs.snmp.field]]
    name = "out_discards"
    oid = ".1.3.6.1.2.1.2.2.1.19.<IFINDEX>"
  [[inputs.snmp.field]]
    name = "oper_status"
    oid = ".1.3.6.1.2.1.2.2.1.8.<IFINDEX>"
  [[inputs.snmp.field]]
    name = "discontinuity"
    oid = ".1.3.6.1.2.1.31.1.1.1.19.<IFINDEX>"
</syntaxhighlight>
</syntaxhighlight>
표준 정의: [https://www.rfc-editor.org/rfc/rfc2863.html RFC 2863]. 각 열의 지원 여부는 실장비에서 확인한다.


1초 수집은 해당 블록의 <code>interval = &quot;1s&quot;</code>로 변경한다. 먼저 카운터 갱신과 수집 소요시간을 확인한다. 800ms는 개별 요청 timeout이지 전체 수집 완료 보장이 아니다. 실패를 반복 재시도하지 않고 결측으로 남긴다. DB에 5초마다 묶어 써도 원래 수집 timestamp는 유지된다.
예시 대상:


=== 7.4 응답시간·손실 보조 수집 ===
<syntaxhighlight lang="bash" line>
203.0.113.11
203.0.113.12
203.0.113.13
</syntaxhighlight>


<code>/etc/telegraf/telegraf.d/30-ipt-ping.conf</code>:
권장 설정:


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
[[inputs.ping]]
[[inputs.ping]]
   urls = ["<SWITCH_IP>", "<CALLSERVER_IP>"]
   urls = [
    "203.0.113.11",
    "203.0.113.12",
    "203.0.113.13"
  ]
 
   method = "native"
   method = "native"
  privileged = true
   count = 3
   count = 1
   deadline = 2.0
   deadline = "1s"
   interval = 10.0
   interval = "5s"
</syntaxhighlight>
</syntaxhighlight>
승인된 시험 전화기 IP를 필요 시 추가한다. ICMP 비응답 장비는 제외한다. 콜서버 Ping은 SIP 서비스 감시가 아니며 VM에서의 경로도 전화기에서의 경로와 같다고 보장할 수 없다.


<code>/etc/systemd/system/telegraf.service.d/10-ipt-monitor.conf</code>의 같은 <code>[Service]</code>에 추가한다:
native ping은 CAP_NET_RAW가 필요하다.
 
<syntaxhighlight lang="bash" line>
systemctl edit telegraf
</syntaxhighlight>


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
[Service]
CapabilityBoundingSet=CAP_NET_RAW
CapabilityBoundingSet=CAP_NET_RAW
AmbientCapabilities=CAP_NET_RAW
AmbientCapabilities=CAP_NET_RAW
</syntaxhighlight>
</syntaxhighlight>
Telegraf를 root로 실행하는 설정은 아니다. [https://docs.influxdata.com/telegraf/v1/input-plugins/ping/ 공식 Ping 권한 설명]. 기존 입력이 있으면 필요한 capability를 제거하지 않도록 확인한다.


=== 7.5 검사·기동 ===
<syntaxhighlight lang="bash" line>
systemctl daemon-reload
systemctl restart telegraf
</syntaxhighlight>
 
== 8. Telegraf Test 및 서비스 시작 ==
 
설정 검사:
 
<syntaxhighlight lang="bash" line>
telegraf \
  --config /etc/telegraf/telegraf.conf \
  --config-directory /etc/telegraf/telegraf.d \
  --test
</syntaxhighlight>
 
실제 서비스 계정으로 확인:


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
chown root:telegraf /etc/telegraf/telegraf.conf /etc/telegraf/telegraf.d/*-*.conf
chmod 640 /etc/telegraf/telegraf.conf /etc/telegraf/telegraf.d/*-*.conf
systemctl daemon-reload
systemctl daemon-reload
systemd-run --unit=ipt-telegraf-check --wait --pipe --collect \
 
    -p User=telegraf -p Group=telegraf \
systemd-run \
    -p EnvironmentFile=/etc/telegraf/ipt-monitor.env \
  --unit=telegraf-config-check \
    -p CapabilityBoundingSet=CAP_NET_RAW \
  --wait \
    -p AmbientCapabilities=CAP_NET_RAW \
  --pipe \
    /usr/bin/telegraf --config /etc/telegraf/telegraf.conf \
  --collect \
    --config-directory /etc/telegraf/telegraf.d --test
  -p User=telegraf \
  -p Group=telegraf \
  -p EnvironmentFile=/etc/telegraf/monitor.env \
  -p CapabilityBoundingSet=CAP_NET_RAW \
  -p AmbientCapabilities=CAP_NET_RAW \
  /usr/bin/telegraf \
  --config /etc/telegraf/telegraf.conf \
  --config-directory /etc/telegraf/telegraf.d \
  --test
</syntaxhighlight>
</syntaxhighlight>
실제로 SNMP/Ping을 한 번 수집하는 검사다. 출력의 IP 등은 사내 정보로 취급한다. <code>--test</code>는 InfluxDB 쓰기를 검증하지 않는다. 수집 검사 성공 후:
 
서비스 시작:


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
systemctl enable --now telegraf
systemctl enable --now telegraf
systemctl restart telegraf
systemctl restart telegraf
systemctl status telegraf --no-pager
systemctl status telegraf --no-pager
journalctl -u telegraf -n 100 --no-pager
journalctl -u telegraf -n 100 --no-pager
</syntaxhighlight>
</syntaxhighlight>
설정 근거: [https://docs.influxdata.com/telegraf/v1/input-plugins/snmp/ SNMP Input Plugin], [https://docs.influxdata.com/telegraf/v1/configuration/ Telegraf 설정]. 입력과 DB 출력이 모두 성공하는지 확인하고 unauthorized, timeout, collection interval 초과, buffer/drop 메시지를 조사한다.


== 8. Grafana 구성과 계산 ==
== 9. Grafana 설치 및 InfluxDB 연결 ==


=== 8.1 서비스 ===
=== 9.1 Grafana 서비스 ===
 
먼저 백업한 뒤 <code>/etc/grafana/grafana.ini</code>의 기존 섹션을 수정한다. 중복 섹션을 추가하지 않는다.


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
cp -a /etc/grafana/grafana.ini "$MON_BACKUP/grafana.ini.initial"
vi /etc/grafana/grafana.ini
vi /etc/grafana/grafana.ini
</syntaxhighlight>
</syntaxhighlight>
내부망 운영 환경에 맞게 listen 주소를 설정한다.
예시:
<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
[server]
[server]
http_addr = 127.0.0.1
http_addr = 0.0.0.0
http_port = 3000
http_port = 3000


629번째 줄: 626번째 줄:
enabled = false
enabled = false
</syntaxhighlight>
</syntaxhighlight>
<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
systemctl enable --now grafana-server
systemctl enable --now grafana-server
systemctl restart grafana-server
systemctl restart grafana-server
systemctl status grafana-server --no-pager
systemctl status grafana-server --no-pager
curl -fsS http://127.0.0.1:3000/api/health
curl -fsS http://127.0.0.1:3000/api/health
ss -lntp | grep ':3000'
</syntaxhighlight>
</syntaxhighlight>
초기 관리자로 로그인 후 즉시 강한 비밀번호로 변경한다. SSH 터널의 <code>http://127.0.0.1:3000</code>을 사용한다. 공유 접속이 필요하면 인증·TLS reverse proxy와 접근제어를 별도 구성하며 평문 HTTP를 전사망에 개방하지 않는다.


=== 8.2 InfluxDB 데이터소스 ===
=== 9.2 InfluxDB Datasource ===


Connections → Data sources → Add data source → InfluxDB:
Grafana:
 
<syntaxhighlight lang="bash" line>
Connections
  → Data sources
  → Add data source
  → InfluxDB
</syntaxhighlight>


{| class="wikitable"
{| class="wikitable"
|-
! 설정
! 설정
! 값
! 값
|-
| Name
| <code>IPT-SNMP</code>
|-
|-
| Query language
| Query language
654번째 줄: 655번째 줄:
|-
|-
| URL
| URL
| <code>http://127.0.0.1:8086</code>
| http://127.0.0.1:8086
|-
|-
| Organization
| Organization
| <code>network</code>
| network
|-
|-
| Token
| Token
| grafana-read 토큰
| grafana-read
|-
|-
| Default Bucket
| Default Bucket
| <code>snmp_raw</code>
| snmp_raw
|-
|-
| Min time interval
| Min time interval
| <code>1s</code> (실제 수집 간격은 패널별로 반영)
| 1s
|}
|}


Save &amp; test 성공 후 Explore에서 measurement를 조회한다. [https://grafana.com/docs/grafana/latest/datasources/influxdb/configure/ Grafana 공식 연결 설정].
== 10. SNMP Dashboard 기본 구성 ==


=== 8.3 필드와 태그 확인 ===
=== 10.1 Traffic ===
 
Flux:


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
from(bucket: "snmp_raw")
from(bucket: "snmp_raw")
   |> range(start: -10m)
   |> range(start: v.timeRangeStart, stop: v.timeRangeStop)
   |> filter(fn: (r) => r._measurement == "switch_port_fast")
   |> filter(fn: (r) => r._measurement == "switch_port")
   |> limit(n: 5)
  |> filter(fn: (r) => r.source == "203.0.113.11")
  |> filter(fn: (r) =>
      r._field == "in_octets" or
      r._field == "out_octets"
  )
  |> derivative(unit: 1s, nonNegative: true)
   |> map(fn: (r) => ({
      r with _value: r._value * 8.0
  }))
</syntaxhighlight>
</syntaxhighlight>
실제 <code>source</code> 태그를 확인해 아래 <code>&lt;SOURCE_TAG_VALUE&gt;</code>에 넣는다. 일반적으로 장비 IP지만 출력으로 확인한다. 먼저 변수 없이 한 대·한 포트의 표시를 완성한다.


=== 8.4 In/Out bps ===
Grafana Unit:
 
<syntaxhighlight lang="bash" line>
bits/sec
</syntaxhighlight>


패널: Time series / Unit: bits/sec / 결측 구간 연결 안 함 / 시간대 Asia/Seoul / 범위 최근 15분 / refresh 5s.
=== 10.2 Error / Discard ===


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
from(bucket: "snmp_raw")
   |> filter(fn: (r) =>
  |> range(start: v.timeRangeStart, stop: v.timeRangeStop)
      r._field == "in_errors" or
   |> filter(fn: (r) => r._measurement == "switch_port_fast")
      r._field == "out_errors" or
  |> filter(fn: (r) => r.source == "<SOURCE_TAG_VALUE>" and r.index == "<IFINDEX>")
      r._field == "in_discards" or
  |> filter(fn: (r) => r._field == "in_octets" or r._field == "out_octets")
      r._field == "out_discards"
  )
   |> derivative(unit: 1s, nonNegative: true)
   |> derivative(unit: 1s, nonNegative: true)
  |> map(fn: (r) => ({r with _value: r._value * 8.0}))
</syntaxhighlight>
</syntaxhighlight>
<code>derivative</code>로 실제 timestamp 차이를 사용해 초당 값으로 바꾼 뒤 8배한다. 고정 간격 5로 나누지 않는다. 장기간 표시에서 집계한다면 '''rate 계산 후''' <code>aggregateWindow(every: v.windowPeriod, fn: max, createEmpty: false)</code>를 추가하고 ‘구간 내 최대 수집간격 평균’으로 표시한다. 누적 카운터를 먼저 평균내지 않는다.


<code>nonNegative</code>만으로 모든 리셋을 판별할 수 없다. discontinuity 변경·sysUpTime 감소·링크 변경 구간은 무효 처리한다. 첫 점은 이전 값이 없어 표시되지 않는다. 결측은 0으로 채우지 않으며 긴 결측을 가로지르는 rate는 장애 증거에 사용하지 않는다. [https://docs.influxdata.com/flux/v0/stdlib/universe/derivative/ Flux derivative].
=== 10.3 Broadcast / Multicast ===
 
Broadcast/Multicast는 별도 패널로 표시하여 폭주 및 비정상 증가를 확인한다.
 
== 11. Prometheus 설치 ==


=== 8.5 PPS、Broadcast、Discard ===
실제 구축에서는 Prometheus 3.13.3을 사용하였다.


위 쿼리의 <code>_field</code> 조건을 다음으로 바꾸고 8배하는 <code>map</code> 행은 제거한다. 각 계열을 별도로 표시한다.
=== 11.1 사용자 및 디렉터리 ===


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
   |> filter(fn: (r) => r._field == "in_ucast" or r._field == "in_mcast" or r._field == "in_bcast")
useradd \
   --system \
  --no-create-home \
  --shell /sbin/nologin \
  prometheus
 
install -d -o prometheus -g prometheus \
  /etc/prometheus \
  /var/lib/prometheus
</syntaxhighlight>
</syntaxhighlight>
합계 In PPS는 세 계열의 합이다. Out PPS는 대응하는 <code>out_</code> 세 계열로 만든다. Broadcast/Multicast 전용 패널도 구성해 정상 시 기준과 비교한다.
 
Prometheus 바이너리는 공식 릴리스에서 다운로드하고 checksum 확인 후 설치한다.


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
  |> filter(fn: (r) => r._field == "in_discards" or r._field == "out_discards")
install -m 0755 prometheus /usr/local/bin/prometheus
install -m 0755 promtool /usr/local/bin/promtool
</syntaxhighlight>
</syntaxhighlight>
위 값은 폐기 packets/sec다. 누적값이 아니라 증가율을 표시한다. 일반 수집 measurement의 <code>in_errors/out_errors</code>도 같은 방식으로 별도 패널을 만든다.


=== 8.6 상태·손실·대시보드 ===
=== 11.2 Prometheus 설정 ===


* <code>oper_status</code>는 rate로 바꾸지 않는다. 1=up, 2=down 등으로 state timeline에 표시한다.
<syntaxhighlight lang="bash" line>
* <code>speed_mbps</code>는 Mbps 단위다. In 사용률=In bps÷(speed_mbps×1,000,000)×100. Out은 따로 계산하며 전이중의 In/Out을 합쳐 100%로 판정하지 않는다.
vi /etc/prometheus/prometheus.yml
* <code>ping</code>의 <code>average_response_ms</code>, <code>percent_packet_loss</code>는 이미 계산된 값이므로 derivative 불필요. 1패킷/5초 설정에서는 각 관측의 손실이 0% 또는 100%다.
</syntaxhighlight>
* LAG 논리 포트와 물리 멤버를 별도로 표시하고 합산 중복 집계를 피한다.
* 초기 패널: 중요 포트 bps, Broadcast/Multicast PPS, Discard/Error 증가율, operStatus, Ping 지연·손실, VM CPU·메모리·디스크, Telegraf 오류·결측.
* 스위치 CPU·PoE는 모델별 OID 확인 후 추가한다. 미설정을 0/정상으로 표시하지 않는다.


대시보드를 저장하고 JSON으로 내보낸다. 이번 Syslog 구성은 파일 조회용으로 Grafana에서 로그를 검색할 수는 없다. 별도 로그 검색 UI를 추가하기 전에는 §11의 CLI로 대조한다.
<syntaxhighlight lang="bash" line>
global:
  scrape_interval: 30s


== 9. Zyxel Syslog 수신 ==
scrape_configs:


=== 9.1 포트 중복 확인 ===
  - job_name: prometheus
    static_configs:
      - targets:
          - "127.0.0.1:9090"
        labels:
          server_name: monitoring-server


<syntaxhighlight lang="bash" line>
  - job_name: node
grep -RniE 'imudp|imtcp|port="514"|UDPServerRun' /etc/rsyslog.conf /etc/rsyslog.d
    static_configs:
ss -lunp | grep ':514'
      - targets:
          - "127.0.0.1:9100"
        labels:
          server_name: monitoring-server
 
      - targets:
          - "192.0.2.20:9100"
        labels:
          server_name: linux-server
 
  - job_name: windows
    static_configs:
      - targets:
          - "192.0.2.30:9182"
        labels:
          server_name: windows-server
 
  - job_name: qnap
    static_configs:
      - targets:
          - "198.51.100.10:9100"
        labels:
          server_name: nas-01
</syntaxhighlight>
</syntaxhighlight>
기존 입력이 있으면 module/input을 중복 생성하지 말고 기존 설정에 통합한다. 아래는 미구성 신규 VM용이다.
 
=== 11.3 systemd ===


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
install -d -m 750 /var/log/zyxel
vi /etc/systemd/system/prometheus.service
restorecon -RF /var/log/zyxel
vi /etc/rsyslog.d/30-zyxel-remote.conf
</syntaxhighlight>
</syntaxhighlight>
<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
module(load="imudp")
[Unit]
Description=Prometheus
Wants=network-online.target
After=network-online.target


template(name="ZyxelRemotePath" type="string"
[Service]
  string="/var/log/zyxel/%fromhost-ip:::secpath-replace%/events.log")
User=prometheus
Group=prometheus


template(name="ZyxelRemoteLine" type="string"
ExecStart=/usr/local/bin/prometheus \
   string="received=%timegenerated:::date-rfc3339% reported=%timereported:::date-rfc3339% src=%fromhost-ip% host=%hostname% facility=%syslogfacility-text% severity=%syslogseverity-text% tag=%syslogtag% msg=%msg%\n")
  --config.file=/etc/prometheus/prometheus.yml \
   --storage.tsdb.path=/var/lib/prometheus \
  --storage.tsdb.retention.time=30d \
  --storage.tsdb.retention.size=20GB \
  --web.listen-address=127.0.0.1:9090


ruleset(name="ZyxelRemoteRules") {
Restart=always
  action(type="omfile"
    dynaFile="ZyxelRemotePath"
    template="ZyxelRemoteLine"
    createDirs="on"
    dirCreateMode="0750"
    fileCreateMode="0640")
  stop
}


input(type="imudp" address="<COLLECTOR_IP>" port="514"
[Install]
  ruleset="ZyxelRemoteRules")
WantedBy=multi-user.target
</syntaxhighlight>
</syntaxhighlight>
송신 IP별 파일에 수신 시각과 장비 보고 시각을 함께 기록한다. 장비 Syslog에 시간대 정보가 없으면 완전한 보정은 불가능하므로 실장비 테스트로 대조한다.


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
rsyslogd -N1
chown -R prometheus:prometheus \
systemctl enable --now rsyslog
  /etc/prometheus \
systemctl restart rsyslog
  /var/lib/prometheus
ss -lunp | grep ':514'
 
journalctl -u rsyslog -n 100 --no-pager
promtool check config /etc/prometheus/prometheus.yml
 
systemctl daemon-reload
systemctl enable --now prometheus
 
systemctl status prometheus --no-pager
</syntaxhighlight>
</syntaxhighlight>
재시작 중 짧은 수신 공백에 주의한다. 실패하면 변경 전 파일을 복원하고 <code>rsyslogd -N1</code> 검사 후 재시작한다. SELinux 거부는 다음으로 확인하며 비활성화하지 않는다.
 
== 12. Node Exporter 설치 ==
 
실제 구축에서는 Node Exporter 1.12.1 환경을 사용하였다.
 
=== 12.1 Linux Node Exporter ===


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
ausearch -m AVC -ts recent
useradd \
ls -Zd /var/log/zyxel
  --system \
  --no-create-home \
  --shell /sbin/nologin \
  node_exporter
 
install -m 0755 node_exporter \
  /usr/local/bin/node_exporter
</syntaxhighlight>
</syntaxhighlight>
근거: [https://docs.rsyslog.com/doc/configuration/modules/imudp.html imudp], [https://docs.rsyslog.com/doc/configuration/modules/omfile.html omfile]. 참조 문서는 최신판을 포함하므로 설치된 rsyslog의 구문 검사를 반드시 수행한다.
=== 9.2 로그 회전과 보존 ===
<code>/etc/logrotate.d/zyxel-remote</code>:


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
/var/log/zyxel/*/events.log {
vi /etc/systemd/system/node_exporter.service
    daily
    rotate 30
    missingok
    notifempty
    compress
    delaycompress
    create 0640 root root
    sharedscripts
    postrotate
        /usr/bin/systemctl kill -s HUP --kill-who=main rsyslog.service >/dev/null 2>&1 || true
    endscript
}
</syntaxhighlight>
</syntaxhighlight>
보존은 ‘30회전분’이며 매일 회전하면 약 30일이다. 정확히 30일이 지난 순간 삭제되는 방식은 아니다. 오래된 파일 삭제를 수반하므로 장애 로그는 별도 보관한다. 신규 VM의 rsyslog root 실행 기준이며 privilege drop을 사용한다면 소유자를 맞춘다.


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
logrotate -d /etc/logrotate.d/zyxel-remote
[Unit]
systemctl status logrotate.timer --no-pager
Description=Node Exporter
systemctl list-timers --all | grep logrotate
After=network-online.target
Wants=network-online.target
 
[Service]
User=node_exporter
Group=node_exporter
 
ExecStart=/usr/local/bin/node_exporter
 
Restart=always
 
[Install]
WantedBy=multi-user.target
</syntaxhighlight>
</syntaxhighlight>
일일 실행이 비활성이라면 패키지 timer/cron을 확인해 활성화한다. 일일 회전만으로 하루 동안의 폭증에 따른 디스크 고갈은 막을 수 없다. 디스크 80% 예고 감시와 증가량 점검을 수행한다. 근거 없이 중요 로그를 제한해 버리지 않는다.


=== 9.3 Zyxel 측 송신 ===
<syntaxhighlight lang="bash" line>
systemctl daemon-reload
systemctl enable --now node_exporter


Syslog 목적지 <code>&lt;COLLECTOR_IP&gt;</code> / UDP514. Informational을 포함하는 범위로 시작하고 상시 Debug는 피한다. Link Up/Down, STP 변경, 루프 탐지, PoE 이상·전력 부족, 재시작 등을 모델 지원 범위에서 활성화한다. 포트별 이벤트 설정이 필요하면 함께 적용한다.
curl -s http://127.0.0.1:9100/metrics | head
</syntaxhighlight>


== 10. firewalld와 접속 확인 ==
외부 서버의 node_exporter는 Prometheus 서버 주소만 9100/tcp에 접근하도록 방화벽을 제한한다.


기본 웹 접속은 SSH 터널만 사용한다. 3000/8086과 서버 수신 UDP161/162를 열지 않는다. SNMP는 서버에서 스위치 UDP161로 조회하고 응답을 받는다. 외향 제한이 있는 경우에만 기존 정책에 필요 허용을 추가한다.
== 13. Windows Exporter ==


=== 10.1 Syslog만 허용 ===
Windows Server에서는 windows_exporter를 사용한다.


영향: 지정 송신원에서 UDP514 수신을 추가한다. zone 변경·SSH 규칙 삭제·광범위 허용은 하지 않는다. 먼저 확인:
기본 포트:


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
firewall-cmd --get-active-zones
9182/tcp
firewall-cmd --zone=<FW_ZONE> --list-all
</syntaxhighlight>
</syntaxhighlight>
실제 값으로 설정한다. 송신자를 엄격히 제한하려면 장비별 /32 규칙을 사용한다.
 
Prometheus Target:


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
MON_ZONE='<FW_ZONE>'
192.0.2.30:9182
MON_SYSLOG_RULE='rule family="ipv4" source address="<SWITCH_SOURCE_CIDR>" port port="514" protocol="udp" accept'
firewall-cmd --zone="$MON_ZONE" --add-rich-rule="$MON_SYSLOG_RULE" --timeout=600
</syntaxhighlight>
</syntaxhighlight>
10분 안에 장비 송신과 SSH 접속 유지를 확인한다. 임시 규칙은 자동 만료된다. 승인 후 동일 규칙만 영구 및 runtime에 적용:
 
성능 카운터 이상 시 다음 명령으로 복구한 사례가 있다.


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
firewall-cmd --permanent --zone="$MON_ZONE" --add-rich-rule="$MON_SYSLOG_RULE"
lodctr /R
firewall-cmd --zone="$MON_ZONE" --remove-rich-rule="$MON_SYSLOG_RULE"
winmgmt /resyncperf
firewall-cmd --zone="$MON_ZONE" --add-rich-rule="$MON_SYSLOG_RULE"
</syntaxhighlight>
</syntaxhighlight>
이미 만료됐다면 remove의 not enabled를 확인하고 다음으로 진행한다. 전체 runtime 저장이나 reload로 다른 작업 변경을 포함하지 않는다.


롤백 — 이번 작업에서 새로 추가한 규칙만 제거:
PowerShell:


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
firewall-cmd --zone="$MON_ZONE" --remove-rich-rule="$MON_SYSLOG_RULE"
Restart-Service windows_exporter
firewall-cmd --permanent --zone="$MON_ZONE" --remove-rich-rule="$MON_SYSLOG_RULE"
</syntaxhighlight>
</syntaxhighlight>
=== 10.2 실장비 도달 확인 ===
 
확인:


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
timeout 20 tcpdump -ni any -nn 'udp dst port 514'
Invoke-WebRequest http://127.0.0.1:9182/metrics
tail -n 50 /var/log/zyxel/<SWITCH_SYSLOG_SOURCE_IP>/events.log
</syntaxhighlight>
</syntaxhighlight>
VM 자신의 logger는 파일 저장 검사용이며 스위치에서의 경로·ACL 검증을 대신하지 못한다.
 
== 14. Grafana Prometheus Datasource ==
 
Grafana:


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
logger --udp --server <COLLECTOR_IP> --port 514 --tag IPT_TEST 'collector receiver test'
Connections
  → Data sources
  → Prometheus
</syntaxhighlight>
</syntaxhighlight>
== 11. 시험 운영·장애 확인 ==


=== 11.1 초기 24시간 ===
URL:


# 한 대·중요 포트 1~2개로 시작한다.
<syntaxhighlight lang="bash" line>
# 실제 SNMP 카운터와 InfluxDB 값 증가를 대조한다.
http://127.0.0.1:9090
# bps/PPS 변동과 operStatus 표시를 확인한다.
</syntaxhighlight>
# 정상 발신 테스트의 시각·단말 IP·Call-ID를 기록한다. 본망에서 부하용 iperf나 루프를 만들지 않는다.
# 별도 수집 중인 콜서버 SIP와 전후 2분의 네트워크 기록을 비교한다.
# 5060 외 SIP/TLS와 RTP는 해당 캡처 범위 밖임을 기록한다.
# 수집시간·스위치 CPU·VM CPU/IO·결측·로그 증가량이 허용 범위면 대상 장비를 늘린다.
# 중요 포트 1초화는 카운터 갱신을 실측한 후 진행한다.


=== 11.2 운영 확인 명령 ===
Save & Test 후 다음 PromQL로 확인한다.


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
systemctl is-active influxdb telegraf grafana-server rsyslog chronyd
up
journalctl -u telegraf --since '-10 min' --no-pager
chronyc tracking
df -h
du -sh /var/log/zyxel
vmstat 1 5
iostat -xz 1 5
nstat -az | grep -E 'Udp(InErrors|RcvbufErrors|InDatagrams)'
</syntaxhighlight>
</syntaxhighlight>
UDP 통계는 누적값이므로 시간 차분으로 본다. VM/호스트의 수신 드롭을 스위치 송신 중단으로 오해하지 않는다. <code>inputs.internal</code>의 수집시간·오류와 DB 출력 drop도 대시보드에 추가한다.


=== 11.3 장애 기록 양식 ===
== 15. Server Dashboard 구성 ==
 
Linux Dashboard:
 
* Uptime
* CPU Usage
* Memory Usage
* Load Average
* Filesystem
* Disk Read / Write
* Network RX / TX
* Network Error / Drop
* Service 상태
 
Windows Dashboard:
 
* CPU
* Memory
* Disk
* Network
* Uptime
* Exporter 상태
 
Disk 용량은 실제 GB/GiB 값으로 표시하며 불필요한 Bar Gauge는 제거한다.
 
== 16. QNAP TS-264 모니터링 ==
 
QNAP은 세 종류의 데이터를 함께 사용한다.


{| class="wikitable"
{| class="wikitable"
! 데이터
! 수집 방법
|-
|-
! 항목
| CPU / Memory / Network / Filesystem / Disk I/O
! 기록
| node_exporter → Prometheus
|-
|-
| 장애 일시·시간대
| HDD / RAID / Storage / Volume
| 초 단위, KST
| SNMP → Telegraf → InfluxDB
|-
|-
| 발신·착신
| Event / Access Log
| 내선, 전화기 IP, 등록 상태
| Syslog → rsyslog → Alloy → Loki
|-
| 연결 위치
| L2 이름, 포트, ifIndex, Voice VLAN
|-
| PC 상태
| 동시 정상·불통·미확인
|-
| SIP
| INVITE 도착, 응답 코드, Call-ID, 재전송
|-
| SNMP
| In/Out bps、PPS、Broadcast、Discard、operStatus
|-
| 로그
| Link, STP, PoE, 재부팅 등
|-
| 수집 상태
| NTP, 결측, VM 수신 drop, tcpdump drop
|}
|}


로그 검색:
=== 16.1 QNAP SNMP ===
 
QNAP SNMPv3:
 
* SHA
* DES
 
Measurement:


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
grep -Ei 'link|stp|loop|poe|power|reboot|restart|topology' \
QNAP_TS264
    /var/log/zyxel/<SWITCH_SYSLOG_SOURCE_IP>/events.log
qnap_disk
qnap_raid
qnap_storage_pool
qnap_volume
</syntaxhighlight>
</syntaxhighlight>
메시지의 수신·보고 시각을 장애 전후와 비교한다. 키워드가 없다는 이유만으로 이벤트가 없다고 판단하지 않는다. 장비 문구는 모델에 따라 다르다.


=== 11.4 초기 알림 제안 ===
Disk:


아래는 운영 후 Grafana 관리형 알림으로 구성할 제안값이며 자동 생성된 설정은 아니다.
* disk_id
* manufacturer
* model
* disk_type
* disk_status
* temperature
* capacity_bytes


* 중요 포트 한 방향 사용률 80% 이상이 30초 지속: 혼잡 예고이며 원인 확정은 아니다.
RAID:
* 중요 포트 Discard 증가: 조사 이벤트. QoS·버퍼·정책 확인.
* Ping 3회 연속 손실: VM에서의 도달성 이상이며 SIP 장애 확정은 아니다.
* 고주기 데이터 15초 이상 결측: No Data를 정상으로 처리하지 않는다.
* 디스크 사용량 80% 이상: 보존 공간 경고.
* Broadcast PPS는 24시간 정상값을 확보한 뒤 임계값을 결정한다. 모든 현장에 공통 고정값을 적용하지 않는다.


전화 연결 실패와 높은 사용률·Discard 증가가 동반되면 혼잡이 유력하다. INVITE 불착만으로 단말 미송신과 경로 손실은 구분할 수 없다. 등록·PoE·Voice VLAN 상태와 함께 판단한다.
* RAID ID
* RAID Name
* RAID Status
* RAID Level
* Capacity


== 12. sFlow 추가 — 조건부 절차 ==
=== 16.2 Volume 단위 보정 ===


=== 12.1 착수 조건 ===
QNAP qnap_volume의 capacity/free 값은 실제 장비에서 KiB 형태로 반환되는 것으로 확인되었다.


'''Zyxel이라는 제조사명만으로 sFlow 지원을 확정할 수 없다.''' 정확한 모델·펌웨어의 공식 매뉴얼에서 exporter, 대상 인터페이스, sampling 범위, source/agent IP 설정을 확인한다. 확인 전 컨테이너 전체를 기동할 필요는 없다.
원시값을 Grafana에서 byte로 바로 해석하면 약 11.3 GiB로 잘못 표시된다.
 
Flux:
 
<syntaxhighlight lang="bash" line>
|> map(fn: (r) => ({
    r with
    capacity_bytes:
      uint(v: r.capacity_bytes) * uint(v: 1024),
 
    free_bytes:
      uint(v: r.free_bytes) * uint(v: 1024),
 
    used_bytes:
      (
        uint(v: r.capacity_bytes)
        - uint(v: r.free_bytes)
      ) * uint(v: 1024),
 
    used_percent:
      if float(v: r.capacity_bytes) > 0.0 then
        (
          float(v: r.capacity_bytes)
          - float(v: r.free_bytes)
        )
        / float(v: r.capacity_bytes)
        * 100.0
      else 0.0
}))
</syntaxhighlight>
 
=== 16.3 실제 Data Volume ===
 
Node Exporter 기준 실제 사용자 Data Volume:
 
<syntaxhighlight lang="bash" line>
mountpoint="/share/CACHEDEV1_DATA"
device="/dev/mapper/cachedev1"
fstype="ext4"
</syntaxhighlight>
 
Snapshot:
 
<syntaxhighlight lang="bash" line>
/mnt/snapshot/...
</syntaxhighlight>
 
위 Snapshot 경로는 Dashboard Filesystem 패널에서 제외한다.
 
DataVol1 사용률:
 
<syntaxhighlight lang="bash" line>
100 * (
  1 -
  node_filesystem_avail_bytes{
    instance="198.51.100.10:9100",
    mountpoint="/share/CACHEDEV1_DATA"
  }
  /
  node_filesystem_size_bytes{
    instance="198.51.100.10:9100",
    mountpoint="/share/CACHEDEV1_DATA"
  }
)
</syntaxhighlight>
 
=== 16.4 Network는 bond0만 표시 ===
 
RX:
 
<syntaxhighlight lang="bash" line>
rate(
  node_network_receive_bytes_total{
    instance="198.51.100.10:9100",
    device="bond0"
  }[$__rate_interval]
) * 8
</syntaxhighlight>
 
TX:
 
<syntaxhighlight lang="bash" line>
rate(
  node_network_transmit_bytes_total{
    instance="198.51.100.10:9100",
    device="bond0"
  }[$__rate_interval]
) * 8
</syntaxhighlight>
 
=== 16.5 Network Error / Drop ===
 
<syntaxhighlight lang="bash" line>
rate(node_network_receive_errs_total{
  instance="198.51.100.10:9100",
  device="bond0"
}[$__rate_interval])
</syntaxhighlight>
 
<syntaxhighlight lang="bash" line>
rate(node_network_transmit_errs_total{
  instance="198.51.100.10:9100",
  device="bond0"
}[$__rate_interval])
</syntaxhighlight>
 
<syntaxhighlight lang="bash" line>
rate(node_network_receive_drop_total{
  instance="198.51.100.10:9100",
  device="bond0"
}[$__rate_interval])
</syntaxhighlight>
 
<syntaxhighlight lang="bash" line>
rate(node_network_transmit_drop_total{
  instance="198.51.100.10:9100",
  device="bond0"
}[$__rate_interval])
</syntaxhighlight>
 
=== 16.6 Disk I/O ===
 
Read:
 
<syntaxhighlight lang="bash" line>
rate(node_disk_read_bytes_total{
  instance="198.51.100.10:9100",
  device=~"sd[a-z]+"
}[$__rate_interval])
</syntaxhighlight>
 
Write:
 
<syntaxhighlight lang="bash" line>
rate(node_disk_written_bytes_total{
  instance="198.51.100.10:9100",
  device=~"sd[a-z]+"
}[$__rate_interval])
</syntaxhighlight>
 
Read IOPS:
 
<syntaxhighlight lang="bash" line>
rate(node_disk_reads_completed_total{
  instance="198.51.100.10:9100",
  device=~"sd[a-z]+"
}[$__rate_interval])
</syntaxhighlight>
 
Write IOPS:
 
<syntaxhighlight lang="bash" line>
rate(node_disk_writes_completed_total{
  instance="198.51.100.10:9100",
  device=~"sd[a-z]+"
}[$__rate_interval])
</syntaxhighlight>
 
== 17. rsyslog 구성 ==
 
=== 17.1 Network Syslog ===
 
예시 로그 파일:
 
<syntaxhighlight lang="bash" line>
/var/log/network-syslog/events.log
</syntaxhighlight>
 
네트워크 장비:
 
<syntaxhighlight lang="bash" line>
UDP/TCP 514
</syntaxhighlight>
 
=== 17.2 Server Syslog ===
 
<syntaxhighlight lang="bash" line>
/var/log/server-syslog/events.log
</syntaxhighlight>
 
예시 포트:
 
<syntaxhighlight lang="bash" line>
5514/tcp
</syntaxhighlight>
 
=== 17.3 QNAP Event / Access 분리 ===
 
QNAP 로그는 Event와 Access를 별도 포트로 분리한다.


{| class="wikitable"
{| class="wikitable"
! 로그
! Port
! File
|-
|-
! 장비 설정
| Event
! 계획
| TCP 5515
| /var/log/qnap/event.log
|-
|-
| Collector
| Access
| <code>&lt;COLLECTOR_IP&gt;</code>
| TCP 5516
|-
| /var/log/qnap/access.log
| UDP port
| 6343, 수집기와 일치
|-
| Agent / source
| SNMP로 도달 가능한 고정 관리 IP
|-
| Sampling
| 지원 범위에서 1:1024 전후를 시험 시작값으로 검토
|-
| Counter polling
| 10~30초 중 장비 지원값
|-
| 관측 대상
| 중요 L2/L3 포트, 가능하면 ingress 중심
|}
|}


패킷 샘플링과 counter polling은 별개다. sFlow는 SIP 전체 패킷 캡처를 대체하지 않는다. IP별 사용률은 추정이고 L3만으로 같은 L2 내부 트래픽을 놓칠 수 있다. 여러 스위치·입출력의 동일 패킷을 중복 합산하지 않는다.
Template:
 
<syntaxhighlight lang="bash" line>
template(name="QnapSyslogLine" type="string"
  string="%timegenerated:::date-rfc3339% src=%fromhost-ip% severity=%syslogseverity-text% host=%hostname% %syslogtag%%msg:::sp-if-no-1st-sp%%msg%\n")
</syntaxhighlight>
 
Event:
 
<syntaxhighlight lang="bash" line>
ruleset(name="QnapEventLog") {
    action(
        type="omfile"
        file="/var/log/qnap/event.log"
        template="QnapSyslogLine"
        fileOwner="root"
        fileGroup="alloy"
        fileCreateMode="0640"
        dirOwner="root"
        dirGroup="alloy"
        dirCreateMode="0750"
        createDirs="on"
    )
    stop
}
 
input(
    type="imtcp"
    port="5515"
    ruleset="QnapEventLog"
)
</syntaxhighlight>
 
Access:
 
<syntaxhighlight lang="bash" line>
ruleset(name="QnapAccessLog") {
    action(
        type="omfile"
        file="/var/log/qnap/access.log"
        template="QnapSyslogLine"
        fileOwner="root"
        fileGroup="alloy"
        fileCreateMode="0640"
        dirOwner="root"
        dirGroup="alloy"
        dirCreateMode="0750"
        createDirs="on"
    )
    stop
}
 
input(
    type="imtcp"
    port="5516"
    ruleset="QnapAccessLog"
)
</syntaxhighlight>
 
검증:
 
<syntaxhighlight lang="bash" line>
rsyslogd -N1
systemctl restart rsyslog
 
ss -lntp | grep -E ':5514|:5515|:5516'
</syntaxhighlight>
 
== 18. Loki / Alloy 구성 ==
 
Loki Endpoint:
 
<syntaxhighlight lang="bash" line>
http://127.0.0.1:3100
</syntaxhighlight>
 
Alloy UI:
 
<syntaxhighlight lang="bash" line>
127.0.0.1:12345
</syntaxhighlight>
 
=== 18.1 Alloy 기본 설정 ===
 
<syntaxhighlight lang="bash" line>
logging {
  level = "info"
}
</syntaxhighlight>
 
=== 18.2 Network Syslog ===
 
<syntaxhighlight lang="bash" line>
loki.source.file "network_syslog" {
  targets = [
    {
      __path__ = "/var/log/network-syslog/events.log",
      job      = "network-syslog",
    },
  ]
 
  forward_to = [loki.process.network_syslog.receiver]
}
 
loki.process "network_syslog" {
  stage.regex {
    expression = `^(?P<received_at>\S+) src=(?P<device_ip>\S+) severity=(?P<severity>\S+) host=(?P<device_host>\S+) (?P<message>.*)$`
  }
 
  stage.timestamp {
    source            = "received_at"
    format            = "RFC3339Nano"
    action_on_failure = "skip"
  }
 
  stage.labels {
    values = {
      device_ip = "",
      severity  = "",
    }
  }
 
  forward_to = [loki.write.local.receiver]
}
</syntaxhighlight>
 
=== 18.3 Server Syslog ===
 
<syntaxhighlight lang="bash" line>
loki.source.file "server_syslog" {
  targets = [
    {
      __path__ = "/var/log/server-syslog/events.log",
      job      = "server-syslog",
    },
  ]
 
  forward_to = [loki.process.server_syslog.receiver]
}
 
loki.process "server_syslog" {
  stage.regex {
    expression = `^(?P<received_at>\S+) src=(?P<server_ip>\S+) severity=(?P<severity>\S+) host=(?P<server_host>\S+) (?P<source>[^:\s\[]+)(?:\[\d+\])?:?\s+(?P<message>.*)$`
  }
 
  stage.timestamp {
    source            = "received_at"
    format            = "RFC3339Nano"
    action_on_failure = "skip"
  }
 
  stage.labels {
    values = {
      server_ip  = "",
      server_host = "",
      severity    = "",
      source      = "",
    }
  }
 
  forward_to = [loki.write.local.receiver]
}
</syntaxhighlight>
 
=== 18.4 QNAP Event ===
 
<syntaxhighlight lang="bash" line>
loki.source.file "qnap_event" {
  targets = [
    {
      __path__ = "/var/log/qnap/event.log",
      job      = "qnap-event",
    },
  ]
 
  forward_to = [loki.process.qnap_event.receiver]
}
 
loki.process "qnap_event" {
  stage.regex {
    expression = `^(?P<received_at>\S+) src=(?P<nas_ip>\S+) severity=(?P<severity>\S+) host=(?P<nas_host>\S+) (?P<message>.*)$`
  }
 
  stage.timestamp {
    source            = "received_at"
    format            = "RFC3339Nano"
    action_on_failure = "skip"
  }
 
  stage.labels {
    values = {
      nas_ip  = "",
      nas_host = "",
      severity = "",
      log_type = "event",
    }
  }
 
  forward_to = [loki.write.local.receiver]
}
</syntaxhighlight>
 
=== 18.5 QNAP Access ===
 
<syntaxhighlight lang="bash" line>
loki.source.file "qnap_access" {
  targets = [
    {
      __path__ = "/var/log/qnap/access.log",
      job      = "qnap-access",
    },
  ]
 
  forward_to = [loki.process.qnap_access.receiver]
}
 
loki.process "qnap_access" {
  stage.regex {
    expression = `^(?P<received_at>\S+) src=(?P<nas_ip>\S+) severity=(?P<severity>\S+) host=(?P<nas_host>\S+) (?P<message>.*)$`
  }
 
  stage.timestamp {
    source            = "received_at"
    format            = "RFC3339Nano"
    action_on_failure = "skip"
  }
 
  stage.labels {
    values = {
      nas_ip  = "",
      nas_host = "",
      severity = "",
      log_type = "access",
    }
  }
 
  forward_to = [loki.write.local.receiver]
}
</syntaxhighlight>
 
=== 18.6 Loki Write ===
 
<syntaxhighlight lang="bash" line>
loki.write "local" {
  endpoint {
    url = "http://127.0.0.1:3100/loki/api/v1/push"
  }
}
</syntaxhighlight>
 
검증:
 
<syntaxhighlight lang="bash" line>
alloy validate /etc/alloy/config.alloy
 
systemctl restart alloy
 
systemctl status alloy --no-pager
</syntaxhighlight>
 
Loki Job 확인:
 
<syntaxhighlight lang="bash" line>
curl -s \
  'http://127.0.0.1:3100/loki/api/v1/label/job/values' \
  | jq
</syntaxhighlight>
 
정상 예:
 
<syntaxhighlight lang="bash" line>
network-syslog
server-syslog
qnap-event
qnap-access
</syntaxhighlight>
 
=== 18.7 Alloy Permission 문제 ===
 
다음과 같은 오류가 발생할 수 있다.
 
<syntaxhighlight lang="bash" line>
failed to tail file
stat failed
permission denied
</syntaxhighlight>
 
확인:
 
<syntaxhighlight lang="bash" line>
namei -l /var/log/qnap/access.log
 
systemctl show alloy \
  -p User \
  -p Group
 
sudo -u alloy \
  head /var/log/qnap/access.log
</syntaxhighlight>
 
권장 권한:
 
<syntaxhighlight lang="bash" line>
drwxr-x--- root alloy /var/log/qnap
-rw-r----- root alloy /var/log/qnap/access.log
-rw-r----- root alloy /var/log/qnap/event.log
</syntaxhighlight>
 
SELinux 확인:
 
<syntaxhighlight lang="bash" line>
getenforce
 
ausearch -m AVC -ts recent \
  | grep -Ei 'alloy|qnap'
 
ls -Zd /var/log/qnap
ls -Z /var/log/qnap/access.log
</syntaxhighlight>
 
== 19. Loki Query ==
 
Network:
 
<syntaxhighlight lang="bash" line>
{job="network-syslog"}
</syntaxhighlight>
 
Server:
 
<syntaxhighlight lang="bash" line>
{job="server-syslog"}
</syntaxhighlight>
 
QNAP Event:
 
<syntaxhighlight lang="bash" line>
{job="qnap-event"}
</syntaxhighlight>
 
QNAP Access:
 
<syntaxhighlight lang="bash" line>
{job="qnap-access"}
</syntaxhighlight>
 
Critical Network Syslog:
 
<syntaxhighlight lang="bash" line>
{job="network-syslog",severity=~"emerg|alert|crit|err"}
</syntaxhighlight>
 
== 20. Grafana Syslog Alert ==
 
Syslog Level 3 이상:
 
<syntaxhighlight lang="bash" line>
sum by (device_ip, severity) (
  count_over_time(
    {
      job="network-syslog",
      severity=~"emerg|alert|crit|err"
    }[1m]
  )
)
</syntaxhighlight>
 
권장 Alert 구조:
 
<syntaxhighlight lang="bash" line>
A = Loki Instant Query
B = Threshold
    A IS ABOVE 0
</syntaxhighlight>
 
Summary:
 
<syntaxhighlight lang="bash" line>
[Syslog 경고] {{ $labels.device_ip }} - {{ $labels.severity }}
</syntaxhighlight>
 
Description:
 
<syntaxhighlight lang="bash" line>
장비 {{ $labels.device_ip }} 에서 Syslog Level 3(Error) 이상의 로그가 발생했습니다.
최근 1분 발생 건수: {{ $values.A.Value }}
Severity: {{ $labels.severity }}
</syntaxhighlight>


=== 12.2 Akvorado 도입 방안 ===
Range Query를 그대로 Alert 조건으로 사용할 경우 다음 오류가 발생할 수 있다.


[https://github.com/akvorado/akvorado Akvorado 공식 프로젝트]는 NetFlow/IPFIX/sFlow, SNMP 메타데이터 보강, ClickHouse 저장과 웹 분석을 제공한다. 공식 Docker Compose는 본문의 RPM 운영과 별개다. '''이전 설명의 Podman Compose 호환을 전제하지 않는다.''' Docker와 Podman 호환 패키지도 혼용하지 않는다.
<syntaxhighlight lang="bash" line>
invalid format of evaluation results for the alert definition A:
looks like time series data, only reduced data can be alerted on.
</syntaxhighlight>


공식 <code>.env</code>는 여러 <code>docker/docker-compose-*.yml</code>을 조합한다. main은 개발 구성을 포함하므로 그대로 운영 배포하지 않는다. [https://github.com/akvorado/akvorado/tree/main/docker 공식 구성 디렉터리].
Range Query를 유지할 경우 Reduce Expression을 추가해야 한다.


지원 확인 후 KVM 스냅샷과 방화벽 백업을 확보하고 Docker Engine의 RHEL/CentOS계 설치 방법을 검토한다. Rocky가 Docker의 공식 지원 대상과 동일하다고 가정하지 않는다. [https://docs.docker.com/engine/install/centos/ Docker CentOS 절차].
== 21. ICMP Alert ==


신규 전용 VM이며 패키지 충돌이 없다고 확인된 경우의 준비 명령:
Flux:


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
rpm -q podman-docker docker-ce docker-ce-cli containerd.io
from(bucket: "snmp_raw")
dnf config-manager --add-repo https://download.docker.com/linux/centos/docker-ce.repo
  |> range(start: -5m)
dnf list --showduplicates docker-ce docker-compose-plugin
  |> filter(fn: (r) =>
dnf install docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin
    r._measurement == "ping" and
    r._field == "percent_packet_loss"
  )
  |> group(columns: ["url"])
  |> last()
  |> keep(columns: ["_time", "_value", "url"])
</syntaxhighlight>
</syntaxhighlight>
충돌 패키지를 자동 삭제하지 않는다. Docker는 전송·패킷 필터 규칙을 변경할 수 있으므로 KVM 게스트에서만 적용하고 관리 경로를 검증한다. KVM 호스트에서는 실행하지 않는다.
 
Alert:


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
systemctl enable --now docker
A = Flux Query
docker version
B = Reduce / Last / Strict
docker compose version
C = Threshold > 99
dnf install -y git
Pending = 2m
git clone https://github.com/akvorado/akvorado.git /opt/akvorado
cd /opt/akvorado
git tag --sort=-version:refname | head -20
</syntaxhighlight>
</syntaxhighlight>
기존 파일이 있는 <code>/opt/akvorado</code>에는 clone하지 않는다. 공식 안정 릴리스의 태그를 선택하고 해당 Changelog·문서를 읽는다. 본문에서 태그는 미확정이다.
 
No Data:


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
git switch --detach <VERIFIED_STABLE_TAG>
Keep Last State
git rev-parse HEAD
docker compose config --services
docker compose config --images
docker compose config --quiet
</syntaxhighlight>
</syntaxhighlight>
<code>docker compose config --quiet</code>는 Compose 구문만 검사한다. 앱 설정 검증과 다르다. 아래 환경값을 '''선택한 릴리스의 schema로 구성한 후''' 기동한다. 추정한 키 이름으로 작성한 YAML은 제공하지 않는다.
 
== 22. 최종 Dashboard 구성안 ==
 
=== 22.1 Network Switch Dashboard ===


{| class="wikitable"
{| class="wikitable"
! Row
! Panels
|-
|-
! 필수 설정
| 상태
! 확인 사항
| Device / Uptime / CPU / Memory / Ping
|-
|-
| sFlow 입력
| Port 상태
| UDP6343 수신·바인딩
| ifName / Alias / Speed / OperStatus
|-
|-
| SNMP metadata
| Traffic
| agent/source 주소, 장비별 인증값, 포트명 보강
| RX bps / TX bps
|-
|-
| 내부 네트워크
| Packet
| Voice/Data CIDR과 사설 IP 분류
| Unicast / Broadcast / Multicast PPS
|-
|-
| 불필요 서비스
| Error
| demo exporter 비활성, Grafana 중복 배치 금지, GeoIP 의존성 확인
| RX/TX Error / Discard
|-
|-
| 인증·공개
| Syslog
| 웹은 loopback+SSH 또는 인증 TLS, DB/Kafka는 외부 공개 금지
| Warning / Error / Critical Event
|-
| 보존
| 원시·집계 데이터 기간, ClickHouse TTL, 용량 계획
|-
| 영속화·SELinux
| named volume 또는 전용 bind mount 라벨, 공유 OS 경로에 <code>:Z</code> 사용 금지
|-
| 재시작
| 서비스 restart policy, Docker 자동 기동
|}
|}


'''Docker 공개 포트를 firewalld INPUT 규칙만으로 제한할 수 있다고 가정하지 않는다.''' 실제 백엔드의 전달 규칙/DOCKER-USER 상당 경로, 상위 ACL과 관리 IP 바인딩을 확인하고 비허가 단말에서 차단되는지 시험한다. 관리 IP 바인딩은 송신원 제한이 아니다. §10 rich rule만으로 Docker의 UDP6343 제한을 완료 처리하지 않는다.
=== 22.2 Linux Server Dashboard ===
 
* Uptime
* CPU
* Memory
* Load
* Disk Capacity
* Disk I/O
* Network RX/TX
* Network Error/Drop
* Service Log
* Security Event


위 조건을 확정한 뒤 실행할 기동·검증 명령:
=== 22.3 Windows Dashboard ===
 
* Uptime
* CPU
* Memory
* Logical Disk
* Disk I/O
* Network
* Windows Exporter 상태
 
=== 22.4 QNAP Dashboard ===
 
최종 구성:
 
<syntaxhighlight lang="bash" line>
Node Exporter
Uptime
CPU
Memory
HDD 최고 온도
 
CPU / Memory / Load
 
Network RX / TX
  - bond0 only
 
Disk 상태
 
Volume
  - DataVol1
  - SNMP capacity/free x1024 보정
 
Filesystem 사용률
  - /share/CACHEDEV1_DATA only
  - Snapshot 제외
 
Network Errors / Drops
  - bond0 RX Error
  - bond0 TX Error
  - bond0 RX Drop
  - bond0 TX Drop
 
Disk I/O
  - Physical sd* only
  - Read B/s
  - Write B/s
  - Read IOPS
  - Write IOPS
 
QNAP Event Log
 
QNAP Access Log
</syntaxhighlight>
 
QNAP Dashboard에서 제거한 항목:
 
* 상단 RAID 상태
* Storage Pool 상태
* Storage Pool 사용률
* RAID 상세 Table
* Storage Pool 상세 Table
 
RAID/Storage Pool 정보는 필요 시 별도 상세 Dashboard에서 조회한다.
 
== 23. 운영 점검 ==
 
전체 서비스:


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
cd /opt/akvorado
systemctl is-active \
docker compose pull
  influxdb \
docker compose up -d
  telegraf \
docker compose ps
  grafana-server \
docker compose logs --tail=100
  prometheus \
  loki \
  alloy \
  rsyslog \
  chronyd
</syntaxhighlight>
 
Listening Port:
 
<syntaxhighlight lang="bash" line>
ss -lntup
</syntaxhighlight>
</syntaxhighlight>
태그·설정이 미정인 상태에서 일괄 실행하는 명령이 아니다. 같은 릴리스의 앱 설정 검사 방법을 따르고 healthy 표시뿐 아니라 실제 데이터 표시도 확인한다.


=== 12.3 수신·분석 합격 기준 ===
Prometheus:


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
timeout 20 tcpdump -ni any -nn 'udp dst port 6343'
curl -s http://127.0.0.1:9090/-/healthy
ss -lunp | grep ':6343'
nstat -az | grep -E 'Udp(InErrors|RcvbufErrors|InDatagrams)'
</syntaxhighlight>
</syntaxhighlight>
Docker의 포트 구현에 따라 호스트 <code>ss</code>에 예상 소켓이 안 보일 수 있다. 실제 Compose 포트 매핑과 수신 메트릭을 함께 확인한다.


합격: 스위치 source에서 패킷 도착, 수집·디코드 카운터 증가, 오류 급증 없음, 웹에서 exporter·입력 포트·단말 IP·시간대 표시, 포트 번호의 실제 이름 매핑. tcpdump 도착만으로 성공 판정하지 않는다.
Loki:


미지원 장비는 SNMP/Syslog를 유지하고 필요한 사건만 승인된 포트 미러링으로 분석한다. 미러링은 별도 vNIC·물리 연결·대역폭 검토가 필요하며 현재 관리 NIC를 임의로 미러 수신 포트로 바꾸지 않는다.
<syntaxhighlight lang="bash" line>
curl -s http://127.0.0.1:3100/ready
</syntaxhighlight>


== 13. 백업·롤백 ==
Telegraf:


=== 13.1 일반 백업 ===
<syntaxhighlight lang="bash" line>
journalctl -u telegraf \
  --since '-10 min' \
  --no-pager
</syntaxhighlight>


Influx CLI 관리자 인증으로 여유 공간이 있는 새 디렉터리에 저장한다. Grafana 읽기 전용·Telegraf 쓰기 전용 토큰으로는 백업할 수 없다. [https://docs.influxdata.com/influxdb/v2/reference/cli/influx/backup/ InfluxDB backup].
Alloy:


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
umask 077
journalctl -u alloy \
MON_ARCHIVE="/root/ipt-monitor-save-$(date +%Y%m%d-%H%M%S)"
  --since '-10 min' \
install -d -m 700 "$MON_ARCHIVE"
  --no-pager
influx backup "$MON_ARCHIVE/influxdb"
</syntaxhighlight>
tar -czf "$MON_ARCHIVE/configs.tar.gz" \
 
    /etc/telegraf /etc/grafana /etc/rsyslog.conf /etc/rsyslog.d \
Filesystem:
    /etc/logrotate.d/zyxel-remote /etc/chrony.conf /etc/firewalld \
 
    /etc/systemd/system/telegraf.service.d \
<syntaxhighlight lang="bash" line>
    /etc/systemd/system/influxdb.service.d
df -hT
</syntaxhighlight>
</syntaxhighlight>
백업에는 SNMP 비밀과 토큰이 포함된다. root 전용 권한을 유지하고 접근제어된 다른 매체에 암호화 복제한다. root 디스크에만 둔 것은 백업 완료로 간주하지 않는다.


Grafana 기본 SQLite의 실행 중 파일만 단순 복사하지 않는다. 화면 중단이 허용되는 시간에 다음을 수행한다:
System I/O:


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
systemctl stop grafana-server
vmstat 1 5
tar -czf "$MON_ARCHIVE/grafana-data.tar.gz" /var/lib/grafana
iostat -xz 1 5
systemctl start grafana-server
curl -fsS http://127.0.0.1:3000/api/health
</syntaxhighlight>
</syntaxhighlight>
이 동안 Telegraf 수집은 계속된다. 외부 DB를 사용하는 Grafana는 별도 DB 정합 백업이 필요하다. 장애 구간 Syslog·콜서버 pcap도 보존한다. 기록 중인 파일은 변경될 수 있으므로 닫힌 파일이나 정합 스냅샷으로 증거를 확보한다.


=== 13.2 설정 변경 롤백 ===
== 24. 장애 점검 순서 ==


* 고주기 SNMP 부하: <code>20-zyxel-fast.conf</code>의 interval을 5/10초로 복귀 후 검사·재시작. 긴급 시 <code>systemctl stop telegraf</code>를 수행하고 수집 중단을 기록한다.
=== Prometheus Target Down ===
* Telegraf 오류: 해당 설정의 변경 전 버전 복원 → 검사 → 재시작. DB는 삭제하지 않는다.
* 신규 Syslog 설정만 취소하려면 아래처럼 비활성화한다. 수신 중단을 알리고 기존 파일을 수정했던 경우에는 변경 전 버전으로 복원한다.


<syntaxhighlight lang="bash" line>
<syntaxhighlight lang="bash" line>
mv /etc/rsyslog.d/30-zyxel-remote.conf /etc/rsyslog.d/30-zyxel-remote.conf.disabled
curl http://TARGET_IP:9100/metrics
rsyslogd -N1
 
systemctl restart rsyslog
journalctl -u prometheus \
  -n 100 \
  --no-pager
</syntaxhighlight>
 
=== SNMP Timeout ===
 
<syntaxhighlight lang="bash" line>
snmpget -v3 \
  -t 5 \
  -r 1 \
  -On \
  SWITCH_IP \
  .1.3.6.1.2.1.1.3.0
</syntaxhighlight>
</syntaxhighlight>
* firewalld는 §10에서 추가한 정확히 일치하는 규칙만 제거한다.
* Akvorado 중지는 <code>/opt/akvorado</code>에서 <code>docker compose stop</code>. <code>down -v</code>, volume 삭제, ClickHouse DROP은 하지 않는다.
* 장비의 sFlow는 추가한 export 설정만 되돌린다. Voice VLAN·PoE·STP·QoS는 일괄 초기화하지 않는다.


=== 13.3 데이터 복원·업그레이드 ===
확인 항목:


운영 InfluxDB에 full restore를 바로 수행하면 덮어쓰기 위험이 있다. 격리된 새 VM에 같은 버전으로 공식 restore 절차를 시험한다. Grafana는 중지 후 동일 구성의 데이터·설정을 복원하고 소유자/SELinux 라벨을 확인해 기동한다. 새 운영 데이터와의 통합은 별도 계획한다.
* SNMP User
* SHA/DES/AES 조합
* ACL
* Source IP
* Timeout
* SNMP View
* 장비 CPU


업그레이드: RPM/이미지 버전 기록 → 백업 → 시험 VM 검증 → 운영 적용. 메이저 변경·DB schema 변경 후 바이너리만 내려서 롤백하지 않는다.
=== Syslog 미수신 ===


== 14. 장애 확인표·인수 체크리스트 ==
<syntaxhighlight lang="bash" line>
ss -lntup \
  | grep -E ':514|:5514|:5515|:5516'


{| class="wikitable"
tcpdump -ni any \
|-
  host DEVICE_IP
! 이상
 
! 첫 확인
tail -f /var/log/qnap/access.log
|-
</syntaxhighlight>
| SNMP timeout
| 같은 조건 CLI, 관리 경로·ACL·CPU·인증·암호화
|-
| No Such Object/Instance
| view·모델 지원·ifIndex·인터페이스 선택
|-
| Telegraf 값은 있는데 Grafana가 비어 있음
| DB 쓰기 토큰·bucket·measurement/tag·UTC/KST·시간 범위
|-
| 계단형 그래프·이상 피크
| 카운터 갱신 주기·리셋·결측·중복 수집
|-
| 재부팅 후 다른 포트 표시
| ifIndex 재조회, 고정 OID·태그 수정
|-
| Syslog는 도착하지만 파일 없음
| ruleset·구문·권한·SELinux·실제 송신 IP
|-
| sFlow 패킷만 도착
| 포트 매핑·디코딩·metadata·DB·시간·필터
|-
| PoE/CPU 없음
| 본 구성 미포함. 모델 전용 MIB 필요
|-
| 전화 장애 때 전체 데이터 결측
| VM/KVM 부하·관리 경로·NTP·수신 drop
|}


=== 초기 구성 완료 조건 ===
=== Loki 미표시 ===


* ☐ OS/RPM 버전과 IP·포트 대응표를 저장했다.
<syntaxhighlight lang="bash" line>
* ☐ SNMPv3 읽기 전용으로 대상 조회 및 ifIndex를 확인했다.
alloy validate \
* ☐ 실데이터의 bytes·PPS·Broadcast·Discard가 확인된다.
  /etc/alloy/config.alloy
* ☐ Grafana가 누적값 대신 rate를 표시하고 상태는 원래 값으로 표시한다.
* ☐ DB 저장·화면 표시와 재부팅 후 수집 지속을 확인했다.
* ☐ 실장비 Syslog 저장과 logrotate dry-run을 확인했다.
* ☐ 콜서버·스위치와 NTP가 일치한다.
* ☐ DB/API 비공개, SSH·송신원 허용을 확인했다.
* ☐ 결측을 0/정상으로 표시하지 않는다.
* ☐ sFlow·PoE·CPU 미확인 항목을 완료 처리하지 않는다.
* ☐ 보존 만료 전 장애 증거 보관과 복원 방침을 정했다.


== 15. 공식 자료·대상 버전 ==
journalctl -u alloy \
  -n 100 \
  --no-pager


공식 자료 확인일: 2026-09-10. 지속 갱신 문서는 고정 게시일이 없으며 실제 RPM 버전을 §4 출력으로 기록한다.
curl -s \
  'http://127.0.0.1:3100/loki/api/v1/label/job/values' \
  | jq
</syntaxhighlight>


{| class="wikitable"
== 25. 운영 원칙 ==
|-
! 대상
! 공식 자료
|-
| Telegraf 1.x
| [https://docs.influxdata.com/telegraf/v1/install/ 설치] / [https://docs.influxdata.com/telegraf/v1/configuration/ 설정] / [https://docs.influxdata.com/telegraf/v1/input-plugins/snmp/ SNMP] / [https://docs.influxdata.com/telegraf/v1/input-plugins/ping/ Ping]
|-
| InfluxDB OSS 2.x
| [https://docs.influxdata.com/influxdb/v2/install/?t=Linux RPM 설치] / [https://docs.influxdata.com/influxdb/v2/reference/cli/influx/backup/ backup]
|-
| Flux
| [https://docs.influxdata.com/flux/v0/stdlib/universe/derivative/ derivative]
|-
| Grafana OSS
| [https://grafana.com/docs/grafana/latest/setup-grafana/installation/redhat-rhel-fedora/ RPM 설치] / [https://grafana.com/docs/grafana/latest/datasources/influxdb/configure/ InfluxDB 연결]
|-
| Net-SNMP
| [https://www.net-snmp.org/docs/man/snmp.conf.html snmp.conf]
|-
| IF-MIB
| [https://www.rfc-editor.org/rfc/rfc2863.html RFC 2863, 2000-06]
|-
| rsyslog 8.x
| [https://docs.rsyslog.com/doc/configuration/modules/imudp.html imudp] / [https://docs.rsyslog.com/doc/configuration/modules/omfile.html omfile]
|-
| Akvorado
| [https://github.com/akvorado/akvorado 공식 프로젝트] / [https://github.com/akvorado/akvorado/tree/main/docker Compose]
|-
| Docker Engine
| [https://docs.docker.com/engine/install/centos/ CentOS 설치 참고]
|}


추가로 필요한 정보: Zyxel L3/L2 모델·펌웨어·대수·관리 IP/Voice VLAN 구조. 확인 후 장비별 SNMPv3/Syslog/sFlow 명령과 CPU·PoE·큐 OID를 확정할 수 있다.
* SNMP Metric은 Telegraf → InfluxDB로 저장한다.
* Host Metric은 Prometheus로 저장한다.
* Syslog는 rsyslog → Alloy → Loki로 저장한다.
* Grafana는 세 데이터소스를 통합한다.
* SNMP Counter는 누적값을 그대로 표시하지 않고 rate/derivative 계산한다.
* No Data를 0 또는 정상으로 강제 표시하지 않는다.
* QNAP Snapshot filesystem은 Data Volume 사용률에서 제외한다.
* QNAP Volume SNMP capacity/free 값은 장비 특성상 ×1024 보정한다.
* QNAP Network는 실제 활성 bond0만 표시한다.
* Loki message 전체를 label로 만들지 않는다.
* 장비 Syslog Alert는 device_ip / severity와 발생 건수를 메일에 포함한다.
* Alert의 Range Query는 Reduce 또는 Instant Query 구조로 구성한다.
* 모든 Token 및 SNMP 비밀번호는 문서에 실제 값을 기록하지 않는다.

2026년 9월 23일 (수) 20:22 판

Rocky Linux 9 통합 모니터링 서버 구축 가이드

Telegraf / InfluxDB / Grafana / Prometheus / Node Exporter / Windows Exporter / Loki / Alloy / rsyslog

작성 기준: 실제 구성 및 장애 처리 기록 기반 문서 형식: MediaWiki IP 주소: 문서용 예시 주소로 치환 운영환경 적용 전 실제 장비 주소, 계정, 토큰, OID를 확인한다.

0. 목적 및 최종 구성

본 문서는 Rocky Linux 9 기반 모니터링 서버를 처음 구축하는 단계부터 네트워크 장비 SNMP, 서버 Metric, Syslog, QNAP NAS 모니터링 및 Grafana Dashboard 구성까지 정리한다.

최종 구조는 다음과 같다.

Network Switch
    |
    +-- SNMPv3 UDP/161
    |      |
    |      +--> Telegraf
    |              |
    |              +--> InfluxDB OSS 2.x
    |
    +-- Syslog TCP/UDP
           |
           +--> rsyslog
                    |
                    +--> Log File
                             |
                             +--> Grafana Alloy
                                      |
                                      +--> Loki

Linux / QNAP
    |
    +-- node_exporter :9100
             |
             +--> Prometheus

Windows
    |
    +-- windows_exporter :9182
             |
             +--> Prometheus

Prometheus + InfluxDB + Loki
             |
             +--> Grafana

0.1 문서용 IP 주소

본 문서에서는 실제 운영 IP를 노출하지 않고 RFC 문서용 주소를 사용한다.

대상 예시 주소 용도
Monitoring Server 192.0.2.10 Grafana / Prometheus / InfluxDB / Telegraf / Loki / Alloy / rsyslog
Linux Server 192.0.2.20 node_exporter
Windows Server 192.0.2.30 windows_exporter
QNAP NAS 198.51.100.10 node_exporter / SNMP / Syslog
Switch-01 203.0.113.11 SNMP / Syslog
Switch-02 203.0.113.12 SNMP / Syslog
Switch-03 203.0.113.13 SNMP / Syslog

1. Rocky Linux 기본 준비

1.1 시스템 상태 확인

cat /etc/rocky-release
uname -m
ip -br address
ip route
df -hT
free -h
getenforce
ss -lntup

SELinux와 firewalld는 초기부터 비활성화하지 않는다.

1.2 백업 디렉터리 생성

기존 서버에 추가 설치하는 경우 설정 파일 백업을 먼저 수행한다.

umask 077

MON_BACKUP="/root/monitor-backup-$(date +%Y%m%d-%H%M%S)"

install -d -m 700 "$MON_BACKUP"

cp -a /etc/rsyslog.conf "$MON_BACKUP/" 2>/dev/null || true
cp -a /etc/rsyslog.d "$MON_BACKUP/" 2>/dev/null || true
cp -a /etc/chrony.conf "$MON_BACKUP/" 2>/dev/null || true
cp -a /etc/firewalld "$MON_BACKUP/" 2>/dev/null || true

rpm -qa | sort > "$MON_BACKUP/packages-before.txt"

1.3 기본 패키지 설치

dnf install -y \
    curl \
    ca-certificates \
    gnupg2 \
    vim-enhanced \
    net-snmp-utils \
    rsyslog \
    logrotate \
    chrony \
    tcpdump \
    iputils \
    sysstat \
    policycoreutils-python-utils \
    dnf-plugins-core \
    jq

1.4 시간 동기화

모니터링에서는 장비와 서버 간 시간이 맞아야 장애 시각을 비교할 수 있다.

timedatectl set-timezone Asia/Seoul

systemctl enable --now chronyd
systemctl restart chronyd

chronyc tracking
chronyc sources -v
date -Ins

2. InfluxDB / Telegraf / Grafana 설치

2.1 InfluxData Repository

curl -fL https://repos.influxdata.com/influxdata-archive.key \
    -o /etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata

gpg --show-keys --with-fingerprint \
    /etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata

rpm --import /etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata

Repository:

vi /etc/yum.repos.d/influxdata.repo
[influxdata]
name=InfluxData Repository - Stable
baseurl=https://repos.influxdata.com/stable/$basearch/main
enabled=1
gpgcheck=1
gpgkey=file:///etc/pki/rpm-gpg/RPM-GPG-KEY-influxdata
sslverify=1

2.2 Grafana Repository

vi /etc/yum.repos.d/grafana.repo
[grafana]
name=Grafana OSS repository
baseurl=https://rpm.grafana.com
repo_gpgcheck=1
enabled=1
gpgcheck=1
gpgkey=https://rpm.grafana.com/gpg.key
sslverify=1

2.3 패키지 설치

dnf makecache

dnf list --showduplicates \
    telegraf \
    influxdb2 \
    influxdb2-cli \
    grafana

dnf install -y \
    telegraf \
    influxdb2 \
    influxdb2-cli \
    grafana

버전 확인:

rpm -q telegraf influxdb2 influxdb2-cli grafana rsyslog

telegraf --version
influxd version
influx version

실제 구축 과정에서는 Telegraf 1.40.0 환경에서 동작을 확인하였다.

3. InfluxDB 초기 구성

3.1 Localhost Bind

InfluxDB API는 Grafana 및 Telegraf가 같은 서버에 있으므로 localhost만 수신하도록 구성한다.

install -d -m 755 /etc/systemd/system/influxdb.service.d

vi /etc/systemd/system/influxdb.service.d/10-listen.conf
[Service]
Environment="INFLUXD_HTTP_BIND_ADDRESS=127.0.0.1:8086"
systemctl daemon-reload
systemctl enable --now influxdb
systemctl restart influxdb

systemctl status influxdb --no-pager

ss -lntp | grep ':8086'

curl -fsS http://127.0.0.1:8086/health

3.2 Initial Setup

influx setup

예시:

설정 값
Organization network
Bucket snmp_raw
Retention 720h

Bucket 확인:

influx bucket list --org network

3.3 Token 분리

서비스별 Token 권한을 분리한다.

  • telegraf-write : snmp_raw Write
  • grafana-read : snmp_raw Read
  • Operator Token : 관리자용

운영 문서에 실제 Token 값을 기록하지 않는다.

4. Telegraf 기본 구성

4.1 환경 변수 파일

vi /etc/telegraf/monitor.env
INFLUX_WRITE_TOKEN='<WRITE_TOKEN>'

ZYXEL_SNMP_USER='<SNMP_USER>'
ZYXEL_SNMP_AUTH='<AUTH_PASSWORD>'
ZYXEL_SNMP_PRIV='<PRIV_PASSWORD>'

QNAP_SNMP_USER='<QNAP_SNMP_USER>'
QNAP_SNMP_AUTH='<QNAP_AUTH_PASSWORD>'
QNAP_SNMP_PRIV='<QNAP_PRIV_PASSWORD>'
chown root:root /etc/telegraf/monitor.env
chmod 600 /etc/telegraf/monitor.env

install -d -m 755 /etc/systemd/system/telegraf.service.d

vi /etc/systemd/system/telegraf.service.d/10-monitor-env.conf
[Service]
EnvironmentFile=/etc/telegraf/monitor.env

4.2 Main Configuration

cp -a /etc/telegraf/telegraf.conf \
    /etc/telegraf/telegraf.conf.orig

vi /etc/telegraf/telegraf.conf
[agent]
  interval = "30s"
  round_interval = true
  flush_interval = "5s"
  precision = "1ms"
  metric_batch_size = 1000
  metric_buffer_limit = 20000
  omit_hostname = false
  snmp_translator = "gosmi"

[[outputs.influxdb_v2]]
  urls = ["http://127.0.0.1:8086"]
  token = "${INFLUX_WRITE_TOKEN}"
  organization = "network"
  bucket = "snmp_raw"

[[inputs.internal]]

[[inputs.cpu]]
  percpu = false
  totalcpu = true

[[inputs.mem]]

[[inputs.disk]]
  mount_points = ["/"]

[[inputs.net]]

5. SNMPv3 사전 확인

SNMP는 v3 authPriv 사용을 기본으로 한다.

Net-SNMP 테스트 파일:

install -d -m 700 /root/.snmp

vi /root/.snmp/snmp.conf
defVersion 3
defSecurityName <SNMP_USER>
defSecurityLevel authPriv
defAuthType SHA
defAuthPassphrase <SNMP_AUTH_PASSWORD>
defPrivType AES
defPrivPassphrase <SNMP_PRIV_PASSWORD>
chmod 600 /root/.snmp/snmp.conf

snmpget -v3 -t 2 -r 0 -On \
    203.0.113.11 \
    .1.3.6.1.2.1.1.3.0

snmpwalk -v3 -t 2 -r 0 -On \
    203.0.113.11 \
    .1.3.6.1.2.1.31.1.1.1.1

ifName OID 마지막 값이 ifIndex이다.

6. Telegraf SNMP 구성

설정 파일은 장비 모델별로 분리한다.

/etc/telegraf/telegraf.d/
├── 10-zyxel-gs1900.conf
├── 11-zyxel-gs1920.conf
├── 20-zyxel-es3128.conf
├── 30-icmp_check.conf
├── 40-qnap_nas.conf
└── sflow.conf

6.1 GS1920 계열

예시 대상:

203.0.113.21
203.0.113.22
203.0.113.23

SNMPv3:

  • SHA
  • DES

CPU OID:

.1.3.6.1.4.1.890.1.15.3.49.1.7.0

Memory:

Total   .1.3.6.1.4.1.890.1.15.3.50.1.1.1.3.1
Used    .1.3.6.1.4.1.890.1.15.3.50.1.1.1.4.1
Percent .1.3.6.1.4.1.890.1.15.3.50.1.1.1.5.1

6.2 GS1900 계열

CPU:

.1.3.6.1.4.1.890.1.15.3.2.4.0

Memory:

.1.3.6.1.4.1.890.1.15.3.2.5.0

실제 구축에서는 2초 timeout / retries 0에서 응답 누락이 발생하여 다음과 같이 완화하였다.

timeout = "5s"
retries = 1

6.3 ES-3128GP

SNMPv3:

  • SHA
  • AES

sysObjectID:

.1.3.6.1.4.1.7800.1.190

CPU / Memory private OID는 장비에서 정상 확인되지 않아 Dashboard에서는 N/A 처리한다.

6.4 IF-MIB 주요 항목

수집 대상:

  • ifName
  • ifAlias
  • ifSpeed
  • ifHCInOctets
  • ifHCOutOctets
  • ifOperStatus
  • ifInErrors
  • ifOutErrors
  • ifInDiscards
  • ifOutDiscards
  • Broadcast
  • Multicast

누적 Counter는 Grafana에서 그대로 표시하지 않고 derivative/rate 계산 후 표시한다.

7. ICMP 수집

Telegraf ping input을 사용한다.

vi /etc/telegraf/telegraf.d/30-icmp_check.conf

예시 대상:

203.0.113.11
203.0.113.12
203.0.113.13

권장 설정:

[[inputs.ping]]
  urls = [
    "203.0.113.11",
    "203.0.113.12",
    "203.0.113.13"
  ]

  method = "native"
  count = 3
  deadline = 2.0
  interval = 10.0

native ping은 CAP_NET_RAW가 필요하다.

systemctl edit telegraf
[Service]
CapabilityBoundingSet=CAP_NET_RAW
AmbientCapabilities=CAP_NET_RAW
systemctl daemon-reload
systemctl restart telegraf

8. Telegraf Test 및 서비스 시작

설정 검사:

telegraf \
  --config /etc/telegraf/telegraf.conf \
  --config-directory /etc/telegraf/telegraf.d \
  --test

실제 서비스 계정으로 확인:

systemctl daemon-reload

systemd-run \
  --unit=telegraf-config-check \
  --wait \
  --pipe \
  --collect \
  -p User=telegraf \
  -p Group=telegraf \
  -p EnvironmentFile=/etc/telegraf/monitor.env \
  -p CapabilityBoundingSet=CAP_NET_RAW \
  -p AmbientCapabilities=CAP_NET_RAW \
  /usr/bin/telegraf \
  --config /etc/telegraf/telegraf.conf \
  --config-directory /etc/telegraf/telegraf.d \
  --test

서비스 시작:

systemctl enable --now telegraf
systemctl restart telegraf

systemctl status telegraf --no-pager
journalctl -u telegraf -n 100 --no-pager

9. Grafana 설치 및 InfluxDB 연결

9.1 Grafana 서비스

vi /etc/grafana/grafana.ini

내부망 운영 환경에 맞게 listen 주소를 설정한다.

예시:

[server]
http_addr = 0.0.0.0
http_port = 3000

[users]
allow_sign_up = false

[auth.anonymous]
enabled = false
systemctl enable --now grafana-server
systemctl restart grafana-server

systemctl status grafana-server --no-pager

curl -fsS http://127.0.0.1:3000/api/health

9.2 InfluxDB Datasource

Grafana:

Connections
  → Data sources
  → Add data source
  → InfluxDB
설정 값
Query language Flux
URL http://127.0.0.1:8086
Organization network
Token grafana-read
Default Bucket snmp_raw
Min time interval 1s

10. SNMP Dashboard 기본 구성

10.1 Traffic

Flux:

from(bucket: "snmp_raw")
  |> range(start: v.timeRangeStart, stop: v.timeRangeStop)
  |> filter(fn: (r) => r._measurement == "switch_port")
  |> filter(fn: (r) => r.source == "203.0.113.11")
  |> filter(fn: (r) =>
      r._field == "in_octets" or
      r._field == "out_octets"
  )
  |> derivative(unit: 1s, nonNegative: true)
  |> map(fn: (r) => ({
      r with _value: r._value * 8.0
  }))

Grafana Unit:

bits/sec

10.2 Error / Discard

  |> filter(fn: (r) =>
      r._field == "in_errors" or
      r._field == "out_errors" or
      r._field == "in_discards" or
      r._field == "out_discards"
  )
  |> derivative(unit: 1s, nonNegative: true)

10.3 Broadcast / Multicast

Broadcast/Multicast는 별도 패널로 표시하여 폭주 및 비정상 증가를 확인한다.

11. Prometheus 설치

실제 구축에서는 Prometheus 3.13.3을 사용하였다.

11.1 사용자 및 디렉터리

useradd \
  --system \
  --no-create-home \
  --shell /sbin/nologin \
  prometheus

install -d -o prometheus -g prometheus \
  /etc/prometheus \
  /var/lib/prometheus

Prometheus 바이너리는 공식 릴리스에서 다운로드하고 checksum 확인 후 설치한다.

install -m 0755 prometheus /usr/local/bin/prometheus
install -m 0755 promtool /usr/local/bin/promtool

11.2 Prometheus 설정

vi /etc/prometheus/prometheus.yml
global:
  scrape_interval: 30s

scrape_configs:

  - job_name: prometheus
    static_configs:
      - targets:
          - "127.0.0.1:9090"
        labels:
          server_name: monitoring-server

  - job_name: node
    static_configs:
      - targets:
          - "127.0.0.1:9100"
        labels:
          server_name: monitoring-server

      - targets:
          - "192.0.2.20:9100"
        labels:
          server_name: linux-server

  - job_name: windows
    static_configs:
      - targets:
          - "192.0.2.30:9182"
        labels:
          server_name: windows-server

  - job_name: qnap
    static_configs:
      - targets:
          - "198.51.100.10:9100"
        labels:
          server_name: nas-01

11.3 systemd

vi /etc/systemd/system/prometheus.service
[Unit]
Description=Prometheus
Wants=network-online.target
After=network-online.target

[Service]
User=prometheus
Group=prometheus

ExecStart=/usr/local/bin/prometheus \
  --config.file=/etc/prometheus/prometheus.yml \
  --storage.tsdb.path=/var/lib/prometheus \
  --storage.tsdb.retention.time=30d \
  --storage.tsdb.retention.size=20GB \
  --web.listen-address=127.0.0.1:9090

Restart=always

[Install]
WantedBy=multi-user.target
chown -R prometheus:prometheus \
  /etc/prometheus \
  /var/lib/prometheus

promtool check config /etc/prometheus/prometheus.yml

systemctl daemon-reload
systemctl enable --now prometheus

systemctl status prometheus --no-pager

12. Node Exporter 설치

실제 구축에서는 Node Exporter 1.12.1 환경을 사용하였다.

12.1 Linux Node Exporter

useradd \
  --system \
  --no-create-home \
  --shell /sbin/nologin \
  node_exporter

install -m 0755 node_exporter \
  /usr/local/bin/node_exporter
vi /etc/systemd/system/node_exporter.service
[Unit]
Description=Node Exporter
After=network-online.target
Wants=network-online.target

[Service]
User=node_exporter
Group=node_exporter

ExecStart=/usr/local/bin/node_exporter

Restart=always

[Install]
WantedBy=multi-user.target
systemctl daemon-reload
systemctl enable --now node_exporter

curl -s http://127.0.0.1:9100/metrics | head

외부 서버의 node_exporter는 Prometheus 서버 주소만 9100/tcp에 접근하도록 방화벽을 제한한다.

13. Windows Exporter

Windows Server에서는 windows_exporter를 사용한다.

기본 포트:

9182/tcp

Prometheus Target:

192.0.2.30:9182

성능 카운터 이상 시 다음 명령으로 복구한 사례가 있다.

lodctr /R
winmgmt /resyncperf

PowerShell:

Restart-Service windows_exporter

확인:

Invoke-WebRequest http://127.0.0.1:9182/metrics

14. Grafana Prometheus Datasource

Grafana:

Connections
  → Data sources
  → Prometheus

URL:

http://127.0.0.1:9090

Save & Test 후 다음 PromQL로 확인한다.

up

15. Server Dashboard 구성

Linux Dashboard:

  • Uptime
  • CPU Usage
  • Memory Usage
  • Load Average
  • Filesystem
  • Disk Read / Write
  • Network RX / TX
  • Network Error / Drop
  • Service 상태

Windows Dashboard:

  • CPU
  • Memory
  • Disk
  • Network
  • Uptime
  • Exporter 상태

Disk 용량은 실제 GB/GiB 값으로 표시하며 불필요한 Bar Gauge는 제거한다.

16. QNAP TS-264 모니터링

QNAP은 세 종류의 데이터를 함께 사용한다.

데이터 수집 방법
CPU / Memory / Network / Filesystem / Disk I/O node_exporter → Prometheus
HDD / RAID / Storage / Volume SNMP → Telegraf → InfluxDB
Event / Access Log Syslog → rsyslog → Alloy → Loki

16.1 QNAP SNMP

QNAP SNMPv3:

  • SHA
  • DES

Measurement:

QNAP_TS264
qnap_disk
qnap_raid
qnap_storage_pool
qnap_volume

Disk:

  • disk_id
  • manufacturer
  • model
  • disk_type
  • disk_status
  • temperature
  • capacity_bytes

RAID:

  • RAID ID
  • RAID Name
  • RAID Status
  • RAID Level
  • Capacity

16.2 Volume 단위 보정

QNAP qnap_volume의 capacity/free 값은 실제 장비에서 KiB 형태로 반환되는 것으로 확인되었다.

원시값을 Grafana에서 byte로 바로 해석하면 약 11.3 GiB로 잘못 표시된다.

Flux:

|> map(fn: (r) => ({
    r with
    capacity_bytes:
      uint(v: r.capacity_bytes) * uint(v: 1024),

    free_bytes:
      uint(v: r.free_bytes) * uint(v: 1024),

    used_bytes:
      (
        uint(v: r.capacity_bytes)
        - uint(v: r.free_bytes)
      ) * uint(v: 1024),

    used_percent:
      if float(v: r.capacity_bytes) > 0.0 then
        (
          float(v: r.capacity_bytes)
          - float(v: r.free_bytes)
        )
        / float(v: r.capacity_bytes)
        * 100.0
      else 0.0
}))

16.3 실제 Data Volume

Node Exporter 기준 실제 사용자 Data Volume:

mountpoint="/share/CACHEDEV1_DATA"
device="/dev/mapper/cachedev1"
fstype="ext4"

Snapshot:

/mnt/snapshot/...

위 Snapshot 경로는 Dashboard Filesystem 패널에서 제외한다.

DataVol1 사용률:

100 * (
  1 -
  node_filesystem_avail_bytes{
    instance="198.51.100.10:9100",
    mountpoint="/share/CACHEDEV1_DATA"
  }
  /
  node_filesystem_size_bytes{
    instance="198.51.100.10:9100",
    mountpoint="/share/CACHEDEV1_DATA"
  }
)

16.4 Network는 bond0만 표시

RX:

rate(
  node_network_receive_bytes_total{
    instance="198.51.100.10:9100",
    device="bond0"
  }[$__rate_interval]
) * 8

TX:

rate(
  node_network_transmit_bytes_total{
    instance="198.51.100.10:9100",
    device="bond0"
  }[$__rate_interval]
) * 8

16.5 Network Error / Drop

rate(node_network_receive_errs_total{
  instance="198.51.100.10:9100",
  device="bond0"
}[$__rate_interval])
rate(node_network_transmit_errs_total{
  instance="198.51.100.10:9100",
  device="bond0"
}[$__rate_interval])
rate(node_network_receive_drop_total{
  instance="198.51.100.10:9100",
  device="bond0"
}[$__rate_interval])
rate(node_network_transmit_drop_total{
  instance="198.51.100.10:9100",
  device="bond0"
}[$__rate_interval])

16.6 Disk I/O

Read:

rate(node_disk_read_bytes_total{
  instance="198.51.100.10:9100",
  device=~"sd[a-z]+"
}[$__rate_interval])

Write:

rate(node_disk_written_bytes_total{
  instance="198.51.100.10:9100",
  device=~"sd[a-z]+"
}[$__rate_interval])

Read IOPS:

rate(node_disk_reads_completed_total{
  instance="198.51.100.10:9100",
  device=~"sd[a-z]+"
}[$__rate_interval])

Write IOPS:

rate(node_disk_writes_completed_total{
  instance="198.51.100.10:9100",
  device=~"sd[a-z]+"
}[$__rate_interval])

17. rsyslog 구성

17.1 Network Syslog

예시 로그 파일:

/var/log/network-syslog/events.log

네트워크 장비:

UDP/TCP 514

17.2 Server Syslog

/var/log/server-syslog/events.log

예시 포트:

5514/tcp

17.3 QNAP Event / Access 분리

QNAP 로그는 Event와 Access를 별도 포트로 분리한다.

로그 Port File
Event TCP 5515 /var/log/qnap/event.log
Access TCP 5516 /var/log/qnap/access.log

Template:

template(name="QnapSyslogLine" type="string"
  string="%timegenerated:::date-rfc3339% src=%fromhost-ip% severity=%syslogseverity-text% host=%hostname% %syslogtag%%msg:::sp-if-no-1st-sp%%msg%\n")

Event:

ruleset(name="QnapEventLog") {
    action(
        type="omfile"
        file="/var/log/qnap/event.log"
        template="QnapSyslogLine"
        fileOwner="root"
        fileGroup="alloy"
        fileCreateMode="0640"
        dirOwner="root"
        dirGroup="alloy"
        dirCreateMode="0750"
        createDirs="on"
    )
    stop
}

input(
    type="imtcp"
    port="5515"
    ruleset="QnapEventLog"
)

Access:

ruleset(name="QnapAccessLog") {
    action(
        type="omfile"
        file="/var/log/qnap/access.log"
        template="QnapSyslogLine"
        fileOwner="root"
        fileGroup="alloy"
        fileCreateMode="0640"
        dirOwner="root"
        dirGroup="alloy"
        dirCreateMode="0750"
        createDirs="on"
    )
    stop
}

input(
    type="imtcp"
    port="5516"
    ruleset="QnapAccessLog"
)

검증:

rsyslogd -N1
systemctl restart rsyslog

ss -lntp | grep -E ':5514|:5515|:5516'

18. Loki / Alloy 구성

Loki Endpoint:

http://127.0.0.1:3100

Alloy UI:

127.0.0.1:12345

18.1 Alloy 기본 설정

logging {
  level = "info"
}

18.2 Network Syslog

loki.source.file "network_syslog" {
  targets = [
    {
      __path__ = "/var/log/network-syslog/events.log",
      job      = "network-syslog",
    },
  ]

  forward_to = [loki.process.network_syslog.receiver]
}

loki.process "network_syslog" {
  stage.regex {
    expression = `^(?P<received_at>\S+) src=(?P<device_ip>\S+) severity=(?P<severity>\S+) host=(?P<device_host>\S+) (?P<message>.*)$`
  }

  stage.timestamp {
    source            = "received_at"
    format            = "RFC3339Nano"
    action_on_failure = "skip"
  }

  stage.labels {
    values = {
      device_ip = "",
      severity  = "",
    }
  }

  forward_to = [loki.write.local.receiver]
}

18.3 Server Syslog

loki.source.file "server_syslog" {
  targets = [
    {
      __path__ = "/var/log/server-syslog/events.log",
      job      = "server-syslog",
    },
  ]

  forward_to = [loki.process.server_syslog.receiver]
}

loki.process "server_syslog" {
  stage.regex {
    expression = `^(?P<received_at>\S+) src=(?P<server_ip>\S+) severity=(?P<severity>\S+) host=(?P<server_host>\S+) (?P<source>[^:\s\[]+)(?:\[\d+\])?:?\s+(?P<message>.*)$`
  }

  stage.timestamp {
    source            = "received_at"
    format            = "RFC3339Nano"
    action_on_failure = "skip"
  }

  stage.labels {
    values = {
      server_ip   = "",
      server_host = "",
      severity    = "",
      source      = "",
    }
  }

  forward_to = [loki.write.local.receiver]
}

18.4 QNAP Event

loki.source.file "qnap_event" {
  targets = [
    {
      __path__ = "/var/log/qnap/event.log",
      job      = "qnap-event",
    },
  ]

  forward_to = [loki.process.qnap_event.receiver]
}

loki.process "qnap_event" {
  stage.regex {
    expression = `^(?P<received_at>\S+) src=(?P<nas_ip>\S+) severity=(?P<severity>\S+) host=(?P<nas_host>\S+) (?P<message>.*)$`
  }

  stage.timestamp {
    source            = "received_at"
    format            = "RFC3339Nano"
    action_on_failure = "skip"
  }

  stage.labels {
    values = {
      nas_ip   = "",
      nas_host = "",
      severity = "",
      log_type = "event",
    }
  }

  forward_to = [loki.write.local.receiver]
}

18.5 QNAP Access

loki.source.file "qnap_access" {
  targets = [
    {
      __path__ = "/var/log/qnap/access.log",
      job      = "qnap-access",
    },
  ]

  forward_to = [loki.process.qnap_access.receiver]
}

loki.process "qnap_access" {
  stage.regex {
    expression = `^(?P<received_at>\S+) src=(?P<nas_ip>\S+) severity=(?P<severity>\S+) host=(?P<nas_host>\S+) (?P<message>.*)$`
  }

  stage.timestamp {
    source            = "received_at"
    format            = "RFC3339Nano"
    action_on_failure = "skip"
  }

  stage.labels {
    values = {
      nas_ip   = "",
      nas_host = "",
      severity = "",
      log_type = "access",
    }
  }

  forward_to = [loki.write.local.receiver]
}

18.6 Loki Write

loki.write "local" {
  endpoint {
    url = "http://127.0.0.1:3100/loki/api/v1/push"
  }
}

검증:

alloy validate /etc/alloy/config.alloy

systemctl restart alloy

systemctl status alloy --no-pager

Loki Job 확인:

curl -s \
  'http://127.0.0.1:3100/loki/api/v1/label/job/values' \
  | jq

정상 예:

network-syslog
server-syslog
qnap-event
qnap-access

18.7 Alloy Permission 문제

다음과 같은 오류가 발생할 수 있다.

failed to tail file
stat failed
permission denied

확인:

namei -l /var/log/qnap/access.log

systemctl show alloy \
  -p User \
  -p Group

sudo -u alloy \
  head /var/log/qnap/access.log

권장 권한:

drwxr-x--- root alloy /var/log/qnap
-rw-r----- root alloy /var/log/qnap/access.log
-rw-r----- root alloy /var/log/qnap/event.log

SELinux 확인:

getenforce

ausearch -m AVC -ts recent \
  | grep -Ei 'alloy|qnap'

ls -Zd /var/log/qnap
ls -Z /var/log/qnap/access.log

19. Loki Query

Network:

{job="network-syslog"}

Server:

{job="server-syslog"}

QNAP Event:

{job="qnap-event"}

QNAP Access:

{job="qnap-access"}

Critical Network Syslog:

{job="network-syslog",severity=~"emerg|alert|crit|err"}

20. Grafana Syslog Alert

Syslog Level 3 이상:

sum by (device_ip, severity) (
  count_over_time(
    {
      job="network-syslog",
      severity=~"emerg|alert|crit|err"
    }[1m]
  )
)

권장 Alert 구조:

A = Loki Instant Query
B = Threshold
    A IS ABOVE 0

Summary:

[Syslog 경고] {{ $labels.device_ip }} - {{ $labels.severity }}

Description:

장비 {{ $labels.device_ip }} 에서 Syslog Level 3(Error) 이상의 로그가 발생했습니다.
최근 1분 발생 건수: {{ $values.A.Value }}
Severity: {{ $labels.severity }}

Range Query를 그대로 Alert 조건으로 사용할 경우 다음 오류가 발생할 수 있다.

invalid format of evaluation results for the alert definition A:
looks like time series data, only reduced data can be alerted on.

Range Query를 유지할 경우 Reduce Expression을 추가해야 한다.

21. ICMP Alert

Flux:

from(bucket: "snmp_raw")
  |> range(start: -5m)
  |> filter(fn: (r) =>
    r._measurement == "ping" and
    r._field == "percent_packet_loss"
  )
  |> group(columns: ["url"])
  |> last()
  |> keep(columns: ["_time", "_value", "url"])

Alert:

A = Flux Query
B = Reduce / Last / Strict
C = Threshold > 99
Pending = 2m

No Data:

Keep Last State

22. 최종 Dashboard 구성안

22.1 Network Switch Dashboard

Row Panels
상태 Device / Uptime / CPU / Memory / Ping
Port 상태 ifName / Alias / Speed / OperStatus
Traffic RX bps / TX bps
Packet Unicast / Broadcast / Multicast PPS
Error RX/TX Error / Discard
Syslog Warning / Error / Critical Event

22.2 Linux Server Dashboard

  • Uptime
  • CPU
  • Memory
  • Load
  • Disk Capacity
  • Disk I/O
  • Network RX/TX
  • Network Error/Drop
  • Service Log
  • Security Event

22.3 Windows Dashboard

  • Uptime
  • CPU
  • Memory
  • Logical Disk
  • Disk I/O
  • Network
  • Windows Exporter 상태

22.4 QNAP Dashboard

최종 구성:

Node Exporter
Uptime
CPU
Memory
HDD 최고 온도

CPU / Memory / Load

Network RX / TX
  - bond0 only

Disk 상태

Volume
  - DataVol1
  - SNMP capacity/free x1024 보정

Filesystem 사용률
  - /share/CACHEDEV1_DATA only
  - Snapshot 제외

Network Errors / Drops
  - bond0 RX Error
  - bond0 TX Error
  - bond0 RX Drop
  - bond0 TX Drop

Disk I/O
  - Physical sd* only
  - Read B/s
  - Write B/s
  - Read IOPS
  - Write IOPS

QNAP Event Log

QNAP Access Log

QNAP Dashboard에서 제거한 항목:

  • 상단 RAID 상태
  • Storage Pool 상태
  • Storage Pool 사용률
  • RAID 상세 Table
  • Storage Pool 상세 Table

RAID/Storage Pool 정보는 필요 시 별도 상세 Dashboard에서 조회한다.

23. 운영 점검

전체 서비스:

systemctl is-active \
  influxdb \
  telegraf \
  grafana-server \
  prometheus \
  loki \
  alloy \
  rsyslog \
  chronyd

Listening Port:

ss -lntup

Prometheus:

curl -s http://127.0.0.1:9090/-/healthy

Loki:

curl -s http://127.0.0.1:3100/ready

Telegraf:

journalctl -u telegraf \
  --since '-10 min' \
  --no-pager

Alloy:

journalctl -u alloy \
  --since '-10 min' \
  --no-pager

Filesystem:

df -hT

System I/O:

vmstat 1 5
iostat -xz 1 5

24. 장애 점검 순서

Prometheus Target Down

curl http://TARGET_IP:9100/metrics

journalctl -u prometheus \
  -n 100 \
  --no-pager

SNMP Timeout

snmpget -v3 \
  -t 5 \
  -r 1 \
  -On \
  SWITCH_IP \
  .1.3.6.1.2.1.1.3.0

확인 항목:

  • SNMP User
  • SHA/DES/AES 조합
  • ACL
  • Source IP
  • Timeout
  • SNMP View
  • 장비 CPU

Syslog 미수신

ss -lntup \
  | grep -E ':514|:5514|:5515|:5516'

tcpdump -ni any \
  host DEVICE_IP

tail -f /var/log/qnap/access.log

Loki 미표시

alloy validate \
  /etc/alloy/config.alloy

journalctl -u alloy \
  -n 100 \
  --no-pager

curl -s \
  'http://127.0.0.1:3100/loki/api/v1/label/job/values' \
  | jq

25. 운영 원칙

  • SNMP Metric은 Telegraf → InfluxDB로 저장한다.
  • Host Metric은 Prometheus로 저장한다.
  • Syslog는 rsyslog → Alloy → Loki로 저장한다.
  • Grafana는 세 데이터소스를 통합한다.
  • SNMP Counter는 누적값을 그대로 표시하지 않고 rate/derivative 계산한다.
  • No Data를 0 또는 정상으로 강제 표시하지 않는다.
  • QNAP Snapshot filesystem은 Data Volume 사용률에서 제외한다.
  • QNAP Volume SNMP capacity/free 값은 장비 특성상 ×1024 보정한다.
  • QNAP Network는 실제 활성 bond0만 표시한다.
  • Loki message 전체를 label로 만들지 않는다.
  • 장비 Syslog Alert는 device_ip / severity와 발생 건수를 메일에 포함한다.
  • Alert의 Range Query는 Reduce 또는 Instant Query 구조로 구성한다.
  • 모든 Token 및 SNMP 비밀번호는 문서에 실제 값을 기록하지 않는다.