Four single-node scripts. Pick one (or two, for monitor + processor). See README.md for full details.
| Script | Purpose |
|---|---|
perlmutter-ersap-monitor.slurm |
monitor stack only, on a normal exclusive compute node (j_dpe + exporter + Prometheus + Grafana) |
perlmutter-ersap-monitor-longrun.slurm |
the same monitor stack, but on a workflow-QOS long-running node — currently sized for 30 days |
perlmutter-ersap-processor.slurm |
one pipeline container, reports to a remote monitor |
perlmutter-ersap-allinone.slurm |
monitor and one pipeline on the same node |
No build of ersap-java needed: git clone/pull this repo for the scripts above plus perlmutter-setup/ — nothing else here is read at run time. ERSAP_HOME (a pre-built ERSAP install, used by monitor*.slurm/allinone.slurm) and the podman-hpc pipeline image (used by processor.slurm/allinone.slurm, default docker.io/gurjyan/pet-sro:v1) are separate dependencies, not part of this repo.
# Install Prometheus and Grafana binaries under $HOME (see README §Perlmutter)
# Deploy their configs:
bash ~/ersap-java/perlmutter-setup/deploy.shEdit each .slurm file's #SBATCH --account= line if your NERSC repo is not amsc016.
Two ways to run the monitor stack, depending on what you need:
perlmutter-ersap-monitor.slurm— a normalregular-QOS,cpu-constraint job that grabs an exclusive compute node (64 cpus) for up to 8h by default. Good for ad hoc testing/debugging of the monitor + Grafana/Prometheus stack.perlmutter-ersap-monitor-longrun.slurm— targets NERSC'sworkflowQOS with thecronconstraint, the mechanism intended for persistent, lightweight, long-running services rather than exclusive compute nodes. It currently requests--time=30-00:00:00(30 days) and--dependency=singleton(so re-submitting under the same job name won't start a second copy). Requires NERSC to have approvedworkflowQOS access for your account first.
cd ~/ersap-java && mkdir -p logs
sbatch perlmutter-ersap-monitor-longrun.slurm # note the JOB_ID
OR
sbatch --qos=debug --time=00:30:00 perlmutter-ersap-monitor.slurmWait for RUNNING, then:
cat logs/monitor-<JOB_ID>/monitor-info.txt # discover node + endpoints
tail -f logs/monitor-<JOB_ID>.out # readiness bannerTunnel from your laptop (also printed in monitor-info.txt):
ssh -N -L 3000:<monitor-node>:3000 -L 9090:<monitor-node>:9090 \
<user>@perlmutter.nersc.gov
# Grafana http://localhost:3000 (admin/changeme)Cancel: scancel <JOB_ID>
cd ~/ersap-java && mkdir -p logs
export MONITOR_ENV_FILE=$PWD/logs/monitor-<monitor-JOB_ID>/monitor.env
sbatch perlmutter-ersap-processor.slurmAlternative: export ERSAP_MONITOR_FE='<ip>%19000_java' instead of MONITOR_ENV_FILE.
Follow: tail -f logs/processor-<JOB_ID>/pipeline.log
Cancel: scancel <JOB_ID> (stops the container cleanly)
cd ~/ersap-java && mkdir -p logs
sbatch perlmutter-ersap-allinone.slurmJob exits when the pipeline finishes; monitor is torn down automatically.
cat logs/allinone-<JOB_ID>/monitor-info.txt for the tunnel command and log paths.
| Var | Default | Applies to |
|---|---|---|
ERSAP_HOME |
/global/homes/g/gurjyan/work/ersap_installation/ersap_home |
monitor, allinone |
MONITOR_PORT |
19000 |
monitor, allinone, processor |
SESSION |
test |
all |
IMAGE |
docker.io/gurjyan/pet-sro:v1 |
processor, allinone |
DATA_DIR |
/global/cfs/cdirs/amsc016/haidis/ersap-data |
processor, allinone |
SERVICES_FILE |
'$ERSAP_USER_DATA/config/pet_services.yaml' (container-resolved) |
processor, allinone |
PIPELINE_PULL |
1 (set 0 to skip podman-hpc pull) |
processor, allinone |
squeue -u $USER
scontrol show job <JOB_ID>
scontrol show hostnames "$(squeue -h -j <JOB_ID> -o '%N')"
squeue --start -j <JOB_ID>
sprio -j <JOB_ID>
scancel <JOB_ID>