Weave on self-managed infrastructure is currently in Private Preview.프로덕션 환경의 경우, W&B는 강력하게 다음을 권장합니다 W&B Dedicated Cloud, Weave가 정식 출시되어 있는 곳입니다.프로덕션 수준의 자체 관리형 인스턴스를 배포하려면 다음으로 문의하세요 support@wandb.com.
이 가이드는 자체 관리형 환경에서 W&B Weave를 실행하는 데 필요한 모든 구성 요소를 배포하는 방법을 설명합니다.자체 관리형 Weave 배포의 핵심 구성 요소는 ClickHouseDB로, Weave 애플리케이션 백엔드가 이에 의존합니다.배포 프로세스가 완전히 기능하는 ClickHouseDB 인스턴스를 설정하지만, 프로덕션 환경에서의 신뢰성과 고가용성을 보장하기 위해 추가 단계가 필요할 수 있습니다.
이 문서의 ClickHouse 배포는 Bitnami ClickHouse 패키지를 사용합니다.Bitnami Helm 차트는 기본 ClickHouse 기능, 특히 ClickHouse Keeper의 사용에 대한 좋은 지원을 제공합니다.Clickhouse를 구성하려면 다음 단계를 완료하세요:
Helm 구성의 가장 중요한 부분은 XML 형식으로 제공되는 ClickHouse 구성입니다. 아래는 필요에 맞게 사용자 정의할 수 있는 매개변수가 있는 예제 values.yaml 파일입니다.
구성 프로세스를 쉽게 하기 위해 관련 섹션에 {/* COMMENT */} 형식으로 주석을 추가했습니다.다음 매개변수를 수정하세요:
clusterName
auth.username
auth.password
S3 버킷 관련 구성
W&B는 clusterName 값을 values.yaml에서 weave_cluster로 유지하는 것을 권장합니다. 이는 W&B Weave가 데이터베이스 마이그레이션을 실행할 때 예상되는 클러스터 이름입니다. 다른 이름을 사용해야 하는 경우 Setting clusterName 섹션에서 자세한 정보를 참조하세요.
## @param clusterName ClickHouse cluster nameclusterName: weave_cluster## @param shards Number of ClickHouse shards to deployshards: 1## @param replicaCount Number of ClickHouse replicas per shard to deploy## if keeper enable, same as keeper count, keeper cluster by shards.replicaCount: 3persistence: enabled: true size: 30G # this size must be larger than cache size.## ClickHouse resource requests and limitsresources: requests: cpu: 0.5 memory: 500Mi limits: cpu: 3.0 memory: 6Gi## Authenticationauth: username: weave_admin password: "weave_123" existingSecret: "" existingSecretKey: ""## @param logLevel Logging levellogLevel: information## @section ClickHouse keeper configuration parameterskeeper: enabled: true## @param extraEnvVars Array with extra environment variables to add to ClickHouse nodes##extraEnvVars: - name: S3_ENDPOINT value: "https://s3.us-east-1.amazonaws.com/bucketname/$(CLICKHOUSE_REPLICA_ID)"## @param defaultConfigurationOverrides [string] Default configuration overrides (evaluated as a template)defaultConfigurationOverrides: | <clickhouse> {/* Macros */} <macros> <shard from_env="CLICKHOUSE_SHARD_ID"></shard> <replica from_env="CLICKHOUSE_REPLICA_ID"></replica> </macros> {/* Log Level */} <logger> <level>{{ .Values.logLevel }}</level> </logger> {{- if or (ne (int .Values.shards) 1) (ne (int .Values.replicaCount) 1)}} <remote_servers> <{{ .Values.clusterName }}> {{- $shards := $.Values.shards | int }} {{- range $shard, $e := until $shards }} <shard> <internal_replication>true</internal_replication> {{- $replicas := $.Values.replicaCount | int }} {{- range $i, $_e := until $replicas }} <replica> <host>{{ printf "%s-shard%d-%d.%s.%s.svc.%s" (include "common.names.fullname" $ ) $shard $i (include "clickhouse.headlessServiceName" $) (include "common.names.namespace" $) $.Values.clusterDomain }}</host> <port>{{ $.Values.service.ports.tcp }}</port> </replica> {{- end }} </shard> {{- end }} </{{ .Values.clusterName }}> </remote_servers> {{- end }} {{- if .Values.keeper.enabled }} <keeper_server> <tcp_port>{{ $.Values.containerPorts.keeper }}</tcp_port> {{- if .Values.tls.enabled }} <tcp_port_secure>{{ $.Values.containerPorts.keeperSecure }}</tcp_port_secure> {{- end }} <server_id from_env="KEEPER_SERVER_ID"></server_id> <log_storage_path>/bitnami/clickhouse/keeper/coordination/log</log_storage_path> <snapshot_storage_path>/bitnami/clickhouse/keeper/coordination/snapshots</snapshot_storage_path> <coordination_settings> <operation_timeout_ms>10000</operation_timeout_ms> <session_timeout_ms>30000</session_timeout_ms> <raft_logs_level>trace</raft_logs_level> </coordination_settings> <raft_configuration> {{- $nodes := .Values.replicaCount | int }} {{- range $node, $e := until $nodes }} <server> <id>{{ $node | int }}</id> <hostname from_env="{{ printf "KEEPER_NODE_%d" $node }}"></hostname> <port>{{ $.Values.service.ports.keeperInter }}</port> </server> {{- end }} </raft_configuration> </keeper_server> {{- end }} {{- if or .Values.keeper.enabled .Values.zookeeper.enabled .Values.externalZookeeper.servers }} <zookeeper> {{- if or .Values.keeper.enabled }} {{- $nodes := .Values.replicaCount | int }} {{- range $node, $e := until $nodes }} <node> <host from_env="{{ printf "KEEPER_NODE_%d" $node }}"></host> <port>{{ $.Values.service.ports.keeper }}</port> </node> {{- end }} {{- else if .Values.zookeeper.enabled }} {{- $nodes := .Values.zookeeper.replicaCount | int }} {{- range $node, $e := until $nodes }} <node> <host from_env="{{ printf "KEEPER_NODE_%d" $node }}"></host> <port>{{ $.Values.zookeeper.service.ports.client }}</port> </node> {{- end }} {{- else if .Values.externalZookeeper.servers }} {{- range $node :=.Values.externalZookeeper.servers }} <node> <host>{{ $node }}</host> <port>{{ $.Values.externalZookeeper.port }}</port> </node> {{- end }} {{- end }} </zookeeper> {{- end }} {{- if .Values.metrics.enabled }} <prometheus> <endpoint>/metrics</endpoint> <port from_env="CLICKHOUSE_METRICS_PORT"></port> <metrics>true</metrics> <events>true</events> <asynchronous_metrics>true</asynchronous_metrics> </prometheus> {{- end }} <listen_host>0.0.0.0</listen_host> <listen_host>::</listen_host> <listen_try>1</listen_try> <storage_configuration> <disks> <s3_disk> <type>s3</type> <endpoint from_env="S3_ENDPOINT"></endpoint> {/* AVOID USE CREDENTIALS CHECK THE RECOMMENDATION */} <access_key_id>xxx</access_key_id> <secret_access_key>xxx</secret_access_key> {/* AVOID USE CREDENTIALS CHECK THE RECOMMENDATION */} <metadata_path>/var/lib/clickhouse/disks/s3_disk/</metadata_path> </s3_disk> <s3_disk_cache> <type>cache</type> <disk>s3_disk</disk> <path>/var/lib/clickhouse/s3_disk_cache/cache/</path> {/* THE CACHE SIZE MUST BE LOWER THAN PERSISTENT VOLUME */} <max_size>20Gi</max_size> </s3_disk_cache> </disks> <policies> <s3_main> <volumes> <main> <disk>s3_disk_cache</disk> </main> </volumes> </s3_main> </policies> </storage_configuration> <merge_tree> <storage_policy>s3_main</storage_policy> </merge_tree> </clickhouse>## @section Zookeeper subchart parameterszookeeper: enabled: false
이 정보를 사용하여 다음 구성을 추가하여 W&B Platform Custom Resource(CR)를 업데이트하세요:
apiVersion: apps.wandb.com/v1kind: WeightsAndBiasesmetadata: labels: app.kubernetes.io/name: weightsandbiases app.kubernetes.io/instance: wandb name: wandb namespace: defaultspec: values: global: [...] clickhouse: host: <release-name>-headless.<namespace>.svc.cluster.local port: 8123 password: <password> user: <username> database: wandb_weave # `replicated` must be set to `true` if replicating data across multiple nodes # This is in preview, use the env var `WF_CLICKHOUSE_REPLICATED` replicated: true weave-trace: enabled: true [...] weave-trace: install: true extraEnv: WF_CLICKHOUSE_REPLICATED: "true" [...]
둘 이상의 복제본을 사용할 때(W&B는 최소 3개의 복제본을 권장함), Weave Traces에 대해 다음 환경 변수가 설정되어 있는지 확인하세요.
extraEnv: WF_CLICKHOUSE_REPLICATED: "true"
이는 replicated: true와 동일한 효과가 있으며 미리보기 중입니다.
다음을 설정하세요 clusterName 다음에서 values.yaml를 weave_cluster로 설정하세요. 그렇지 않으면 데이터베이스 마이그레이션이 실패합니다.또는 다른 클러스터 이름을 사용하는 경우, 아래 예시와 같이 WF_CLICKHOUSE_REPLICATED_CLUSTER 환경 변수를 weave-trace.extraEnv에서 선택한 이름과 일치하도록 설정하세요.
[...] clickhouse: host: <release-name>-headless.<namespace>.svc.cluster.local port: 8123 password: <password> user: <username> database: wandb_weave # `replicated` must be set to `true` if replicating data across multiple nodes # This is in preview, use the env var `WF_CLICKHOUSE_REPLICATED` replicated: true weave-trace: enabled: true[...]weave-trace: install: true extraEnv: WF_CLICKHOUSE_REPLICATED: "true" WF_CLICKHOUSE_REPLICATED_CLUSTER: "different_cluster_name"[...]