This model is deprecated. We only continue performance benchmarking for the default 10k input token workload. Results for other workloads are historical and no longer updated.
NVIDIA has launched a newer model, Llama Nemotron Super 49B v1.5. We suggest considering it instead.
For more information, see comparison of Llama Nemotron Super 49B v1.5 to other models and API provider benchmarks for Llama Nemotron Super 49B v1.5.
Llama 3.3 Nemotron Super 49B v1 (Non-reasoning) API 제공업체 벤치마킹 및 분석
이 분석은 사용 사례에 가장 적합한 Llama 3.3 Nemotron Super 49B v1 (Non-reasoning) API 제공업체를 선택하는 데 도움을 드립니다.
가장 빠름
출력 속도
총 제공업체 0곳
가장 낮은 지연 시간
첫 토큰까지 걸린 시간
총 제공업체 0곳
가장 낮은 가격
혼합 가격(토큰 100만 개당)
총 제공업체 0곳
현재 Llama 3.3 Nemotron Super 49B을 제공하는 API 제공업체가 없습니다.
모델 세부 정보와 다른 모델 대비 지능은 Llama 3.3 Nemotron Super 49B v1 (Non-reasoning) 모델 페이지에서 확인하세요.
주요 내용
가격
Pricing: Cache Hit, Input, and Output
Pricing: Blended Price
Output Speed vs. Price
속도
출력 속도(초당 토큰 수)로 측정
Output Speed: Llama 3.3 Nemotron Super 49B v1 (Non-reasoning)
Latency vs. Output Speed
지연 시간
첫 토큰까지 걸린 시간(초)으로 측정
첫 토큰까지 걸린 시간
종단 간 응답 시간
Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed
종단 간 응답 시간
주요 비교 지표 및 API 기능
| No results. | |||||||||
자주 묻는 질문
Llama 3.3 Nemotron Super 49B v1 (Non-reasoning) 제공업체에 관한 일반적인 질문
현재 벤치마킹 대상 API 제공업체 중 Llama 3.3 Nemotron Super 49B v1 (Non-reasoning)을 제공하는 곳은 없습니다. 오픈 웨이트 모델이므로 자체 호스팅할 수 있습니다.