Great work! I would like to ask how the Selection-to-Tuning Ratio (STR) metric is calculated in the paper. In my experiment, the time for a single run of data selection is approximately over two hours, while fine-tuning the model takes around twenty minutes. Is this a normal phenomenon? I am using a single H-series GPU in both cases.
Great work! I would like to ask how the Selection-to-Tuning Ratio (STR) metric is calculated in the paper. In my experiment, the time for a single run of data selection is approximately over two hours, while fine-tuning the model takes around twenty minutes. Is this a normal phenomenon? I am using a single H-series GPU in both cases.