Implement prompt-token loss masking given (prompt_len, total_len) per sample

SFT, Instruction Tuning & PEFT1 / 6