我有兩個要合并的DF。但我需要根據字符串包含并使用多個列來合并它們
df_1
IN Start_Time Description Per_Extr
0 IN7305517 2022-07-24 00:06:59 ABEND JOB PP_BRAI_VAR_CARTAO_IND_IBI_D and JOB_STREAM_NAME P26_BRAI_RS2... FROM : 2022/01/08 TO : 2022/12/09
1 IN7305465 2022-07-24 00:09:49 ABEND JOB PP_AAAR_4898_POUP_MOV_TDCH_D and JOB_STREAM_NAME P26_AAAR_006_TSA... FROM : 2022/01/08 TO : 2022/12/09
2 IN7305466 2022-07-24 00:10:16 ABEND JOB PP_AAAR_4898_POUPMOV_D and JOB_STREAM_NAME P26_AAAR_006_TSA... FROM : 2022/01/08 TO : 2022/12/09
3 IN7305493 2022-07-24 00:20:27 ABEND JOB PP_BGDTPRODHBACMS102020_01_M and JOB_STREAM_NAME P26_BGDTDCHF_PUM... FROM : 2022/01/08 TO : 2022/12/09
df_2
JOB_STREAM_NAME JOB_NAME
NaN P26_BRAI_RS2 PP_BRAI_VAR_CARTAO_IND_IBI_D
NaN P26_BRAI_VAR_TOD PP_BRAI_VAR_CARTAO_IND_IBI_D
NaN P26_AAAR_006_TSA PP_AAAR_4898_POUP_MOV_TDCH_D
NaN P26_AAAR_006_TSA PP_AAAR_4898_POUPMOV_D
NaN P26_BGDTDCHF_PUM PP_BGDTPRODHBACMS102020_01_M
描述列中有JOB_NAME和JOB_STREAM_NAME
我的目標是這樣一個df:merged_df
IN JOB_STREAM_NAME JOB_NAME Start_Time Description Per_Extr
0 IN7305517 P26_BRAI_RS2 PP_BRAI_VAR_CARTAO_IND_IBI_D 2022-07-24 00:06:59 ABEND JOB PP_BRAI_VAR_CARTAO_IND_IBI_D and JOB_STREAM_NAME P26_BRAI_RS2... FROM : 2022/01/08 TO : 2022/12/09
1 NaN P26_BRAI_VAR_TOD PP_BRAI_VAR_CARTAO_IND_IBI_D NaN NaN NaN
2 IN7305465 P26_AAAR_006_TSA PP_AAAR_4898_POUP_MOV_TDCH_D 2022-07-24 00:10:16 ABEND JOB PP_AAAR_4898_POUPMOV_D and JOB_STREAM_NAME P26_AAAR_006_TSA... FROM : 2022/01/08 TO : 2022/12/09
3 IN7305466 P26_AAAR_006_TSA PP_AAAR_4898_POUPMOV_D 2022-07-24 00:10:16 ABEND JOB PP_AAAR_4898_POUPMOV_D and JOB_STREAM_NAME P26_AAAR_006_TSA... FROM : 2022/01/08 TO : 2022/12/09
4 IN7305493 P26_AAAR_006_TSA PP_AAAR_4898_POUPMOV_D 2022-07-24 00:20:27 ABEND JOB PP_BGDTPRODHBACMS102020_01_M and JOB_STREAM_NAME P26_BGDTDCHF_PUM... FROM : 2022/01/08 TO : 2022/12/09
請注意,作業PP_BRAI_VAR_CARTAO_IND_IBI_D位于2 JOB_STREAM_NAME中,其中一個作業沒有in,這就是為什么在merged_df中JOB_STREAM_NAME=P26_BRAI_VAR_TOD中的作業沒有in(NaN)的原因
我被指示對一個列執行此操作,但對多個列執行相同的操作。
在一篇專欄文章中,我使用了這種方法:
jobs_list= "|".join(map(str, df_2['JOB_NAME']))
new_df.insert(0, 'merge_key', df_1['Description'].str.extract("("+jobs_list+")", expand=False))
df_merged = new_df.merge(df_1, how='right', left_on='merge_key', right_on='JOB_NAME').drop('merge_key', axis=1)
你們能幫我嗎?
您需要一個鍵來合并這兩者,所以我們提取這些鍵并使用它們進行合并。