A journal of IEEE and CAA , publishes high-quality papers in English on original theoretical/experimental research and development in all areas of automation
Volume 13 Issue 6
Jun.  2026

IEEE/CAA Journal of Automatica Sinica

  • JCR Impact Factor: 18.3, Top 1 (SCI Q1)
    CiteScore: 28.2, Top 1% (Q1)
    Google Scholar h5-index: 95, TOP 5
Turn off MathJax
Article Contents
G. Zhang, S. Kan, W. Xu, Y. Jin, Y. Li, and Y. Cen, “Bimodal predicate refinement with decoupled entity-predicate representations for scene graph generation,” IEEE/CAA J. Autom. Sinica, vol. 13, no. 6, pp. 1470–1491, Jun. 2026. doi: 10.1109/JAS.2026.125786
Citation: G. Zhang, S. Kan, W. Xu, Y. Jin, Y. Li, and Y. Cen, “Bimodal predicate refinement with decoupled entity-predicate representations for scene graph generation,” IEEE/CAA J. Autom. Sinica, vol. 13, no. 6, pp. 1470–1491, Jun. 2026. doi: 10.1109/JAS.2026.125786

Bimodal Predicate Refinement With Decoupled Entity-Predicate Representations for Scene Graph Generation

doi: 10.1109/JAS.2026.125786
Funds:  This work was supported in part by the Fundamental Research Funds for the Central Universities (2025YJS051), the National Natural Science Foundation of China (62473033, 62576028), the Beijing Natural Science Foundation (L231012), and the Hunan Provincial Natural Science Foundation of China (2022JJ40632)
More Information
  • Scene graph generation is crucial for visual understanding, which can provide a structural foundation for image captioning, image comprehension, and visual question answering. However, datasets with significant long-tail distributions adversely impact model performance. Existing approaches have proposed various fusion and training strategies to address the long-tail distribution issue. Unfortunately, many strategies often lead to the “entanglement” of predicate and entity representations. This means that the representation space of predicates becomes confused with the representations of subjects and objects, resulting in suboptimal model performance. To tackle this challenge, we introduce a novel approach called bimodal predicate refinement with decoupled entity-predicate representations (BiRef). Specifically, we initialize predicate representations using Gaussian distribution and then progressively refine effective predicate representations in both visual and semantic spaces. Then, the refined predicate information is adaptively weighted and fused from the visual and semantic branches to obtain the predicate representation for the triplet. Extensive experiments conducted on the visual genome, GQA, and open images datasets demonstrated that our method achieved state-of-the-art scene graph generation performance and effectively mitigates the prediction bias problem associated with long-tail distributions. Our code has been made open source on https://github.com/gavin-gqzhang/BiRef.

     

  • loading
  • 1 Pretrained weights are available at: https://huggingface.co/openai/clip-vit-base-patch32

    2 Pretrained weights are available at: https://huggingface.co/openai/clip-vit-base-patch32

  • [1]
    K. Nguyen, S. Tripathi, B. Du, T. Guha, and T. Q. Nguyen, “In defense of scene graphs for image captioning,” in Proc. IEEE/CVF Int. Conf. Computer Vision, Montreal, Canada, 2021, pp. 1387−1396.
    [2]
    L. Gao, B. Wang, and W. Wang, “Image captioning with scene-graph based semantic concepts,” in Proc. 10th Int. Conf. Machine Learning and Computing, Macau, China, 2018, pp. 225−229.
    [3]
    J. Johnson, R. Krishna, M. Stark, L.-J. Li, D. A. Shamma, M. S. Bernstein, and L. Fei-Fei, “Image retrieval using scene graphs,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, Boston, USA, 2015, pp. 3668−3678.
    [4]
    X. Wang, Y. Jin, H. Yu, Y. Cen, and Y. Li, “Visual perception-inspired 3D point cloud sampling,” Pattern Recognit., vol. 169, Art. no. 111883, Jan. 2026. doi: 10.1016/j.patcog.2025.111883
    [5]
    T. Qian, J. Chen, S. Chen, B. Wu, and Y.-G. Jiang, “Scene graph refinement network for visual question answering,” IEEE Trans. Multimedia, vol. 25, pp. 3950–3961, Jan. 2023. doi: 10.1109/TMM.2022.3169065
    [6]
    S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “VQA: Visual question answering,” in Proc. IEEE Int. Conf. Computer Vision, Santiago, Chile, 2015, pp. 2425−2433.
    [7]
    T. He, L. Gao, J. Song, and Y.-F. Li, “Exploiting scene graphs for human-object interaction detection,” in Proc. IEEE/CVF Int. Conf. Computer Vision, Montreal, Canada, 2021, pp. 15964−15973.
    [8]
    G. Gkioxari, R. Girshick, P. Dollár, and K. He, “Detecting and recognizing human-object interactions,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018, pp. 8359−8367.
    [9]
    G. Zhang, S. Kan, L. Shi, W. Xu, G. An, and Y. Cen, “Cross-scene visual context parsing with large vision-language model,” Pattern Recognit., vol. 166, Art. no. 111641, Oct. 2025. doi: 10.1016/j.patcog.2025.111641
    [10]
    R. Zellers, M. Yatskar, S. Thomson, and Y. Choi, “Neural motifs: Scene graph parsing with global context,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018, pp. 5831−5840.
    [11]
    R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al., “Visual Genome: Connecting language and vision using crowdsourced dense image annotations,” Int. J. Comput. Vis., vol. 123, no. 1, pp. 32–73, May 2017. doi: 10.1007/s11263-016-0981-7
    [12]
    K. Tang, Y. Niu, J. Huang, J. Shi, and H. Zhang, “Unbiased scene graph generation from biased training,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Seattle, USA, 2020, pp. 3713−3722.
    [13]
    J. Yu, Y. Chai, Y. Wang, Y. Hu, and Q. Wu, “CogTree: Cognition tree loss for unbiased scene graph generation,” in Proc. 30th Int. Joint Conf. Artificial Intelligence, Montreal, Canada, 2021, pp. 1274−1280.
    [14]
    H. Kang and C. D. Yoo, “Skew class-balanced re-weighting for unbiased scene graph generation,” Mach. Learn. Knowl. Extr., vol. 5, no. 1, pp. 287–303, Mar. 2023. doi: 10.3390/make5010018
    [15]
    S. Yan, C. Shen, Z. Jin, J. Huang, R. Jiang, Y. Chen, and X.-S. Hua, “PCPL: Predicate-correlation perception learning for unbiased scene graph generation,” in Proc. 28th ACM Int. Conf. Multimedia, Seattle, USA, 2020, pp. 265−273.
    [16]
    Y. Cong, M. Y. Yang, and B. Rosenhahn, “RelTR: Relation transformer for scene graph generation,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 9, pp. 11169–11183, Sep. 2023. doi: 10.1109/TPAMI.2023.3268066
    [17]
    R. Li, S. Zhang, and X. He, “SGTR: End-to-end scene graph generation with transformer,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, New Orleans, USA, 2022, pp. 19464−19474.
    [18]
    C. Zheng, X. Lyu, L. Gao, B. Dai, and J. Song, “Prototype-based embedding network for scene graph generation,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Vancouver, Canada, 2023, pp. 22783−22792.
    [19]
    G. Zhang, S. Kan, Y. Zhang, Y. Cen, W. Xu, Y. Jin, and Y. Li, “Query-guided predicate decoupling and prototype approximation learning for scene graph generation,” Expert Syst. Appl., vol. 297, Art. no. 129525, Feb. 2026. doi: 10.1016/j.eswa.2025.129525
    [20]
    L. van der Maaten and G. Hinton, “Visualizing data using t-SNE,” J. Mach. Learn. Res., vol. 9, no. 86, pp. 2579–2605, Nov. 2008.
    [21]
    S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” in Proc. 29th Int. Conf. Neural Information Processing Systems, Montreal, Canada, 2015, pp. 91−99.
    [22]
    N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in Proc. 16th European Conf. Computer Vision, Glasgow, UK, 2020, pp. 213−229.
    [23]
    D. Xu, Y. Zhu, C. B. Choy, and L. Fei-Fei, “Scene graph generation by iterative message passing,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, Honolulu, USA, 2017, pp. 3097−3106.
    [24]
    K. Tang, H. Zhang, B. Wu, W. Luo, and W. Liu, “Learning to compose dynamic tree structures for visual contexts,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Long Beach, USA, 2019, pp. 6612−6621.
    [25]
    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. 31st Int. Conf. Neural Information Processing Systems, Long Beach, USA, 2017, pp. 6000−6010.
    [26]
    S. Khandelwal and L. Sigal, “Iterative scene graph generation,” in Proc. 36th Int. Conf. Neural Information Processing Systems, New Orleans, USA, 2022, Art. no. 1764.
    [27]
    L. Li, C. Wang, Y. Qin, W. Ji, and R. Liang, “Biased-predicate annotation identification via unbiased visual predicate representation,” in Proc. 31st ACM Int. Conf. Multimedia, Ottawa, Canada, 2023, pp. 4410−4420.
    [28]
    J. Li, Y. Wang, X. Guo, R. Yang, and W. Li, “Leveraging predicate and triplet learning for scene graph generation,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Seattle, USA, 2024, pp. 28369−28379.
    [29]
    G. Zhang, S. Kan, F. Zhang, W. Xu, Y. Zhang, and Y. Cen, “Noise-guided predicate representation extraction and diffusion-enhanced discretization for scene graph generation,” in Proc. 42nd Int. Conf. Machine Learning, Vancouver, Canada, 2025, pp. 75234−75252.
    [30]
    H. Zhou, J. Zhang, T. Luo, Y. Yang, and J. Lei, “Debiased scene graph generation for dual imbalance learning,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 4, pp. 4274–4288, Apr. 2023. doi: 10.1109/tpami.2022.3198965
    [31]
    H. Zhou, T. Luo, J. Zhang, and L. Liu, “Exploring the essence of relationships for scene graph generation via causal features enhancement network,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 47, no. 8, pp. 6616–6630, Aug. 2025. doi: 10.1109/TPAMI.2025.3559995
    [32]
    X. Lin, C. Ding, J. Zeng, and D. Tao, “GPS-Net: Graph property sensing network for scene graph generation,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Seattle, USA, 2020, pp. 3743−3752.
    [33]
    A.-A. Liu, H. Tian, N. Xu, W. Nie, Y. Zhang, and M. Kankanhalli, “Toward region-aware attention learning for scene graph generation,” IEEE Trans. Neural Netw. Learn. Syst., vol. 33, no. 12, pp. 7655–7666, Dec. 2022. doi: 10.1109/TNNLS.2021.3086066
    [34]
    M. Khademi and O. Schulte, “Deep generative probabilistic graph neural networks for scene graph generation,” in Proc. 34th AAAI Conf. Artificial Intelligence, New York, USA, 2020, pp. 11237−11245.
    [35]
    A. Desai, T.-Y. Wu, S. Tripathi, and N. Vasconcelos, “Single-stage visual relationship learning using conditional queries,” in Proc. 36th Int. Conf. Neural Information Processing Systems, New Orleans, USA, 2022, Art. no. 949.
    [36]
    C. Wang, R. Xu, Y. Huang, J. Pei, C. Huang, W. Zhu, and J. Yang, “Limited-data SAR ATR causal method via dual-invariance intervention,” IEEE Trans. Geosci. Remote Sens., vol. 63, Art. no. 5203319, Jan. 2025.
    [37]
    C. Jing, Y. Wu, X. Zhang, Y. Jia, and Q. Wu, “Overcoming language priors in VQA via decomposed linguistic representations,” in Proc. 34th AAAI Conf. Artificial Intelligence, New York, USA, 2020, pp. 11181−11188.
    [38]
    Y. Wang and T. Derr, “Tree decomposed graph neural network,” in Proc. 30th ACM Int. Conf. Information & Knowledge Management, Australia, 2021, pp. 2040−2049.
    [39]
    A. Agrawal, D. Batra, D. Parikh, and A. Kembhavi, “Don’t just assume; look and answer: Overcoming priors for visual question answering,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018, pp. 4971−4980.
    [40]
    I. Higgins, L. Matthey, A. Pal, C. Burgess, X. Glorot, M. Botvinick, S. Mohamed, and A. Lerchner, “Beta-VAE: Learning basic visual concepts with a constrained variational framework,” in Proc. 5th Int. Conf. Learning Representations, Toulon, France, 2017, pp. 1−22.
    [41]
    H. Kim and A. Mnih, “Disentangling by factorising,” in Proc. 35th Int. Conf. Machine Learning, Stockholm, Sweden, 2018, pp. 2654−2663.
    [42]
    K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask R-CNN,” in Proc. IEEE Int. Conf. Computer Vision, Venice, Italy, 2017, pp. 2980−2988.
    [43]
    J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Proc. 34th Int. Conf. Neural Information Processing Systems, Vancouver, Vancouver, 2020, Art. no. 574.
    [44]
    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al., “Learning transferable visual models from natural language supervision,” in Proc. 38th Int. Conf. Machine Learning, 2021, pp. 8748−8763.
    [45]
    J. Pennington, R. Socher, and C. D. Manning, “GloVe: Global vectors for word representation,” in Proc. Conf. Empirical Methods in Natural Language Processing, Doha, Qatar, 2014, pp. 1532−1543.
    [46]
    T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in Proc. IEEE Int. Conf. Computer Vision, Venice, Italy, 2017, pp. 2999−3007.
    [47]
    D. A. Hudson and C. D. Manning, “GQA: A new dataset for real-world visual reasoning and compositional question answering,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Long Beach, USA, 2019, pp. 6693−6702.
    [48]
    A. Kuznetsova, H. Rom, N. Alldrin, J. Uijlings, I. Krasin, J. Pont-Tuset, S. Kamali, S. Popov, M. Malloci, A. Kolesnikov, et al., “The Open Images dataset V4: Unified image classification, object detection, and visual relationship detection at scale,” Int. J. Comput. Vis., vol. 128, no. 7, pp. 1956–1981, Jul. 2020. doi: 10.1007/s11263-020-01316-z
    [49]
    X. Dong, T. Gan, X. Song, J. Wu, Y. Cheng, and L. Nie, “Stacked hybrid-attention and group collaborative learning for unbiased scene graph generation,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, New Orleans, USA, 2022, pp. 19405−19414.
    [50]
    R. Li, S. Zhang, B. Wan, and X. He, “Bipartite graph network with adaptive message passing for unbiased scene graph generation,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Nashville, USA, 2021, pp. 11104−11114.
    [51]
    J. Zhang, K. J. Shih, A. Elgammal, A. Tao, and B. Catanzaro, “Graphical contrastive losses for scene graph parsing,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Long Beach, USA, 2019, pp. 11527−11535.
    [52]
    A. Zhang, Y. Yao, Q. Chen, W. Ji, Z. Liu, M. Sun, and T.-S. Chua, “Fine-grained scene graph generation with data transfer,” in Proc. 17th European Conf. Computer Vision, Tel Aviv, Israel, 2022, pp. 409−424.
    [53]
    X. Han, X. Dong, X. Song, T. Gan, Y. Zhan, Y. Yan, and L. Nie, “Divide-and-conquer predictor for unbiased scene graph generation,” IEEE Trans. Circuits Syst. Video Technol., vol. 32, no. 12, pp. 8611–8622, Dec. 2022. doi: 10.1109/TCSVT.2022.3193857
    [54]
    K. Kim, K. Yoon, Y. In, J. Moon, D. Kim, and C. Park, “Adaptive self-training framework for fine-grained scene graph generation,” in Proc. 12th Int. Conf. Learning Representations, Vienna, Austria, 2024, pp. 1−25.
    [55]
    T. Jin, F. Guo, Q. Meng, S. Zhu, X. Xi, W. Wang, Z. Mu, and W. Song, “Fast contextual scene graph generation with unbiased context augmentation,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Vancouver, Canada, 2023, pp. 6302−6311.
    [56]
    C. Zhang, S. Stepputtis, J. Campbell, K. Sycara, and Y. Xie, “HiKER-SGG: Hierarchical knowledge enhanced robust scene graph generation,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Seattle, USA, 2024, pp. 28233−28243.
    [57]
    J. Im, J. Y. Nam, N. Park, H. Lee, and S. Park, “EGTR: Extracting graph from transformer for scene graph generation,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Seattle, USA, 2024, pp. 24229−24238.
    [58]
    M.-J. Chiou, H. Ding, H. Yan, C. Wang, R. Zimmermann, and J. Feng, “Recovering the unbiased scene graphs from the biased ones,” in Proc. 29th ACM Int. Conf. Multimedia, China, 2021, pp. 1581−1590.
    [59]
    X. Han, X. Song, X. Dong, Y. Wei, M. Liu, and L. Nie, “DBiased-P: Dual-biased predicate predictor for unbiased scene graph generation,” IEEE Trans. Multimedia, vol. 25, pp. 5319–5329, Jan. 2023. doi: 10.1109/TMM.2022.3190135
    [60]
    L. Tao, L. Mi, N. Li, X. Cheng, Y. Hu, and Z. Chen, “Predicate correlation learning for scene graph generation,” IEEE Trans. Image Process., vol. 31, pp. 4173–4185, Jun. 2022. doi: 10.1109/TIP.2022.3181511
    [61]
    L. Li, L. Chen, Y. Huang, Z. Zhang, S. Zhang, and J. Xiao, “The devil is in the labels: Noisy label correction for robust scene graph generation,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, New Orleans, USA, 2022, pp. 18847−18856.
    [62]
    Y. Deng, Y. Li, Y. Zhang, X. Xiang, J. Wang, J. Chen, and J. Ma, “Hierarchical memory learning for fine-grained scene graph generation,” in Proc. 17th European Conf. Computer Vision, Tel Aviv, Israel, 2022, pp. 266−283.
    [63]
    B. A. Biswas and Q. Ji, “Probabilistic debiasing of scene graphs,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Vancouver, Canada, 2023, pp. 10429−10438.
    [64]
    L. Li, J. Xiao, H. Shi, W. Wang, J. Shao, A.-A. Liu, Y. Yang, and L. Chen, “Label semantic knowledge distillation for unbiased scene graph generation,” IEEE Trans. Circuits Syst. Video Technol., vol. 34, no. 1, pp. 195–206, Jan. 2024. doi: 10.1109/TCSVT.2023.3282349
    [65]
    Z. Wang, X. Xu, G. Wang, Y. Yang, and H. T. Shen, “Quaternion relation embedding for scene graph generation,” IEEE Trans. Multimedia, vol. 25, pp. 8646–8656, Jan. 2023. doi: 10.1109/TMM.2023.3239229
    [66]
    G. Yang, J. Zhang, Y. Zhang, B. Wu, and Y. Yang, “Probabilistic modeling of semantic ambiguity for scene graph generation,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Nashville, USA, 2021, pp. 12522−12531.
    [67]
    A. Goel, B. Fernando, F. Keller, and H. Bilen, “Not all relations are equal: Mining informative labels for scene graph generation,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, New Orleans, USA, 2022, pp. 15575−15585.
    [68]
    Y. Guo, L. Gao, X. Wang, Y. Hu, X. Xu, X. Lu, H. T. Shen, and J. Song, “From general to specific: Informative scene graph generation via balance adjustment,” in Proc. IEEE/CVF Int. Conf. Computer Vision, Montreal, Canada, 2021, pp. 16363−16372.
    [69]
    W. Li, H. Zhang, Q. Bai, G. Zhao, N. Jiang, and X. Yuan, “PPDL: Predicate probability distribution based loss for unbiased scene graph generation,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, New Orleans, USA, 2022, pp. 19425−19434.
    [70]
    T. He, L. Gao, J. Song, and Y.-F. Li, “State-aware compositional learning toward unbiased training for scene graph generation,” IEEE Trans. Image Process., vol. 32, pp. 43–56, 2023. doi: 10.1109/TIP.2022.3224872
    [71]
    J. Wang, W. Zhang, Y. Zang, Y. Cao, J. Pang, T. Gong, K. Chen, Z. Liu, C. C. Loy, and D. Lin, “Seesaw loss for long-tailed instance segmentation,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Nashville, USA, 2021, pp. 9690−9699.
    [72]
    C. Chen, Y. Zhan, B. Yu, L. Liu, Y. Luo, and B. Du, “Resistance training using prior bias: Toward unbiased scene graph generation,” in Proc. 36th AAAI Conf. Artificial Intelligence, 2022, pp. 212−220.
    [73]
    J. Tan, C. Wang, B. Li, Q. Li, W. Ouyang, C. Yin, and J. Yan, “Equalization loss for long-tailed object recognition,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Seattle, USA, 2020, pp. 11659−11668.
    [74]
    D. Liu, M. Bober, and J. Kittler, “Neural belief propagation for scene graph generation,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 8, pp. 10161–10172, Aug. 2023. doi: 10.1109/TPAMI.2023.3243306
    [75]
    D. Liu, M. Bober, and J. Kittler, “Constrained structure learning for scene graph generation,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 10, pp. 11588–11599, Oct. 2023. doi: 10.1109/TPAMI.2023.3282889
    [76]
    Y. Li, W. Ouyang, B. Zhou, K. Wang, and X. Wang, “Scene graph generation from objects, phrases and region captions,” in Proc. IEEE Int. Conf. Computer Vision, Venice, Italy, 2017, pp. 1270−1279.
    [77]
    C. Zheng, L. Gao, X. Lyu, P. Zeng, A. El Saddik, and H. T. Shen, “Dual-branch hybrid learning network for unbiased scene graph generation,” IEEE Trans. Circuits Syst. Video Technol., vol. 34, no. 3, pp. 1743–1756, Mar. 2024. doi: 10.1109/TCSVT.2023.3297842
    [78]
    H. Zhang, Z. Kyaw, S.-F. Chang, and T.-S. Chua, “Visual translation embedding network for visual relation detection,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, Honolulu, USA, 2017, pp. 3107−3115.
    [79]
    J. Yang, J. Lu, S. Lee, D. Batra, and D. Parikh, “Graph R-CNN for scene graph generation,” in Proc. 15th European Conf. Computer Vision, Munich, Germany, 2018, pp. 690−706.
    [80]
    T. He, T. Wu, D. Zhang, G. Duan, K. Qin, and Y.-F. Li, “Towards lifelong scene graph generation with knowledge-ware in-context prompt learning,” arXiv preprint arXiv: 2401.14626, 2024.
    [81]
    M. Suhail, A. Mittal, B. Siddiquie, C. Broaddus, J. Eledath, G. Medioni, and L. Sigal, “Energy-based learning for scene graph generation,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Nashville, USA, 2021, pp. 13931−13940.
    [82]
    X. Lyu, L. Gao, Y. Guo, Z. Zhao, H. Huang, H. T. Shen, and J. Song, “Fine-grained predicates learning for scene graph generation,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, New Orleans, USA, 2022, pp. 19445−19453.
    [83]
    S. Khandelwal, M. Suhail, and L. Sigal, “Segmentation-grounded scene graph generation,” in Proc. IEEE/CVF Int. Conf. Computer Vision, Montreal, Canada, 2021, pp. 15859−15869.
    [84]
    T. Chen, W. Yu, R. Chen, and L. Lin, “Knowledge-embedded routing network for scene graph generation,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Long Beach, USA, 2019, pp. 6156−6164.
    [85]
    A. Zareian, S. Karaman, and S.-F. Chang, “Bridging knowledge graphs to generate scene graphs,” in Proc. 16th European Conf. Computer Vision, Glasgow, UK, 2020, pp. 606−623.
    [86]
    T.-J. J. Wang, S. Pehlivan, and J. Laaksonen, “Tackling the unannotated: Scene graph generation with bias-reduced models,” in Proc. 31st British Machine Vision Conf., UK, 2020, pp. 1−13.
    [87]
    D. Chen, X. Liang, Y. Wang, and W. Gao, “Soft transfer learning via gradient diagnosis for visual relationship detection,” in Proc. IEEE Winter Conf. Applications of Computer Vision, Waikoloa, USA, 2019, pp. 1118−1126.
    [88]
    H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in Proc. 37th Int. Conf. Neural Information Processing Systems, New Orleans, USA, 2023, Art. no. 1516.
    [89]
    S. Liu, H. Cheng, H. Liu, H. Zhang, F. Li, T. Ren, X. Zou, J. Yang, H. Su, J. Zhu, et al., “LLaVA-Plus: Learning to use tools for creating multimodal agents,” in Proc. 18th European Conf. Computer Vision, Milan, Italy, 2025, pp. 126−142.
    [90]
    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. Conf. North American Chapter of the Association for Comput. Linguistics: Human Language Tech., Minneapolis, USA, 2019, vol. 1, pp. 4171−4186.
    [91]
    J. Li, D. Li, S. Savarese, and S. Hoi, “BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” in Proc. 40th Int. Conf. Machine Learning, Honolulu, USA, 2023, pp. 19730−19742.

Catalog

    通讯作者: 陈斌, bchen63@163.com
    • 1. 

      沈阳化工大学材料科学与工程学院 沈阳 110142

    1. 本站搜索
    2. 百度学术搜索
    3. 万方数据库搜索
    4. CNKI搜索

    Figures(12)  / Tables(9)

    Article Metrics

    Article views (265) PDF downloads(10) Cited by()

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return