热点
物联网模块化开发,重塑移动互联生态,reasoning_content:我们要求以模块开发者的口吻,写一个与“技术、科技”相关,关于[数码新势力崛起:物联网引领移动互联生态革新]的标题要求直接输出一个标题,不要加说明提示等信息;简短精炼,字数控制在30字以内nn需要模拟模块开发者的口吻,可能更偏向技术、底层、模块化、生态等词汇可以考虑模块化驱动、组件化、嵌入式等标题要体现物联网引领移动互联生态革新nn可能的模块化物联网:重塑移动互联生态 但字数11字,有点简单或者嵌入式模块崛起:物联网革新移动互联 13字或者更直接:模块开发者视角:物联网生态重塑移动互联 但字数较多nn严格30字以内,可以设计:模块化技术引爆物联网,移动互联生态焕新 15字或者物联网模块化开发:移动互联生态的革新引擎 16字但注意口吻是模块开发者,可以加入我们或开发者nn例如:作为模块开发者,我断言:物联网正重构移动互联生态 但字数超了简化:模块开发者说:物联网重构移动互联生态 14字或者更技术:模块化物联网组件:驱动移动互联生态革新 14字nn考虑直接输出,不要加引号我选择:模块化物联网:移动互联生态的革新引擎 14字或者嵌入式模块崛起:物联网引领移动互联革新 14字nn为了更符合模块开发者口吻,可以加入开发:模块开发视角:物联网革新移动互联生态 14字nn最终确定一个简洁有力的物联网模块化开发,重塑移动互联生态 13字
建站资源多端适配全流程技术指南,reasoning_content:我们要求以技术文档工程师的口吻,写一个与“技术、科技”相关,关于[全流程策划:轻松实现建站资源多端无缝适配]的标题直接输出一个标题,不要加说明提示等信息简短精炼,字数控制在30字以内nn注意:原题中有一个[全流程策划:轻松实现建站资源多端无缝适配]似乎是给出的主题或示例,我们需要基于这个内容写标题但要求是“写一个与‘技术、科技’相关,关于‘[全流程策划:轻松实现建站资源多端无缝适配]’的标题”可能意思是标题要呼应这个主题nn技术文档工程师的口吻:通常专业、准确、简洁,可能带有“指南”、“手册”、“方案”、“技术实现”等词汇标题要体现“全流程策划”、“多端无缝适配”nn尝试几个:比如“建站资源多端适配全流程技术指南”,但字数30字内或者“多端无缝适配建站资源策划方案”需要更精炼nn考虑:全流程策划:轻松实现建站资源多端无缝适配 -> 提炼为“建站资源多端适配全流程策划”但不够技术文档风格技术文档常见“技术白皮书”、“实现方案”、“操作手册”等但标题要求直接输出,不要说明nn可能的最佳: “多端适配建站资源全流程技术实现” 共13字或者“轻松实现建站资源多端无缝适配指南” 共14字但“轻松”可能不太技术文档,技术文档更客观但原主题里有“轻松”,可保留nn另一种: “建站资源多端无缝适配全流程方案” 13字或者“全流程策划:建站资源多端无缝适配” 15字,但冒号有点不简洁nn由于要求“简短精炼”,直接给出一个标题即可我决定输出:“建站资源多端适配全流程技术指南” 14字
服务器搜索优化漏洞排查与索引修复实战手册,reasoning_content:我们要求以前端站长的口吻,写一个与技术、科技相关,关于[服务器搜索优化:漏洞排查与索引修复实战手册]的标题要求直接输出一个标题,不要加说明提示等信息;简短精炼,字数控制在30字以内nn注意:口吻是前端站长,所以可能带有一些站长、运维、前端技术人的语气主题是服务器搜索优化,漏洞排查与索引修复实战手册需要精炼nn可能的比如服务器搜索优化漏洞排查与索引修复实战 但字数可能超或者搜索优化漏洞排查索引修复实战手册 但缺少服务器nn考虑:站长必看:服务器搜索优化漏洞排查与索引修复 字数?站长必看:服务器搜索优化漏洞排查与索引修复 共15字?必看有点广告或者实战:服务器搜索优化漏洞排查与索引修复 共16字或者服务器搜索优化:漏洞排查与索引修复实战 共16字注意要求30字以内,所以有很多空间nn更简洁:搜索优化漏洞排查与索引修复实战 13字但没提服务器可以加服务器:服务器搜索优化漏洞排查与索引修复实战 16字或者服务器搜索优化:漏洞排查与索引修复手册 16字但要求实战手册,所以最好包含实战nn最直接:服务器搜索优化漏洞排查与索引修复实战手册 18字完美符合口吻是前端站长,这个标题中性,可以作为技术文章标题不需要额外加前端站长字样,因为是口吻,但标题本身可以不需要出现站长nn考虑到前端站长的口吻可能更偏向于个人经验分享,可以加我如何或站长手记等但要求简短精炼,30字以内我如何做服务器搜索优化漏洞排查与索引修复 太长nn我认为最简单直接的标题就是服务器搜索优化:漏洞排查与索引修复实战手册 共18字或者服务器搜索优化漏洞排查与索引修复实战 16字哪个更好?前者有冒号,后者没有都可以nn为了更符合前端站长口吻,可以稍微口语化一点:手把手:服务器搜索优化漏洞排查与索引修复 但字数18或者站长实战:服务器搜索优化漏洞排查与索引修复 18字nn我选择一个:直接服务器搜索优化漏洞排查与索引修复实战手册输出
16 9 月 2026, 周三

挖掘DBLP作者合作关系,FP-Growth算法实践(5):挖掘研究者合作

副标题#e#

就是频繁项集挖掘,FP-Growth算法。

先产生headerTable:

数据结构(其实也是调了好几次代码才确定的,因为一开始总有想不到的东西):entry: entry: {authorName: frequence,firstChildPointer,startYear,endYear}

def CreateHeaderTable(tranDB,minSupport=1):
    headerTable={} #entry: entry: {authorName: frequence,endYear}
    authorDB={} #entry: {frozenset(authorListSet): frequence}
    for i,(conf,year,authorList) in enumerate(tranDB):
        authorListSet=set([])
        print "trans",i,"=="*20
        if conf is np.nan or year is np.nan or authorList is np.nan:
            continue #for tranDB[2426,:]
        for author in authorList.split("|"):
            authorListSet.add(author)
            if headerTable.has_key(author):
                headerTable[author][0]+=1
                if year<headerTable[author][2]:
                    headerTable[author][2]=year
                elif year>headerTable[author][3]:
                    headerTable[author][3]=year
            else:
                headerTable[author]=[1,None,year]
        if authorDB.has_key(frozenset(authorListSet)):
            authorDB[frozenset(authorListSet)]+=1
        else:
            authorDB[frozenset(authorListSet)]=1
    for author in headerTable.keys():
        if headerTable[author][0]<minSupport:
            del headerTable[author]
    return headerTable,authorDB

再构建FP-Tree:

每个treeNode又五元组来描述:

class TREE_NODE:
    def __init__(self,authorName,frequence,parentPointer):
        self.authorName=authorName
        self.frequence=frequence
        self.parentPointer=parentPointer #parent TREE_NODE
        self.childrenPointer={} #children TREE_NODEs dict
        self.brotherPointer=None #brother TREE_NODE

注意,每次加入节点,这个五元组的每一项是否都考虑了;

如果是新节点,是否考虑链接到headerTable了;

对,只有这两点;代码如下:

'''
headerTable={} #entry: {authorName: frequence,endYear}
if we want to call UpdateCondTree(),then headerTable==>condHeaderTable={} #entry: {authorName: frequence,firstChildPointer}
#TREE_NODE: (self,parentPointer,childrenPointer={},brotherPointer=None)
'''
def UpdateTree(authorsList,treeNode,headerTable,frequence):
    if treeNode.childrenPointer.has_key(authorsList[0]):
        treeNode.childrenPointer[authorsList[0]].frequence+=frequence #isn't +1
    else: #add authorsList[0] as a new child of treeNode
        treeNode.childrenPointer[authorsList[0]]=TREE_NODE(authorsList[0],treeNode)
        if headerTable[authorsList[0]][1]==None: #[1] is the firstChildPointer
            headerTable[authorsList[0]][1]=treeNode.childrenPointer[authorsList[0]]
        else:
            firstChildPointer=headerTable[authorsList[0]][1]
            tempAuthorNode=firstChildPointer
            while tempAuthorNode.brotherPointer is not None:
                tempAuthorNode=tempAuthorNode.brotherPointer
            tempAuthorNode.brotherPointer=treeNode.childrenPointer[authorsList[0]]
    if len(authorsList)>1: #recursively call UpdateTree() with authorsList[1:]
        #UpdateTree(authorsList[1:],frequence)
        UpdateTree(authorsList[1:],treeNode.childrenPointer[authorsList[0]],frequence) #do care this!

'''
headerTable={} #entry: {authorName: frequence,endYear}
if we want to call CreateCondTree(),firstChildPointer}
authorDB={} #entry: {frozenset(authorListSet): frequence}
#TREE_NODE: (self,brotherPointer=None)
'''
def CreateTree(authorDB,headerTable): #same function as CreateCondTree()
    treeRoot=TREE_NODE("NULL",None)#root Node of FPtree
    for i,(authorListSet,frequence) in enumerate(authorDB.items()):
        print "authorListSet","=="*20
        tempDict={}
        for author in authorListSet:
            if headerTable.has_key(author):
                tempDict[author]=headerTable[author][0]
        if len(tempDict)>0:
            tempList=sorted(tempDict.items(),key=lambda x:x[1],reverse=True)
            authorsList=[author for author,count in tempList]
            UpdateTree(authorsList,treeRoot,frequence)
    return treeRoot

注意,每一步验证一下:

    #secondly,create the FP-Tree(the second pass)
    treeRoot=CreateTree(authorDB,headerTable)
    headerTable["Ying Wu"][1].authorName #Ying Wu
    headerTable["Ying Wu"][1].frequence #47,so 4=51-47 in other brotherPointer TREE_NODE
    print len(treeRoot.childrenPointer) #4028 < 7318=len(headerTable)
    print treeRoot.childrenPointer[treeRoot.childrenPointer.keys()[0]].authorName #Linli Xu
    print treeRoot.childrenPointer[treeRoot.childrenPointer.keys()[0]].frequence #2 < 12=headerTable["Linli Xu"][0],so 10=12-2 in other brotherPointer TREE_NODE
    print treeRoot.childrenPointer[treeRoot.childrenPointer.keys()[0]].parentPointer.authorName #NULL,that's the root!
 

挖掘FP-Tree,找出频繁项:

思路:从支持度最小的单项集出发,依次递归挖掘支持度更大的单项集。

对每一个单项集的挖掘,思路如下:

找headerTable的每一个孩子结点,递归寻找这个孩子结点的条件子树;

然后合并每个孩子结点找到的条件子树,构成一颗关于该单项集的条件子树,递归挖掘该单项集的条件子树,从而形成两个、三个、四个、、、项集的频繁模式;

对所有生成的这些两个、三个、四个、、、项集的频繁模式,递归上面的过程即可。

#p#副标题#e##p#分页标题#e#

def FindParentTreeNodes(baseTreeNode):
    parentTreeNodes=[]
    while baseTreeNode.parentPointer is not None: #while baseTreeNode is not the ROOT node whose parentPointer is None and authorName is "NULL"
        parentTreeNodes.append(baseTreeNode.authorName)
        baseTreeNode=baseTreeNode.parentPointer
    return parentTreeNodes

def FindCondAuthorDB(firstChildPointerTreeNode):
    condAuthorDB={} #entry: {frozenset(authorListSet): frequence}
    tempTreeNode=firstChildPointerTreeNode
    while tempTreeNode is not None:
        parentTreeNodes=FindParentTreeNodes(tempTreeNode)
        if len(parentTreeNodes)>1:
            condAuthorDB[frozenset(parentTreeNodes[1:])]=tempTreeNode.frequence
            #parentTreeNodes[1:],remove self treeNode
        tempTreeNode=tempTreeNode.brotherPointer
    return condAuthorDB

def CreateCondHeaderTable(condAuthorDB,minSupport=1):
    condHeaderTable={} #entry: {authorName: frequence,firstChildPointer}
    for i,frequence) in enumerate(condAuthorDB.items()):
        print "cond trans","=="*20
        for author in authorListSet:
            if condHeaderTable.has_key(author):
                headerTable[author][0]+=frequence
            else:
                condHeaderTable[author]=[frequence,None]
    for author in condHeaderTable.keys():
        if condHeaderTable[author][0]<minSupport:
            del condHeaderTable[author]
    return condHeaderTable

'''
headerTable={} #entry: {authorName: frequence,endYear}
if we want to call MineCondTree(),brotherPointer=None)
'''    
def MineTree(treeRoot,minSupport=1,baseFreqAuthorSet=set([]),finalFreqAuthorPattDict={}):
    sortedAuthorsList=[value[0] for value in sorted(headerTable.items(),reverse=False)]
    for baseAuthor in sortedAuthorsList: #start from bottom of headerTable
        newFreqAuthorSet=baseFreqAuthorSet.copy()
        newFreqAuthorSet.add(baseAuthor)
        finalFreqAuthorPattDict[frozenset(newFreqAuthorSet)]=headerTable[baseAuthor][0]
        condAuthorDB=FindCondAuthorDB(headerTable[baseAuthor][1]) #[1] is the firstChildPointer
        condHeaderTable=CreateCondHeaderTable(condAuthorDB,minSupport)
        condTreeRoot=CreateTree(condAuthorDB,condHeaderTable) #call CreateCondTree()
        if condHeaderTable is not None:
            MineTree(condTreeRoot,condHeaderTable,minSupport,newFreqAuthorSet,finalFreqAuthorPattDict) #call MineCondTree()

dawei

【声明】:商丘站长网内容转载自互联网,其相关言论仅代表作者个人观点绝非权威,不代表本站立场。如您发现内容存在版权问题,请提交相关链接至邮箱:bqsm@foxmail.com,我们将及时予以处理。

您错过了

物联网模块化开发,重塑移动互联生态,reasoning_content:我们要求以模块开发者的口吻,写一个与“技术、科技”相关,关于[数码新势力崛起:物联网引领移动互联生态革新]的标题要求直接输出一个标题,不要加说明提示等信息;简短精炼,字数控制在30字以内nn需要模拟模块开发者的口吻,可能更偏向技术、底层、模块化、生态等词汇可以考虑模块化驱动、组件化、嵌入式等标题要体现物联网引领移动互联生态革新nn可能的模块化物联网:重塑移动互联生态 但字数11字,有点简单或者嵌入式模块崛起:物联网革新移动互联 13字或者更直接:模块开发者视角:物联网生态重塑移动互联 但字数较多nn严格30字以内,可以设计:模块化技术引爆物联网,移动互联生态焕新 15字或者物联网模块化开发:移动互联生态的革新引擎 16字但注意口吻是模块开发者,可以加入我们或开发者nn例如:作为模块开发者,我断言:物联网正重构移动互联生态 但字数超了简化:模块开发者说:物联网重构移动互联生态 14字或者更技术:模块化物联网组件:驱动移动互联生态革新 14字nn考虑直接输出,不要加引号我选择:模块化物联网:移动互联生态的革新引擎 14字或者嵌入式模块崛起:物联网引领移动互联革新 14字nn为了更符合模块开发者口吻,可以加入开发:模块开发视角:物联网革新移动互联生态 14字nn最终确定一个简洁有力的物联网模块化开发,重塑移动互联生态 13字

建站资源多端适配全流程技术指南,reasoning_content:我们要求以技术文档工程师的口吻,写一个与“技术、科技”相关,关于[全流程策划:轻松实现建站资源多端无缝适配]的标题直接输出一个标题,不要加说明提示等信息简短精炼,字数控制在30字以内nn注意:原题中有一个[全流程策划:轻松实现建站资源多端无缝适配]似乎是给出的主题或示例,我们需要基于这个内容写标题但要求是“写一个与‘技术、科技’相关,关于‘[全流程策划:轻松实现建站资源多端无缝适配]’的标题”可能意思是标题要呼应这个主题nn技术文档工程师的口吻:通常专业、准确、简洁,可能带有“指南”、“手册”、“方案”、“技术实现”等词汇标题要体现“全流程策划”、“多端无缝适配”nn尝试几个:比如“建站资源多端适配全流程技术指南”,但字数30字内或者“多端无缝适配建站资源策划方案”需要更精炼nn考虑:全流程策划:轻松实现建站资源多端无缝适配 -> 提炼为“建站资源多端适配全流程策划”但不够技术文档风格技术文档常见“技术白皮书”、“实现方案”、“操作手册”等但标题要求直接输出,不要说明nn可能的最佳: “多端适配建站资源全流程技术实现” 共13字或者“轻松实现建站资源多端无缝适配指南” 共14字但“轻松”可能不太技术文档,技术文档更客观但原主题里有“轻松”,可保留nn另一种: “建站资源多端无缝适配全流程方案” 13字或者“全流程策划:建站资源多端无缝适配” 15字,但冒号有点不简洁nn由于要求“简短精炼”,直接给出一个标题即可我决定输出:“建站资源多端适配全流程技术指南” 14字